Novel cas13b orthologs crispr enzymes and systems
Patent Information
- Application Number
- JP2023045024
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-10-02
- Filing Date
- 2023-03-22
- Publication Date
- 2025-07-25
AI Technical Summary
Existing CRISPR-Cas systems lack efficient and specific RNA-targeting tools, particularly for RNA interference and editing, due to the high diversity and complexity of CRISPR-Cas systems, limiting their application in gene expression control and genome engineering.
Development of novel Cas13b orthologs and engineered compositions, including CRISPR complexes with enhanced or suppressed activity through accessory proteins like csx27 and csx28, and optimized guide sequences for targeted RNA binding and editing.
The novel Cas13b orthologs and engineered systems provide precise and efficient RNA targeting and editing capabilities, enabling effective gene expression control and genome engineering in both prokaryotic and eukaryotic cells, with reduced off-target effects.
Smart Images

Figure 00000329_0000 
Figure 00000329_0001 
Figure 00000330_0000
Abstract
Description
Technical field
[0001] Related Applications and Incorporation by Reference Priority is claimed to U.S. Provisional Application No. 62 / 471,710 filed March 15, 2017 and U.S. Provisional Application No. 62 / 566,829 filed October 2, 2017.
[0002] Reference is made to the PCT applications, including application PCT / US Patent Application Publication No. 2016 / 058302, filed October 21, 2016, specifically designating the United States. U.S. Provisional Patent Application No. 62 / 245,270, filed October 22, 2015; U.S. Provisional Patent Application No. 62 / 296,548, filed February 17, 2016; See U.S. Provisional Patent Application Nos. 62 / 376,367 and 62 / 376,382 filed in U.S.A. See also U.S. Provisional Patent Application No. 62 / 471,792 filed March 15, 2017 and U.S. Provisional Patent Application No. 62 / 484,786 filed April 12, 2017. Smargon et al.(2017),“Cas13b Is a Type VI-B CRISPR-Associated RNA-Guided RNase Differentially Regulated by Accessory Proteins Csx27 and Csx28”,Molecular Cell 65,618-630(Feb.16,2017)doi:10.1016 / j .molcel.2016.12.023.Epub Jan 5,2017 and Smargon et al.(2017),“Cas13b Is a Type VI-B CRISPR-Associated RNA-Guided RNase Differentially Regulated by Accessory Proteins Csx27 and Csx28”,bioRxiv 092577;doi : https: / / doi.org / 10.1101 / 092577.Posted December 9,2017. Each of the aforementioned applications and literature citations is hereby incorporated herein by reference.
[0003] Indeed, all documents cited or referenced in this specification and documents cited herein are in any document mentioned herein or incorporated by reference herein. Any manufacturer's instructions, descriptions, product specifications and product sheets for any product, together with which are hereby incorporated by reference herein, may be used in the practice of the present invention. More specifically, all documents referenced are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.
[0004] STATEMENT ON FEDERALLY SPONSORED RESEARCH This invention was made with federal support under Grant Nos. MH100706 and MH110049 awarded by the National Institutes of Health. The federal government has certain rights in this invention.
[0005] The present invention generally relates to the control of gene expression, including sequence targeting, which can use vector systems for clustered regularly spaced short palindromic repeats (CRISPR) and their components, such as perturbation of gene transcripts or nucleic acid editing. It relates to systems, methods and compositions for use in [Background technology]
[0006] Bacterial and archaeal CRISPR-CRISPR-associated (Cas) adaptive immune systems are some such systems that exhibit extremely high diversity in protein composition and genomic locus organization. The CRISPR-Cas system loci have over 50 gene families and no strictly universal genes, suggesting rapid evolution and a very high degree of diversity in locus organization. To date, approximately 395 profiles of cas genes for 93 Cas proteins have been comprehensively identified by taking a multidirectional approach. Included in the classification is the signature of the signature gene profile + locus organization. A new classification of CRISPR-Cas systems has been proposed, in which these systems broadly fall into two classes, class 1 with multi-subunit effector complexes and single-subunit effector modules exemplified in the Cas9 protein. are divided into classes 2 and 2 with Novel effector proteins associated with class 2 CRISPR-Cas systems can be developed as powerful genome engineering tools, and prediction of putative novel effector proteins and their engineering and optimization are important. Novel Cas13b orthologs and uses thereof are desirable.
[0007] Citation or identification of any document in this application is not an admission that such document is available as prior art to the present invention. [Outline of the invention] [Means for solving the problem]
[0008] Effector proteins comprise two subgroups, types VI-B1 and VI-B2, which include members that are RNA-programmable nucleases, RNA-interfering, and can participate in bacterial adoptive immunity against RNA phages. . The Cas13b system contains a single large effector (approximately 1100 amino acids long) and one or neither of two small putative accessory proteins (approximately 200 amino acids long) in the vicinity of the CRISPR array. can be Based on this nearby small protein, the system divides into two loci, A and B. Up to 25 kilobase pairs upstream or downstream from the array, no additional proteins are conserved between species with each locus. With few exceptions, the CRISPR array contains direct repeat sequences 36 nucleotides long and spacer sequences 30 nucleotides long. Direct repeats are generally well conserved, especially at the ends, where the 5'-terminal GTTG / GUUG is reverse complementary to the 3'-terminal CAAC. This conservation suggests strong base pairing towards RNA loop structures that may interact with proteins within the locus. Motif searches complementary to direct repeats did not reveal candidate tracrRNAs near the array, possibly representing a single crRNA such as that found in the Cpf1 locus.
[0009] In embodiments of the invention, the Type VI-B system comprises a novel Cas13b effector protein and optionally small accessory proteins encoded upstream or downstream of the Cas13b effector protein. In certain embodiments, small accessory proteins enhance the ability of Cas13b effectors to target RNA.
[0010] The present invention provides non-naturally occurring or engineered compositions comprising:
[0011] i) certain novel Cas13b effector proteins, and
[0012] ii) crRNAs,
[0013] wherein the crRNA comprises a) a guide sequence capable of hybridizing to a target RNA sequence and b) a direct repeat sequence,
[0014] A CRISPR complex is thereby formed containing the Cas13b effector protein complexed with a guide sequence that is hybridized to the target RNA sequence. Complexes may be formed in vitro or ex vivo and introduced into cells, or contacted with RNA, or formed in vivo.
[0015] In some embodiments, the non-naturally occurring or engineered compositions of the invention can include accessory proteins that enhance type VI-B CRISPR-Cas effector protein activity.
[0016] In certain such embodiments, the accessory protein that enhances Cas13b effector protein activity is csx28 protein. In such embodiments, the Type VI-B CRISPR-Cas effector protein and the Type VI-B CRISPR-Cas accessory protein may be from the same source or from different sources.
[0017] In some embodiments, the non-naturally occurring or engineered compositions of the invention comprise accessory proteins that suppress Cas13b effector protein activity.
[0018] In certain such embodiments, the accessory protein that suppresses Cas13b effector protein activity is the csx27 protein. In such embodiments, the Type VI-B CRISPR-Cas effector protein and the Type VI-B CRISPR-Cas accessory protein may be from the same source or from different sources. In certain embodiments of the invention, the Type VI-B CRISPR-Cas effector protein is from Table 1. In certain embodiments, the Type VI-B CRISPR-Cas accessory protein is from Table 1.
[0019] In some embodiments, the non-naturally occurring or engineered compositions of the invention comprise two or more crRNAs.
[0020] In some embodiments, the non-naturally occurring or engineered compositions of the invention comprise a guide sequence that hybridizes to a prokaryotic target RNA sequence.
[0021] In some embodiments, the non-naturally occurring or engineered compositions of the invention comprise a guide sequence that hybridizes to a eukaryotic target RNA sequence.
[0022] In some embodiments, the Cas13b effector protein comprises one or more nuclear localization signals (NLS).
[0023] A Cas13b effector protein of the invention is, is, or comprises, consists essentially of, or consists of, or is such a protein from or as shown in Table 1. to be concerned with or to relate to The invention provides, relates to, involves, or uses a protein from or as set forth in Table 1, including mutations or modifications thereof as set forth herein. intended to comprise or consist essentially of or consist of; Table 1 Cas13b effector proteins are discussed in further detail herein in conjunction with Table 1.
[0024] In some embodiments of the non-naturally occurring or engineered compositions of the invention, the Cas13b effector protein is associated with one or more functional domains. Association can be by direct linking of the effector protein to the functional domain or by association with crRNA. In a non-limiting example, a crRNA contains additions or insertions of sequences that can associate with a functional domain of interest, including, for example, nucleotides that bind to aptamers or nucleic acid binding adapter proteins.
[0025] In certain non-limiting embodiments, the non-naturally occurring or engineered compositions of the invention comprise functional domains that cleave target RNA sequences.
[0026] In certain non-limiting embodiments, the non-naturally occurring or engineered compositions of the invention comprise functional domains that alter the transcription or translation of target RNA sequences.
[0027] In some embodiments of the compositions of the invention, the Cas13b effector protein is associated with one or more functional domains and the effector protein comprises one or more mutations within the HEPN domain, thereby , the complex can deliver epigenetic modifiers or transcriptional or translational activation or repression signals. Complexes may be formed in vitro or ex vivo and introduced into cells, or contacted with RNA, or formed in vivo.
[0028] In some embodiments of the non-naturally occurring or engineered compositions of the invention, the Cas13b effector protein and accessory protein are derived from the same organism.
[0029] In some embodiments of the non-naturally occurring or engineered compositions of the invention, the Cas13b effector protein and accessory protein are derived from different organisms.
[0030] The present invention provides a type VI-B CRISPR-Cas vector system, a first regulatory element operably linked to a nucleotide sequence encoding a Cas13b effector protein; and a second regulatory element operably linked to a nucleotide sequence encoding crRNA Also provided is a Type VI-B CRISPR-Cas vector system comprising one or more vectors comprising
[0031] In certain embodiments, the vector system of the invention further comprises a regulatory element operably linked to the nucleotide sequence of the type VI-B CRISPR-Cas accessory protein.
[0032] Where appropriate, the nucleotide sequence encoding the type VI-B CRISPR-Cas effector protein and / or the nucleotide sequence encoding the type VI-B CRISPR-Cas accessory protein is codon-optimized for expression in eukaryotic cells.
[0033] In some embodiments of the vector systems of the invention, the nucleotide sequences encoding the Cas13b effector and accessory proteins are codon-optimized for expression in eukaryotic cells.
[0034] In some embodiments, the vector system of the invention comprises a single vector.
[0035] In some embodiments of the vector system of the invention, the one or more vectors comprises a viral vector.
[0036] In some embodiments of the vector system of the invention, the one or more vectors comprise one or more retroviral, lentiviral, adenoviral, adeno-associated or herpes simplex viral vectors.
[0037] The present invention i) a Cas13b effector protein, and ii) crRNAs, A delivery system configured to deliver a Cas13b effector protein and one or more nucleic acid components of a non-naturally occurring or engineered composition comprising The crRNA contains a) a guide sequence that hybridizes to the target RNA sequence of the cell and b) a direct repeat sequence, The Cas13b effector protein forms a complex with crRNA and the guide sequence directs sequence-specific binding to the target RNA sequence; A delivery system is provided whereby a CRISPR complex is formed comprising a Cas13b effector protein complexed with a guide sequence hybridized to a target RNA sequence. Complexes may be formed in vitro or ex vivo and introduced into cells, or contacted with RNA, or formed in vivo.
[0038] In some embodiments of the delivery system of the invention, the system comprises one or more vectors or one or more polynucleotide molecules, wherein the one or more vectors or polynucleotide molecules are non-naturally occurring or It comprises one or more polynucleotide molecules encoding the Cas13b effector protein and one or more nucleic acid components of the engineered composition.
[0039] In some embodiments, the delivery systems of the invention comprise delivery vehicles comprising liposomes, particles, exosomes, microvesicles, gene guns or one or more viral vectors.
[0040] In some embodiments, the non-naturally occurring or engineered compositions of the invention are for use in therapeutic treatment methods or research programmes.
[0041] In some embodiments, the non-naturally occurring or engineered vector systems of the invention are for use in therapeutic treatment methods or research programmes.
[0042] In some embodiments, the non-naturally occurring or engineered delivery systems of the invention are for use in therapeutic treatment methods or research programmes.
[0043] The present invention provides a method of altering the expression of a target gene of interest, comprising: i) a Cas13b effector protein, and ii) crRNAs, comprising contacting with one or more non-naturally occurring or engineered compositions comprising The crRNA contains a) a guide sequence that hybridizes to the target RNA sequence of the cell and b) a direct repeat sequence, The Cas13b effector protein forms a complex with crRNA and the guide sequence directs sequence-specific binding to the target RNA sequence in the cell; thereby forming a CRISPR complex comprising the Cas13b effector protein complexed with a guide sequence hybridized to the target RNA sequence, Methods are provided whereby the expression of a target locus of interest is altered. Complexes may be formed in vitro or ex vivo and introduced into cells, or contacted with RNA, or formed in vivo.
[0044] In some embodiments, the method of altering expression of a target gene of interest further comprises contacting the target RNA with an accessory protein that enhances Cas13b effector protein activity.
[0045] In some embodiments of methods of altering expression of a target gene of interest, the accessory protein that enhances Cas13b effector protein activity is csx28 protein.
[0046] In some embodiments, the method of altering expression of a target gene of interest further comprises contacting the target RNA with an accessory protein that suppresses Cas13b effector protein activity.
[0047] In some embodiments of methods of altering expression of a target gene of interest, the accessory protein that suppresses Cas13b effector protein activity is csx27 protein.
[0048] In some embodiments, the method of altering expression of a target gene of interest comprises cleaving the target RNA.
[0049] In some embodiments, methods of altering expression of a target gene of interest comprise increasing or decreasing expression of a target RNA.
[0050] In some embodiments of methods of altering expression of a target gene of interest, the target gene is in a prokaryotic cell.
[0051] In some embodiments of methods of altering expression of a target gene of interest, the target gene is in a eukaryotic cell.
[0052] The invention provides cells comprising a modified target of interest, wherein the target of interest has been modified by any of the methods disclosed herein.
[0053] In some embodiments of the invention, the cells are prokaryotic cells.
[0054] In some embodiments of the invention, the cells are eukaryotic cells.
[0055] In some embodiments, modification of a target of interest in a cell comprises a cell containing altered expression of at least one gene product; a cell comprising altered expression of at least one gene product, wherein expression of the at least one gene product is increased, or A cell comprising altered expression of at least one gene product, wherein expression of the at least one gene product is decreased bring.
[0056] In some embodiments, the cells are mammalian or human cells.
[0057] The invention provides cell lines of or containing the cells disclosed herein or cells or progeny thereof modified by any of the methods disclosed herein.
[0058] The invention provides multicellular organisms comprising one or more cells disclosed herein or one or more cells modified by any of the methods disclosed herein.
[0059] The invention provides plant or animal models comprising one or more cells disclosed herein or one or more cells modified by any of the methods disclosed herein.
[0060] The invention provides gene products from the cells, or cell lines, or organisms, or plant or animal models disclosed herein.
[0061] In some embodiments, the amount of gene product expressed is greater or less than the amount of gene product from cells that do not have altered expression.
[0062] The present invention provides an isolated Cas13b effector protein comprising, consisting essentially of, or consisting of, or as set forth in Table 1. The Cas13b effector proteins of Table 1 are as discussed in further detail herein in conjunction with Table 1. The invention provides an isolated nucleic acid encoding a Cas13b effector protein. In some embodiments of the invention, the isolated nucleic acid comprises a DNA sequence and further comprises a sequence encoding crRNA. The invention provides an isolated eukaryotic cell containing a nucleic acid encoding a Cas13b effector protein. Thus, as used herein, "Cas13b effector protein" or "effector protein" or "Cas" or "Cas protein" or "RNA targeting effector protein" or "RNA targeting protein" or similar expressions are , should be understood in relation to Table 1 and can be read as Cas13b effector protein in Table 1; expressions such as "RNA-targeting CRISPR system" should be understood in relation to Table 1, Table 1 can be read as Cas13b effector protein CRISPR system; and references to guide RNAs or sgRNAs should be read in conjunction with the discussion herein of Cas13b-based crRNAs, e.g., sgRNAs in other systems. Some may be considered crRNAs in the present invention or as such.
[0063] The present invention provides a method of identifying suitable guide sequence requirements for a Cas13b effector protein of the invention (e.g., Table 1), comprising: (a) selecting a set of essential genes in an organism; (b) designing a library of targeting guide sequences capable of hybridizing to the regions of these genes, the coding regions and the 5' and 3' UTRs of these genes; (c) generating randomized guide sequences as control guides that do not hybridize to any region within the organism's genome; (d) preparing a plasmid comprising an RNA targeting protein and a first resistance gene and a guide plasmid library comprising said library of targeting guides, said control guide and a second resistance gene; (e) co-introducing said plasmid into a host cell; (f) introducing said host cell into a selective medium for said first and second resistance genes; (g) sequencing essential genes of the growing host cell; (h) determining the significance of depletion of cells transformed with the targeting guide by comparing depletion of cells with the control guide; and (i) Determining Suitable Guide Sequence Requirements Based on Depleted Guide Sequences to provide a method comprising:
[0064] In one embodiment of such methods, determination of the PFS sequence against a suitable guide sequence of an RNA targeting protein is by comparison of the guide's target sequence in depleted cells. In one aspect of such methods, the method further comprises comparing guide abundances for different conditions in different replicate experiments. In one aspect of such methods, a control guide is selected in that it is determined to exhibit limited deviation for guide depletion in replicate experiments. In one embodiment of such methods, the significance of depletion is determined as (a) depletion higher than the most depleted control guide, or (b) depletion higher than the mean depletion plus twice the standard deviation for the control guide. In one embodiment of such methods, the host cell is a bacterial host cell. In one embodiment of such a method, co-introducing the plasmids is by electroporation and the host cell is an electrocompetent host cell.
[0065] Cas13b The present invention provides methods of altering sequences associated with or at a target locus of interest, which methods include non-naturally occurring or engineered sequences comprising a Cas13b effector protein and one or more nucleic acid moieties. delivering a composition to said genetic locus, wherein the effector protein forms a complex with one or more nucleic acid moieties, and upon binding of said complex to the genetic locus of interest, the effector protein: Inducing alterations in sequences associated with or at the target locus of interest. In preferred embodiments, the modification is the introduction of a strand break. In a preferred embodiment, the sequence associated with or at the target locus of interest comprises or consists of RNA.
[0066] The present invention provides a method of altering sequences associated with or at a target locus of interest, comprising a Cas13b effector protein, optionally a small accessory protein, and one or more nucleic acid components. wherein the effector protein is complexed with one or more nucleic acid moieties, and wherein the complex is the desired Upon binding to the locus, the effector protein induces alteration of sequences associated with or at the target locus of interest. In preferred embodiments, the modification is the introduction of a strand break. In a preferred embodiment, the sequence associated with or at the target locus of interest comprises or consists of RNA.
[0067] The present invention provides a method of altering sequences associated with or at a target locus of interest, which method comprises a non-naturally occurring or delivering an engineered composition to said sequence associated with or at a genetic locus, wherein said Cas13b effector protein forms a complex with one or more nucleic acid moieties, said complex comprising Upon binding to the locus of interest, the effector protein induces alteration of sequences associated with or at the target locus of interest. In preferred embodiments, the modification is the introduction of a strand break. In a preferred embodiment, the Cas13b effector protein forms a complex with one nucleic acid component, preferably an engineered or non-naturally occurring nucleic acid component. Induction of sequence alterations associated with or at a target locus of interest can be under Cas13b effector protein-nucleic acid guidance. In a preferred embodiment, one nucleic acid component is CRISPR RNA (crRNA). In a preferred embodiment, one nucleic acid component is a mature crRNA or guide RNA, wherein the mature crRNA or guide RNA comprises a spacer sequence (or guide sequence) and direct repeat (DR) sequences or derivatives thereof. In a preferred embodiment, the spacer sequence or derivative thereof comprises a seed sequence, wherein the seed sequence is critical for recognition and / or hybridization with the sequence at the target locus. In preferred embodiments of the invention, the crRNA is a short crRNA that can be associated with a short DR sequence. In another embodiment of the invention, the crRNA is a long crRNA that can be associated with a long DR sequence (or double DR). Aspects of the invention relate to Cas13b effector protein complexes having one or more non-naturally occurring or engineered or modified or optimized nucleic acid components. In preferred embodiments, the nucleic acid component comprises RNA. In preferred embodiments, the nucleic acid component of the complex may comprise a guide sequence linked to a direct repeat sequence, where the direct repeat sequence comprises one or more stem loops or optimized secondary structures. In a preferred embodiment of the invention, direct repeats may be short DRs or long DRs (double DRs). In preferred embodiments, direct repeats may be modified to contain one or more protein-binding RNA aptamers. In preferred embodiments, direct repeats may be modified to contain one or more protein-binding RNA aptamers. In preferred embodiments, one or more aptamers may be included as part of an optimized secondary structure. Such aptamers may have the ability to bind to bacteriophage coat proteins. Bacteriophage coat proteins are Qβ, F2, GA, fr, JP501, MS2, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, It may be selected from the group comprising φCb5, φCb8r, φCb12r, φCb23r, 7s and PRR1. In preferred embodiments, the bacteriophage coat protein is MS2. The invention also provides nucleic acid components of complexes that are 30 or more, 40 or more, or 50 or more nucleotides in length.
[0068] The present invention provides a method for genome editing or modifying sequences associated with or at a target locus of interest, which method allows any desired cell type, prokaryotic or eukaryotic cell to contain the Cas13b complex. so that the Cas13b effector protein complex effectively functions to interact with eukaryotic or prokaryotic RNA. In preferred embodiments, the cell is a eukaryotic cell and the RNA is transcribed from the mammalian genome or is present in the mammalian cell. In preferred methods of RNA editing or genome editing in human cells, Cas13b effector proteins may include, but are not limited to, the specific species of Cas13b effector proteins disclosed herein.
[0069] The present invention also provides a method of modifying a target locus of interest, comprising producing a non-naturally occurring or engineered composition comprising a Cas13b locus effector protein and one or more nucleic acid components at said locus. wherein the Cas13b effector protein forms a complex with one or more nucleic acid moieties, and when said complex binds to the locus of interest, the effector protein is delivered to the target locus of interest Induce modification. In preferred embodiments, the modification is the introduction of a strand break.
[0070] In such methods, the target locus of interest can be contained within an RNA molecule. In such methods, the target locus of interest can be contained in the RNA in vitro.
[0071] In such methods, the target locus of interest can be contained in an RNA molecule within the cell. Cells can be prokaryotic or eukaryotic. A cell can be a mammalian cell. Modifications introduced into cells by the present invention may be such as to alter the cells and their progeny for enhanced production of biological products such as antibodies, starches, alcohols or other desired cellular products. . Modifications introduced into a cell by the present invention may be such that the cell and progeny of the cell contain changes that alter the biological product produced.
[0072] Mammalian cells are cells of non-human mammals, such as primate, bovine, ovine, porcine, canine, rodent, Leporidae, e.g. monkey, cow, ovine, porcine, canine, rabbit, rat or mouse cells. could be. The cells may be non-mammalian eukaryotic cells, such as poultry avian (eg chicken), vertebrate fish (eg salmon) or crustacean (eg oyster, claim, lobster, shrimp) cells. Cells can also be plant cells. The plant cell may be of a monocotyledonous or dicotyledonous plant or of a crop or cereal plant such as cassava, maize, sorghum, soybean, wheat, oats or rice. Plant cells may be algal, arboreal or producing plants, fruit or vegetable (e.g. citrus trees such as orange, grapefruit or lemon trees; peach or nectarine trees; apple or pear trees; almonds or Nut trees such as walnut or pistachio trees; Solanaceae plants; Brassica plants; Lactuca plants; Spinacia plants; Capsicum plants; Tobacco, asparagus, carrots, cabbage, broccoli, cauliflower, tomatoes, eggplant, pepper, lettuce, spinach, strawberries, blueberries, raspberries, blackberries, grapes, coffee, cocoa, etc.).
[0073] The present invention provides a method of modifying a target locus of interest, wherein the method delivers a non-naturally occurring or engineered composition comprising a Cas13b effector protein and one or more nucleic acid moieties to said locus. wherein the effector protein forms a complex with one or more nucleic acid components, and when said complex binds to the locus of interest, the effector protein induces modification of the target locus of interest do. In preferred embodiments, the modification is the introduction of a strand break.
[0074] In such methods, the target locus of interest can be contained within an RNA molecule. In a preferred embodiment, the target locus of interest comprises or consists of RNA.
[0075] The invention also provides a method of modifying a target locus of interest, which method delivers a non-naturally occurring or engineered composition comprising a Cas13b effector protein and one or more nucleic acid moieties to said locus. wherein the Cas13b effector protein forms a complex with one or more nucleic acid components, and when said complex binds to the locus of interest, the effector protein modifies the target locus of interest Induce. In preferred embodiments, the modification is the introduction of a strand break.
[0076] Preferably, in such methods, the target locus of interest may be contained in an RNA molecule in vitro. Also preferably, in such methods the target locus of interest may be contained in an intracellular RNA molecule. Cells can be prokaryotic or eukaryotic. A cell can be a mammalian cell. The cells can be rodent cells. The cells can be mouse cells.
[0077] In any of the methods described, the target locus of interest can be a genomic or epigenomic locus of interest. In any of the methods described, the complex can be delivered with multiple guides for multiplexed use. More than one protein can be used in any of the methods described.
[0078] In a further aspect of the invention, the nucleic acid component may comprise CRISPR RNA (crRNA) sequences. With the effector protein being a Cas13b effector protein, the nucleic acid component may contain CRISPR RNA (crRNA) sequences and generally not contain any transactivating crRNA (tracr RNA) sequences.
[0079] In any of the methods described, the effector protein and nucleic acid components may be provided by one or more polynucleotide molecules encoding the protein and / or nucleic acid components, wherein the one or more polynucleotide molecules are , is operably configured to express protein and / or nucleic acid components. One or more polynucleotide molecules may contain one or more regulatory elements operably configured to express protein and / or nucleic acid components. One or more polynucleotide molecules can be contained within one or more vectors. In any of the methods described, the target locus of interest can be a genomic, epigenomic or transcriptome locus of interest. In any of the methods described, the complexes can be delivered with multiple guides for multiplexed applications. More than one protein can be used in any of the methods described.
[0080] In any of the methods described, the strand break can be a single strand break or a double strand break. In a preferred embodiment, a double-strand break is defined as two sections of RNA formed when a single-stranded RNA molecule folds on itself or an RNA molecule comprising a self-complementary sequence folds and self-folds a portion of that RNA. It can refer to the cleavage of two sections of RNA, such as the putative double helix formed by allowing pairing.
[0081] Regulatory elements can include inducible promoters. Polynucleotide and / or vector systems may include inducible systems.
[0082] In any of the methods described, one or more polynucleotide molecules can be included in the delivery system, or one or more vectors can be included in the delivery system.
[0083] In any of the methods described, the non-naturally occurring or engineered composition can be delivered by liposomes, particles including nanoparticles, exosomes, microvesicles, gene guns or one or more viral vectors.
[0084] The invention also provides non-naturally occurring or engineered compositions having characteristics as discussed herein or compositions defined by any of the methods described herein. .
[0085] In certain embodiments, therefore, the present invention provides non-naturally occurring or engineered compositions, such as compositions that are specifically capable of or configured to modify a target locus of interest. and said composition comprises a Cas13b effector protein and one or more nucleic acid moieties, wherein said effector protein is complexed with said one or more nucleic acid moieties and said complex is at a locus of interest. Upon binding, the effector protein induces modification of the target locus of interest. In certain embodiments, the effector protein can be a Cas13b effector protein.
[0086] The invention also provides, in a further aspect, non-naturally occurring or engineered compositions, such as compositions specifically capable of or configured to modify a target locus of interest, said The composition comprises (a) a guide RNA molecule (or a combination of guide RNA molecules, such as a first guide RNA molecule and a second guide RNA molecule) or a nucleic acid encoding a guide RNA molecule (or a combination of guide RNA molecules) (one or more nucleic acids that do), (b) a Cas13b effector protein. In certain embodiments, the effector protein can be a Cas13b effector protein.
[0087] The present invention provides, in a further aspect, (I.) one comprising (a) a guide sequence capable of hybridizing to a target sequence of a polynucleotide locus, (b) a tracr mate sequence, and (c) a tracrRNA sequence. Also provided is a non-naturally occurring or engineered composition comprising the above CRISPR-Cas system polynucleotide sequence and (II.) a second polynucleotide sequence encoding a Cas13b effector protein, wherein upon transcription: The tracr mate sequence hybridizes to the tracr RNA sequence, and the guide sequence directs sequence-specific binding to the target sequence by the CRISPR complex, wherein the CRISPR complex is (1) hybridized to the target sequence It contains a guide sequence and (2) a Cas13b effector protein complexed with a tracr mate sequence that hybridizes to a tracr RNA sequence. In certain embodiments, the effector protein can be a Cas13b effector protein.
[0088] In certain embodiments, tracrRNA may not be required. Accordingly, the present invention provides, in certain embodiments, (I.) one or more CRISPRs comprising (a) a guide sequence capable of hybridizing to a target sequence of a polynucleotide locus, and (b) a direct repeat sequence. - A non-naturally occurring or engineered composition comprising a Cas system polynucleotide sequence and (II.) a second polynucleotide sequence encoding a Cas13b effector protein, wherein, upon transcription, the guide sequence is , directs sequence-specific binding to the target sequence by the CRISPR complex, where the CRISPR complex is complexed with (1) a guide sequence hybridized to the target sequence, and (2) a direct repeat sequence. Contains Cas13b effector protein. Preferably, the effector protein may be the Cas13b effector protein. Without limitation, Applicants hypothesize that in such cases, the direct repeat sequence may contain sufficient secondary structure to charge the crRNA to the effector protein. By way of example and without limitation, such secondary structure may comprise, consist essentially of, or consist of stem loops (such as one or more stem loops) within direct repeats.
[0089] The present invention also provides a vector system comprising one or more vectors, wherein the one or more vectors are naturally occurring compositions of matter having characteristics as defined in the methods described herein. comprising one or more polynucleotide molecules encoding components of the non- or engineered composition.
[0090] The invention also provides delivery systems comprising one or more vectors or one or more polynucleotide molecules, wherein the one or more vectors or polynucleotide molecules have characteristics as discussed herein or A composition defined by any of the methods described herein, comprising one or more polynucleotide molecules encoding components of the non-naturally occurring or engineered composition.
[0091] The present invention encodes a non-naturally occurring or engineered composition, or one or more polynucleotides encoding components of said composition, or components of said composition, for use in therapeutic methods of treatment. Also provided are vectors or delivery systems comprising one or more polynucleotides. Therapeutic treatment methods may include gene or gene editing or gene therapy.
[0092] The present invention provides methods and compositions wherein one or more amino acid residues of an effector protein can be modified, e.g. Alternatively, non-naturally occurring Cas13b effector proteins are also provided. In certain embodiments, the modification may involve mutation of one or more amino acid residues of the effector protein. One or more mutations can be in one or more catalytically active domains of the effector protein. The effector protein may have reduced or no nuclease activity compared to an effector protein lacking said one or more mutations. This effector protein may not induce cleavage of one RNA strand at the target locus of interest. In preferred embodiments, the one or more mutations may include two mutations. In preferred embodiments, one or more amino acid residues are modified in a Cas13b effector protein, eg, an engineered or non-naturally occurring Cas13b effector protein. In certain embodiments of the invention, the effector protein comprises one or more HEPN domains. In preferred embodiments, the effector protein comprises two HEPN domains. In another preferred embodiment, the effector protein contains one HEPN domain at the C-terminus of the protein and another HEPN domain at the N-terminus. In certain embodiments, the one or more mutations or the two or more mutations may be in a catalytically active domain of an effector protein comprising the HEPN domain or a catalytically active domain homologous to the HEPN domain. In certain embodiments, the effector protein comprises one or more of the following mutations: R116A, H121A, R1177A, H1182A (wherein the amino acid position is group 29 from Bergeyella zoohelcum ATCC 43767). corresponding to amino acid positions in the protein). Those skilled in the art will appreciate that corresponding amino acid positions in different Cas13b proteins can be mutated to have the same effect. In certain embodiments, one or more mutations completely or partially abolish the protein's catalytic activity (eg, alter cleavage rate, alter specificity, etc.). In certain embodiments, the effector protein as described herein is a "dead" effector protein, such as a dead Casl3b effector protein (ie dCasl3b). In certain embodiments, the effector protein has one or more mutations in HEPN domain 1. In certain embodiments, the effector protein has one or more mutations in HEPN domain 2. In certain embodiments, the effector protein has one or more mutations in HEPN domain 1 and HEPN domain 2. Effector proteins may contain one or more heterologous functional domains. One or more heterologous functional domains may comprise one or more nuclear localization signal (NLS) domains. One or more heterologous functional domains may comprise at least two or more NLS domains. One or more NLS domains may be located at, near or adjacent to the terminus of an effector protein (e.g., a Cas13b effector protein), and in the case of two or more NLS, each of the two It can be located at or near or close to the terminus (eg, the Cas13b effector protein). One or more heterologous functional domains may comprise one or more transcriptional activation domains. In preferred embodiments, the transcriptional activation domain may comprise VP64. One or more heterologous functional domains may comprise one or more transcriptional repression domains. In preferred embodiments, the transcription repression domain may comprise a KRAB domain or a SID domain (eg SID4X). One or more heterologous functional domains may comprise one or more nuclease domains. In preferred embodiments, the nuclease domain comprises Fok1.
[0093] The present invention provides that one or more heterologous functional domains have the following activities: methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription terminator activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, It is also provided to have one or more of double stranded RNA cleaving activity, single stranded DNA cleaving activity, double stranded DNA cleaving activity and nucleic acid binding activity. The at least one or more heterologous functional domains may be at or near the amino terminus of the effector protein and / or wherein the at least one or more heterologous functional domains are at the carboxy terminus of the effector protein or near it. One or more heterologous functional domains can be fused to the effector protein. One or more heterologous functional domains can be tethered to the effector protein. One or more heterologous functional domains can be linked to the effector protein via a linker moiety.
[0094] In certain embodiments, the Cas13b effector protein as contemplated herein is 30-40 bp long, more typically 34-38 bp long, even more typically 36-37 bp long, e.g. , 32, 33, 34, 35, 36, 37, 38, 39 or 40 bp long loci containing short CRISPR repeats. In certain embodiments, the CRISPR repeats are 80-350 bp long, such as 80-200 bp long, even more typically 86-88 bp long, such as 80, 81, 82, 83, 84, 85, 86, 87, 88. , 89 or 90 bp long or double repeats.
[0095] In certain embodiments, a protospacer adjacent motif (PAM) or PAM-like motif directs binding to a target locus of interest by an effector protein (e.g., Cas13b effector protein) complex as disclosed herein. . In some embodiments, the PAM can be a 5' PAM (ie located upstream of the 5' end of the protospacer). In other embodiments, the PAM can be the 3'PAM (ie located downstream of the 5' end of the protospacer). In other embodiments, both 5'PAM and 3'PAM are required. In certain embodiments of the invention, PAM or PAM-like motifs may be unnecessary to direct binding of effector proteins (eg, Cas13b effector proteins). In certain embodiments, 5'PAM is D (ie, A, G or U). In certain embodiments, 5'PAM is D for Casl3b effectors. In certain embodiments of the invention, truncation in the repeat sequence results in short nucleotides (e.g. 5, 6, 7, 8, 9 or 10 nt or more if it is a double repeat) at the 5' end. A crRNA (eg, short or long crRNA) can be generated that includes a repeat sequence (which can be referred to as a crRNA “tag”) and a complete spacer sequence flanked at the 3′ end by the remainder of the repeat. In certain embodiments, targeting by the effector proteins described herein may require lack of homology between the crRNA tag and the target 5'flanking sequences. This requirement can be similar to that further described in Samai et al. "Co-transcriptional DNA and RNA Cleavage during Type III CRISPR-Cas Immunity" Cell 161, 1164-1174, May 21, 2015, where , the requirement is thought to distinguish the real target on the invading nucleic acid from the CRISPR array itself, and the presence of repeat sequences would lead to perfect homology with the crRNA tag and prevent autoimmunity.
[0096] In certain embodiments, the Cas13b effector protein may be engineered to contain one or more mutations that reduce or eliminate nuclease activity and thereby reduce or eliminate RNA interference activity. Mutations can be made in adjacent residues, such as amino acids near those involved in nuclease activity. In some embodiments, one or more putative catalytic nuclease domains are inactivated and the effector protein complex lacks cleavage activity and function as an RNA binding complex. In preferred embodiments, the resulting RNA binding complex may be linked to one or more functional domains as described herein.
[0097] In certain embodiments, one or more functional domains are regulatable, ie inducible.
[0098] In certain embodiments of the invention, the guide RNA or mature crRNA comprises, consists essentially of, or consists of a direct repeat sequence and a guide or spacer sequence. In certain embodiments, the guide RNA or mature crRNA comprises, consists essentially of, or consists of a direct repeat sequence linked to a guide or spacer sequence. In preferred embodiments of the invention, the mature crRNA comprises a stem-loop or an optimized stem-loop structure or an optimized secondary structure. In preferred embodiments, the mature crRNA comprises a stem-loop or optimized stem-loop structure in the direct repeat sequence, where the stem-loop or optimized stem-loop structure is important for cleavage activity. In certain embodiments, the mature crRNA preferably contains a single stem-loop. In certain embodiments, the direct repeat sequence preferably contains a single stem loop. In certain embodiments, the cleavage activity of the effector protein complex is altered by introducing mutations that affect stem-loop RNA duplex structure. In a preferred embodiment, mutations can be introduced that maintain the RNA duplex of the stem-loop, thereby maintaining the cleaving activity of the effector protein complex. In other preferred embodiments, mutations can be introduced that disrupt the RNA duplex structure of the stem-loop, thereby completely abolishing the cleavage activity of the effector protein complex.
[0099] CRISPR systems as provided herein can utilize crRNA or similar polynucleotides containing guide sequences, wherein the polynucleotide is RNA, DNA, or a mixture of RNA and DNA; and / or wherein the polynucleotide comprises one or more nucleotide analogues. The sequence can comprise any structure, including but not limited to structures of native crRNA such as bulge, hairpin or stem-loop structures. In certain embodiments, a polynucleotide comprising a guide sequence forms a duplex with a second polynucleotide sequence, which can be an RNA or DNA sequence.
[0100] In certain embodiments, the method utilizes chemically modified guide RNAs. Examples of guide RNA chemical modifications include, without limitation, 2'-O-methyl (M), 2'-O-methyl 3' phosphorothioate (MS) or 2'-O-methyl 3 at one or more terminal nucleotides. 'ThioPACE (MSP) incorporation. Such chemically modified guide RNAs can contain increased stability and increased activity when compared to unmodified guide RNAs, but on-target versus off-target specificity is unpredictable. . (See Hendel, 2015, Nat Biotechnol. 33(9):985-9, doi:10.1038 / nbt.3290, published online 29 June 2015). Chemically modified guide RNAs further include, without limitation, RNAs with phosphorothioate linkages and locked nucleic acid (LNA) nucleotides containing a methylene bridge between the 2' and 4' carbons of the ribose ring.
[0101] The invention also provides nucleotide sequences encoding effector proteins that are codon-optimized for expression in eukaryotes or eukaryotic cells in any of the methods or compositions described herein. In certain embodiments of the invention, the codon-optimized effector protein is any Cas13b effector protein discussed herein, eukaryotic cells or organisms, e.g. are codon-optimized for operability in various cells or organisms such as, without limitation, yeast cells or mammalian cells or organisms such as mouse cells, rat cells and human cells or non-human eukaryotes such as plants.
[0102] In certain embodiments of the invention, at least one nuclear localization signal (NLS) is added to the nucleic acid sequence encoding the Cas13b effector protein. In a preferred embodiment, at least one or more C-terminal or N-terminal NLS is added (thus, a nucleic acid molecule encoding a Cas13b effector protein may contain coding for an NLS such that the NLS is added or added to the expressed product). connected). In a preferred embodiment, a C-terminal NLS is added for optimal expression and nuclear targeting in eukaryotic cells, preferably human cells. The invention also encompasses methods of delivering multiple nucleic acid components, wherein each nucleic acid component is specific for a different target locus of interest, thereby modifying multiple target loci of interest. The nucleic acid component of the complex can include one or more protein-binding RNA aptamers. One or more aptamers may be capable of binding bacteriophage coat proteins. Bacteriophage coat proteins are Qβ, F2, GA, fr, JP501, MS2, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, It may be selected from the group comprising φCb5, φCb8r, φCb12r, φCb23r, 7s and PRR1. In preferred embodiments, the bacteriophage coat protein is MS2. The invention also provides nucleic acid components of complexes that are 30 or more, 40 or more, or 50 or more nucleotides in length.
[0103] In a further aspect, the invention provides a eukaryotic cell comprising a modified target locus of interest, wherein the target locus of interest is modified by any of the methods described herein. ing. A further aspect provides a cell line of said cells. Another aspect provides a multicellular organism comprising one or more of said cells.
[0104] In certain embodiments, the modification of the target locus of interest is a eukaryotic cell comprising altered expression of at least one gene product, a eukaryotic cell comprising altered expression of at least one gene product, wherein at least one expression of one gene product is increased in a eukaryotic cell, a eukaryotic cell comprising altered expression of at least one gene product, wherein expression of at least one gene product is decreased in a eukaryotic cell, or can result in eukaryotic cells containing the edited genome.
[0105] In certain embodiments, eukaryotic cells can be mammalian or human cells.
[0106] In further embodiments, the non-naturally occurring or engineered composition, vector system or delivery system as described herein is used for site-specific gene knockout, site-specific genome editing, RNA sequence-specific interference Or it can be used for multiplex genome engineering.
[0107] Also provided are gene products from the cells, cell lines or organisms as described herein. In certain embodiments, the amount of gene product expressed may be greater or less than the amount of gene product from cells that do not have altered expression or edited genomes. In certain embodiments, the gene product may be altered relative to the gene product from cells that do not have altered expression or the edited genome.
[0108] In another aspect, the invention provides a method of identifying novel nucleic acid modification effectors, comprising at least one nucleic acid modification effector that is within a defined distance from a conserved genomic element of a locus and exceeds a defined size limit. identifying a putative nucleic acid modification locus from a set of nucleic acid sequences that contain a protein or encode both putative nucleic acid modification enzyme loci; combining the identified putative nucleic acid modification loci into subsets containing homologous proteins , one or more of the following: a subset containing loci with putative effector proteins with low domain homology matches to known protein domains compared to loci in the other subset, conserved compared to loci in the other subset the subset containing the putative protein with the smallest distance to the genomic element in the other subset; Subsets containing putative effector proteins with lower existing nucleic acid modification taxonomy than other subsets, subsets containing loci with lower proximity to known nucleic acid modification loci compared to other subsets, and the total number of candidate loci in each subset identifying a final set of candidate nucleic acid modification loci by selecting nucleic acid modification loci from one or more subsets based on.
[0109] In one embodiment, the set of nucleic acid sequences is obtained from a genomic or metagenomic database, such as a genomic or metagenomic database containing prokaryotic genomic or metagenomic sequences.
[0110] In one embodiment, the constant distance from conserved genomic elements is 1 kb to 25 kb.
[0111] In one embodiment, conserved genomic elements include repetitive elements such as CRISPR arrays. In a specific embodiment, the constant distance from the conserved genomic elements is within 10 kb of the CRISPR array.
[0112] In one embodiment, the defined size limit for proteins contained within the putative nucleic acid modification (effector) locus is greater than 200 amino acids, or more particularly the defined size limit is greater than 700 amino acids. In one embodiment, the putative nucleic acid modification locus is 900-1800 amino acids.
[0113] In one embodiment, conserved genomic elements are identified using repeat or pattern discovery analysis of nucleic acid sets, such as PILER-CR.
[0114] In one embodiment, the summarizing step of the methods described herein is based, at least in part, on the results of a domain homology search or HHpred protein domain homology search.
[0115] In one embodiment, the constant threshold is the BLAST nearest neighbor cutoff value between 0 and 1e-7.
[0116] In one embodiment, the methods described herein further comprise a filtering step to include only loci with putative proteins of 900-1800 amino acids.
[0117] In one embodiment, the methods described herein involve generating a set of nucleic acid constructs encoding nucleic acid modification effectors and performing PAM validation in bacterial colonies, in vitro cleavage assays, Surveyor methods, experiments in mammalian cells, PFS Further includes experimental verification of the nucleic acid modifying function of the candidate nucleic acid modifying effector, including performing one or more biochemical validation assays, such as using validation or a combination thereof.
[0118] In one embodiment, the methods described herein further comprise preparing a non-naturally occurring or engineered composition comprising one or more proteins from the identified nucleic acid modification loci.
[0119] In one embodiment, the identified locus comprises a class 2 CRISPR effector, or the identified locus lacks Cas1 or Cas2, or the identified locus comprises a single effector.
[0120] In one embodiment, the single large effector protein is greater than 900 amino acids long, or greater than 1100 amino acids long, or comprises at least one HEPN domain.
[0121] In one embodiment, at least one HEPN domain is near the N-terminus or C-terminus of the effector protein, or is located at an internal position of the effector protein.
[0122] In one embodiment, a single large effector protein contains HEPN domains at the N-terminus and C-terminus, and two HEPN domains within the protein.
[0123] In one embodiment, the identified locus further comprises one or two small putative accessory proteins within 2-10 kb from the CRISPR array.
[0124] In one embodiment, the small accessory protein is less than 700 amino acids. In one embodiment, the small accessory protein is 50-300 amino acids long.
[0125] In one embodiment, the small accessory protein comprises multiple predicted transmembrane domains, or four predicted transmembrane domains, or at least one HEPN domain.
[0126] In one embodiment, the small accessory protein comprises at least one HEPN domain and at least one transmembrane domain.
[0127] In one embodiment, the locus does not contain additional proteins up to 25 kb from the CRISPR array.
[0128] In one embodiment, the CRISPR array comprises direct repeat sequences comprising about 36 nucleotides in length. In a specific embodiment, the direct repeat comprises a GTTG / GUUG at the 5' end that is reverse complementary to CAAC at the 3' end.
[0129] In one embodiment, the CRISPR array comprises spacer sequences comprising about 30 nucleotides in length.
[0130] In one embodiment, the identified locus lacks small accessory proteins.
[0131] The present invention provides a method of identifying novel CRISPR effectors, the method comprising: a) identifying, in a genome or metagenomic database, a sequence encoding a CRISPR array; b) in said selected sequence, Identifying one or more open reading frames (ORFs) within 10 kb, c) selecting loci based on the presence of putative CRISPR effector proteins of size 900-1800 amino acids, d) 50-300 amino acids. and e) identifying loci encoding putative CRISPR effectors and CRISPR accessory proteins and optionally classifying them based on structural analysis.
[0132] In one embodiment, the CRISPR effector is a type VI CRISPR effector. In one embodiment, step (a) comprises: i) comparing sequences in a genome and / or metagenomic database with at least one pre-identified seed sequence encoding a CRISPR array, and a sequence comprising said seed sequence; or ii) identifying CRISPR arrays based on the CRISPR algorithm.
[0133] In some embodiments, step (d) comprises identifying the nuclease domain. In some embodiments, step (d) comprises identifying RuvC, HPN and / or HEPN domains.
[0134] In certain embodiments, ORFs encoding Cas1 or Cas2 are not within 10 kb of the CRISPR array.
[0135] In certain embodiments, the ORF in step (b) encodes a putative accessory protein of 50-300 amino acids.
[0136] In certain embodiments, the putative novel CRISPR effector obtained in step (d) is obtained by further comparison of genomic and / or metagenomic sequences and subsequent gene of interest as described in steps a) to d) of claim 1. Used as seed sequence for locus selection. In certain embodiments, pre-identified seed sequences are used by (a) identifying CRISPR motifs in a genome or metagenomic database, (b) extracting multiple features in said identified CRISPR motifs, (c) a teacher Classifying the CRISPR loci using pone learning, (d) identifying conserved locus elements based on said classification, and (e) selecting therefrom suitable putative CRISPR effectors as seed sequences. obtained by a method including
[0137] In certain embodiments, features include protein elements, repeat structures, repeat sequences, spacer sequences and spacer mapping. In certain embodiments, the genome and metagenomic databases are bacterial and / or archaeal genomes. In certain embodiments, genomic and metagenomic sequences are obtained from the Ensembl and / or NCBI genomic databases. In certain embodiments, structural analysis in step (d) is based on secondary structure prediction and / or sequence alignment. In one embodiment, step (d) is achieved by clustering the remaining loci based on the protein they encode and manually curating the resulting clusters.
[0138] Accordingly, the object of the present invention is to cover any previously known product, product, as Applicants reserve the right and hereby disclose a disclaimer of any previously known product, process, or method. The process of making or using the product is also not within the scope of this invention. Further, the present invention is subject to any U.S. patents to which applicants have reserved and which hereby disclose disclaimers of any previously described product, process of making the product, or method of use of the product. Any product, process of making a product or product that does not meet the description and enablement requirements of the Trademark Office (USPTO) (Section 112, first paragraph of the United States Patent Act) or the European Patent Office (EPO) (Article 83 EPC) It is noted that methods of use are also not intended to be encompassed within the scope of the present invention. In the practice of the present invention, it may be advantageous to comply with Article 53(c) EPC and Regulation 28(b) and (c) EPC. Nothing in this specification should be construed as prospective.
[0139] In the present disclosure and particularly in the claims and / or paragraphs, the terms "comprising," "contains," "comprising," and the like may have the meanings ascribed to them under United States patent law. For example, these can mean "includes," "includes," "contains," etc., and terms such as "consisting essentially of" and "consisting essentially of" , have the meaning ascribed to them in United States patent law, e.g., these terms allow elements not explicitly recited, but may be used to describe elements found in the prior art or elements fundamental or novel to the invention. It is noted that excluding factors that affect characteristics.
[0140] These and other embodiments are disclosed or are apparent from and encompassed by the following detailed description.
[0141] The novel features of the invention are pointed out with particularity in the appended claims. A further understanding of the features and advantages of the present invention may be realized by reference to the following detailed description and accompanying drawings, which illustrate illustrative embodiments in which the principles of the invention are employed. [Brief description of the drawing]
[0142] [Fig. 1A] Figures 1A-1B show tree alignments of Cas13b orthologs. [Fig. 1B] Figures 1A-1B show tree alignments of Cas13b orthologs.
[0143] [Figure 2A] Figures 2A-2C show tree alignments of C2c2 and Cas13b orthologues. [Fig.2B] Figures 2A-2C show tree alignments of C2c2 and Cas13b orthologues. [Figure 2C] Figures 2A-2C show tree alignments of C2c2 and Cas13b orthologues.
[0144] [Fig.3] FIG. 3 shows exemplary test results of Cas13b orthologs for activity in E. coli, comparing introduction of Cas13b from B. zoohelcum with introduction of an empty vector.
[0145] [Figure 4] FIG. 4 shows a general comparison of the specific RNA cleavage activities obtained with the different orthologs of Table 1.
[0146] [Fig.5] FIG. 5 shows an alignment of different Cas13b orthologs as provided in FIGS.
[0147] [Figure 6A-6B] (A) shows the protospacer design for the MS2 phage drop plaque assay testing RNA interference to identify PFS. (B) shows an RNA interference assay schematic. The target sequence is placed in-frame at the start of the transcribed bla gene that confers ampicillin resistance or in a non-transcribed region on the opposite strand of the same target plasmid. Target plasmids were co-transformed with the bzcas13b plasmid, which confers chloramphenicol resistance, or an empty vector, and plated on double-selection antibiotic plates. Depleted colonies were identified and the corresponding targets were sequenced to identify PFS.
[0148] [Fig.7] Figure 7 shows a heatmap of normalized PFS scores from safely depleted spacers for orthologs 1, 13 and 16 in the absence and presence of csx27 accessory proteins.
[0149] [Fig.8] Figure 8 shows a heatmap of normalized PFS scores from safely depleted spacers for orthologs 2, 3, 8, 9, 14, 19 and 21 in the absence and presence of csx28 accessory proteins.
[0150] [Fig.9] FIG. 9 shows a heatmap of normalized PFS scores from safely depleted spacers for orthologs 5, 6, 7, 10, 12 and 15.
[0151] [Fig. 10A-10B] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10C-10D] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10E-10F] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10G-10H] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10I-10J] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10K-10L] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10M-10N] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10O-10P] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10Q-10R] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10S-10T] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10U-10V] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10W-10X] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10Y-10Z] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs. [Fig. 10AA-10BB] Figures 10A-10BB show heatmaps of normalized PFS scores and derived PFS from safely depleted spacers for different orthologs.
[0152] [Fig. 11] Figure 11 provides a summary of luciferase interference data for Cas13b orthologs that were less active in mammalian cells in the guides tested.
[0153] [Fig. 12] Figure 12 provides a summary of luciferase interference data for Cas13b orthologues that showed low to moderate activity in mammalian cells in the guide tested.
[0154] [Fig. 13] Figure 13 provides a summary of luciferase interference data for some of the Cas13b orthologs that showed significant activity in mammalian cells in the Guide tested.
[0155] [Fig. 14A] Figures 14A-14B provide a summary of luciferase interference data for some of the Cas13b orthologs that showed significant activity in mammalian cells in the Guide tested. [Fig. 14B] Figures 14A-14B provide a summary of luciferase interference data for some of the Cas13b orthologs that showed significant activity in mammalian cells in the Guide tested.
[0156] [Fig. 15] Figure 15 provides a summary of luciferase interference data for the Cas13b orthologues that showed significant activity in mammalian cells in the guides tested and a comparison with C2c2 activity.
[0157] [Fig. 16] FIG. 16 shows synthetic data from orthologues with significant activity in eukaryotic cells with the same guide sequences.
[0158] [Figure 17A-17G] Figures 17A-17G show collateral effects of Cas13b orthologs in mammalian cells using two different reporter genes, G-luciferase and C-luciferase. [Fig. 17C-17D] Figures 17A-17G show collateral effects of Cas13b orthologs in mammalian cells using two different reporter genes, G-luciferase and C-luciferase. [Fig. 17E-17F] Figures 17A-17G show collateral effects of Cas13b orthologs in mammalian cells using two different reporter genes, G-luciferase and C-luciferase. [Fig. 17G]Figures 17A-17G show collateral effects of Cas13b orthologs in mammalian cells using two different reporter genes, G-luciferase and C-luciferase.
[0159] [Fig. 18] Figure 18 shows the characterization of highly active Cas13b orthologues for RNA knockdown. A) Schematic representation of the stereotypic Cas13 loci and corresponding crRNA structures. B) Determination of 19 Cas13a, 15 Cas13b and 7 Cas13c orthologues for luciferase knockdown using two different guides. The host organism name is displayed for orthologs for which knockdown is efficient using both guides. C) Compare the knockdown activity of PspCas13b and LwaCas13a by tiling guides against Gluc and measuring luciferase expression. D) Compare the knockdown activity of PspCas13b and LwaCas13a by tiling guides against Cluc and measuring luciferase expression. E) log2 of all genes detected in the RNA-seq library (in million transcripts) of the non-targeting control (x-axis) compared to the Gluc-targeting conditions (y-axis) for LwaCas13a (red) and shRNA (black). expression levels by percentage (TPM) values. Mean values of three biological replicates are shown. Gluc transcript data points are labeled. F) log2 of all genes detected in the RNA-seq library (in million transcripts) of the non-targeting control (x-axis) compared to the Gluc-targeting conditions (y-axis) for PspCas13b (blue) and shRNA (black). expression levels by percentage (TPM) values. Mean values of three biological replicates are shown. Gluc transcript data points are labeled. G) Number of significant off-targets from Gluc knockdown for LwaCasl3a, PspCasl3b and shRNA from transcriptome-wide analysis in E and F.
[0160] [Fig. 19] Figure 19 shows engineering of dCas13b-ADAR fusions for RNA editing. A) Schematic of RNA editing by the dCas13b-ADAR fusion protein. B) Schematic of Cypridina luciferase W85X target and targeting guide design. C) Quantification of luciferase activity recovery for Cas13b-dADAR1 (left) and Cas13b-ADAR2-cd (right) tiling guides of length 30, 50, 70 or 84 nt. D) Schematic representation of target sites for targeting Cypridina luciferase W85X. E) Sequencing quantification of A→I editing for a 50nt guide targeting Cypridina luciferase W85X.
[0161] [Fig.20] FIG. 20 shows sequence flexibility measurements for RNA editing by REPAIRv1. A) Schematic of the screen to determine protospacer flanking site (PFS) preference for RNA editing by REPAIRv1. B) Distribution of RNA editing efficiencies for all combinations of 4-N PFS at two different editing sites. C) Quantification of percent editing rate of REPAIRv1 in Cluc W85 at all possible 3-nucleotide motifs. D) Heatmap of 5' and 3' base preferences for RNA editing in Cluc W85 for all possible 3 base motifs.
[0162] [Fig.21] FIG. 21 shows correction of disease-associated mutations by REPAIRv1. A) Schematic of target and guide design for targeting AVPR2 878G>A. B) The 878G>A mutation in AVPR2 is corrected to varying percentages using REPAIRv1 with three different guide designs. C) Schematic of target and guide design for targeting FANCC 1517G>A. D) The 1517G>A mutation in FANCC is corrected to varying percentages using REPAIRv1 with three different guide designs. E) Quantification of percent editability of 34 different disease-associated G>A mutations using REPAIRv1. F) Analysis of all possible G>A mutations that can be corrected as annotated by the ClinVar database. G) Distribution of editing motifs for all G>A mutations in ClinVar versus editing efficiency by REPAIRv1 per motif as quantified on Gluc transcripts.
[0163] [Fig.22] Figure 22 shows the characterization of the specificity of REPAIRv1. A) Schematic representation of KRAS target sites and guide design. B) Quantification of percent edit rate for tiled KRAS targeting guides. The percentage editability at on-target and flanking adenosine sites is shown. For each guide, the region of double-stranded RNA is outlined in red. C) Transcriptome-wide sites of significant RNA editing by REPAIRv1 in Cluc targeting guides. The on-target site Cluc site (254 A>G) is highlighted in orange. D) Transcriptome-wide sites of significant RNA editing by REPAIRv1 in non-targeting guides.
[0164] [Fig.23] FIG. 23 shows rational mutagenesis of ADAR2 to improve the specificity of REPAIRv1. A) Quantification of luciferase signal recovery by various dCas13-ADAR2 mutants and their specificity scores plotted along a schematic for contacts between key ADAR2 deaminase residues and dsRNA targets. A specificity score is defined as the luciferase signal ratio between targeting and non-targeting guide conditions. B) Quantification of luciferase signal recovery by various dCas13-ADAR2 mutants relative to their specificity scores. C) Measurement of on-target editing rates and number of significant off-targets for each dCas13-ADAR2 mutant by transcriptome-wide sequencing of mRNA. D) Transcriptome-wide sites of significant RNA editing by REPAIRv1 and REPAIRv2 at guides targeting premature termination sites in Cluc. The on-target Cluc site (254 A>G) is highlighted in orange. E) RNA-sequencing reads around the on-target Cluc editing site (254 A>G) highlighting off-target editing differences between REPAIRv1 and REPAIRv2. All A>G edits are highlighted in red, while sequencing errors are highlighted in blue. F) RNA editing by REPAIRv1 and REPAIRv2 in guides targeting out-of-frame UAG sites of endogenous KRAS and PPIB transcripts. The percentage of on-target edits for each condition row is shown as a horizontal bar graph on the right. The duplex region formed by the guide RNA is indicated by a red box.
[0165] [Fig.24] Figure 24 shows bacterial screening and PFS determination of Cas13b orthologues for in vivo efficiency. A) Schematic of bacterial assay for determining PFS of Cas13b orthologs. A Cas13b ortholog with a β-lactamase targeting spacer is co-transformed with a β-lactamase expression plasmid and subjected to double selection. B) Quantification of the interfering activity of Cas13b ortholog targeting β-lactamase as measured by colony forming units (cfu). C) PFS logos of Cas13b orthologs as determined by depletion sequences from bacterial assays.
[0166] [Fig.25] Figure 25 shows optimization of Casl3b knockdown and further characterization of mismatch specificity. A) Gluc knockdown by two different guides is measured using the top two Cas13a and top four Cas13b orthologues fused to various nuclear localization and nuclear export tags. B) KRAS knockdown is measured for LwaCas13a, RanCas13b, PguCas13b and PspCas13b at four different guides and compared to four position-matched shRNA controls. C) Schematic of the single and double mismatch plasmid libraries used to assess the specificity of LwaCas13a and PspCas13b knockdown. All possible single and double mismatches are present at the three positions immediately adjacent to the target sequence and the 5' and 3' ends of the target site. D) Transcript depletion levels at the indicated single mismatches are plotted as a heatmap for both LwaCasl3a and PspCasl3b conditions. E) Transcript depletion levels at the indicated double mismatches are plotted as a heatmap for both LwaCasl3a and PspCasl3b conditions.
[0167] [Fig.26]Figure 26 shows the characterization of design parameters for dCas13-ADAR2 RNA editing. A) Knockdown efficiency of Gluc targeting for wild-type Cas13b and catalytically inactive H133A / H1058A Cas13b (dCas13b). B) Quantification of luciferase activity restoration by dCasl3b fused to either the wild-type ADAR2 catalytic domain or the highly active E488Q mutant ADAR2 catalytic domain, tested by tiling the Cluc targeting guide. C) Guide design and sequencing quantification of A→I editing for a 30nt guide targeting Cypridina luciferase W85X. D) Guide design and sequencing quantification of A→I edits for 50nt guides targeting PPIB. E) Effect of linker selection on recovery of luciferase activity by REPAIRv1. F) Effect of opposite base identity of target adenosine on recovery of luciferase activity by REPAIRv1.
[0168] [Fig.27] Figure 27 shows the clinVar motif distribution for G>A mutations. Number of each possible triplet motif observed in the ClinVar database for all G>A mutations.
[0169] [Figure 28A-28B] Figures 28A-28B show RNA binding by truncation of dCas13b. Illustrated various N- and C-terminal truncations of dCasl3b. Comparing activity using targeting and non-targeting guides shows RNA binding when there is ADAR dependent RNA editing as measured by restoration of luciferase signal. Amino acid positions correspond to those of the Prevotella sp. P5-125 Cas13b protein.
[0170] [Fig.29] Figure 29 shows a comparison of other programmable ADAR systems with the dCas13-ADAR2 compilation. A) Schematic of two programmable ADAR schemes: BoxB-based targeting and full-length ADAR2 targeting. In the BoxB scheme (top), the ADAR2 deaminase domain (ADAR2DD(E488Q)) is fused to a small bacterial viral protein called lambda N (λN), which is specific for a small RNA sequence called BoxB-λ. bind to A guide RNA containing two BoxB-λ hairpins can then guide ADAR2DD(E488Q), -λN for site-specific editing. In the full-length ADAR2 scheme (bottom), the dsRNA-binding domain of ADAR2 binds to a guide RNA hairpin, allowing programmable ADAR2 editing. B) Transcriptome-wide sites of significant RNA editing by BoxB-ADAR2DD(E488Q) in Cluc-targeted and non-targeted guides. The on-target Cluc site (254 A>G) is highlighted in orange. C) Transcriptome-wide sites of significant RNA editing by ADAR2 in Cluc-targeted and non-targeted guides. The on-target Cluc site (254 A>G) is highlighted in orange. D) Transcriptome-wide sites of significant RNA editing by REPAIRv1 in Cluc-targeted and non-targeted guides. The on-target Cluc site (254 A>G) is highlighted in orange. E) Quantification of on-target editability percentages of BoxB-ADAR2DD (E488Q), ADAR2 and REPAIRv1 for targeting guides to Cluc. F) Overlap of off-target sites between different targeting and non-targeting conditions for the programmable ADAR system.
[0171] [Fig.30] Figure 30 shows the efficiency and specificity of the dCas13b-ADAR2 mutants. A) Quantification of luciferase activity recovery by the dCas13b-ADAR2DD(E488Q) mutant for Cluc targeting and non-targeting guides. B) Relationship between the ratio of targeting and non-targeting guides and the number of RNA editing off-targets as quantified by transcriptome-wide sequencing. C) Quantification of the number of transcriptome-wide off-target RNA editing sites versus on-target Cluc editing efficiency for the dCas13b-ADAR2DD(E488Q) mutant.
[0172] [Fig.31] Figure 31 shows the transcriptome-wide specificity of RNA editing by the dCas13b-ADAR2DD(E488Q) mutant. A) Transcriptome-wide sites of significant RNA editing by the dCas13b-ADAR2DD(E488Q) mutant in Cluc-targeted guides. The on-target Cluc site (254 A>G) is highlighted in orange. B) Transcriptome-wide sites of significant RNA editing by the dCas13b-ADAR2DD(E488Q) mutant on non-targeting guides.
[0173] [Fig.32] Figure 32 shows the characterization of motif bias in the off-target of dCasl3b-ADAR2DD(E488Q) editing. A) Motifs present across all A>G off-target edits in the transcriptome are shown for each dCas13b-ADAR2DD(E488Q) mutant. B) Distribution of off-target A>G editing by motif identity for REPAIRv1 with targeting and non-targeting guides. C) Distribution of off-target A>G editing by motif identity for REPAIRv2 with targeting and non-targeting guides.
[0174] [Fig.33] Figure 33 shows further characterization of REPAIRv1 and REPAIRv2 off-targets. A) Histogram of off-target numbers per transcript for REPAIRv1. B) Histogram of off-target numbers per transcript for REPAIRv2. C) Mutation effect prediction of REPAIRv1 off-target. D) Distribution of potential oncogenic effects of REPAIRv1 off-targets. E) Mutation effect prediction of REPAIRv2 off-targets. F) Distribution of potential oncogenic effects of REPAIRv2 off-targets.
[0175] [Fig. 34] Figure 34 shows the RNA editing efficiency and specificity of REPAIRv1 and REPAIRv2. A) Quantification of percent editing of KRAS by KRAS targeting guide 1 at target adenosine and flanking sites for REPAIRv1 and REPAIRv2. B) Quantification of percent editing of KRAS by KRAS targeting guide 3 at target adenosine and flanking sites for REPAIRv1 and REPAIRv2. C) Quantification of percent editability of PPIB with PPIB targeting guide 2 at target adenosine and flanking sites for REPAIRv1 and REPAIRv2.
[0176] [Fig.35] Figure 35 shows the demonstration of all potential codon changes with A>G RNA edits. A) Table of all potential codon conversions made possible by A>I editing. B) Codon table demonstrating all potential codon transitions made possible by A>I editing. C).
[0177] [Figure 36A-36B] Figures 36A-36B show the effect of Csx protein on RNA interference. A) Comparison of Cas13b and Cas13b+Csx27 interference; B) Comparison of Cas13b and Cas13b+Csx28 interference.
[0178] [Figure 37A-37C] Figures 37A-37F show a comparison of transcript knockdown by Cas13b and Cas13b+Csx28 using a luciferase assay. A) Pin Cas13b(WP_036860899); B) Pbu Cas13b(WP_004343973); C) Rin Cas13b(WP_004919755); D) Pau Cas13b(WP025000926); E) Pgu Cas13b(WP_039434803); F) Pig Cas13 b(WP_053444417). [Figure 37D-37F] Figures 37A-37F show a comparison of transcript knockdown by Cas13b and Cas13b+Csx28 using a luciferase assay. A) Pin Cas13b(WP_036860899); B) Pbu Cas13b(WP_004343973); C) Rin Cas13b(WP_004919755); D) Pau Cas13b(WP025000926); E) Pgu Cas13b(WP_039434803); F) Pig Cas13 b(WP_053444417). [Mode for carrying out the invention]
[0179] The figures herein are for illustrative purposes only and are not necessarily drawn to scale.
[0180] In general, CRISPR-Cas or the CRISPR system includes sequences encoding the Cas gene, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr mate sequences (“direct repeats" and partially direct repeats processed by tracrRNA), guide sequences (also referred to as "spacers" in the context of the endogenous CRISPR system) or " RNA" (e.g., RNA that guides Cas such as Cas9, e.g., CRISPR RNA and transactivating (tracr) RNA or single guide RNA (sgRNA) (chimeric RNA)) or other sequences and transcripts from the CRISPR locus. Collectively, it refers to transcripts and other elements involved in the expression or directing the activity of CRISPR-associated (“Cas”) genes, including. In general, CRISPR systems are characterized by elements that facilitate the formation of CRISPR complexes at the site of target sequences (also called protospacers in the context of endogenous CRISPR systems). When the CRISPR protein is a class 2 VI-B type effector, tracrRNA is unnecessary. In the engineered system of the invention, direct repeats may comprise naturally occurring or non-naturally occurring sequences. The direct repeats of the invention are not limited to naturally occurring lengths and sequences. Direct repeats can be 36 nt long, but there can be a variety of longer or shorter direct repeats. For example, direct repeats can be 30 nt or more, such as 30-100 nt or more. For example, direct repeats can be 30nt, 40nt, 50nt, 60nt, 70nt, 70nt, 80nt, 90nt, 100nt or more in length. In some embodiments, a direct repeat of the invention may comprise a synthetic nucleotide sequence inserted between the 5' and 3' ends of a naturally occurring direct repeat. In certain embodiments, the inserted sequence may be self-complementary, eg, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% self-complementary. In addition, the direct repeats of the invention may contain insertions of nucleotides (for association with functional domains), such as sequences that bind aptamers or adapter proteins. In certain embodiments, one end of a direct repeat containing such an insertion is approximately the front half of the short DR and this end is approximately the rear half of the short DR.
[0181] In the context of CRISPR complex formation, "target sequence" refers to a sequence to which the guide sequence is designed to be complementary, where hybridization between the target and guide sequences promotes CRISPR complex formation. A target sequence may comprise any polynucleotide, such as a DNA or RNA polynucleotide. In some embodiments, the target sequence is located in the nucleus or cytoplasm of the cell. In some embodiments, direct repeats can be identified in silico by searching for repeat motifs that satisfy some or all of the following criteria: 1. A 2 Kb window of genomic sequence flanking the type II CRISPR locus. present; spanning 2.20-50 bp; and spaced 3.20-50 bp. In some embodiments, two of these criteria may be used, eg, 1 and 2, 2 and 3, or 1 and 3. In some embodiments, all three criteria may be used.
[0182] In an embodiment of the present invention, the terms guide sequence and guide RNA, i.e. RNA having the ability to guide the Cas13b effector protein to target loci, are defined in WO2014 / 093622 (PCT / U.S. Patent Application Publication No. 2013 / 074667), etc., are used interchangeably as in the references cited herein. In general, the guide sequence (or spacer sequence) is any polynucleotide having sufficient complementarity with the target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding to the target sequence by the CRISPR complex. is an array. In some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence is about 50%, 60%, 75%, 80% when optimally aligned using a suitable alignment algorithm. , 85%, 90%, 95%, 97.5%, 99% or higher. Optimal alignment can be determined using any algorithm suitable for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, and algorithms based on the Burroughs-Wheeler transformation. (e.g. Burrows Wheeler Aligners), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (soap.genomics.org.cn) (available at maq.sourceforge.net) and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence (or spacer sequence) is about 5,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25 , 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 nucleotides in length or longer. In some embodiments, the guide sequence is less than or less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12 nucleotides in length. Preferably, the guide sequence is 10-40 nucleotides long, such as 20-30 or 20-40 nucleotides long or more, such as 30 nucleotides long or about 30 nucleotides long. In certain embodiments, the guide sequence for the Cas13b effector is 10-30 nucleotides in length, such as 20-30 or 20-40 nucleotides in length or more, such as 30 nucleotides in length or about 30 nucleotides in length. In certain embodiments, the guide sequence is 10-30 nucleotides long, such as 20-30 nucleotides long, such as 30 nucleotides long. The ability of a guide sequence to direct sequence-specific binding of a CRISPR complex to a target sequence can be assessed by any suitable assay. A host cell in which sufficient components of the CRISPR system to form a CRISPR complex have corresponding target sequences, such as by transfection of vectors encoding the components of the CRISPR sequences, including the guide sequence to be tested. , followed by assessment of preferential cleavage within the target sequence, such as by Surveyor assays as described herein. Similarly, cleavage of a target polynucleotide sequence provides a control guide sequence that is different from the target sequence, the components of the CRISPR complex including the guide sequence to be tested, and the test guide sequence, and controls the test guide sequence. It can be determined in vitro by comparing the rate of binding or cleavage at the target sequence during reactions with the guide sequence. Other assays are possible and will occur to those skilled in the art.
[0183] The present invention provides specific Ca13b effectors, nucleic acids, systems, vectors and methods of use. Any VI-B is distinguishable from VI-A (Cas13b) by structure and also by the position of the HEPN domain). There appears to be little that can separate Cas13b that lacks small accessory proteins from Cas13b that does, with the exception of the VI-B2 locus, which contains a much more highly conserved small accessory protein. appears to have
[0184] As used herein, the terms Ca13b-s1 accessory protein, Cas13b-s1 protein, Cas13b-s1, Csx27 and Csx27 protein are used interchangeably and the terms Cas13b-s2 accessory protein, Cas13b-s2 protein , Cas13b-S2, Csx28 and csx28 proteins are used interchangeably.
[0185] In classical CRISPR-Cas systems, the degrees of complementarity between guide sequences and their corresponding target sequences are approximately 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%. , 99% or 100% or more, or more; , 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 nucleotides long or longer; or the guide or RNA or crRNA is about 75, 50, It can be less than or shorter than 45, 40, 35, 30, 25, 20, 15, 12 nucleotides in length; and advantageously the tracr RNA is 30 or 50 nucleotides in length. However, one aspect of the invention is to reduce off-target interactions, eg, guide interaction with less complementary target sequences. Indeed, in examples, a distinction is made between target sequences and off-target sequences that have greater than 80% to about 95% complementarity, such as 83%-84% or 88-89% or 94-95% complementarity ( For example, mutations that result in a CRISPR-Cas system capable of distinguishing a target with 18 nucleotides from an off-target of 18 nucleotides with 1, 2 or 3 mismatches are shown to be relevant to the present invention. Thus, in the context of the present invention, the degree of complementarity between a guide sequence and its corresponding target sequence is 94.5% or 95% or 95.5% or 96% or 96.5% or 97% or 97.5% or 98% or greater than 98.5% or 99% or 99.5% or 99.9% or 100%. Off-target is 100% or 99.9% or 99.5% or 99% or 99% or 98.5% or 98% or 97.5% or 97% or 96.5% or 96% or 95.5% or 95% or 94.5% or 94% or 93% or 92% or 91% or 90% or 89% or 88% or 87% or 86% or 85% or 84% or 83% or 82% or 81% or less than 80% complementary Advantageously, the off-target is 100% or 99.9% or 99.5% or 99% or 99% or 98.5% or 98% or 97.5% or 97% or 96.5% or 96% or 95.5% or 95% or 94.5% complementarity.
[0186] In particularly preferred embodiments of the present invention, the guide RNA (having the ability to guide Cas to a target locus) comprises (1) a eukaryotic target locus (polynucleotide target locus such as an RNA target locus); (2) a single RNA, ie sgRNA (arranged in 5' to 3' direction) or direct repeat (DR) sequence in crRNA).
[0187] In particular embodiments, the wild-type Cas13b effector protein has RNA binding and cleavage functions.
[0188] In particular embodiments, the Cas13b effector protein may have a DNA-cleaving function. In these embodiments, methods may be provided based on the effector proteins provided herein, which include delivering a vector as discussed herein to a cell. Inducing one or more mutations in a eukaryotic cell as is (in vitro, ie in an isolated eukaryotic cell). This mutation may involve the introduction, deletion or substitution of one or more nucleotides in each target sequence of the cell via the guide RNA or sgRNA or crRNA. This mutation may involve the introduction, deletion or substitution of 1-75 nucleotides in each target sequence of the cell via guide RNA, or sgRNA, or crRNA. The mutations include 1, 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 1, 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 in each target sequence of the cell via guide RNA, or sgRNA, or crRNA; Introductions, deletions or substitutions of 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or 75 nucleotides may be included. 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 14, 15, 16, 17, 18, 19, 20, 21, in each target sequence of the cell via guide RNA, or sgRNA, or crRNA, Introductions, deletions or substitutions of 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or 75 nucleotides may be included. 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, in each target sequence of the cell via guide RNA, or sgRNA, or crRNA. Introductions, deletions or substitutions of 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or 75 nucleotides may be included. The mutations include guide RNA, or sgRNA, or crRNA-mediated target sequences in the cell, Introductions, deletions or substitutions of 45, 50 or 75 nucleotides may be included. The mutation includes the introduction, deletion or substitution of 40, 45, 50, 75, 100, 200, 300, 400 or 500 nucleotides in each target sequence of said cell via guide RNA, sgRNA or crRNA. can be included.
[0189] To minimize toxicity and off-target effects, it may be important to control the concentration of Cas13b mRNA and guide RNA delivered. Optimal concentrations of Cas13b mRNA and guide RNA are determined by testing various concentrations in cellular or non-human eukaryotic animal models and analyzing the degree of alteration at potential off-target genomic loci using deep sequencing. be able to. Alternatively, to minimize toxicity levels and off-target effects, a Cas13b nickase mRNA (e.g., S. pyogenes Cas9 with a D10A mutation) is combined with a pair of guide RNAs targeting sites of interest. can be delivered. Guide sequences and strategies to minimize toxicity and off-target effects can be as in WO2014 / 093622 (PCT / US2013 / 074667), or Mutations as described herein may be used.
[0190] Typically, in the context of the endogenous CRISPR system, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in a target (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50 base pairs or more from the disconnection occurs.
[0191] Nucleic acid molecules encoding Cas13b are advantageously codon-optimized. Examples of codon-optimized sequences in this case are optimized for eukaryotic, e.g., human expression (i.e., optimized for human expression), or are discussed herein. is another eukaryotic, animal or mammalian optimized sequence such as; See the standardization array. Although this is preferred, it is understood that other examples are possible, codon optimization for non-human host species or codon optimization for specific organs are known. In some embodiments, an enzyme-encoding sequence encoding Cas is codon-optimized for expression in a particular cell, eg, a eukaryotic cell. Eukaryotic cells may be derived from certain organisms, e.g., mammals, including but not limited to humans, or non-human eukaryotes or animals or mammals as discussed herein, e.g., mice, rats, rabbits, dogs. , domestic or non-human mammals or primates. In some embodiments, methods of altering human germline genetic identity and / or methods of altering animal genetic identity that may cause distress to humans or animals without any substantial medical benefit to them. Additionally, animals obtained from such methods may be excluded. Generally, codon-optimization involves making at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of a native sequence Refers to a method of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing the codons used more frequently or most frequently in the host cell's genes while maintaining the amino acid sequence. Various species exhibit specific biases for certain codons for particular amino acids. Codon bias (differences in codon usage between organisms) is often correlated with the translation efficiency of messenger RNA (mRNA), which in turn determines, among other things, the characteristics and specific characteristics of the codons translated. It is believed to depend on the availability of transfer RNA (tRNA) molecules. The predominance of the tRNA of choice in the cell generally reflects the codons most frequently used in peptide synthesis. Thus, genes can be tuned for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example in the "Codon Usage Database" available at www.kazusa.orjp / codon / , and these tables can be adapted in a number of ways. See Nakamura, Y., et al. "Codon usage tabulated from the international DNA sequence databases:status for the year 2000" Nucl. Acids Res. 28:292 (2000). Computer algorithms are also available for codon-optimizing a particular sequence for expression in a particular host cell, such as Gene Forge (Aptagen; Jacobus, PA). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more, or all codons) in the sequence encoding Cas are corresponds to the most frequently used codons for the amino acids of
[0192] In certain embodiments, the methods as described herein comprise one or more guide RNAs encoding one or more guide RNAs operably linked in a cell to regulatory elements, including promoters of one or more genes of interest. It may comprise providing Cas13b transgenic cells into which one or more nucleic acids are provided or introduced. As used herein, the term "Cas13b transgenic cell" refers to a cell, such as a eukaryotic cell, into which the Cas13b gene has been genomically integrated. The nature, type or origin of the cells is not particularly limited according to the invention. Also, the method of introducing the Cas13b transgene into the cell can vary and can be any method as known in the art. In certain embodiments, Cas13b transgenic cells are obtained by introducing a Cas13b transgene into an isolated cell. In certain other embodiments, Cas13b transgenic cells are obtained by isolating cells from a Cas13b transgenic organism. By way of example and without limitation, a 13bCas transgenic cell as referred to herein may be derived from a Cas13b transgenic eukaryote, such as a Cas13b knock-in eukaryote. Reference is made to WO2014 / 093622 (PCT / US Patent Application Publication No. 13 / 74667), which is incorporated herein by reference. The methods of US Patent Application Publication Nos. 20120017290 and 20110265198 assigned to Sangamo BioSciences, Inc. for targeting the Rosa locus can be modified to take advantage of the CRISPR Cas system of the present invention. The method of US Patent Application Publication No. 20130236946 assigned to Cellectis for targeting the Rosa locus can also be modified to take advantage of the CRISPR Cas system of the present invention. As a further example, see Platt et. al. (Cell; 159(2):440-455 (2014)) (incorporated herein by reference) describing Cas9 knock-in mice. The Cas13b transgene may further contain a Lox-Stop-polyA-Lox (LSL) cassette, thereby rendering Cas13b expression inducible by Cre recombinase. Alternatively, Cas13b transgenic cells can be obtained by introducing a Cas13b transgene into isolated cells. Transgene delivery systems are well known in the art. By way of example, the Cas13b transgene may be used for vectors (e.g., AAV, adenovirus, lentivirus) and / or particles and / or particle delivery, e.g., in eukaryotic cells, as also described elsewhere herein. can be delivered using
[0193] One skilled in the art will appreciate that cells such as Cas13b transgenic cells as referred to herein are, for example and without limitation, Platt et al. (2014), Chen et al., (2014) or Kumar et al.. (2009), in addition to having the Cas13b gene integrated, it contains additional genomic alterations or Cas13b when complexed with RNA capable of guiding Cas13b to target loci. It will be understood that mutations resulting from the sequence-specific action of , such as one or more oncogenic mutations, can be included.
[0194] In some embodiments, the Cas13b sequence comprises one or more nuclear localization sequences (NLS), such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more, or more. Fused with many NLS or NES. In some embodiments, the Cas13b comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the amino terminus; about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs in the vicinity thereof, or combinations thereof (e.g., zero or at least one NLS at the amino terminus); or more NLSs or zero or more NLSs at the carboxy terminus). When more than one NLS is present, each may be selected independently of the others, thus a single NLS may be present in two or more copies and / or one present in one or more copies. It can exist in combination with other NLS above. In a preferred embodiment of the invention, Casl3b contains at most 6 NLSs. In some embodiments, the NLS is about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40 along the polypeptide chain from the N-terminus or C-terminus to the nearest amino acid of the NLS. , is considered to be near the N-terminus or C-terminus when within 50 or more amino acids. Non-limiting examples of NLSs include the NLS of SV40 viral large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO:X); the NLS of nucleoplasmin (e.g., sequence KRPAATKKAGQAKKKK) (SEQ ID NO:X); NLS; c-myc NLS with amino acid sequence PAAKRVKLD (SEQ ID NO: X) or RQRRNELKRSP (SEQ ID NO: X); hRNPA1 M9 NLS with sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: X); ); the sequences VSRKRPRP (SEQ ID NO:X) and PPPKARED (SEQ ID NO:X) of the fibroid T protein; the sequence POPKKKPL (SEQ ID NO:X) of human p53; the sequence SALIKKKKKMAP (SEQ ID NO:X) of mouse c-abl IV; sequence DLRRR (SEQ ID NO:X) and PKQKKRK (SEQ ID NO:X); sequence RKLKKKIKKL (SEQ ID NO:X) of hepatitis virus delta antigen; sequence REKKKFLKRR (SEQ ID NO:X) of mouse Mx1 protein; sequence KRKGDEVDGVDEVAKKKSKK of human poly(ADP-ribose) polymerase (SEQ ID NO:X); and NLS sequences derived from the steroid hormone receptor (human) glucocorticoid sequence RKCLQAGMNLEARKTKK (SEQ ID NO:X). In general, one or more NLSs are strong enough to drive the accumulation of detectable amounts of Cas in the nucleus of eukaryotic cells. In general, the strength of nuclear localization activity can be derived from the number of NLSs within Cas, the specific NLS used, or a combination of these factors. Detection of nuclear accumulation may be performed by any suitable technique. For example, a detectable marker can be fused to Cas to visualize its location within the cell, such as in combination with a means of detecting the location of the nucleus (e.g., a nuclear-specific stain such as DAPI). . Cell nuclei can also be isolated from cells and then their contents analyzed by any suitable protein detection method such as immunohistochemistry, Western blot or enzyme activity assay. Accumulation in the nucleus is assayed for the effects of CRISPR complex formation (e.g., assays for DNA breaks or mutations in target sequences or assays for altered gene expression activity under the influence of CRISPR complex formation and / or Cas enzymatic activity). ) compared to controls not exposed to Cas or complexes or to controls exposed to Cas lacking one or more NLSs.
[0195] In certain embodiments, the present invention provides, for example, for delivering or introducing into cells Cas13b and / or RNA having the ability to guide Cas13b to a target locus (i.e., guide RNA), and for introducing these components (e.g., It relates to vectors, for propagating in prokaryotic cells). As used herein, a "vector" is a tool that enables or facilitates the transfer of an entity from one environment to another. This is a replicon, such as a plasmid, phage or cosmid, into which another DNA segment can be inserted, resulting in replication of the inserted segment. In general, vectors are replication competent when associated with appropriate regulatory elements. In general, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors include, but are not limited to, single-stranded, double-stranded or partially double-stranded nucleic acid molecules; nucleic acid molecules without free ends (e.g., circular) containing one or more free ends; DNA, RNA; or both; and various other polynucleotides known in the art. Certain vectors are "plasmids," which refer to circular double-stranded DNA loops into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, wherein vectors include packaging into viruses (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses and adeno-associated viruses (AAV)). There are viral DNA or RNA sequences for A viral vector also includes a virus-carrying polynucleotide for transfection into a host cell. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (eg, bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (eg, non-episomal mammalian vectors) integrate into the genome of the host cell upon introduction into the host cell, thereby replicating along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors". Common expression vectors of utility in recombinant DNA techniques are often in the form of plasmids.
[0196] A recombinant expression vector can contain a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, i.e., the recombinant expression vector is operably linked to the nucleic acid sequence to be expressed. It is meant to contain one or more regulatory elements, which can be selected based on the host cell used for expression. Within the scope of recombinant expression vectors, "operably linked" refers to sequences capable of expression (e.g., in an in vitro transcription / translation system or in a host cell if the vector is introduced into the host cell). is intended to mean that the nucleotide sequence of interest is linked to the regulatory element in any manner. U.S. Patent Application No. 10 / 815,730, published September 2, 2004 as U.S. Patent Application Publication No. 2004-0171156 A1, the contents of which are incorporated herein in their entirety, for recombination and cloning methods. incorporated by reference as).
[0197] A vector may include regulatory elements such as promoters. The vector may contain a Cas13b coding sequence and / or a single but possibly a guide RNA (e.g. crRNA) coding sequence, which may comprise at least 3 or 8 or 16 or 32 or 48 or 50 e.g. ~3, 1~4, 1~5, 3~6, 3~7, 3~8, 3~9, 3~10, 3~8, 3~16, 3~30, 3~32, 3~48 , can contain 3-50 RNAs (eg, crRNAs). There may be a promoter for each RNA (e.g. crRNA), advantageously if there are no more than about 16 RNAs (e.g. crRNAs) in a single vector; crRNA), one or more promoters can drive the expression of two or more of those RNAs (e.g. crRNAs), e.g. if there are 32 RNAs (e.g. sgRNAs or crRNAs), If each promoter can drive the expression of 2 RNAs (e.g. sgRNA or crRNA) and there are 48 RNAs (e.g. sgRNA or crRNA), then each promoter can drive 3 RNAs (e.g. sgRNA or crRNA). ) expression. Simple arithmetic, well-established cloning protocols and the teachings of the present disclosure will enable one skilled in the art to generate RNA, e.g., sgRNA or crRNA, and a suitable promoter, such as the U6 promoter, for suitable exemplary vectors such as AAV The invention can be readily implemented with U6-sgRNA or -crRNA. For example, the packaging limit for AAV is approximately 4.7 kb. A person skilled in the art can easily fit about 12-16, for example 13, U6-sgRNA or crRNA cassettes in a single vector. This can be assembled by any suitable means, such as the golden gate strategy used for TALE assembly (http: / / www.genome-engineering.org / taleffectors / ). A person skilled in the art can increase the number of U6-sgRNAs or -crRNAs by about 1.5-fold, for example, from 12-16, such as 13, to about 18-24, such as about 19 U6-sgRNAs or -crRNAs, in tandem. Guided strategies can also be used. Thus, one skilled in the art can easily arrive at about 18-24, such as about 19 promoter-RNAs, such as U6-sgRNA or -crRNA, in a single vector, such as an AAV vector. A further means of increasing the number of promoters and RNAs, such as sgRNAs or crRNAs, in a vector is to use a single promoter (e.g., U6) with RNAs, such as sgRNAs or crRNAs, separated by cleavable sequences. to express the array. And yet another means of increasing the number of promoter-RNAs, such as sgRNAs or crRNAs, in a vector is to express an array of promoter-RNAs, such as sgRNAs or crRNAs, separated by cleavable sequences in the coding sequence or introns of the gene. and in this case it may be advantageous to use a polymerase II promoter, which may increase expression and allow transcription of long RNAs in a tissue-specific manner (e.g. nar.oxfordjournals.org / content / 34 / 7 / e53.short, see www.nature.com / mt / journal / v16 / n9 / abs / mt2008144a.html). In advantageous embodiments, AAV can package U6 tandem sgRNAs targeting up to about 50 genes. Thus, from the knowledge in the art and the teachings of this disclosure, one skilled in the art will be able to identify a plurality of RNAs under the control of, or operably or functionally linked to, one or more promoters, or guides, or sgRNAs or crRNA-expressing vectors, such as single vectors, are easily made and used without any undue experimentation (especially with respect to the number of RNAs or guides or sgRNAs or crRNAs discussed herein) be able to.
[0198] Guide RNAs, such as sgRNA or crRNA coding sequences and / or Cas13b coding sequences, may be functionally or operably linked to regulatory elements, which thus drive expression. The promoter can be a constitutive promoter and / or a conditional promoter and / or an inducible promoter and / or a tissue specific promoter. Promoters include RNA polymerase, pol I, pol II, pol III, T7, U6, H1, retrovirus Rous sarcoma virus (RSV) LTR promoter, cytomegalovirus (CMV) promoter, SV40 promoter, dihydrofolate reductase promoter, β- It may be selected from the group consisting of actin promoter, phosphoglycerol kinase (PGK) promoter and EF1α promoter. An advantageous promoter is the promoter U6.
[0199] In one aspect of the invention, the novel RNA targeting system, also called RNA- or RNA-targeting CRISPR system of the present application, is based on the Cas13b protein identified herein, which targets specific RNA sequences. Rather, a single enzyme can be programmed to recognize a specific RNA target by an RNA molecule; can be mobilized to
[0200] In some embodiments, one or more elements of the nucleic acid targeting system are derived from a particular organism that contains an endogenous CRISPR RNA targeting system. In certain embodiments, this CRISPR RNA targeting system is found in the genera Eubacterium and Ruminococcus. In certain embodiments, the effector protein comprises targeted and collateral ssRNA cleavage activity. In certain embodiments, the effector protein comprises dual HEPN domains. In certain embodiments, the effector protein lacks a counterpart to the helical-1 domain of Casl3a. In certain embodiments, the effector protein is smaller than previously characterized Class 2 CRISPR effectors, with a median size of 928 aa. The median size is 190 aa (17%) smaller than Casl3c, >200 aa (18%) smaller than Casl3b, and >300 aa (26%) smaller than Casl3a. In certain embodiments, effector proteins have no requirement for flanking sequences (eg, PFS, PAM).
[0201] In certain embodiments, the effector protein locus structure comprises WYL domains containing accessory proteins (so named after the three conserved amino acids within these domains originally identified; e.g. , see WYL domain IPR026881). In certain embodiments, the WYL domain accessory protein comprises at least one helix-turn-helix (HTH) or ribbon-helix-helix (RHH) DNA binding domain. In certain embodiments, the WYL domain containing accessory protein increases both the targeted and collateral ssRNA cleavage activity of the RNA targeting effector protein. In certain embodiments, the WYL domain containing the accessory protein contains a pattern of predominantly hydrophobic conserved residues, including the N-terminal RHH domain as well as an invariant tyrosine-leucine doublet corresponding to the original WYL motif. including groups. In certain embodiments, the WYL domain containing accessory protein is WYL1. WYL1 is a single WYL domain protein primarily related to the genus Ruminococcus.
[0202] In another exemplary embodiment, the Type VI RNA-targeting Cas enzyme is Cas 13d. In certain embodiments, Cas13d is Eubacterium siraeum DSM 15702 (EsCas13d) or Ruminococcus sp. N15.MGS-57 (RspCas13d) (eg, Yan et al. , “Cas13d Is a Compact RNA-Targeting Type VI CRISPR Effector Positively Modulated by a WYL-Domain-Containing Accessory Protein”, Molecular Cell (2018), doi.org / 10.1016 / j.molcel.2018.02.028) . RspCasl3d and EsCasl3d do not have flanking sequence requirements (eg PFS, PAM).
[0203] Nucleic acid targeting systems, vector systems, vectors and compositions described herein are useful for a variety of nucleic acid targeting applications, alteration or modification of the synthesis of gene products such as proteins, nucleic acid cleavage, nucleic acid editing, splicing of nucleic acids; target nucleic acids; , tracking of target nucleic acids, isolation of target nucleic acids, visualization of target nucleic acids, and the like.
[0204] In an advantageous embodiment, the invention encompasses Cas13b effector proteins associated with Table 1. The Cas13b effector proteins of Table 1 are as discussed in further detail herein in conjunction with Table 1.
[0205] Cas13b nuclease A Cas13b effector protein of the invention is, is, or comprises, consists essentially of, or consists of, or is such a protein from or as shown in Table 1. to be concerned with or to relate to The invention provides, relates to, involves, or uses a protein from or as set forth in Table 1, including mutations or modifications thereof as set forth herein. intended to comprise or consist essentially of or consist of; The Cas13b effector proteins of Table 1 are as discussed in further detail herein in conjunction with Table 1.
[0206] Thus, in some embodiments, the effector protein can be an RNA binding protein, such as a dead Cas-type effector protein, which is optionally a transcriptional activator or repressor, as described herein. domain, NLS or other functional domains. In some embodiments, the effector protein can be an RNA binding protein that cleaves a single strand of RNA. If the binding RNA is ssRNA, the ssRNA is completely cleaved. In some embodiments, the effector protein can be an RNA binding protein that double-strands cleaves RNA, eg, if it contains two RNase domains. If the binding RNA is dsRNA, the dsRNA is completely cleaved. In some embodiments, the effector protein can be an RNA binding protein with nickase activity, ie it binds to dsRNA but cleaves only one of the RNA strands.
[0207] The RNase function in the CRISPR system is well known, and for example, mRNA targeting has been reported for certain type III CRISPR-Cas systems (Hale et al., 2014, Genes Dev, vol. 28, 2432-2443; Hale et al., 2009, Cell, vol. 139, 945-956; Peng et al., 2015, Nucleic acids research, vol. 43, 406-417), offer great advantages. Accordingly, CRISPR-Cas systems, compositions or methods for targeting RNA by the present effector proteins are provided.
[0208] The target RNA, or RNA of interest, is the RNA to be targeted by the present invention that leads to the recruitment and binding of the effector protein to the target site of interest on the target RNA. Target RNA can be any suitable form of RNA. This may include mRNA in some embodiments. In other embodiments, target RNA may include tRNA or rRNA.
[0209] Interfering RNA (RNAi) and microRNA (miRNA) In other embodiments, target RNAs may include interfering RNAs, ie, RNAs involved in the RNA interference pathway, such as shRNAs, siRNAs, and the like. In other embodiments, target RNAs may include microRNAs (miRNAs). Regulation of interfering RNAs or miRNAs can help reduce off-target effects (OTEs) seen in approaches by shortening the life span of interfering RNAs or miRNAs in vivo or in vitro.
[0210] If the effector protein and suitable guide are selectively expressed (e.g. under the control of a suitable promoter, e.g. tissue-specific or cell-cycle-specific promoter and / or enhancer) in space or time, this may be used. could "protect" a cell or system (in vivo or in vitro) from RNAi in that cell. This is useful for comparison in adjacent tissues or cells where RNAi is not required, or in cells or tissues in which effector proteins and suitable guides are expressed and not expressed (i.e. RNAi unregulated and RNAi regulated, respectively). could be. Effector proteins may be used to regulate or bind to molecules comprising or consisting of RNA, such as ribozymes, ribosomes or riboswitches. In embodiments of the invention, the RNA guide recruits effector proteins to such molecules, thereby allowing the effector proteins to bind to them.
[0211] Ribosomal RNA (rRNA) For example, azalide antibiotics such as azithromycin are well known. It targets and destroys the 50S ribosomal subunit. The effector protein, along with a suitable guide RNA that targets the 50S ribosomal subunit, can in some embodiments be recruited to and bind to the 50S ribosomal subunit. Accordingly, the subject effector proteins are provided together with suitable guides directed to the ribosome (particularly the 50s ribosomal subunit). This use effector protein use in conjunction with suitable guides directed to the ribosome (particularly the 50s ribosomal subunit) may include use as an antibiotic. Specifically, its use as an antibiotic mimics the action of azalide antibiotics such as azithromycin. In some embodiments, prokaryotic ribosomal subunits may be targeted, such as the prokaryotic 70S subunit, the 50S subunit, 30S subunit and 16S and 5S subunits described above. In other embodiments, eukaryotic ribosomal subunits may be targeted, such as eukaryotic 80S, 60S, 40S and 28S, 18S, 5.8S and 5S subunits.
[0212] The effector protein can be an RNA binding protein, optionally functionalized as described herein. In some embodiments, the effector protein can be an RNA binding protein that cleaves a single strand of RNA. In either case, however, ribosome function may be modulated, particularly reduced or disrupted, particularly when the RNA binding protein cleaves a single strand of RNA. This can be applied to any ribosomal RNA and any ribosomal subunit and the sequences of rRNA are well known.
[0213] Therefore, regulation of ribosomal activity through the use of the subject effector proteins in conjunction with suitable guides to ribosomal targeting is envisioned. This may be due to cleavage or binding to the ribosome. Specifically, a reduction in ribosomal activity is envisaged. It can be useful in in vivo or in vitro assays of ribosome function, and also as a means of controlling in vivo or in vitro therapies based on ribosomal activity. Further, regulation (ie, reduction) of protein synthesis in in vivo or in vitro systems is envisioned, including use as antibiotics and research and diagnostic uses.
[0214] riboswitch Riboswitches (also known as aptazymes) are regulatory segments of messenger RNA molecules that bind small molecules. This typically results in altered production of the protein encoded by the mRNA. Thus, regulation of riboswitch activity through the use of the subject effector proteins in conjunction with suitable guides to riboswitch targets is therefore envisioned. This may be due to cleavage of the riboswitch or binding to it. Specifically, a reduction in riboswitch activity is envisioned. This may be useful in in vivo or in vitro assays of riboswitch function, and also as a means of controlling in vivo or in vitro therapeutics based on riboswitch activity. Furthermore, regulation (ie, reduction) of protein synthesis in in vivo or in vitro systems is envisioned. This control, as far as rRNA is concerned, can include antibiotic use as well as research and diagnostic uses.
[0215] Ribozyme Ribozymes are RNA molecules that have catalytic properties, similar to enzymes (which, of course, are proteins). Ribozymes, both naturally-occurring and engineered, comprise or consist of RNA and thus can be targets of the present RNA-binding effector proteins as well. In some embodiments, the effector protein can be an RNA binding protein that cleaves the ribozyme, thereby rendering it ineffective. Thus, regulation of ribozyme activity through the use of the subject effector proteins in conjunction with suitable guides to ribozyme targets is envisioned. This may be due to cleavage of the ribozyme or binding to it. Specifically, a reduction in ribozyme activity is envisioned. This can be useful in in vivo or in vitro assays of ribozyme function, and also as a means of controlling in vivo or in vitro therapies based on ribozyme activity.
[0216] Gene expression including RNA processing Effector proteins, along with suitable guides, can also be used to target gene expression, including by regulating RNA processing. Control of RNA processing includes RNA splicing, including alternative splicing by targeting RNApol; viral replication, including plant viroids (particularly satellite viruses, bacteriophages and retroviruses such as HBV, HBC and HIV and this those of other viruses listed herein); and RNA processing reactions such as tRNA biogenesis. Effector proteins and suitable guides can also be used to control RNA activation (RNAa). RNAa leads to stimulation of gene expression, so regulation of gene expression can be achieved by disrupting or reducing RNAa, thus reducing stimulation of gene expression.
[0217] RNAi screen RNAi screens can identify gene products whose knockdown is associated with phenotypic changes to probe biological pathways and identify components. Effector proteins and suitable guides are used to remove or reduce the activity of RNAi in the screen and thus the activity of the (previously interfered) gene product (by removing or reducing interference / suppression). Restoration may exert control over or between these screens.
[0218] Satellite RNA (satRNA) and satellite viruses can also be processed.
[0219] Regulation herein in relation to RNase activity generally means reduction, negative disruption or knockdown or knockout.
[0220] In vivo RNA application Inhibition of gene expression The target-specific RNases provided herein are capable of highly specific cleavage of target RNA. Interference at the RNA level can be regulated both spatially and temporally and in a non-invasive manner since the genome is not modified.
[0221] A number of diseases have been demonstrated to be treatable by mRNA targeting. Although most of those studies concern administration of siRNA, it is clear that the RNA targeting effector proteins provided herein are equally applicable.
[0222] Examples of mRNA targets (and corresponding disease treatments) are VEGF, VEGF-R1 and RTP801 (for treating AMD and / or DME), Caspase 2 (for treating Naion), ADRB2 (for treating intraocular pressure), TRPVI (for the treatment of dry eye syndrome, Syk kinase (for the treatment of asthma), Apo B (for the treatment of hypercholesterolemia), PLK1, KSP and VEGF (for the treatment of solid tumors), Ber-Abl (for the treatment of CML) ) (Burnett and Rossi Chem Biol. 2012, 19(1):60-71)). Similarly, RNA targeting has been demonstrated to be effective in treating RNA virus-mediated diseases such as HIV (HIV Tet and Rev targeting), RSV (RSV nucleocapsid targeting) and HCV (miR-122 targeting). (Burnett and Rossi Chem Biol. 2012, 19(1):60-71).
[0223] Further, it is envisioned that the RNA targeting effector proteins of the invention may be used for mutation-specific or allele-specific knockdown. Guide RNAs can be designed to specifically target sequences of transcribed mRNAs containing mutations or allele-specific sequences. Such specific knockdown is particularly suitable for therapeutic applications for disorders associated with mutated or allele-specific gene products. For example, most cases of familial hypobetalipoproteinemia (FHBL) are caused by mutations in the ApoB gene. This gene encodes two versions of the apolipoprotein B protein: a short version (ApoB-48) and a longer version (ApoB-100). Several ApoB gene mutations leading to FHBL make both versions of ApoB abnormally short. Specific targeting and knockdown of mutated ApoB mRNA transcripts with the RNA targeting effector proteins of the invention may be beneficial in the treatment of FHBL. As another example, Huntington's disease (HD) is caused by expansion of the CAG triplet repeat of the gene encoding the huntingtin protein, resulting in an abnormal protein. Specific targeting and knockdown of mutated or allele-specific mRNA transcripts encoding huntingtin protein with the RNA targeting effector proteins of the invention may be beneficial in the treatment of HD.
[0224] It is noted that for various applications in this context and more generally as described herein, the use of split versions of RNA targeting effector proteins may be envisioned. Indeed, this may not only allow for increased specificity, but may also be advantageous for delivery. Cas13b is split in the sense that the two parts of the Cas13b enzyme are essentially functional Cas13b. Ideally, the split should always be such that the catalytic domain remains unaffected. This Cas13b may function as a nuclease, or it may be dead Cas13b, an RNA-binding protein with essentially little or no catalytic activity, typically due to mutations in its catalytic domain.
[0225] Each half of the split Cas13b can be fused to a dimerization partner. By way of example and not limitation, the use of a rapamycin-sensitive dimerization domain allows the creation of chemically inducible split Cas13b to temporally control Cas13b activity. Thus, Casl3b can be made chemically inducible by splitting it into two fragments, and this rapamycin-sensitive dimerization domain can be used to reassemble Casl3b in a controlled manner. The two parts of split Cas13b can be considered as the N'-terminal part and the C'-terminal part of split Cas13b. This fusion is typically the split point of Cas13b. In other words, the C'-terminus of the N'-terminal part of the split Cas13b is fused to one half-dimer, while the N'-terminus of the C'-terminal part is fused to the other half-dimer.
[0226] Cas13b need not be split in the sense that the breakpoint is newly created. The split point is typically designed in silico and cloned into the construct. Together, the two parts of the split Cas13b, the N'-terminal part and the C'-terminal part, preferably contain at least 70% or more of the wild-type amino acids (or the nucleotides encoding them), preferably the wild-type amino acids (or forming a complete Casl3b comprising at least 80% or more, preferably at least 90% or more, preferably at least 95% or more and most preferably at least 99% or more of the nucleotides encoding it. It is possible that there is some trimming and mutants are envisioned. Non-functional domains can be completely removed. The important point is that the two parts can be brought together and the desired Cas13b function restored or reverted. Dimers can be homodimers or heterodimers.
[0227] In certain embodiments, Cas13b effectors as described herein may be used for mutation-specific or allele-specific targeting, such as mutation-specific or allele-specific knockdown.
[0228] The RNA targeting effector protein may also synergize with another functional RNase domain such as a non-specific RNase or Argonaute 2 to increase RNase activity or ensure further degradation of the message. can be fused.
[0229] Regulation of gene expression by regulation of RNA function Apart from the direct effects on gene expression through mRNA cleavage, RNA targeting can also be used to influence specific aspects of RNA processing within the cell, thereby making gene expression more sensitive. It may be possible to adjust to In general, modulation can be mediated, for example, by interfering with protein binding to RNA, eg, by blocking protein binding or recruiting RNA binding proteins. Indeed, regulation can be ensured at various levels, such as splicing, transport, localization, translation and turnover of mRNA. In the therapeutic context as well, it can be envisioned to address (pathogenic) dysfunctions at each of these levels by using RNA-specific targeting molecules. In these embodiments, often the RNA targeting protein has lost the ability to cleave the RNA target, but retains its ability to bind to it, such as the mutant forms of Cas13b described herein. preferably "dead" Cas13b.
[0230] a) alternative splicing Many human genes express multiple mRNAs as a result of alternative splicing. A variety of diseases have been shown to be associated with aberrant splicing leading to loss-of-function or gain-of-function of expressed genes. Some of these diseases are caused by mutations that cause splicing defects, but many others are not. One therapeutic option is to directly target the splicing machinery. The RNA targeting effector proteins described herein can, for example, block or promote slicing, influence exon inclusion or exclusion and expression of specific isoforms and / or stimulate expression of alternative protein products. can be used for Such applications are described in more detail below.
[0231] Binding of the RNA targeting effector protein to the target RNA can sterically block access of splicing factors to the RNA sequence. An RNA targeting effector protein that targets a splice site can block splicing at that site and optionally redirect splicing to an adjacent site. For example, binding of an RNA targeting effector protein that binds to the 5' splice site can block recruitment of the U1 component of the spliceosome, favoring its exon skipping. Alternatively, RNA targeting effector proteins that target splicing enhancers or silencers can prevent binding of trans-acting regulatory splicing factors at target sites, effectively blocking or promoting splicing. In addition, exon exclusion can be achieved by recruiting ILF2 / 3 near exons of precursor mRNAs by RNA targeting effector proteins as described herein. As yet another example, a glycine-rich domain can be added for hnRNP A1 recruitment and exon exclusion (Del Gatto-Konczak et al. Mol Cell Biol. 1999 Jan;19(1):251-60 ).
[0232] In certain embodiments, proper selection of the gRNA will allow certain splice variants to be targeted, while others will not be targeted.
[0233] In some cases, RNA targeting effector proteins can be used to promote slicing (eg, if splicing is deficient). For example, an RNA targeting effector protein can be associated with an effector capable of stabilizing a splicing regulatory stem-loop for further splicing. Linking an RNA targeting effector protein to the consensus binding site sequence of a particular splicing factor allows the protein to be recruited to target DNA.
[0234] Examples of diseases associated with aberrant splicing include, but are not limited to, paraneoplastic oculoclonus myoclonic ataxia (or POMA) and non-function caused by loss of Nova proteins that regulate the splicing of proteins that function at synapses. Cystic fibrosis caused by defective splicing of the cystic fibrosis transmembrane conductance regulator, which leads to the production of sex chloride channels. In other diseases, aberrant RNA splicing leads to gain-of-function. This applies, for example, to myotonic dystrophy caused by CUG triplet repeat expansions (50->1500 repeats) in the 3'UTR of mRNAs that cause splicing defects.
[0235] RNA targeting effector proteins can be used to exclude exons by recruiting splicing factors (such as U1) to the 5' splice site to facilitate excision of introns surrounding the desired exon. Such recruitment may be mediated by fusion with an arginine / serine-rich domain that functions as a splicing activator (Gravely BR and Maniatis T, Mol Cell. 1998(5):765-71).
[0236] It is envisioned that RNA targeting effector proteins could be used to block the splicing machinery at the desired locus, thereby preventing exon recognition and expression of another protein product. An example of a treatable disorder is Duchenne muscular dystrophy (DMD) caused by mutations in the gene encoding the dystrophin protein. Almost all DMD mutations lead to frameshifts and impaired dystrophin translation. Combining RNA targeting effector proteins with splice junctions or exon splicing enhancers (ESEs), thereby preventing exon recognition, can result in translation of partially functional proteins. This converts the lethal Duchenne phenotype to the less severe Becker phenotype.
[0237] b) RNA modification RNA editing is a natural process by which small modifications of RNA increase the diversity of gene products for a given sequence. Typically, this modification involves conversion of adenosine (A) to inosine (I) resulting in an RNA sequence different from that encoded by the genome. RNA modifications are generally ensured by ADAR enzymes, whereby pre-RNA targets form imperfect double-stranded RNAs by base-pairing between exons containing edited adenosines and non-coding intronic elements. do. A classic example of A-I editing is the glutamate receptor GluR-B mRNA, according to which this change alters the conductance properties of the channel (Higuchi M, et al. Cell. 1993;75:1361). -70).
[0238] In humans, heterozygous functional null mutations in the ADAR1 gene lead to a skin disease, human hereditary pigmentary dermatosis (Miyamura Y, et al. Am J Hum Genet. 2003;73:693-9). It is envisioned that the RNA targeting effector proteins of the invention may be used to correct dysfunctional RNA modifications.
[0239] Further, it is envisioned that RNA adenosine methylase (N(6)-methyladenosine) may be fused to the RNA targeting effector proteins of the invention to target transcripts of interest. This methylase causes reversible methylation, has a regulatory role, and can influence gene expression and cell fate decisions by regulating multiple RNA-associated cellular pathways (Fu et al Nat Rev Genet.2014 15(5):293-306).
[0240] c) Polyadenylation Polyadenylation of mRNA is important for nuclear transport, translational efficiency and stability of mRNA, and all of these and also polyadenylation processes are dependent on specific RBPs. Many eukaryotic mRNAs receive a 3' poly(A) tail of approximately 200 nucleotides after transcription. Polyadenylation involves a variety of RNA-binding protein complexes that stimulate the activity of poly(A) polymerase (Minvielle-Sebastia L et al. Curr Opin Cell Biol. 1999;11:352-7). It is envisioned that the RNA targeting effector proteins provided herein may be used to interfere with or facilitate the interaction between RNA binding proteins and RNA.
[0241] An example of a disease that has been associated with defective proteins involved in polyadenylation is oculopharyngeal muscular dystrophy (OPMD) (Brais B, et al. Nat Genet. 1998;18:164-7).
[0242] d) RNA nuclear export After pre-mRNA processing, mRNA is transported from the nucleus to the cytoplasm. This is ensured by cellular mechanisms involving the generation of carrier complexes, which then translocate through the nuclear pore and release the mRNA in the cytoplasm, followed by recycling of the carrier.
[0243] Overexpression of proteins that play a role in nuclear export of RNA (such as TAP) has been found to increase nuclear export of transcripts that are originally inefficiently exported in Xenopus laevis (Katahira J, et al.EMBO J.1999;18:2593-609).
[0244] e) mRNA localization mRNA localization ensures spatially regulated protein production. Localization of a transcript to a particular region of the cell can be ensured by a localization element. In particular embodiments, it is envisioned that the effector proteins described herein may be used to target localization elements to RNAs of interest. Effector proteins can be designed to bind a target transcript and shuttle it to a location within the cell determined by its peptide signal tag. For example, more specifically, an RNA targeting effector protein fused to a nuclear localization signal (NLS) can be used to alter RNA localization.
[0245] Further examples of localization signals include the zipcode binding protein (ZBP1) that ensures cytoplasmic localization of β-actin in some asymmetric cell types, the KDEL retention sequence (endoplasmic reticulum localization ), nuclear export signal (localization to the cytoplasm), mitochondrial targeting signal (localization to mitochondria), peroxisomal targeting signal (localization to peroxisomes) and m6A-labeled / YTHDF2 (localization to p-body conversion). Another approach envisioned is the fusion of RNA targeting effector proteins with proteins of known localization (eg membrane, synapse).
[0246] Alternatively, effector proteins according to the invention can be used, for example, for localization-dependent knockdown. By fusing the effector protein with an appropriate localization signal, the effector is targeted to a specific subcellular compartment. Only target RNAs in this compartment will be effectively targeted, while targets in other identical but different subcellular compartments are not, thus establishing localization-dependent knockdown. can be done.
[0247] f) translation The RNA targeting effector proteins described herein can be used to enhance or repress translation. Translational upregulation is envisioned as a highly robust method of regulating cellular circuits. Furthermore, for functional studies, protein translation screens may be advantageous compared to transcriptional upregulation screens, which suffer from the drawback that upregulation of transcripts does not lead to increased protein production.
[0248] The RNA targeting effector proteins described herein can be used to bring a translation initiation factor such as EIF4G into close proximity to the 5' untranslated repeat (5'UTR) of the messenger RNA of interest to drive translation. is assumed (as described in De Gregorio et al. EMBO J. 1999;18(17):4865-74 for non-reprogrammable RNA binding proteins). As another example, the cytoplasmic poly(A) polymerase GLD2 can be recruited to target mRNAs by RNA targeting effector proteins. This may allow directed polyadenylation of target mRNAs and thereby stimulation of translation.
[0249] Similarly, the RNA targeting effector proteins envisioned herein can be used to block translational repressors of mRNA, such as ZBP1 (Huttelmaier S, et al. Nature. 2005;438:512-5). . Binding of the target RNA to the translation initiation site can directly affect translation.
[0250] Additionally, fusing the RNA targeting effector protein to a protein that stabilizes the mRNA, eg, by preventing its degradation, such as an RNase inhibitor, can increase protein production from the transcript of interest.
[0251] It is envisioned that the RNA targeting effector proteins described herein can be used to bind to the 5'UTR region of RNA transcripts and inhibit translation by preventing ribosome formation and translation initiation.
[0252] Furthermore, recruitment of Caf1, a component of the CCR4-NOT deadenylase complex, to target mRNAs using RNA targeting effector proteins can result in deadenylation of target transcripts and inhibition of protein translation.
[0253] For example, the RNA targeting effector proteins of the invention can be used to increase or decrease translation of therapeutically relevant proteins. Examples of therapeutic applications in which RNA targeting effector proteins may be used to downregulate or upregulate translation are amyotrophic lateral sclerosis (ALS) and cardiovascular disorders. Decreased levels of the glial glutamate transporter EAAT2 have been reported in ALS motor cortex and spinal cord, and multiple abnormal EAAT2 mRNA transcripts have also been reported in ALS brain tissue. Loss of EAAT2 protein and function is believed to be the major cause of excitotoxicity in ALS. Restoration of EAAT2 protein levels and function may provide therapeutic benefit. Thus, RNA targeting effector proteins can be beneficially used to upregulate expression of the EAAT2 protein, for example by blocking translational repressors or stabilizing mRNA as described above. Apolipoprotein A1 is the major protein component of high density lipoprotein (HDL), and ApoA1 and HDL are generally considered to be anti-atherogenic. It is envisioned that RNA targeting effector proteins may be beneficially used to upregulate expression of ApoA1, for example by blocking translational repressors or stabilizing mRNA as described above.
[0254] g) mRNA turnover Translation is closely linked to mRNA turnover and regulatory mRNA stability. Transcript stability has been described to involve specific proteins (ELAV / Hu proteins in neurons, Keene JD, 1999, Proc Natl Acad Sci USA.96:5-7) and tristetraproline (TTP). )Such. These proteins stabilize target mRNAs by protecting the message from degradation in the cytoplasm (Peng SS et al., 1988, EMBO J. 17:3461-70).
[0255] It is envisioned that the RNA targeting effector proteins of the invention may be used to interfere with or promote the activity of proteins that serve to stabilize mRNA transcripts such that mRNA turnover is affected. obtain. For example, recruitment of human TTP to target RNA using an RNA targeting effector protein can enable adenylate-uridylate-rich element (AU-rich element)-mediated translational repression and target degradation. AU-rich elements are found in the 3'UTRs of many mRNAs encoding proto-oncogenes, nuclear transcription factors and cytokines, and promote RNA stability. As another example, an RNA targeting effector protein is fused to another mRNA stabilizing protein, HuR (Hinman MN and Lou H, Cell Mol Life Sci 2008;65:3168-81) to recruit it to target transcripts. can extend its life span or stabilize short-lived mRNAs.
[0256] Further, it is envisioned that the RNA targeting effector proteins described herein may be used to promote degradation of target transcripts. For example, recruitment of m6A methyltransferase to the target transcript and localization of the transcript to P-bodies can result in target degradation.
[0257] As yet another example, fusing an RNA targeting effector protein as described herein to the non-specific endonuclease domain PilT N-terminus (PIN) recruits it to the target transcript and its degradation. can make it possible.
[0258] Patients with paraneoplastic neuropathy (PND)-associated encephalomyelitis and neuropathy develop autoantibodies against the Hu protein in tumors outside the central nervous system (Szabo A et al. 1991, Cell.; 67:325-33; It is the patient who then crosses the blood-brain barrier It can be envisioned that the RNA targeting effector proteins of the invention can be used to interfere with the binding of autoantibodies to mRNA transcripts.
[0259] Patients with dystrophy type 1 (DM1), caused by an expansion of (CUG)n in the 3'UTR of the myotonic dystrophic protein kinase (DMPK) gene, are characterized by accumulation of such transcripts in the nucleus. It is envisioned that an RNA targeting effector protein of the invention fused to an endonuclease targeting (CUG)n repeats could inhibit the accumulation of such aberrant transcripts.
[0260] h) interactions with multifunctional proteins Some RNA-binding proteins bind to multiple sites on many RNAs and function in diverse processes. For example, the hnRNP A1 protein binds to exon splicing silencer sequences to antagonize splicing factors, associates with telomere ends (thereby stimulating telomere activity), and binds to miRNAs to promote Drosha-mediated processing. , which has been shown to affect maturity. It is envisioned that an RNA binding effector protein of the invention may interfere with binding of an RNA binding protein at one or more locations.
[0261] i) RNA folding RNA adopts a defined structure in order to exert its biological activity. Conformational transitions between alternative tertiary structures are critical to many RNA-mediated processes. However, RNA folding can be associated with several problems. For example, RNA may tend to fold into and maintain an inappropriate alternate conformation, and / or the correct tertiary structure may not be sufficiently thermodynamically favorable relative to the alternate conformation. Sometimes I can't. RNA targeting effector proteins of the invention, in particular truncated or dead RNA targeting proteins, may be used to direct the folding of (m)RNA and / or ensure its correct tertiary structure.
[0262] Use of RNA-targeting effector proteins in modulating cellular states In certain embodiments, Casl3b complexed with crRNA is activated upon binding to the target RNA and subsequently cleaves any nearby ssRNA targets (ie, a “collateral” or “bystander” effect). Cas13b can cleave other (non-complementary) RNA molecules when primed by its cognate target. Such promiscuous RNA cleavage can potentially cause cytotoxicity or otherwise affect cell physiology or cell state.
[0263] Thus, in certain embodiments, the non-naturally occurring or engineered compositions, vector systems or delivery systems as described herein are used for or for use in inducing cellular dormancy. belongs to. In certain embodiments, the non-naturally occurring or engineered composition, vector system or delivery system as described herein is used or for use in inducing cell cycle arrest. is. In certain embodiments, the non-naturally occurring or engineered composition, vector system or delivery system as described herein is used for or in reducing cell growth and / or cell proliferation. It is for In certain embodiments, the non-naturally occurring or engineered composition, vector system or delivery system as described herein is used or for use in inducing cellular anergy. be. In certain embodiments, the non-naturally occurring or engineered composition, vector system or delivery system as described herein is used or for use in inducing cell apoptosis. be. In certain embodiments, the non-naturally occurring or engineered composition, vector system or delivery system as described herein is used or for use in inducing cell necrosis. be. In certain embodiments, the non-naturally occurring or engineered composition, vector system or delivery system as described herein is used or for use in inducing cell death. be. In certain embodiments, the non-naturally occurring or engineered composition, vector system or delivery system as described herein is used for or for use in inducing programmed cell death. is.
[0264] In certain embodiments, the invention provides a method of inducing cell dormancy comprising introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein. Regarding. In certain embodiments, the invention provides a method of inducing cell cycle arrest comprising introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein. Regarding. In certain embodiments, the present invention provides cell growth and / or cell growth comprising introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein. It relates to a method of reducing proliferation. In certain embodiments, the invention relates to a method of inducing cellular anergy comprising introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein. . In certain embodiments, the invention relates to a method of inducing cell apoptosis comprising introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein. . In certain embodiments, the invention relates to a method of inducing cell necrosis comprising introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein. . In certain embodiments, the invention relates to a method of inducing cell death comprising introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein. . In certain embodiments, the invention provides a method of inducing programmed cell death comprising introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein. Regarding.
[0265] The methods and uses as described herein may be therapeutic or prophylactic and may target specific cells, cell (sub)populations or cell / tissue types. In particular, the methods and uses as described herein can be therapeutic or prophylactic, expressing one or more target sequences, such as one or more specific target RNAs (e.g., ssRNA). Specific cells, cell (sub)populations or cell / tissue types may be targeted. Without limitation, target cells are e.g. cancer cells expressing a particular transcript, e.g. neurons of a given class, e.g. (immune) cells causing autoimmunity or specific (e.g. viral) pathogens. It can be infected cells and the like.
[0266] Thus, in certain embodiments, the present invention provides for unwanted cells ( The present invention relates to methods of treating pathological conditions characterized by the presence of host cells). In certain embodiments, the present invention provides non-naturally occurring or engineered compositions as described herein for treating pathological conditions characterized by the presence of unwanted cells (host cells). use of the product, vector system or delivery system. In certain embodiments, the present invention provides non-naturally occurring or engineered cells as described herein for use in treating pathological conditions characterized by the presence of unwanted cells (host cells). composition, vector system or delivery system. It should be appreciated that the CRISPR-Cas system preferably targets specific targets to unwanted cells. In certain embodiments, the invention relates to the use of the non-naturally occurring or engineered composition, vector system or delivery system as described herein for treating, preventing or alleviating cancer. In certain embodiments, the invention relates to non-naturally occurring or engineered compositions, vector systems or delivery systems as described herein for use in treating, preventing or alleviating cancer. In certain embodiments, the invention provides a method for treating, preventing or treating cancer comprising introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein. Regarding how to mitigate. Preferably, it should be understood that the CRISPR-Cas system targets cancer cell-specific targets. In certain embodiments, the invention provides a non-naturally occurring or engineered composition, vector system or delivery system as described herein for treating, preventing or reducing infection of cells by pathogens. regarding the use of In certain embodiments, the invention provides a non-naturally occurring or engineered composition, vector system or composition as described herein for use in treating, preventing or alleviating infection of a cell by a pathogen. Regarding delivery systems. In certain embodiments, the invention provides for the infection of cells by pathogens, including introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein. It relates to methods of treatment, prevention or alleviation. It should be appreciated that the CRISPR-Cas system preferably targets a target specific to a cell infected with a pathogen (eg, a pathogen-derived target). In certain embodiments, the invention provides the use of non-naturally occurring or engineered compositions, vector systems or delivery systems as described herein for treating, preventing or ameliorating autoimmune disorders. Regarding. In certain embodiments, the invention provides a non-naturally occurring or engineered composition, vector system or delivery system as described herein for use in treating, preventing or alleviating an autoimmune disorder. Regarding. In certain embodiments, the invention provides for the treatment of autoimmune disorders comprising introducing or inducing a non-naturally occurring or engineered composition, vector system or delivery system as described herein; It relates to methods of prevention or alleviation. It should be appreciated that the CRISPR-Cas system preferably targets targets specific to cells involved in autoimmune disorders (eg, specific immune cells).
[0267] Use of RNA targeting effector proteins in RNA detection Further, it is envisioned that RNA targeting effector proteins may be used in Northern blot assays. Northern blotting involves size separation of RNA samples using electrophoresis. RNA targeting effector proteins can be used to specifically bind and detect target RNA sequences.
[0268] RNA targeting effector proteins can be fused to fluorescent proteins (such as GFP) and used to track RNA localization in living cells. More specifically, an RNA targeting effector protein can be inactivated to the point that it no longer cleaves RNA. In a detailed embodiment, it is envisioned that split RNA targeting effector proteins may be used to ensure more precise visualization, whereby the signal depends on the binding of both subproteins. Alternatively, a split fluorescent protein that reconstitutes upon binding of multiple RNA targeting effector protein complexes to the target transcript can be used. Furthermore, it is envisioned that transcripts may be targeted at multiple binding sites along the mRNA such that the fluorescent signal may amplify the true signal to allow localized discrimination. As yet another alternative, fluorescent proteins can be reconstituted from split inteins.
[0269] RNA targeting effector proteins are suitably used, for example, to determine the localization of RNA or specific splice variants, levels of mRNA transcripts, up- or down-regulation of transcripts and disease-specific diagnostics. RNA targeting effector proteins can be detected using, for example, fluorescence microscopy or flow cytometry, such as fluorescence-activated cell sorting (FACS), which allows high-throughput screening of cells and recovery of viable cells after cell sorting (live). It can be used for visualization of RNA in cells. Furthermore, the expression levels of various transcripts can be assessed simultaneously under stress, eg inhibition of cancer growth using molecular inhibitors or hypoxic conditions for the cells. Another application could be to follow the localization of transcripts to synaptic connections during nerve stimulation using two-photon microscopy.
[0270] In certain embodiments, the components or conjugates of the invention as described herein are subjected to multiplexed error-robust fluorescence in situ hybridization (MERFISH; Chen et al.), e.g. al.Science;2015;348(6233)).
[0271] In vitro APEX labeling Cellular processes depend on networks of molecular interactions between proteins, RNA and DNA. Accurate detection of protein-DNA and protein-RNA interactions is critical to the understanding of such processes. In vitro proximity labeling techniques utilize affinity tags, for example in combination with photoactivatable probes, to label polypeptides and RNA in the vicinity of the protein or RNA of interest in vitro. After UV irradiation, the photoactivatable group reacts with proteins and other molecules in close proximity to the tagging molecule, thereby labeling them. Labeled interacting molecules can then be recovered and identified. The RNA targeting effector proteins of the invention can be used, for example, to target probes to RNA sequences of choice.
[0272] These applications may also be applicable in animal models for disease-related applications or in vivo imaging of difficult-to-culture cell types.
[0273] Use of RNA targeting effector proteins in RNA origami / in vitro assembly lines - Combinatorics RNA origami refers to nanoscale folding structures for creating two- or three-dimensional structures using RNA as an embedded template. The folding structure is encoded in the RNA and thus the shape of the resulting RNA is determined by the synthesized RNA sequence (Geary, et al.2014.Science,345(6198).pp.799-804). RNA origami can serve as a scaffold for arranging other components, such as proteins, into complexes. The RNA targeting effector proteins of the invention can be used, for example, to target proteins of interest to RNA origami using suitable guide RNAs.
[0274] These applications can also be applied in disease relevant applications or in animal models for in vivo imaging of difficult to culture cell types.
[0275] Use of RNA targeting effector proteins in RNA isolation or purification, enrichment or depletion Further, it is envisioned that RNA targeting effector proteins when complexed with RNA may be used to isolate and / or purify RNA. For example, an RNA targeting effector protein can be fused to an affinity tag that can be used to isolate and / or purify the RNA-RNA targeting effector protein complex. Such applications are useful, for example, in analyzing gene expression profiles in cells. In particular embodiments, it can be envisioned that RNA targeting effector proteins can be used to target specific non-coding RNAs (ncRNAs) and thereby block their activity, providing useful functional probes. In certain embodiments, effector proteins as described herein may be used to specifically enrich certain RNAs (including but not limited to increasing stability, etc.) or RNA (eg, without limitation, particular splice variants, isoforms, etc.) can be specifically depleted.
[0276] Examination of lincRNA function and other nuclear RNAs Current RNA knockdown strategies, such as siRNA, have the disadvantage of being mostly limited to targeting cytoplasmic transcripts because the protein machinery is cytoplasmic. An advantage of the RNA targeting effector protein of the present invention, being an extrinsic system that is not essential for cellular function, is that it can be used in any compartment of the cell. By fusing the NLS signal to an RNA targeting effector protein, it can be directed to the nucleus, allowing targeting of nuclear RNA. For example, probing the function of lincRNA is envisioned. Long intergenic noncoding RNA (lincRNA) is a very extensively investigated research area. Many of the lincRNAs have as yet unexplained functions, which may lead to studies using the RNA targeting effector proteins of the present invention.
[0277] Identification of RNA binding proteins Identification of proteins that bind specific RNAs can be useful in understanding the roles of many RNAs. For example, many lincRNAs are associated with transcriptional and epigenetic regulators in the regulation of transcription. Understanding what proteins bind to a given lincRNA can help unravel the components of a given regulatory pathway. The RNA targeting effector proteins of the invention can be designed to recruit biotin ligase to specific transcripts to locally label bound proteins with biotin. The protein can then be pulled down and analyzed by mass spectrometry to identify it.
[0278] Assembly of complexes into RNA and substrate shuttling Additionally, RNA targeting effector proteins of the invention can be used to assemble complexes onto RNA. This can be achieved by functionalizing the RNA targeting effector protein with multiple related proteins (eg, components of specific synthetic pathways). Alternatively, multiple RNA targeting effector proteins can be functionalized with such different related proteins to target the same or adjacent target RNAs. A useful application of assembly of complexes into RNA is, for example, the promotion of substrate shuttling between proteins.
[0279] synthetic biology The development of biological systems has wide utility, including clinical applications. It is envisioned that the programmable RNA targeting effector proteins of the invention can be used to fuse, for example, cancer-associated RNAs to split proteins of toxic domains for targeted cell death using cancer-associated RNAs as target transcripts. Furthermore, in synthetic biological systems, pathways involving protein-protein interactions can be affected by fusion complexes with appropriate effectors, such as kinases or other enzymes.
[0280] protein splicing: inteins Protein splicing is a post-translational process in which an intervening polypeptide, called an intein, catalyzes the excision of itself from its neighboring polypeptides, called exteins, and the subsequent ligation of the exteins. Directing the release of split inteins using assembly of two or more RNA targeting effector proteins as described herein into a target transcript (Topilina and Mills Mob DNA.2014 Feb 4;5(1): 5), which may allow direct calculation of the presence of mRNA transcripts and subsequent release of protein products (for downstream operation of transcription pathways) such as metabolic enzymes or transcription factors. This application can have great implications in synthetic biology (see above) or large-scale bioproduction (producing products only under certain conditions).
[0281] Inducible, dosing and self-deactivating systems In one embodiment, a fusion complex comprising an RNA targeting effector protein of the invention and an effector component is designed to be inducible, eg, light- or chemically-inducible. Such inducibility allows activation of the effector component at a desired time.
[0282] Light inducibility is achieved, for example, by designing fusion complexes in which CRY2PHR / CIBN pairing is used for fusion. This system is particularly useful for light induction of protein interactions in living cells (Konermann S, et al. Nature. 2013;500:472-476).
[0283] Chemoinducibility is provided, for example, by designing fusion complexes where the fusion uses FKBP / FRB (FK506 binding protein / FKBP rapamycin binding) pairing. Using this system, rapamycin is required for protein binding (Zetsche et al. Nat Biotechnol. 2015;33(2):139-42 describes the use of this system for Cas9).
[0284] In addition, the RNA targeting effector proteins of the present invention, when introduced into cells as DNA, can induce hormone-inducible gene expression, such as tetracycline- or doxycycline-controlled transcriptional activation (Tet-On and Tet-Off expression systems), e.g., ecdysone-inducible gene expression systems. It can be regulated by inducible promoters, such as expression systems and arabinose-inducible gene expression systems. When delivered as RNA, the expression of RNA-targeting effector proteins can be regulated by riboswitches, which can detect small molecules such as tetracycline (Goldfless et al. Nucleic Acids Res. 2012;40 (9): as described in e64).
[0285] In one embodiment, modulating the delivery of the RNA targeting effector protein of the invention to alter the amount of protein or crRNA within the cell, thereby altering the magnitude of the desired effect or any undesired off-target effect. can be done.
[0286] In one embodiment, the RNA targeting effector proteins described herein can be designed to self-inactivate. This is because when delivered to cells as RNA, either mRNA or replicating RNA therapeutics (Wrobleska et al Nat Biotechnol.2015 Aug;33(8):839-841), they destroy self-RNA, thereby By reducing extrinsic and potentially unwanted effects, it can self-inactivate manifestations and subsequent effects.
[0287] For further in vivo applications of RNA targeting effector proteins as described herein, see Mackay JP et al (Nat Struct Mol Biol. 2011 Mar;18(3):256-61), Nelles et al (Bioessays. 2015 Jul;37(7):732-9) and Abil Z and Zhao H (Mol Biosyst. 2015 Oct;11(10):2658-65), which are incorporated herein by reference. Specifically, the following applications are envisioned in certain embodiments of the present invention, preferably by using catalytically inactive Cas13b: enhancement of translation (e.g., Cas13b-translational facilitator fusions translational repression (e.g. gRNAs targeting ribosome binding sites); exon skipping (e.g. gRNAs targeting splice donor and / or acceptor sites); exon inclusion (e.g. specific Cas13b fused to or recruiting a gRNA or spliceosomal component (e.g., U1 snRNA) that targets exon splice donor and / or acceptor sites of the target to include it; access to RNA localization (e.g., Cas13b- marker fusions (e.g. EGFP fusions)); changes in RNA localization (e.g. Cas13b-localizing signal fusions (e.g. NLS or NES fusions)); RNA degradation (in this case dependent on the activity of Cas13b). If so, no catalytically inactive Cas13b shall be used, or, for increased specificity, split Cas13b may be used); such as by degradation of gRNAs or binding to functional sites (optionally Cas13b-signal sequence fusions (titrates out at specific sites by relocalization by activator), inhibition of non-coding RNA function (eg, miRNAs).
[0288] As described herein above and demonstrated in the Examples, Casl3b function is robust to 5' or 3' elongation of crRNA and elongation of the crRNA loop. It is therefore envisioned that MS2 loops and other recruitment domains could be added to crRNAs without affecting complex formation and binding to target transcripts. Such modifications to crRNA to recruit different effector domains are applicable in the use of RNA-targeted effector proteins described above.
[0289] Cas13b has the ability to mediate RNA phage resistance. Thus, it is envisioned that Casl3b can be used to immunize, for example, animals, humans and plants against RNA-only pathogens, including but not limited to Ebola virus and Zika virus.
[0290] In certain embodiments, Casl3b can process (cleave) its own array. This applies to both wild-type Cas13b proteins and mutant Cas13b proteins containing one or more mutant amino acid residues as discussed herein. Thus, it is envisioned that multiple crRNAs designed for different target transcripts and / or applications can be delivered as a single pre-crRNA or as a single transcript driven by one promoter. . Such delivery methods have the advantage of being substantially more compact, easier to synthesize, and easier to deliver in viral systems. The exact amino acid positions for Cas13b orthologs herein may vary and are suitably determined by protein alignments, as known in the art and described elsewhere herein. It will be understood to get Aspects of the present invention are directed to genome engineering in vitro, in vivo or ex vivo in prokaryotic or eukaryotic cells, e.g., to alter or manipulate the expression of one or more genes or one or more gene products. Also included are methods and uses of the compositions and systems described therein.
[0291] In one aspect, the invention provides methods and compositions for modulating, eg, reducing, expression of a target RNA in a cell. The method provides a Cas13b system of the invention that interferes with RNA transcription, stability and / or translation.
[0292] In certain embodiments, an effective amount of the Cas13b system is used to cleave RNA or otherwise inhibit RNA expression. In this context, this system has similar uses to siRNA and shRNA and can therefore also replace such methods. This method includes, but is not limited to, the use of the Cas13b system as a substitute for interfering ribonucleic acid (such as siRNA or shRNA) or its transcription template, eg, DNA encoding the shRNA. The Cas13b system is introduced into target cells, eg, by administration to a mammal containing the target cells.
[0293] Advantageously, the Cas13b system of the invention is specific. For example, interfering ribonucleic acid (such as siRNA or shRNA) polynucleotide systems suffer from design and stability issues as well as off-target binding, while the Cas13b system of the present invention can be designed with high specificity.
[0294] destabilizing Cas13b In certain embodiments, the effector proteins of the invention as described herein are associated or fused to a destabilization domain (DD). In some embodiments, DD is ER50. The corresponding stabilizing ligand of this DD is 4HT in some embodiments. Therefore, in some embodiments, one of the at least one DD is ER50 and thus the stabilizing ligand is 4HT or CMP8. In some embodiments, DD is DHFR50. The corresponding stabilizing ligand of this DD is TMP in some embodiments. Therefore, in some embodiments, one of the at least one DD is DHFR50 and thus the stabilizing ligand is TMP. In some embodiments, DD is ER50. The corresponding stabilizing ligand of this DD is CMP8 in some embodiments. Therefore, CMP8 may be an alternative stabilizing ligand for 4HT in the ER50 system. While it is possible that CMP8 and 4HT may / should be used competitively, some cell types may be susceptible to one or the other of these two ligands, and the present disclosure and From knowledge in the art, one skilled in the art can use CMP8 and / or 4HT.
[0295] In some embodiments, one or two DDs may be fused to the N-terminal end of Casl3b and one or two DDs may be fused to the C-terminal end of Casl3b. In some embodiments, at least two DDs are associated with Casl3b and the DDs are the same DD, ie the DDs are homologous. Therefore, both (or more than one) of the DDs can be ER50 DDs. This is preferred in some embodiments. Alternatively, both (or more than one) of the DDs can be DHFR50 DDs. This is also preferred in some embodiments. In some embodiments, at least two DDs are associated with Casl3b and the DDs are different DDs, ie the DDs are heterologous. Thus, one of the DDs can be ER50, while one or more of the DDs or any other DD can be DHFR50. Having two or more DDs that are heterologous can be advantageous as it can provide a higher level of degradation control. Tandem fusions of two or more DDs at the N-terminus or C-terminus can enhance degradation; and such tandem fusions can be, for example, ER50-ER50-Casl3b or DHFR-DHFR-Casl3b. While high levels of degradation can occur if neither stabilizing ligand is present, and intermediate levels of degradation can occur if one stabilizing ligand is absent and the other (or another) stabilizing ligand is present. , it is assumed that low levels of degradation can occur if both (or more than one) of the stabilizing ligands are present. Control can also be provided by having an N-terminal ER50 DD and a C-terminal DHFR50 DD.
[0296] In some embodiments, the fusion of Casl3b and DD includes a linker between DD and Casl3b. In some embodiments, the linker is a GlySer linker. In some embodiments, DD-Casl3b further comprises at least one nuclear export signal (NES). In some embodiments, DD-Casl3b comprises two or more NES. In some embodiments, DD-Casl3b comprises at least one nuclear localization signal (NLS). This could be because it contains in addition to the NES. In some embodiments, Casl3b comprises or consists essentially of a localization (nuclear import or export) signal as or part of a linker between Casl3b and DD; or consists of HA or Flag tags are also within the scope of the invention as linkers. Applicants have used NLS and / or NES as linkers and have also used as short as GS to maximum (GGGGS) 3 Also use glycerin linkers up to .
[0297] Destabilization domains have general utility in conferring instability to a wide range of proteins; (incorporated herein by reference). CMP8 or 4-hydroxy tamoxifen can be the destabilizing domain. More generally, temperature-sensitive mutants of mammalian DHFR (DHFRts), destabilizing residues due to the N-end rule, were found to be stable at the permissive temperature but unstable at 37°C. Addition of methotrexate, a high-affinity ligand for mammalian DHFR, to cells expressing DHFRts partially inhibited protein degradation. This was an important demonstration that small molecule ligands can stabilize proteins that are naturally targeted for degradation in cells. The use of rapamycin derivatives stabilized an unstable mutant of the FRB domain of mTOR (FRB*) and restored the function of the fused kinase GSK-3β6,7. This system demonstrated that ligand-dependent stability represents an attractive strategy to modulate the function of specific proteins in complex biological environments. A regulatory system for protein activity may involve a DD that becomes functional upon ubiquitin complementation by rapamycin-induced dimerization of FK506-binding proteins with FKBP12. Mutants of the human FKBP12 or ecDHFR proteins can be engineered to be metabolically unstable in the absence of their high affinity ligands, Shield-1 or trimethoprim (TMP), respectively. These mutants are part of a possible destabilization domain (DD) useful in the practice of the present invention, and the instability of the DD as a fusion with Cas13b increases the Cas13b of the entire fusion protein by the proteasome. bring decomposition. Shield-1 and TMP bind and stabilize DD in a dose-dependent manner. The estrogen receptor ligand binding domain (ERLBD, residues 305-549 of ERS1) can also be engineered as a destabilizing domain. Because the estrogen receptor signaling pathway is involved in various diseases such as breast cancer, this pathway has been extensively studied and many estrogen receptor agonists and antagonists have been developed. Thus, compatible pairs of ERLBD and drugs are known. There are ligands that bind to the mutant form of ERLBD but not to the wild-type form of ERLBD. By using one of these mutated domains encoding three mutations (L384M, M421G, G521R)12, the stability of ERLBD-derived DD with ligands that do not perturb the endogenous estrogen-sensitive network can be adjusted. An additional mutation (Y537S) can be introduced to further destabilize ERLBD and constitute it as a potential DD candidate. This quadruple mutant is an advantageous DD expansion. This mutant ERLBD can be fused to Casl3b and ligands can be used to modulate or perturb its stability so that Casl3b has a DD. Another DD can be a mutated FKBP protein-based 12 kDa (107 amino acids) tag that is stabilized by Shield1 ligand; see, eg, Nature Methods 5, (2008). For example, the DD can be a synthetic, biologically inert small molecule, modified FK506 binding protein 12 (FKBP12) that binds to and is reversibly stabilized by Shield-1; , Banaszynski LA, Chen LC, Maynard-Smith LA, Ooi AG, Wandless TJ. "A rapid, reversible, and tunable method to regulate protein function in living cells using synthetic small molecules". Cell.2006;126:995-1004; Banaszynski LA, Sellmyer MA, Contag CH, Wandless TJ, Thorne SH.“Chemical control of protein stability and function in living mice”. Nat Med.2008;14:1123-1127;Maynard-Smith LA, Chen LC, Banaszynski LA, Ooi AG,Wandless TJ.“A directed approach for engineering conditional protein stability using biologically silent small molecules”.The Journal of biological chemistry.2007;282:24866-24872;and Rodriguez,Chem Biol.Mar 23,2012;19(3 ):391-398, all of which are incorporated herein by reference, and can be used in the practice of the present invention in selecting DDs to associate with Casl3b. As can be appreciated, the knowledge in the art contains a number of DDs and DDs can be associated, e.g. and can destabilize DD in its absence, thereby destabilizing Cas13b as a whole, or DD can be stabilized in the absence of ligand. and can destabilize DD when a ligand is present; DD allows Cas13b, and thus the CRISPR-Cas13b complex or system, to be modulated or controlled--turned on and off, as it were, thereby Means are provided to modulate or control the system, eg, in an in vivo or in vitro environment. For example, when a protein of interest is expressed as a fusion with a DD tag, it is destabilized intracellularly and rapidly degraded by, for example, the proteasome. Thus, the absence of stabilizing ligands leads to degradation of D-associated Cas. Fusing the novel DD to the protein of interest confers its instability to the protein of interest, resulting in rapid degradation of the entire fusion protein. Cas peak activity is sometimes beneficial in reducing off-target effects. Therefore, short bursts of high activity are preferred. The present invention is capable of providing such peaks. In a sense, this system is inducible. In another sense, the system is repressed in the absence of the stabilizing ligand and derepressed in the presence of the stabilizing ligand.
[0298] Cas13 mutation In certain embodiments, the effector protein of the invention (CRISPR enzyme; Cas13; effector protein) as described herein is a catalytically inactive or dead Cas13 effector protein (dCas13). In some embodiments, the dCasl3 effector comprises a mutation in the nuclease domain. In some embodiments, the dCasl3 effector protein is truncated. To reduce the size of the fusion protein between the Cas13 effector and one or more functional domains, the C-terminus of the Cas13 effector can be truncated while still maintaining its RNA binding function. For example, at least 20 amino acids, at least 50 amino acids, at least 80 amino acids, or at least 100 amino acids, or at least 150 amino acids, or at least 200 amino acids, or at least 250 amino acids, or at least 300 amino acids, or at least 350 amino acids, or up to 120 amino acids, or Up to 140 amino acids, or up to 160 amino acids, or up to 180 amino acids, or up to 200 amino acids, or up to 250 amino acids, or up to 300 amino acids, or up to 350 amino acids, or up to 400 amino acids may be truncated at the C-terminus of the Casl3b effector. Specific examples of Cas13 truncation include C-terminal Δ984-1090, C-terminal Δ1026-1090 and C-terminal Δ1053-1090, C-terminal Δ934-1090, C-terminal Δ884-1090, C-terminal Δ834-1090, C-terminal Δ784- 1090 and C-terminal Δ734-1090 are included, where the amino acid positions correspond to those of the Prevotella sp. P5-125 Cas13b protein. See Figure 28.
[0299] Regulation of Cas13 effector proteins The present invention provides accessory proteins that modulate CRISPR protein function. In certain embodiments, accessory proteins modulate the catalytic activity of CRISPR proteins. In certain embodiments of the invention, accessory proteins modulate targeted or sequence-specific nuclease activity. In certain embodiments of the invention, the accessory protein modulates collateral nuclease activity. In certain embodiments of the invention, the accessory protein modulates binding to the target nucleic acid.
[0300] According to the present invention, the nuclease activity that is modulated can be against nucleic acids comprising or consisting of RNA, including, without limitation, mRNA, miRNA, siRNA, and nucleic acids comprising cleavable RNA linkages, together with nucleotide analogues. . In certain embodiments of the invention, the nuclease activity that is modulated may be on a nucleic acid comprising or consisting of DNA, including without limitation nucleic acids comprising cleavable DNA linkages and nucleic acid analogs.
[0301] In certain embodiments of the invention, accessory proteins enhance the activity of CRISPR proteins. In certain such embodiments, the accessory protein contains a HEPN domain and enhances RNA cleavage. In certain embodiments, the accessory protein inhibits the activity of CRISPR protein. In certain such embodiments, the accessory protein comprises an inactivated HEPN domain or lacks the HEPN domain entirely.
[0302] According to the present invention, naturally occurring accessory proteins of the type VI CRISPR system include small proteins encoded at or near the CRISPR locus that serve to modify the activity of the CRISPR protein. In general, CRISPR loci can be identified that contain putative CRISPR arrays and / or that encode putative CRISPR effector proteins. In certain embodiments, the effector protein can be 800-2000 amino acids, or 900-1800 amino acids, or 950-1300 amino acids. In certain embodiments, the accessory protein may be encoded within 25 kb from the putative CRISPR effector protein or array, or within 20 kb, or within 15 kb, or within 10 kb, or between 2 kb and 10 kb from the putative CRISPR effector protein or array.
[0303] In certain embodiments of the invention, the accessory protein is 50-300 amino acids, or 100-300 amino acids, or 150-250 amino acids, or about 200 amino acids. Non-limiting examples of accessory proteins include the csx27 and csx28 proteins identified herein.
[0304] The identification and use of the CRISPR accessory proteins of the invention are independent of the CRISPR effector protein taxonomy. The accessory proteins of the invention can be found associated with, or engineered to function with, various CRISPR effector proteins. The examples of accessory proteins identified and used herein generally represent CRISPR effector proteins. CRISPR effector protein classes include homology, characteristic positions (e.g., REC domains, NUC domains, HEPN sequence positions), nucleic acid targets (e.g., DNA or RNA), presence or absence of tracr RNA, 5' or 3 of direct repeats. It is understood that the position of the 'side guide / spacer sequence or other criteria may be involved. In embodiments of the invention, the identification and use of accessory proteins transcends such classification.
[0305] In RNA-targeted Type VI CRISPR-Cas systems, Cas proteins usually contain two conserved HEPN domains involved in RNA cleavage. In certain embodiments, the Cas protein processes crRNA to produce mature crRNA. The guide sequence of crRNA recognizes target RNA with complementary sequence and Cas protein degrades the target strand. More specifically, in certain embodiments, upon target binding, the Cas protein undergoes a structural rearrangement that bridges the two HEPN domains together to form an active HEPN catalytic site, followed by cleavage of the target RNA. The catalytic site near the surface of the Cas protein allows non-specific collateral ssRNA cleavage.
[0306] In certain embodiments, accessory proteins serve to increase or decrease targeted and / or collateral RNA cleavage. Without being bound by theory, an accessory protein that activates CRISPR activity (e.g., a csx28 protein or an ortholog or variant containing a HEPN domain) interacts with the Cas protein and combines its HEPN domain with the HEPN domain of the Cas protein. can be assumed to have the ability to form an active HEPN catalytic site, whereas inhibitory accessory proteins (e.g., csx27, which lacks a HEPN domain) interact with Cas proteins and bring the two HEPN domains together. It can be assumed that it has the ability to reduce or block the conformation of the Cas protein obtained.
[0307] According to the present invention, in certain embodiments, enhancing the activity of a type VI Cas protein or complex thereof comprises adding a type VI Cas protein or complex thereof to an accessory protein from the same organism that activates the Cas protein. including contact with In other embodiments, enhancing the activity of a type VI Cas protein or complex thereof comprises activating the type VI Cas protein or complex thereof with an activator from a different organism within the same subclass (e.g., type VI-b). Including contacting with an accessory protein. In other embodiments, enhancing the activity of a type VI Cas protein or complex thereof is to convert the type VI Cas protein or complex thereof to an accessory protein not within its subclass (e.g., a VI other than type VI-b contacting a type Cas protein with a type VI-b accessory protein or vice versa).
[0308] According to the present invention, in certain embodiments, inhibiting the activity of a type VI Cas protein or complex thereof comprises exposing the type VI Cas protein or complex thereof to an accessory protein from the same organism that inhibits the Cas protein. Including contacting. In other embodiments, inhibiting the activity of a type VI Cas protein or complex thereof is performed by suppressing the activity of the type VI Cas protein or complex thereof with a repressor from a different organism within the same subclass (e.g., type VI-b). Including contacting with an accessory protein. In other embodiments, inhibiting the activity of a type VI Cas protein or complex thereof redirects the type VI Cas protein or complex to a repressor accessory protein not within its subclass (e.g., to a repressor accessory protein other than type VI-b to the type VI-b repressor accessory protein or vice versa).
[0309] In certain embodiments in which the type VI Cas protein and the type VI accessory protein are derived from the same organism, these two proteins can function together in an engineered CRISPR system. In certain embodiments, it may be desirable to alter the function of the engineered CRISPR system, for example by altering either or both of the proteins or their expression. In embodiments where the Type VI Cas protein and the Type VI accessory protein are from different organisms, which may be within the same class or within different classes, these proteins are combined in an engineered CRISPR system. Although they can function together, in many cases it will be desirable or necessary to modify either one or both of the proteins to function together.
[0310] Thus, in certain embodiments of the invention, either one or both of the Cas protein and the accessory protein may be modified to modulate aspects of the protein-protein interaction between the Cas protein and the accessory protein. In certain embodiments, either or both the Cas protein and the accessory protein may be modified to modulate aspects of protein-nucleic acid interactions. Methods of modulating protein-protein and protein-nucleic acid interactions include, without limitation, adaptation of molecular surfaces, polar interactions, hydrogen bonding and regulation of van der Waals interactions. In certain embodiments, modulating protein-protein interactions or protein-nucleic acid binding comprises increasing or decreasing binding interactions. In certain embodiments, modulating protein-protein interactions or protein-nucleic acid binding comprises alterations that favor or disadvantage the conformation of the protein or nucleic acid.
[0311] "Adaptation" means, including by automatic or semi-automatic means, between one or more atoms of a Cas13 protein and at least one atom of a Cas13 accessory protein, or one or more atoms of a Cas13 protein and a nucleic acid. or between one or more atoms of a Cas13 accessory protein and a nucleic acid, and calculating how stable such interactions are do. Interactions include attractive and repulsive forces exerted by charges, steric considerations, and the like.
[0312] Three-dimensional structures of type VI CRISPR proteins or complexes thereof or type VI CRISPR accessory proteins or complexes thereof, in the context of the present invention, provide additional tools for identifying additional mutations in orthologues of Cas13. do. The crystal structure can also serve as a basis for designing new specific Cas13 and Cas13 accessory proteins. Various computer-based adaptation methods are further described. The binding interactions of Ca13, accessory proteins and nucleic acids can be investigated through the use of computer modeling with docking programs. Docking programs are known; for example, GRAM, DOCK or AUTODOCK (Walters et al. Drug Discovery Today, vol. , 27-42). This procedure can include computer fitting to ascertain how well the shape and chemical structure of the binding partner are. A computer-assisted manual search of the active site or binding site of the Type VI system can be performed. Programs such as GRID (P. Goodford, J. Med. Chem, 1985, 28, 849-57) - a program that determines potential sites of interaction between molecules with various functional groups - are also useful for partial identification of binding compounds. It can be used for active site or binding site analysis to predict structure. Computer programs can be used to estimate the attraction, repulsion or steric hindrance of two binding partners, eg a component or nucleic acid molecule of the type VI CRISPR system and a component of the type VI CRISPR system.
[0313] Amino acid substitutions may be made on the basis of differences or similarities in amino acid properties (such as polarity, charge, solubility, hydrophobicity, hydrophilicity and / or amphipathic properties of residues), thus dividing amino acids into functional groups. It is useful to summarize Amino acids can be grouped based on the properties of their side chains alone. In comparing orthologs, they are likely residues that are conserved for structural or catalytic reasons. These sets can be described in Venn diagram form (Livingstone C.D. and Barton G.J. (1993) "Protein sequence alignments: a strategy for the hierarchical analysis of residue conservation" Comput. Appl. Biosci. 9:745-756). (Taylor W. R. (1986) "The classification of amino acid conservation" J. Theor. Biol. 119; 205-218). Conservative substitutions may be made, for example according to the table below, which describes the generally accepted Venn diagram grouping of amino acids.
table 1
[0314] In an engineered Ca13 system, the modification may involve modification of one or more amino acid residues of the Cas13 protein and / or modification of one or more amino acid residues of the Cas13 accessory protein.
[0315] In an engineered Ca13 system, modification may involve modification of one or more amino acid residues located in regions containing residues that are positively charged in the unmodified Cas13 protein and / or Cas13 accessory proteins.
[0316] In an engineered Cas13 system, modification may involve modification of one or more amino acid residues that are positively charged in the unmodified Cas13 protein and / or Cas13 accessory proteins.
[0317] In an engineered Cas13 system, modification may involve modification of one or more amino acid residues that are not positively charged in the unmodified Cas13 protein and / or Cas13 accessory proteins.
[0318] Modifications may include modification of one or more uncharged amino acid residues in the unmodified Cas13 protein and / or the Cas13 accessory protein.
[0319] Modifications can include modification of one or more amino acid residues that are negatively charged in the unmodified Cas13 protein and / or the Cas13 accessory protein.
[0320] Modifications can include modification of one or more amino acid residues that are hydrophobic in the unmodified Cas13 protein and / or the Cas13 accessory protein.
[0321] Modifications may include modification of one or more amino acid residues that are polar in the unmodified Cas13 protein and / or the Cas13 accessory protein.
[0322] Modifications may include substitution of hydrophobic or polar amino acids to charged amino acids, which may be negatively or positively charged amino acids. Modifications may include substitution of negatively charged amino acids with positively charged, or polar, or hydrophobic amino acids. Modifications can include substitution of positively charged amino acids for negatively charged, or polar, or hydrophobic amino acids.
[0323] Embodiments of the present invention include possible homologous substitutions (both substitution and substitution are used herein to mean the interchange of existing amino acid residues or nucleotides with alternative residues or nucleotides), i.e. It includes sequences (both polynucleotides and polypeptides) that may contain substitutions of like types, such as between basic, between acidic, and between polar amino acids. non-homologous substitutions, i.e. substitutions from one class of residues to another, or ornithine (hereinafter referred to as Z), ornithine diaminobutyrate (hereinafter referred to as B), norleucine ornithine (hereinafter referred to as O), Substitutions involving the incorporation of unnatural amino acids such as pyrylalanine, thienylalanine, naphthylalanine and phenylglycine are also possible. A variant amino acid sequence may be inserted between any two amino acid residues of the sequence containing amino acid spacers such as glycine or β-alanine residues as well as alkyl groups such as methyl, ethyl or propyl groups. spacer groups. Further forms of variants, which involve the presence of one or more amino acid residues in peptoid form, are well understood by those skilled in the art. For the avoidance of doubt, "peptoid form" is used to refer to variant amino acid residues in which the α-carbon substituent is not on the α-carbon, but rather on the nitrogen atom of the residue. Methods for preparing peptides in peptoid form are known in the art (see, for example, Simon RJ et al., PNAS (1992) 89(20), 9367-9371 and Horwell DC, Trends Biotechnol. (1995) 13(4) ), 132-134).
[0324] Homology modeling: Corresponding residues in other Cas13 orthologues were identified by Zhang et al., 2012 (Nature; 490(7421):556-60) and Chen et al., 2015 (PLoS Comput Biol; 11(5): e1004248) method—computational protein-protein interaction (PPI) methods that predict interactions mediated by domain-motif interfaces. PrePPI (Predictive PPI), a structure-based PPI prediction method, combines structural evidence with non-structural evidence using a Bayesian statistical framework. The method involves pairing query proteins and using structural alignments to identify representative structures that correspond to either their experimentally determined structures or homology models. Structural alignment is further used to identify both near and far neighboring structures by considering global and local geometric relationships. Whenever two neighboring structures of a representative structure form a complex reported in the protein databank, this defines a template for modeling the interaction between two query proteins. A model of the complex is created by superimposing representative structures on corresponding neighboring structures in the template. This technique is in Dey et al., 2013 (Prot Sci;22:359-66).
[0325] Application of RNA-targeting CRISPR system to plants and yeast Definition: In general, the term "plant" includes any of the various photosynthetic organisms, eukaryotes, It relates to unicellular or multicellular organisms. The term plant includes monocotyledonous and dicotyledonous plants. Specifically, plants include, without limitation, acacia, alfalfa, amaranth, apple, apricot, artichoke, ash tree, asparagus, avocado, banana, barley, legumes, sugar beet, birch, beech, blackberry, Blueberries, broccoli, Brussels sprouts, cabbage, canola, cantaloupe, carrots, cassava, cauliflower, cedar, grains, celery, chestnuts, cherries, Chinese cabbage, citrus, clementines, clover, coffee, corn, cotton, cowpea, cucumber, cypress, Eggplant, elm, endive, eucalyptus, fennel, fig, fir, geranium, grape, grapefruit, peanuts, physalis, gum hemlock, hickory, kale, kiwifruit, kohlrabi, larch, lettuce, chive, lemon, lime , black locust, pine, maidenhair, corn, mango, maple, melon, millet, mushroom, mustard, nuts, oak, oat, oil palm, okra, onion, orange, ornamental plant or ornamental flower or ornamental tree, papaya, palm , parsley, parsnip, pea, peach, peanut, pear, peat, pepper, oyster, pigeon pea, pine, pineapple, plantain, plum, pomegranate, potato, pumpkin, red chicory, radish, oilseed rape, raspberry, rice, rye , sorghum, safflower, saliva, soybean, spinach, spruce, pumpkin, strawberry, sugar beet, sugar cane, sunflower, sweet potato, sweet corn, tangerine, tea, tobacco, tomato, tree, triticale, lawn grass, turnip, vine, walnut Angiosperms and gymnosperms such as, watermelon, wheat, yam, yew and zucchini are intended to be included. The term plants also includes algae, which are mostly photoautotrophs, united primarily by their lack of roots, leaves and other organs that characterize higher plants.
[0326] The method of modulating gene expression using the RNA targeting system as described herein can be used to confer desired traits on essentially any plant. Various plants and plant cell lines can be engineered for the desired physiological and agronomic properties described herein using the nucleic acid constructs of the present disclosure and the various transformation methods described above. In preferred embodiments, plants and plant cells targeted for engineering include, but are not limited to, cereal crops (e.g., wheat, corn, rice, millet, barley), fruit crops (e.g., tomatoes, apples, pears, strawberries). , oranges), forage crops (e.g. alfalfa), root crops (e.g. carrots, potatoes, sugar beets, yams), leafy vegetable crops (e.g. lettuce, spinach); flowering plants (e.g. petunias, roses, chrysanthemums) , conifers and pine trees (e.g. fir, spruce); plants used in phytoremediation (e.g. heavy metal accumulating plants); oil crops (e.g. sunflower, rapeseed) and plants used for experimental purposes (e.g. Arabidopsis spp.) (Arabidopsis)), monocotyledonous and dicotyledonous plants. Thus, the method and CRISPR-Cas system span a wide range of plants, e.g. ), Ranunculales, Papeverales, Sarraceniaceae, Trochodendrales, Hamamelidales, Eucomiales, Leitneriales, Myricales, Fagales, Casuarinales, Caryophyllales, Batales, Polygonales, Plumbaginales, Dilleniales, Theales, Malvaceae (Malvales), Urticales, Lecythidales, Violales, Salicales, Capparales, Ericales, Diapensales, Ebenales ), Primulales, Rosales, Fabales, Podostemales, Haloragales, Myrtales, Cornales, Proteales, San tales, Rafflesiales, Celastrales, Euphorbiales, Rhamnales, Sapindales, Juglandales, Geraniales, Coleoptera Polygalales, Umbellales, Gentianales, Polemoniales, Lamiales, Plantaginals, Scrophulariales, Campanulales, Rubiformes Rubiales, Dipsacales, and Asterales, such as dicotyledonous plants; Najadales, Triuridales, Commelinales, Eriocaulales, Restionales, Poales, Juncales, Cyperales, Typhales , Bromeliales, Zingiberales, Arecales, Cyclanthales, Pandanales, Arales, Liliales and Orchid ales plants belonging to the order of the monocotyledonous or gymnospermae, such as those belonging to the order Pinales, Ginkgoales, Cycadale, Araucariales, Cupressales and Gnetum It can be used by those belonging to the order (Gnetales).
[0327] The RNA targeting CRISPR systems and methods of use described herein can be used for a wide range of plant species, including the following non-limiting list of dicotyledonous, monocotyledonous or gymnosperm genera: Wolf Atropa, Alseodaphne, Anacardium, Arachis, Beilschmiedia, Brassica, Carthamus, Cocculus, Croton, Cucumis, Citrus, Citrullus, Capsicum, Catharanthus, Cocos, Coffea, Cucurbita Cucurbita, Daucus, Duguetia, Eschscholzia, Ficus, Fragaria, Glaucium, Glycine, Cotton Gossypium, Helianthus, Hevea, Hyoscyamus, Lactuca, Landolphia, Linum, Litsea, Lycopersicon ), Lupine, Manihot, Majorana, Malus, Medicago, Nicotiana, Olea, Parthenium, Papaver, Persea, Phaseolus, Pistacia, Pisum, Pyrus, Prunus, Raphanus, Castor bean (Ricinus), Senecio, Sinomenium, Stephania, Sinapis, Solanum, Theobroma, Trifolium, Trigonella ), Vicia, Vinca, Vilis and Vigna; and Allium, Andropogon, Aragrostis, Aragrostis ), Avena, Cynodon, Elaeis, Festuca, Festulolium, Heterocallis, Hordeum, Lemna ), Lolium, Musa, Oryza, Panicum, Pannesetum, Phleum, Poa, Secale , Sorghum, Triticum, Zea, Abies, Cunninghamia, Ephedra, Picea, Pinus and Pine Genus (Pseudotsuga).
[0328] RNA-targeting CRISPR systems and methods of use are described, for example, in the phylum Rhodophyta (red algae), Chlorophyta (green algae), Phaeophyta (brown algae), Bacillariophyta. algae selected from several eukaryotic phyla including (Diatoms), Eustigmatophyta and dinoflagellates, and prokaryotic phyla, Cyanobacteria (Cyanobacteria) It can also be used with a wide range of "algae" or "algal cells", including (algea). The term "algae" includes, for example, the genera Amphora, Anabaena, Ankstrodesmis, Botryococcus, Chaetoceros, Chlamydomonas, Chlorella, Chlorococcum, Cyclotella, Cylindrotheca, Dunaliella, Emiliana, Euglena, Hematococcus, Isochrysis Genus Isochrysis, Monochrysis, Monoraphidium, Nannochloris, Nannnochloropsis, Navicula, Nephrochloris, Nephrosermis Genus Nephroselmis, Nitzschia, Nodularia, Nostoc, Oochromonas, Oocystis, Oscillartoria, Pavlova, Phaeodactylum Genus Phaeodactylum, Playtmonas, Pleurochrysis, Porhyra, Pseudoanabaena, Pyramimonas, Stichococcus, Synechococcus, Synechocystis (Synechocystis), Tetraselmis, Thalassiosira and Trichodesmium.
[0329] Parts of plants, or "plant tissues," may be treated according to the methods of the invention to produce improved plants. Plant tissue also includes plant cells. The term "plant cell" as used herein is an intact whole plant or in vitro tissue culture on medium or agar, in suspension in growth medium or buffer, or Refers to an individual unit of a living plant, either in isolated form grown as part of a highly organized unit, such as a plant tissue, plant organ or whole plant.
[0330] A "protoplast" is a machine that reforms its cell wall, multiplies, and regenerates into an intact, biochemically competent unit of a living plant capable of developing into a whole plant under suitable growth conditions. It refers to a plant cell whose protective cell wall has been completely or partially removed, for example using chemical or enzymatic means.
[0331] The term "transformation" broadly refers to methods by which a plant host is genetically modified by introduction of DNA using Agrobacteria or one of a variety of chemical or physical methods. As used herein, the term "plant host" refers to a plant, including any cell, tissue, organ or progeny of a plant. Many suitable plant tissues or plant cells can be transformed, including but not limited to protoplasts, somatic embryos, pollen, leaves, seedlings, stems, callus, stolons, microtubules and shoots. Plant tissue, whether sexually or asexually produced, includes any clones and cuttings or seeds of such plants, seeds, progeny, progeny of any of the foregoing, such as plant stems. Also point
[0332] The term "transformed" as used herein refers to a cell, tissue, organ or organism into which an exogenous DNA molecule such as a construct has been introduced. The introduced DNA molecule can be integrated into the genomic DNA of the recipient cell, tissue, organ or organism such that the introduced DNA molecule is passed on to subsequent progeny. In these embodiments, a "transformed" or "transgenic" cell or plant includes the progeny of that cell or plant and DNA produced and introduced from a breeding program using such transformed plants as parents in crosses. Progeny that exhibit phenotypic changes resulting from the presence of the molecule may also be included. Preferably, the transgenic plants are fertile and capable of transmitting the introduced DNA to their offspring through sexual reproduction.
[0333] The term "progeny", such as the progeny of a transgenic plant, is that which is born from, produced by or derived from a plant or transgenic plant. An introduced DNA molecule may also be transiently introduced into a recipient cell, such that the introduced DNA molecule will not be inherited by subsequent progeny and thus is not considered "transgenic." Thus, as used herein, a "non-transgenic" plant or plant cell is a plant that does not contain foreign DNA stably integrated into its genome.
[0334] The term "plant promoter" as used herein is a promoter capable of initiating transcription in a plant cell, whether or not its origin is a plant cell. Exemplary suitable plant promoters include, but are not limited to, those obtained from bacteria such as Agrobacterium or Rhizobium, including genes expressed in plants, plant viruses and plant cells. .
[0335] As used herein, "fungal cell" refers to any type of eukaryotic cell within the kingdom Fungi. Phylum within the Kingdom of Fungi include the phylum Ascomycota, Basidiomycota, Blastocladiomycota, Chytridiomycota, Glomeromycota, Microsporidia ( Microsporidia) and Neocalimastigomycota. Fungal cells can include yeast, mold and fungi. In some embodiments, the fungal cell is a yeast cell.
[0336] As used herein, the term "yeast cell" refers to any fungal cell within the phylum Ascomycota and Basidiomycota. Yeast cells may include budding yeast cells, fission yeast cells and mold cells. Although not limited to these organisms, many types of yeast used in laboratory and industrial settings are part of the phylum Ascomycota. In some embodiments, the yeast cell is a S. cerervisiae, Kluyveromyces marxianus, or Issatchenkia orientalis cell. Other yeast cells include, without limitation, Candida spp. (e.g. Candida albicans), Yarrowia spp. (e.g. Yarrowia lipolytica). , Pichia spp. (e.g. Pichia pastoris), Kluyveromyces spp. (e.g. Kluyveromyces lactis and Kluyveromyces marcusianus ( Kluyveromyces marxianus), Neurospora spp. (e.g. Neurospora crassa), Fusarium spp. (e.g. Fusarium oxysporum) and Issatchenkia spp. ) (eg Issatchenkia orientalis, also known as Pichia kudriavzevii and Candida acidothermophilum). In some embodiments, the fungal cell is a filamentous fungal cell. As used herein, the term "filamentous fungal cell" refers to any type of fungal cell that grows filamentously, ie, as a hypha or mycelium. Examples of filamentous fungal cells include, without limitation, Aspergillus spp. (e.g. Aspergillus niger), Trichoderma spp. ), Rhizopus spp. (eg Rhizopus oryzae) and Mortierella spp. (eg Mortierella isabellina).
[0337] In some embodiments, the fungal cell is an industrial strain. As used herein, "industrial strain" refers to any strain of fungal cell used in or isolated from an industrial process, such as the production of a product on a commercial or industrial scale. . Industrial strains may refer to fungal species typically used in industrial processes, or it may refer to isolates of fungal species that may also be used for non-industrial purposes (eg, laboratory research). Examples of industrial processes include fermentation (eg, in the production of food or beverage products), distillation, biofuel production, chemical compound production and polypeptide production. Examples of industrial strains include, without limitation, JAY270 and ATCC4124.
[0338] In some embodiments, the fungal cell is a polyploid cell. As used herein, a "polyploid" cell may refer to any cell in which the genome exists in more than one copy. A polyploid cell may refer to a type of cell that is naturally found in the polyploid state, or it may be subject to specific regulation, alteration, It may refer to cells that have been induced (by inactivation, activation or modification). A polyploid cell can refer to a cell whose entire genome is polyploid, or it can refer to a cell that is polyploid at a particular genomic locus of interest. Without wishing to be bound by theory, the abundance of guide RNAs can often be a rate-limiting component in genome engineering of polyploid cells compared to haploid cells, and is therefore described herein. It is conceivable that the Cas13b CRISPR system-based method could take advantage of the use of specific fungal cell types.
[0339] In some embodiments, the fungal cell is a diploid cell. As used herein, a "diploid" cell can refer to any cell in which the genome exists in two copies. A diploid cell may refer to a type of cell found in the diploid state in nature, or it may refer to a cell that exists in a diploid state (e.g., specific regulation of meiosis, cytokinesis or DNA replication). , alteration, inactivation, activation or modification). For example, S. cerevisiae strain S228C can be maintained in a haploid or diploid state. A diploid cell can refer to a cell whose entire genome is diploid, or it can refer to a cell that is diploid at a particular genomic locus of interest. In some embodiments, the fungal cell is a haploid cell. As used herein, a "haploid" cell may refer to any cell in which the genome exists in one copy. A haploid cell may refer to a type of cell that is found in the haploid state in nature, or it may refer to a cell that exists in the haploid state (e.g., specific regulation of meiosis, cytokinesis or DNA replication). , alteration, inactivation, activation or modification). For example, S. cerevisiae strain S228C can be maintained in a haploid or diploid state. A haploid cell can refer to a cell whose entire genome is haploid, or it can refer to a cell that is haploid at a particular genomic locus of interest.
[0340] As used herein, a "yeast expression vector" is a nucleic acid containing one or more sequences encoding RNA and / or polypeptides and any desired vector that controls the expression of that nucleic acid. It refers to a nucleic acid that may further contain elements and any elements that enable the replication and maintenance of an expression vector inside a yeast cell. Many suitable yeast expression vectors and their characteristics are known in the art; for example, various vectors and techniques are described in Yeast Protocols, 2nd edition, Xiao, W., ed. ) and Buckholz, R.G. and Gleeson, M.A. (1991) Biotechnology (NY) 9(11):1067-72. Yeast vectors include, without limitation, a centromere (CEN) sequence, an autonomously replicating sequence (ARS), a promoter such as the RNA polymerase III promoter operably linked to the sequence or gene of interest, a terminator such as the RNA polymerase III terminator, a replication Origin and marker genes (eg, auxotrophs, antibiotics or other selectable markers) may be included. Examples of expression vectors used in yeast include plasmids, yeast artificial chromosomes, 2μ plasmids, yeast integrating plasmids, yeast replicating plasmids, shuttle vectors and episomal plasmids.
[0341] Stable integration of RNA-targeting CRISP system components in the genome of plants and plant cells A detailed embodiment envisions introducing polynucleotides encoding components of an RNA-targeting CRISPR system for stable integration into the genome of a plant cell. In these embodiments, the design of the transformation vector or expression system can be adjusted depending on when, where and under what conditions the guide RNA and / or RNA targeting gene is to be expressed.
[0342] A detailed embodiment envisions stably introducing the components of the RNA-targeting CRISPR system into the genomic DNA of the plant cell. Additionally or alternatively, it is envisioned to introduce components of an RNA-targeting CRISPR system for stable integration into the DNA of plant organelles such as, but not limited to plastids, mitochondria or chloroplasts.
[0343] Expression systems for stable integration into the genome of plant cells may contain one or more of the following elements: promoter elements that may be used to express guide RNAs and / or RNA targeting enzymes in plant cells; 5' untranslated region to enhance; intronic elements to further enhance expression in certain cells, such as monocot cells; for inserting one or more guide RNAs and / or RNA targeting gene sequences and other desired elements; a multiple cloning site that provides convenient restriction sites; and a 3' untranslated region that provides efficient termination of expressed transcripts.
[0344] The elements of an expression system can be on one or more expression constructs that are either circular, such as plasmids or transformation vectors, or non-circular, such as linear double-stranded DNA. In particular embodiments, the RNA targeting CRISPR expression system comprises at least: (a) a nucleotide sequence encoding a guide RNA (gRNA) that hybridizes to a plant target sequence, the guide RNA comprising a guide sequence and a direct repeat sequence; and (b) a nucleotide sequence encoding an RNA targeting protein; including Here, components (a) or (b) are located on the same or different constructs and thereby different nucleotide sequences can be under the control of the same or different regulatory elements operable in plant cells.
[0345] A DNA construct containing the components of an RNA-targeting CRISPR system can be introduced into the genome of a plant, plant part or plant cell by a variety of conventional techniques. This process generally includes the steps of selecting a suitable host cell or host tissue, introducing the construct into the host cell or tissue, and regenerating a plant cell or plant therefrom. In particular embodiments, the DNA construct can be introduced into plant cells using techniques such as, but not limited to, electroporation, microinjection, aerosol beam injection of plant cell protoplasts, or the DNA construct can be introduced into plant cells such as DNA particle bombardment. (see also Fu et al., Transgenic Res. 2000 Feb;9(1):11-9). The basis of particle bombardment is to accelerate particles coated with the gene of interest towards the cell, thereby allowing the particles to enter the cytoplasm and typically obtain stable integration into the genome (e.g. , Klein et al, Nature (1987), Klein et al, Bio / Technology (1992), Casas et al, Proc.
[0346] In a detailed embodiment, DNA constructs containing components of the RNA-targeting CRISPR system can be introduced into plants by Agrobacterium-mediated transformation. The DNA construct can be combined with suitable T-DNA flanking regions and introduced into a conventional Agrobacterium tumefaciens host vector. Foreign DNA is incorporated into the plant's genome by infecting the plant or by incubating plant protoplasts with Agrobacterium bacteria containing one or more Ti (tumor-inducing) plasmids. (See, eg, Fraley et al., (1985), Rogers et al., (1987) and US Pat. No. 5,563,055).
[0347] plant promoter To ensure proper expression in plant cells, the components of the Cas13b CRISPR system described herein are typically placed under the control of plant promoters, ie, promoters operable in plant cells. The use of promoters of various types is envisioned.
[0348] A constitutive plant promoter is capable of causing the open reading frame (ORF) it controls to be expressed in all or nearly all plant tissues at all or nearly all developmental stages of the plant (referred to as "constitutive expression"). possible promoter. One non-limiting example of a constitutive promoter is the cauliflower mosaic virus 35S promoter. The present invention also contemplates methods of modifying RNA sequences and thus modulating the expression of plant biomolecules. Therefore, in particular embodiments of the invention, it is advantageous to place one or more elements of the RNA targeting CRISPR system under the control of a promoter, which may be regulated. "Regulated promoter" refers to a promoter that is not constitutive but directs gene expression in a temporally and / or spatially regulated manner, and includes tissue-specific, tissue-preferred and inducible promoters. Different promoters may direct gene expression in different tissues or cell types, or at different developmental stages, or in response to different environmental conditions. In particular embodiments, one or more of the RNA-targeting CRISPR components are expressed under the control of a constitutive promoter, such as the cauliflower mosaic virus 35S promoter, and utilize tissue-preferred promoters to target specific plant tissues within specific plant tissues. It is possible to target enhanced expression in seed cell types, such as leaf or root vascular cells or seed specific cells. For examples of detailed promoters used in RNA targeting CRISPR systems see Kawamata et al., (1997) Plant Cell Physiol 38:792-803; Yamamoto et al., (1997) Plant J 12:255-65; Hire et al. al, (1992) Plant Mol Biol 20:207-18, Kuster et al, (1995) Plant Mol Biol 29:759-72 and Capana et al., (1994) Plant Mol Biol 25:681-91. . Examples of promoters that are inducible and allow gene editing or spatiotemporal control of gene expression may use some form of energy. Forms of energy may include, but are not limited to, acoustic energy, electromagnetic radiation, chemical energy and / or thermal energy. Examples of inducible systems include tetracycline-inducible promoters (Tet-On or Tet-Off), small molecule two-hybrid transcriptional activation systems (FKBP, ABA, etc.) or light-inducible systems (Phytochrome, LOV domains or Cryptochrome). For example, light-inducible transcriptional effectors (LITEs) that lead to sequence-specific changes in transcriptional activity. Components of a light-inducible system can include the RNA-targeting Cas13b, a light-responsive cytochrome heterodimer (eg, from Arabidopsis thaliana) and a transcriptional activation / repression domain. Further examples of inducible DNA binding proteins and methods of their use are provided in US Provisional Patent Application No. 61 / 736465 and US Provisional Patent Application No. 61 / 721,283, which are hereby incorporated by reference in their entirety. ).
[0349] In particular embodiments, transient or inducible expression can be achieved using, for example, chemically regulated promoters, whereby gene expression is induced upon addition of an exogenous chemical. Modulation of gene expression can also be achieved by chemical repressible promoters, where addition of a chemical represses gene expression. Chemically inducible promoters include, but are not limited to, the maize ln2-2 promoter (De Veylder et al., (1997) Plant Cell Physiol 38:568-77) activated by benzenesulfonamide herbicide antidotes, germination The maize GST promoter (GST-II-27, WO 93 / 01294) activated by hydrophobic electrophilic compounds used as preherbicides and the tobacco PR-1a promoter activated by salicylic acid (Ono et al.) ., (2004) Biosci Biotechnol Biochem 68:803-7). Promoters regulated by antibiotics, such as tetracycline-inducible and tetracycline-repressible promoters (Gatz et al., (1991) Mol Gen Genet 227:229-37; U.S. Pat. Nos. 5,814,618 and 5,789,156) can also be used herein.
[0350] Translocation and / or expression in specific plant organelles Expression systems may include elements for translocation to and / or expression in specific plant organelles.
[0351] Chloroplast targeting Detailed embodiments envision the use of an RNA-targeting CRISPR system to specifically alter the expression and / or translation of chloroplast genes, or ensure expression in the chloroplast. For this purpose, chloroplast transformation methods or compartmentalization of RNA-targeted CRISPR components into the chloroplast are used. For example, introducing genetic modifications into the plastid genome can mitigate biosafety issues such as gene flow through pollen.
[0352] Chloroplast transformation methods are known in the art and include particle bombardment, PEG treatment and microinjection. Additionally, methods involving transfer of the transformation cassette from the nuclear genome to the plastid can be used, as described in WO2010061186.
[0353] Alternatively, targeting one or more of the RNA targeting CRISPR components to the plant chloroplast is envisioned. This is accomplished by incorporating into the expression construct a sequence encoding a chloroplast transit peptide (CTP) or plastid transit peptide operably linked to the 5' region of the sequence encoding the RNA targeting protein. CTP is removed in a processing step during transition to the chloroplast. Chloroplast targeting of expressed proteins is well known to those of skill in the art (see, eg, Protein Transport into Chloroplasts, 2010, Annual Review of Plant Biology, Vol. 61:157-180). In such embodiments, it may also be desirable to target one or more guide RNAs to the plant chloroplast. Methods and constructs that can be used to translocate guide RNAs into chloroplasts using chloroplast localization sequences are described, for example, in US Patent Application Publication No. 20040142476 (incorporated herein by reference). is described in By incorporating such various constructs into the expression system of the present invention, the RNA targeting guide RNA can be transferred efficiently.
[0354] Introduction of a polynucleotide encoding a CRISPR-RNA targeting system in algal cells Transgenic algae (or other plants such as canola) may be particularly useful for the production of vegetable oils or biofuels such as alcohols (particularly methanol and ethanol) or other products. They can be engineered to express or overexpress high levels of oils or alcohols used in the oil or biofuel industry.
[0355] US Pat. No. 8,945,839 describes methods for engineering microalgae (Chlamydomonas reinhardtii cells) species using Cas9. Using similar tools, the RNA-targeting CRISPR-based methods described herein can be applied to Chlamydomonas species and other algae. In a detailed embodiment, the RNA targeting protein and guide RNA are introduced and expressed in algae using vectors that express the RNA targeting protein under the control of a constitutive promoter such as Hsp70A-Rbc S2 or β2-tubulin. Guide RNA is optionally delivered using a vector containing a T7 promoter. Alternatively, RNA targeting mRNA and in vitro transcription guide RNA can be delivered to algal cells. Electroporation protocols are available to those skilled in the art, such as the standard recommended protocols from the GeneArt Chlamydomonas Engineering Kit.
[0356] Introduction of Polynucleotides Encoding RNA Targeting Components in Yeast Cells In particular embodiments, the invention relates to the use of the RNA-targeting CRISPR system for RNA editing in yeast cells. Methods for transforming yeast cells that can be used to introduce polynucleotides encoding RNA targeting CRISPR system components are well known to those skilled in the art and are described in Kawai et al., 2010, Bioeng Bugs. 2010 Nov-Dec;1(6). ):395-403). Non-limiting examples include transformation of yeast cells by lithium acetate treatment (which may further include carrier DNA and PEG treatment), by bombardment or electroporation.
[0357] Transient expression of RNA-targeting CRISP system components in plants and plant cells In particular embodiments, transient expression of the guide RNA and / or RNA targeting gene in the plant cell is envisioned. These embodiments ensure that the RNA-targeting CRISPR system modifies the RNA-target molecule only when both the guide RNA and the RNA-targeting protein are present in the cell, thus further controlling gene expression. Since expression of the RNA targeting enzyme is transient, plants regenerated from such plant cells typically do not contain foreign DNA. In particular embodiments, the RNA targeting enzyme is stably expressed by the plant cell and the guide sequence is transiently expressed.
[0358] In a particularly preferred embodiment, RNA targeting CRISPR system components can be introduced into plant cells using plant viral vectors (Scholthof et al. 1996, Annu Rev Phytopathol. 1996;34:299-323). In a further detailed embodiment, said viral vector is a vector derived from a DNA virus. For example, geminiviruses (e.g., cabbage leaf curl virus, bean dwarf virus, wheat dwarf virus, tomato leaf curl virus, corn streak virus, tobacco leaf curl virus, or tomato golden mosaic virus) or nanoviruses (e.g., broad bean spotted virus). virus). In another detailed embodiment, said viral vector is a vector derived from an RNA virus. For example, Tobravirus (eg, tobacco stem bark virus, tobacco mosaic virus), Potexvirus (eg, potato X virus) or Hordeivirus (eg, wheat spotted leaf mosaic virus). The replicating genomes of plant viruses are non-integrating vectors, which is advantageous in relation to avoiding the production of GMO plants.
[0359] In particular embodiments, the vector used for transient expression of the RNA-targeting CRISPR construct is, for example, the pEAQ vector, which is tailored for Agrobacterium-mediated transient expression in protoplasts. (Sainsbury F. et al., Plant Biotechnol J. 2009 Sep;7(7):682-93). Precise targeting of genomic locations was demonstrated using modified cabbage curl virus (CaLCuV) vectors to express gRNAs in stable transgenic plants expressing Cas13b (Scientific Reports 5, Article number: 14926 (2015) , doi:10.1038 / srep14926).
[0360] In particular embodiments, double-stranded DNA fragments encoding guide RNAs or crRNAs and / or RNA targeting genes can be transiently introduced into plant cells. In such embodiments, the introduced double-stranded DNA fragment is provided in an amount sufficient to modify the RNA molecule in the cell, but not remain after the intended period of time or after one or more cell divisions. be. Methods for directing DNA transfer in plants are known to those of skill in the art (see, eg, Davey et al. Plant Mol Biol. 1989 Sep;13(3):273-85).
[0361] In other embodiments, the RNA polynucleotide encoding the RNA targeting protein is released after a period of time sufficient, but contemplated, to modify the RNA molecule (in the presence of at least one guide RNA). It is introduced into the plant cell in an amount that does not remain after more than one cell division, which is then translated and processed by the host cell to produce protein. Methods for introducing mRNA into plant protoplasts for transient expression are known to those skilled in the art (see, eg, Gallie, Plant Cell Reports (1993), 13; 119-122). Combinations of the different methods described above are also envisioned.
[0362] Delivery of RNA-targeting CRISPR components to plant cells In particular embodiments, it is beneficial to directly deliver one or more components of the RNA targeting CRISPR system to the plant cell. This is particularly useful for producing non-transgenic plants. In particular embodiments, one or more of the RNA targeting components are prepared outside the plant or plant cell and delivered to the cell. For example, in a detailed embodiment, an RNA targeting protein is prepared in vitro and then introduced into a plant cell. RNA targeting proteins can be prepared by a variety of methods known to those skilled in the art, including recombinant production. After expression, the RNA targeting protein is isolated, optionally refolded, purified, and optionally treated to remove any purification tags, such as His tags. Once a crude, partially purified or more completely purified RNA targeting protein is obtained, the protein can be introduced into a plant cell.
[0363] In a detailed embodiment, an RNA targeting protein is mixed with a guide RNA that targets the RNA of interest to form a preassembled ribonucleoprotein.
[0364] Individual components or preassembled ribonucleoproteins may be transported across cell membranes using electroporation, by bombardment of particles coated with RNA targeting-related gene products, by chemical transfection, or by other means. can be introduced into plant cells by any means of For example, transfection of plant protoplasts with pre-assembled CRISPR ribonucleoproteins has been demonstrated to ensure targeted modification of the plant genome (Woo et al. Nature Biotechnology, 2015; DOI:10.1038 / nbt. .3389). These methods can be modified to achieve targeted modification of RNA molecules in plants.
[0365] In a detailed embodiment, RNA targeting CRISPR system components are introduced into plant cells using nanoparticles. Components, either as proteins or nucleic acids, or combinations thereof, can be uploaded onto or packaged in nanoparticles and applied to plants (e.g., WO2008042156 and US Patent such as the description of Japanese Patent Application Publication No. 20130185823). In particular, embodiments of the present invention provide that a DNA molecule encoding an RNA targeting protein, a DNA molecule encoding a guide RNA and / or an isolated guide RNA are uploaded or packaged as described in WO2015089419. containing nano-particles.
[0366] A further means of introducing one or more components of the RNA-targeting CRISPR system into plant cells is through the use of cell penetrating peptides (CPPs). Accordingly, in particular embodiments of the invention include compositions comprising a cell penetrating peptide linked to an RNA targeting protein. In particular embodiments of the invention, RNA targeting proteins and / or guide RNAs are coupled to one or more CPPs to effectively transport them inside plant protoplasts (for Cas9 in human cells see Ramakrishna (2014 Genome Res. 2014 Jun;24(6):1020-7) In other embodiments, RNA targeting genes and / or guide RNAs are coupled to one or more CPPs for plant protoplast delivery. Encoded by one or more circular or non-circular DNA molecules Plant protoplasts are then regenerated into plant cells and even plants CPPs generally transport biomolecules across cell membranes in a receptor-independent manner Described as short peptides of less than 35 amino acids derived either from proteins with the ability to transport or from chimeric sequences, CPPs include cationic peptides, peptides with hydrophobic sequences, amphiphilic peptides, proline-rich They can be peptides with microbial sequences and chimeric or bipartite peptides (Pooga and Langel 2005).CPPs are capable of permeating biological membranes, thus allowing the movement of various biomolecules across cell membranes into the cytoplasm. and can improve their intracellular trafficking, thus facilitating the interaction of biomolecules with targets.Examples of CPPs include, among others, the intranuclear cell required for viral replication by HIV type 1. Tat, which is a transcriptional activation protein, penetratin, Kaposi fibroblast growth factor (FGF) signal peptide sequence, integrin β3 signal peptide sequence; polyarginine peptide Arg sequence, guanine-rich molecular transporter, sweet arrow peptide and the like.
[0367] Target RNA expected to be applied to plants, algae or fungi The target RNA, ie the RNA of interest, is the RNA targeted by the present invention, leading to the recruitment and binding of the RNA targeting protein to and binding to the desired target site on the target RNA. Target RNA can be any suitable form of RNA. This may include mRNA in some embodiments. In other embodiments, target RNA may include transfer RNA (tRNA) or ribosomal RNA (rRNA). In other embodiments, target RNAs can include interfering RNAs (RNAi), microRNAs (miRNAs), microswitches, microzymes, satellite RNAs and RNA viruses. The target RNA can be located in the cytoplasm of the plant cell, or in the cell nucleus, or in plant organelles such as mitochondria, chloroplasts or plastids.
[0368] In particular embodiments, an RNA-targeting CRISPR system is used to cleave RNA or otherwise inhibit RNA expression.
[0369] Use of RNA-targeting CRISPR system for regulation of plant gene expression by RNA regulation RNA targeting proteins can also be used to target gene expression through regulation of RNA processing, along with suitable guide RNAs. Regulation of RNA processing may include RNA splicing, including alternative splicing; viral replication (particularly that of plant viruses, including plant viroids) and RNA processing reactions such as tRNA biogenesis. Suitable guides. RNA targeting proteins in combination with RNA can also be used to regulate RNA activation (RNAa) RNAa leads to enhanced gene expression, so regulation of gene expression results in disruption or reduction of RNAa and thus reduced promotion of gene expression. It can be realized in the form of
[0370] The RNA targeting effector proteins of the invention can also be used for antiviral activity in plants, especially against RNA viruses. Effector proteins can be targeted to viral RNA using a suitable guide RNA selective for the viral RNA sequence of choice. Specifically, the effector protein can be an active nuclease that cleaves RNA, such as single-stranded RNA. Accordingly, use of the RNA targeting effector proteins of the invention as antiviral agents is provided. Examples of viruses that can be antagonized in this manner include, but are not limited to, Tobacco Mosaic Virus (TMV), Tomato Spotted Spot Virus (TSWV), Cucumber Mosaic Virus (CMV), Potato Virus Y (PVY), Cauliflower Mosaic Virus. (CaMV) (RT virus), Plumpox virus (PPV), Brom mosaic virus (BMV) and Potato X virus (PVX).
[0371] Examples of regulation of RNA expression in plants, algae or fungi as an alternative to targeted genetic modification are further described herein.
[0372] Of particular interest is the regulated regulation of gene expression by regulated mRNA cleavage. This can be accomplished by placing the RNA targeting element under the control of a regulated promoter as described herein.
[0373] Use of RNA-targeting CRISPR system to restore function of tRNA molecules Pring et al describe RNA editing in plant mitochondria and chloroplasts that alters mRNA sequences to encode proteins different from DNA. (Plant Mol. Biol. (1993) 21(6):1163-1170.doi:10.1007 / BF00023611). In particular embodiments of the invention, introduction of elements of the RNA-targeting CRISPR system that specifically target mitochondrial and chloroplast mRNAs into plants or plant cells allows expression of different proteins in such plant organelles. can mimic the processes that occur in vivo.
[0374] Using the RNA-targeting CRISPR system as an alternative to RNA interference to inhibit RNA expression RNA-targeting CRISPR systems have similar uses to RNA inhibition or RNA interference and can therefore replace such methods. In particular embodiments, the methods of the invention include, for example, the use of RNA-targeting CRISPR as an alternative to interfering ribonucleic acids (such as siRNA or shRNA or dsRNA). Examples of inhibition of RNA expression in plants, algae or fungi as an alternative to targeted genetic modification are further described herein.
[0375] Using an RNA-targeting CRISPR system to control RNA interference Regulation of interfering RNAs or miRNAs can help reduce off-target effects (OTEs) seen in approaches by shortening the life span of interfering RNAs or miRNAs in vivo or in vitro. In particular embodiments, target RNAs may include interfering RNAs, ie, RNAs involved in the RNA interference pathway, such as shRNAs, siRNAs, and the like. In other embodiments, target RNA may include microRNA (miRNA) or double-stranded RNA (dsRNA).
[0376] In other particular embodiments, an RNA targeting protein and a suitable guide RNA are selected (e.g. under the control of a spatially or temporally regulated promoter, e.g. a tissue-specific or cell cycle-specific promoter and / or enhancer). When specifically expressed, it can be used to "protect" a cell or system (in vivo or in vitro) from RNAi in that cell. This is useful for comparison in adjacent tissues or cells where RNAi is not required, or in cells or tissues in which effector proteins and suitable guides are expressed and not expressed (i.e. RNAi unregulated and RNAi regulated, respectively). could be. RNA targeting proteins can be used to regulate or bind to molecules comprising or consisting of RNA, such as ribozymes, ribosomes or riboswitches. In embodiments of the invention, the guide RNA recruits the RNA targeting protein to such molecules, allowing the RNA targeting protein to bind to them.
[0377] The RNA-targeting CRISPR system of the present invention can be used for pest management, plant disease management and herbicide tolerance without undue experimentation from the present disclosure, as the present application provides the basis for the informed engineering of this system. It can be applied in the area of implanta RNAi technology, including management and in plant assays and for other applications (e.g. Kim et al., Pesticide Biochemistry and Physiology (Impact Factor: 2.01).01 / 2015; 120.DOI:10.1016 / j.pestbp.2015.01.002;Sharma et al.in Academic Journals(2015),Vol.12(18)pp2303-2312);Green J.M,inPest Management Science,Vol 70(9),pp 1351-1357).
[0378] Use of RNA-targeting CRISPR systems to engineer riboswitches and control metabolic regulation in plants, algae and fungi Riboswitches (also known as aptazymes) are regulatory segments of messenger RNA that bind small molecules and thus regulate gene expression. This mechanism allows cells to sense intracellular concentrations of these small molecules. A particular riboswitch typically regulates its neighboring genes by altering the transcription, translation or splicing of that gene. Accordingly, specific embodiments of the present invention contemplate regulation of riboswitch activity by using an RNA targeting protein in combination with a suitable guide RNA to target the riboswitch. This may be due to cleavage of the riboswitch or binding to it. A detailed embodiment envisions a reduction in riboswitch activity. A riboswitch that binds thiamine pyrophosphate (TPP) was recently characterized and found to regulate thiamine biosynthesis in plants and algae. Furthermore, this element appears to be an essential regulator of primary metabolism in plants (Bocobza and Aharoni, Plant J.2014 Aug;79(4):693-703.doi:10.1111 / tpj.12540.Epub 2014 Jun. 17). The TPP riboswitch is also found in certain fungi, such as Neurospora crassa, where it regulates alternative splicing to conditionally produce an upstream open reading frame (uORF), thereby Affects gene expression (Cheah MT et al., (2007) Nature 447(7143):497-500.doi:10.1038 / nature05769). The RNA targeting CRISPR system described herein can be used to manipulate the endogenous riboswitch activity of plants, algae or fungi, thus altering the expression of downstream genes regulated thereby. In particular embodiments, the RNA targeting CRISP system can be used in assays of riboswitch function in vivo or in vitro and studies of its relevance to metabolic networks. In a detailed embodiment, the RNA-targeting CRISPR system could potentially be used to engineer riboswitches as metabolite sensors in plants and genetic control platforms.
[0379] Use of RNA-targeting CRISPR systems in plant, algal or fungal RNAi screens RNAi screens can identify gene products whose knockdown is associated with phenotypic changes to probe biological pathways and identify components. In particular embodiments of the invention, the Guide 29 or Guide 30 proteins and suitable guide RNAs described herein are used to eliminate or reduce the activity of RNAi in the screen, thus (preventing interference). By restoring the activity of the gene product (by removing or reducing the interference / suppression), control may be exerted on or during such screens.
[0380] Use of RNA targeting proteins to visualize RNA molecules in vivo and in vitro In particular embodiments, the invention provides nucleic acid binding systems. In situ hybridization of RNA with complementary probes is a powerful technique. Typically, fluorescent DNA oligonucleotides are used for the detection of nucleic acids by hybridization. Although increased efficiency has been achieved with certain modifications, such as locked nucleic acids (LNA), there remains a need for efficient and versatile alternatives. Thus, labeling elements of RNA targeting systems can be used as an alternative to efficient and adaptable systems for in situ hybridization.
[0381] Further application of the RNA-targeting CRISPR system in plants and yeast Use of RNA-targeting CRISPR system in biofuel production The term "biofuel," as used herein, is an alternative fuel made from plants and plant-derived resources. Renewable biofuels can be extracted from organic matter that has been energized through the process of carbon fixation, or produced by the use or conversion of biomass. This biomass can be used directly for biofuels or converted into convenient energy-bearing substances by thermal, chemical and biochemical conversions. This biomass conversion can yield fuel in solid, liquid or gaseous form. There are two types of biofuels: bioethanol and biodiesel. Bioethanol is produced primarily by the sugar fermentation process of cellulose (starch), mostly derived from corn and sugar cane. Biodiesel, on the other hand, is mainly produced from oil crops such as rapeseed, palm and soybean. Biofuels are mainly used for transportation.
[0382] Enhancing plant properties for biofuel production In a particular embodiment, the cell walls are lysed using methods using RNA-targeted CRISPR as described herein to make key hydrolytic agents accessible for more efficient release of sugars during fermentation. change the characteristics of In particular embodiments, cellulose and / or lignin biosynthesis is modified. Cellulose is the major constituent of cell walls. Cellulose and lignin biosynthesis are co-regulated. By reducing the proportion of lignin in the plant, the proportion of cellulose can be increased. In particular embodiments, the methods described herein are used to downregulate lignin biosynthesis in plants, thereby increasing fermentable carbohydrates. More specifically, 4-coumarate 3-hydroxylase (CH), phenylalanine ammonia lyase (PAL), cinnamon acid-4-hydroxylase (C4H), hydroxycinnamoyltransferase (HCT), caffeic acid O-methyltransferase (COMT), caffeoyl CoA3-O-methyltransferase (CCoAOMT), ferulic acid 5-hydroxylase (F5H), at least a first selected from the group consisting of cinnamyl alcohol dehydrogenase (CAD), cinnamoyl CoA-reductase (CCR), 4-coumarate-CoA ligase (4CL), monolignol-lignin specific glycosyltransferase and aldehyde dehydrogenase (ALDH) lignin biosynthetic genes are downregulated.
[0383] In particular embodiments, the methods described herein are used to produce plant masses that generate lower levels of acetic acid during fermentation (see also WO2010096488).
[0384] Yeast modification for biofuel production In particular embodiments, RNA targeting enzymes provided herein are used for bioethanol production by recombinant microorganisms. For example, RNA targeting enzymes can be used to engineer microorganisms such as yeast to make biofuels or biopolymers from fermentable sugars, and optionally agricultural waste as a source of fermentable sugars. It is possible to degrade plant-derived lignocellulose obtained from More particularly, the present invention provides methods of using RNA-targeting CRISPR complexes to alter the expression of endogenous genes required for biofuel production and / or to alter endogenous genes that may interfere with biofuel synthesis. offer. More specifically, the method involves stimulating expression in a microorganism, such as yeast, of one or more nucleotide sequences encoding enzymes involved in the conversion of pyruvate to ethanol or another desired product. In particular embodiments, the method ensures stimulation of the expression of one or more enzymes that effect microbial degradation of cellulose, such as cellulase. In yet another embodiment, RNA-targeted CRISPR complexes are used to suppress endogenous metabolic pathways that compete with biofuel production pathways.
[0385] Modification of algae and plants for the production of vegetable oils or biofuels Transgenic algae or other plants such as rape can be particularly useful for the production of vegetable oils or biofuels such as alcohols (particularly methanol and ethanol). They can be engineered to express or overexpress high levels of oils or alcohols used in the oil or biofuel industry.
[0386] US Pat. No. 8,945,839 describes methods for engineering microalgae (Chlamydomonas reinhardtii cells) species using Cas9. Using similar tools, the RNA-targeting CRISPR-based methods described herein can be applied to Chlamydomonas species and other algae. In a detailed embodiment, the RNA targeting effector protein and guide RNA are introduced and expressed in algae using vectors that express the RNA targeting effector protein under the control of a constitutive promoter such as Hsp70A-Rbc S2 or β2-tubulin. Let Guide RNA will be delivered using a vector containing a T7 promoter. Alternatively, in vitro transcribed guide RNA can be delivered to algal cells. The electroporation protocol follows the standard recommended protocol of the GeneArt Chlamydomonas Engineering Kit.
[0387] Detailed application of RNA-targeting enzymes in plants In a particular embodiment, the present invention is capable of cleaving viral RNA and thus can be used as a therapeutic agent for virus elimination in plant systems. Previous studies in human systems have demonstrated the successful use of CRISPR in targeting a single-stranded RNA virus, hepatitis C (A.Price, et al., Proc.Natl.Acad.Sci, 2015). These methods can also be adapted for use with RNA-targeting CRISPR systems in plants.
[0388] improved plant The present invention also provides plant and yeast cells obtainable by and obtained by the methods provided herein. The improved plants obtained by the methods described herein are, for example, food or feed by modifying the expression of genes that ensure tolerance to plant pests, herbicides, drought, cold or hot, excess water, etc. It can be useful in manufacturing.
[0389] The improved plants, particularly crops and algae, obtained by the methods described herein are, for example, food or feed products, by expression of higher protein, carbohydrate, nutrient or vitamin levels than would normally be found in the wild type. It can be useful in manufacturing. In this respect, improved plants, especially legumes and tubers, are preferred.
[0390] Improved algae or other plants such as rape can be particularly useful for the production of vegetable oils or biofuels such as alcohols (particularly methanol and ethanol). They can be engineered to express or overexpress high levels of oils or alcohols used in the oil or biofuel industry.
[0391] The invention also provides improved plant parts. Plant parts include, but are not limited to, leaves, stems, roots, tubers, seeds, endosperm, ovules and pollen. Plant parts as contemplated herein may be viable, non-viable, renewable and / or non-renewable.
[0392] Also included herein is the provision of plant cells and plants produced by the methods of the invention. Also included within the scope of the invention are gametes, seeds, embryos (whether zygotic or somatic), progeny or hybrids of plants containing genetic modifications produced by conventional breeding methods. Such plants may contain heterologous or foreign DNA sequences inserted into or in place of the target sequence. Alternatively, such plants may contain only certain changes (mutations, deletions, insertions, substitutions) in one or more nucleotides. Such plants will therefore differ from their progenitor plants only in the presence of certain modifications.
[0393] In one embodiment of the invention, pathogen-resistant plants are engineered using the Cas13b system to create resistance to diseases caused by, for example, bacteria, fungi or viruses. In certain embodiments, pathogen resistance can be achieved by engineering crops to create a Cas13b system that will be ingested by pests and lead to death. In one embodiment of the invention, the Cas13b system is used to engineer abiotic stress tolerance. In another embodiment, the Cas13b system is used to engineer drought stress tolerance or salt stress tolerance or cold or heat stress tolerance. Younis et al. 2014, Int. J. Biol. Are suitable. Some non-limiting target crops include Arabidopsis maize (Zea mays), thaliana, rice (Oryza sativa L), plum (Prunus domestica L.), cotton (Gossypium). hirsutum), Nicotiana rustica, maize (Zea mays), Medicago sativa, Nicotiana benthamiana and Arabidopsis thaliana.
[0394] In one embodiment of the invention, the Cas13b system is used to control crop pests. For example, Cas13b systems operable in crop pests can be expressed from plant hosts or directly transferred to targets using, for example, viral vectors.
[0395] In certain embodiments, the invention provides methods for efficiently generating homozygous organisms from heterozygous non-human starting organisms. In one embodiment, the invention is used in plant breeding. In another embodiment, the invention is used in animal breeding. In such embodiments, homozygous organisms, such as plants or animals, act by preventing or suppressing recombination by interfering with at least one target gene involved in double-strand breaks, chromosomal pairing and / or strand exchange. be done.
[0396] Application of CAS13B protein in an optimized functional RNA targeting system In one aspect, the present invention provides systems for specifically delivering functional components into an RNA environment. This can be ensured using a CRISPR system comprising the RNA targeting effector proteins of the invention that allow specific targeting of various components to the RNA. More specifically, such components include activators or repressors, such as activators or repressors of RNA translation, degradation, etc. Applications of this system are described elsewhere herein.
[0397] According to one aspect, the invention provides a non-naturally occurring or engineered composition comprising a guide RNA comprising a guide sequence capable of hybridizing to a target sequence within a genomic locus of interest in a cell, wherein In , the guide RNA is modified by the insertion of one or more individual RNA sequences that bind to adapter proteins. In particular embodiments, an RNA sequence can bind to two or more adapter proteins (eg, aptamers), and each adapter protein is associated with one or more functional domains. The guide RNA for the Cas13bc2c2 enzyme described herein is shown to be suitable for modification of the guide sequence. In particular embodiments, the guide RNA is modified by insertion of a separate RNA sequence 5' to the direct repeat, within the direct repeat, or 3' to the guide sequence. When there is more than one functional domain, those functional domains may be the same or different, eg two same or two different activators or repressors. In certain embodiments, the invention provides compositions discussed herein, wherein one or more functional domains have been added to an RNA targeting enzyme such that the functional domains are , resulting in a spatial arrangement that enables a functional domain to function in its assigned function; A CRISPR-Cas complex having five functional domains, at least one of which is associated with an RNA targeting enzyme and at least two of which are associated with gRNAs.
[0398] Accordingly, in one aspect, the invention provides a non-naturally occurring or engineered CRISPR-Cas13b complex composition comprising a guide RNA as discussed herein and an RNA targeting enzyme, Cas13b, wherein and optionally, the RNA targeting enzyme has at least one mutation and optionally at least one or more such that the RNA targeting enzyme has 5% or less of the nuclease activity of the enzyme without the at least one mutation. containing one or more nuclear localization sequences of In particular embodiments, the guide RNA may additionally or alternatively still ensure binding of the RNA targeting enzyme, but prevent cleavage by the RNA targeting enzyme (detailed elsewhere herein). as per) is modified.
[0399] In particular embodiments, the RNA targeting enzyme is a Casl3b enzyme that has at least 97% or 100% reduced nuclease activity when compared to a Casl3b enzyme that does not have at least one mutation. In certain embodiments, the invention provides compositions discussed herein, wherein the Cas13b enzyme comprises two or more mutations otherwise discussed herein.
[0400] In particular embodiments, an RNA targeting system as described herein above is provided comprising two or more functional domains. In particular embodiments, the two or more functional domains are heterologous functional domains. In particular embodiments, the system comprises an adapter protein that is a fusion protein comprising a functional domain, the fusion protein optionally comprising a linker between the adapter protein and the functional domain. In particular embodiments, the linker comprises a GlySer linker. Additionally or alternatively, one or more functional domains are attached to the RNA effector protein by a linker, optionally a GlySer linker. In particular embodiments, one or more functional domains are added to the RNA targeting enzyme at one or both of the HEPN domains.
[0401] In certain aspects, the invention provides compositions discussed herein, wherein one or more functional domains associated with an adapter protein or RNA targeting enzyme activate or repress RNA translation. It is a domain that has the ability to In certain embodiments, the invention provides compositions discussed herein, wherein at least one of the one or more functional domains associated with the adapter protein comprises methylase activity, demethylase activity, transcription one or more including activating activity, transcription repressing activity, transcription terminator activity, histone modifying activity, DNA integration activity, RNA cleavage activity, DNA cleavage activity or nucleic acid binding activity or molecular switching activity or chemical or light inducible ability active.
[0402] In certain aspects, the invention provides compositions discussed herein comprising aptamer sequences. In particular embodiments, the aptamer sequences are two or more aptamer sequences specific for the same adapter protein. In certain embodiments, the invention provides compositions discussed herein, wherein the aptamer sequences are two or more aptamer sequences specific for different adapter proteins. In certain embodiments, the invention provides compositions discussed herein, wherein the adapter proteins are MS2, PP7, Qβ, F2, GA, fr, JP501, M12, R17, BZ13, JP34, Including JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, PRR1. Thus, in particular embodiments, the aptamer is selected from binding proteins that specifically bind any one of the adapter proteins listed above. In one aspect, the invention provides compositions discussed herein, wherein the cell is a eukaryotic cell. In certain aspects, the invention provides compositions discussed herein, wherein the eukaryotic cell is a mammalian, plant or yeast cell, whereby the mammalian cell optionally comprises: mouse cells. In certain aspects, the invention provides compositions discussed herein, wherein the mammalian cells are human cells.
[0403] In certain embodiments, the invention provides compositions discussed herein above, wherein there are two or more guide RNAs, or gRNAs, or crRNAs, and which target different sequences, Thereby there is multiplexing when the composition is used. In certain embodiments, the invention provides compositions in which there are two or more guide RNAs, or gRNAs, or crRNAs modified by insertion of individual RNA sequences that bind to one or more adapter proteins.
[0404] In certain embodiments, the invention provides compositions discussed herein, wherein one or more adapter proteins associated with one or more functional domains are present and inserted into the guide RNA Binds to individual RNA sequences.
[0405] In certain embodiments, the invention provides compositions discussed herein, wherein the guide RNA is modified to have at least one non-coding functional loop; Functional loops are inhibitory; for example, at least one non-coding functional loop contains Alu.
[0406] In certain embodiments, the invention provides methods of altering gene expression comprising administering to a host or in vivo expression in a host one or more of the compositions as discussed herein.
[0407] In one aspect, the invention provides a method discussed herein comprising delivery of a composition or a nucleic acid molecule encoding it, wherein said nucleic acid molecule is operably linked to a regulatory sequence, and expressed in vivo. In one aspect, the invention provides the methods discussed herein, wherein the in vivo expression is lentivirus, adenovirus or AAV mediated.
[0408] In one aspect, the invention provides a mammalian cell line of cells as discussed herein, wherein the cell line is optionally a human cell line or a mouse cell line. In certain aspects, the invention provides a transgenic mammalian model, optionally a mouse, wherein the model has been transformed with a composition discussed herein, or are descendants.
[0409] In one aspect, the invention provides a nucleic acid molecule encoding a guide RNA or RNA-targeting CRISPR-Cas13b complex or composition as discussed herein. In one aspect, the invention provides a vector comprising a nucleic acid molecule encoding a guide RNA (gRNA) or crRNA comprising a guide sequence capable of hybridizing to a cellular RNA target sequence, wherein a direct repeat of the gRNA or crRNA is modified by the insertion of distinct RNA sequences that bind to two or more adapter proteins, and each adapter protein is associated with one or more functional domains; modified to have a loop. In certain embodiments, the present invention provides a gRNA or crRNA discussed herein and an RNA targeting enzyme [optionally, the RNA targeting enzyme comprises at least one mutation, such that the RNA targeting enzyme comprises at least one has 5% or less of the nuclease activity of an unmutated RNA targeting enzyme, and optionally contains at least one or more nuclear localization sequences] A vector containing a nucleic acid molecule encoding the modified CRISPR-Cas13b complex composition is provided. In some embodiments, the vector is a eukaryotic protein operably linked to a guide RNA (gRNA) or a nucleic acid molecule encoding a crRNA and / or a nucleic acid molecule encoding an RNA targeting enzyme and / or an optional nuclear localization sequence. It may further comprise regulatory elements operable in cells.
[0410] In one aspect, the invention provides kits comprising one or more of the components described herein. In some embodiments, the kit includes a vector system as described herein and instructions for use of the kit.
[0411] In certain embodiments, the invention provides methods of screening for gain-of-function (GOF) or loss-of-function (LOF) or screening of non-coding RNAs or potential regulatory regions (e.g., enhancers, repressors), which methods is a cell line or model as discussed herein containing or expressing an RNA targeting enzyme and a composition as discussed herein. by which the gRNA or crRNA comprises either an activator or a repressor, and that the introduced gRNA or crRNA comprises an activator with respect to those cells or with respect to those cells Whether the introduced gRNA or crRNA contains a repressor, including monitoring GOF or LOF, respectively.
[0412] In one embodiment, the present invention provides a non-naturally occurring or engineered RNA-targeting CRISPR guide RNA (gRNA) or crRNA, each comprising an RNA-targeting enzyme, comprising a guide sequence capable of hybridizing to a target RNA sequence of interest in a cell. A library of compositions is provided, wherein the RNA targeting enzyme comprises at least one mutation, such that the RNA targeting enzyme has a nuclease activity that is 5% or less of an RNA targeting enzyme that does not have the at least one mutation. and the gRNA or crRNA is modified by the insertion of distinct RNA sequences that bind to one or more adapter proteins, and the adapter proteins are associated with one or more functional domains, the composition comprising one or more or Including two or more adapter proteins, each protein associated with one or more functional domains, and gRNAs or crRNAs comprising a genome-wide library comprising multiple RNA targeting guide RNAs (gRNAs) or crRNAs. In some embodiments, the invention provides a library as discussed herein, wherein the RNA targeting RNA targeting enzyme has at least 97% or have 100% reduced nuclease activity. In one aspect, the invention provides a library as discussed herein, wherein the adapter protein is a fusion protein comprising the functional domain. In one aspect, the invention provides a library as discussed herein, wherein the gRNAs or crRNAs are modified by insertion of individual RNA sequences that bind to one or more adapter proteins. not. In one aspect, the invention provides a library as discussed herein, wherein one or more functional domains are associated with RNA targeting enzymes. In one aspect, the invention provides a library as discussed herein, wherein the cell population of cells is a eukaryotic cell population. In one aspect, the invention provides a library as discussed herein, wherein the eukaryotic cell is a mammalian cell, plant cell or yeast cell. In one aspect, the invention provides a library as discussed herein, wherein the mammalian cells are human cells. In one aspect, the invention provides a library as discussed herein, wherein the population of cells is a population of embryonic stem (ES) cells.
[0413] In one aspect, the invention provides a library as discussed herein, wherein targeting is about 100 or more RNA sequences. In one aspect, the invention provides a library as discussed herein, wherein targeting is about 1000 or more RNA sequences. In one aspect, the invention provides a library as discussed herein, wherein targeting is about 20,000 or more sequences. In one aspect, the invention provides a library as discussed herein, wherein the targeting is the entire transcriptome. In one aspect, the invention provides a library as discussed herein, wherein targeting is a panel of target sequences focused on relevant or desirable pathways. In one aspect, the invention provides a library as discussed herein, wherein the pathway is an immunization pathway. In one aspect, the invention provides a library as discussed herein, wherein the pathway is a cell division pathway.
[0414] In one aspect, the invention provides a method of generating a model eukaryotic cell containing a gene with altered expression. In some embodiments, the disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method comprises (a) introducing into a eukaryotic cell one or more vectors encoding components of the systems described herein above; binding the target polynucleotide to alter expression of the gene, thereby creating a model eukaryotic cell containing altered gene expression.
[0415] The structural information provided herein makes it possible to examine guide RNA or crRNA interactions with target RNAs and RNA targeting enzymes so that the functionality of the entire RNA targeting CRISPR-Cas13b system is optimized. It becomes possible to engineer or change the guide RNA structure. For example, the insertion of an adapter protein capable of binding to RNA can extend the guide RNA or crRNA without conflict with the RNA targeting protein. These adapter proteins can further recruit effector proteins or fusions containing one or more functional domains.
[0416] Certain aspects of the invention are those in which the above elements are included in a single composition or in separate compositions. These compositions can be advantageously applied to the host to induce functional effects at the genomic level.
[0417] One skilled in the art will recognize that modifications to the guide RNA or crRNA will allow binding of the adapter + functional domain, but proper positioning of the adapter + functional domain (e.g., steric hindrance within the three-dimensional structure of the CRISPR-Cas13b complex) You will understand that anything that cannot be done (due to One or more of the modified guide RNAs or crRNAs can be modified by introduction of individual RNA sequences 5' to the direct repeat, within the direct repeat, or 3' to the guide sequence.
[0418] The modified guide RNA or crRNA, the inactivating RNA targeting enzyme (with or without a functional domain), and the binding protein with one or more functional domains are each individually included in the composition and delivered to the host individually or Can be administered together. Alternatively, these components may be provided in a single composition for administration to the host. Administration to a host can be performed using viral vectors (eg, lentiviral vectors, adenoviral vectors, AAV vectors) known to those of skill in the art or described herein for delivery to a host. Different selectable markers (e.g., for lentiviral gRNA or crRNA selection) and gRNA or crRNA concentrations (e.g., depending on whether multiple gRNAs or crRNAs are used) are used as described herein. Using can be advantageous to produce an enhanced effect.
[0419] One skilled in the art can use the provided compositions to advantageously and specifically target single or multiple loci at the same or different functional domains to trigger one or more genomic events. The compositions are useful for screening in libraries of cells and in vivo functional modeling (e.g., identification of gene activation and function of lincRNA; gain-of-function modeling; loss-of-function modeling; cell lines and transgenic animals for optimization and screening purposes). The use of the composition of the present invention to establish a) can be applied in a variety of ways.
[0420] The invention encompasses the use of the compositions of the invention to establish and exploit conditional or inducible CRISPR Cas13b RNA targeting events. (For example, Platt et al., Cell (2014), http: / / dx.doi.org / 10.1016 / j.cell.2014.09.014 or PCT patent publications cited herein, such as WO 2014 / See 093622 (PCT / U.S. Patent Application Publication No. 2013 / 074667) (which are not considered to precede the present invention or this application). For example, the target cell conditionally or inducibly contains the RNA targeting CRISRP enzyme (e.g., in the form of a Cre-dependent construct) and / or the adapter protein conditionally or inducibly contains and is introduced into the target cell. Upon expression of the vector, the vector is expressed to induce or condition the RNA targeting enzyme and / or adapter expression in the target cell. Inducible gene expression influenced by functional domains is also an aspect of the invention by applying the teachings and compositions of the invention in conjunction with known methods of making CRISPR complexes. Alternatively, provision of the adapter protein as a conditional or inducible element together with a conditional or inducible RNA targeting enzyme may provide an effective model for screening purposes, which is advantageous for a variety of applications. and require only minimal design and administration of specific gRNAs.
[0421] Guide RNA according to the present invention containing a dead guide sequence In one aspect, the invention allows for successful formation of the CRISPR-Cas13b complex and binding to the target while allowing or not allowing successful nuclease activity (i.e. no nuclease activity / no indel activity). A guide sequence is provided that is modified to be any of For purposes of explanation, such modified guide sequences are referred to as "dead guides" or "dead guide sequences." These dead guides or dead guide sequences can be considered catalytically inactive or conformationally inactive in terms of nuclease activity. Indeed, dead guide sequences may not fully participate in productive base-pairing in their ability to promote catalytic activity or to discriminate between on-target and off-target binding activity. Briefly, the assay involves synthesizing CRISPR target RNAs and guide RNAs containing mismatches to the target RNAs, combining them with RNA targeting enzymes, and gel-based gels based on the presence of bands produced by the cleavage products. and quantifying cleavage based on relative band intensities.
[0422] Accordingly, in a related aspect, the present invention provides a non-naturally occurring or engineered composition comprising a functional RNA targeting enzyme as described herein and a guide RNA (gRNA) or crRNA targeting CRISPR-Cas RNA. A system is provided wherein the gRNA or crRNA contains a dead guide sequence whereby the gRNA is activated by the RNA-targeting CRISPR-Cas system intracellularly without detectable RNA-cleaving activity of the non-mutated RNA-targeting enzymes of the system. will have the ability to hybridize to a target sequence that directs it to the genomic locus of interest. It should be understood that any of the gRNAs or crRNAs of the present invention as described elsewhere herein can be used as crRNAs containing dead gRNAs / dead guide sequences.
[0423] The ability of dead guide sequences to direct sequence-specific binding of CRISPR complexes to RNA target sequences can be assessed by any suitable assay. For example, a component of the CRISPR Cas13b system sufficient to form a CRISPR Cas13b complex encodes the components of that system to host cells with corresponding target sequences, including dead guide sequences to be tested. can be provided, such as by transfection of a vector, and then assessed for preferential cleavage within the target sequence.
[0424] As explained further herein, it is possible to drive the appropriate framework to such dead guides by several structural parameters. A dead guide sequence can typically be shorter than each guide sequence that provides for active RNA cleavage. In detailed embodiments, the dead guides are 5%, 10%, 20%, 30%, 40%, 50% shorter than the respective guides for the same.
[0425] As explained below and known in the art, one aspect of gRNA or crRNA--RNA targeting specificity--is a direct repeat sequence, which should be appropriately linked to such guides. Specifically, this implies that the direct repeat sequence is designed according to the origin of the RNA targeting enzyme. Structural data available for validated dead-guide sequences can be used to design Cas13b-specific equivalents. For example, structural similarities between the orthologous nuclease domains HEPN of two or more Cas13b effector proteins can be used to transfer dead guides of equivalent design. Accordingly, the dead guides herein can be appropriately modified in length and sequence to reflect such Cas13b-specific equivalents, which allow successful formation of the CRISPR-Cas13b complex and binding to target RNA. At the same time, successful nuclease activity becomes unacceptable.
[0426] Dead guides allow the use of gRNAs or crRNAs as a means of gene targeting without nuclease activity, while providing an inducible means of activation or repression. Guide RNAs or crRNAs, including dead guides, are capable of functionally positioning elements, particularly gene effectors (e.g., activators or repressors of gene activity) in a manner that allows activation or repression of gene activity. It may be modified to further include protein adapters (eg, aptamers) as described elsewhere herein to enable One example is the uptake of aptamers, as described herein and in the state of the art. By engineering gRNAs or crRNAs containing dead guides to incorporate protein-interacting aptamers (Konermann et al., “Genome-scale transcription activation by an engineered CRISPR-Cas9 complex”, doi:10.1038 / nature14136, see incorporated herein), multiple individual effector domains may be assembled. This can be modeled after natural processes.
[0427] general information In embodiments of the present invention, the terms guide sequence and guide RNA and crRNA are as in the aforementioned references, such as WO2014 / 093622 (PCT / US2013 / 074667). Used synonymously. In general, a guide sequence is any polynucleotide sequence having sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of the CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence is about 50%, 60%, 75%, 80% when optimally aligned using a suitable alignment algorithm. , 85%, 90%, 95%, 97.5%, 99% or higher. Optimal alignment can be determined using any algorithm suitable for aligning sequences, non-limiting examples of which include Smith-Waterman algorithm, Needleman-Wunsch algorithm, algorithms based on the Burroughs-Wheeler transformation. (e.g. Burrows Wheeler Aligners), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (soap.genomics.org.cn) (available at maq.sourceforge.net) and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence is about 28, 29, 30, 35, 40, 45, 50, 75 nucleotides in length or longer. In some embodiments, the guide sequence is less than or less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12 nucleotides in length. Preferably, the guide sequence is 10-30 nucleotides long, eg 30 nucleotides long. The ability of a guide sequence to direct sequence-specific binding of a CRISPR complex to a target sequence can be assessed by any suitable assay. For example, enough components of the CRISPR system to form a CRISPR complex have a corresponding target sequence, such as by transfection of vectors encoding components of the CRISPR sequence, including the guide sequence to be tested. It can be provided to a host cell and then assessed for preferential cleavage within the target sequence, such as by a Surveyor assay as described herein. Similarly, cleavage of a target polynucleotide sequence provides a target sequence, a component of a CRISPR complex including the guide sequence to be tested, and a control guide sequence that differs from the test guide sequence. can be determined in vitro by comparing the rate of binding or cleavage at the target sequence between reactions with a control guide sequence. Other assays are possible and will occur to those skilled in the art. Guide sequences can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the genome of the cell. Exemplary target sequences include those that are unique in the target genome.
[0428] The term "vector" generally and throughout this specification refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors include, but are not limited to, single-stranded, double-stranded or partially double-stranded nucleic acid molecules; nucleic acid molecules without free ends (e.g., circular) containing one or more free ends; DNA, RNA; or nucleic acid molecules containing both; and other types of polynucleotides known in the art. Certain vectors are "plasmids," which refer to circular double-stranded DNA loops into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, wherein the vector includes a vector for packaging into a virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses and adeno-associated viruses). Viral DNA or RNA sequences are present. A viral vector also includes a polynucleotide carried by the virus for transfection into a host cell. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (eg, bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (eg, non-episomal mammalian vectors) integrate into the genome of the host cell upon introduction into the host cell, thereby replicating along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors". Vectors for and effecting expression in eukaryotic cells may be referred to herein as "eukaryotic expression vectors." Common expression vectors of utility in recombinant DNA techniques are often in the form of plasmids.
[0429] A recombinant expression vector can contain a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, i.e., the recombinant expression vector is operably linked to the nucleic acid sequence to be expressed. It is meant to contain one or more regulatory elements, which can be selected based on the host cell used for expression. Within the scope of recombinant expression vectors, "operably linked" refers to sequences capable of expression (e.g., in an in vitro transcription / translation system or in a host cell if the vector is introduced into the host cell). is intended to mean that the nucleotide sequence of interest is linked to the regulatory element in any manner.
[0430] The term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES) and other expression control elements (eg, polyadenylation signals and transcription termination signals such as poly U sequences). Such regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of nucleotide sequences in many types of host cells and those that direct expression of nucleotide sequences only in particular host cells (eg, tissue-specific regulatory sequences). A tissue-specific promoter may direct expression primarily in a desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g. liver, pancreas) or specific cell types (e.g. lymphocytes). . Regulatory elements may also direct expression in a time-dependent manner, such as a cell cycle-dependent or developmental stage-dependent manner, and this expression may or may not also be tissue- or cell-type specific. In some embodiments, the vector comprises one or more pol III promoters (eg, 1, 2, 3, 4, 5 or more pol III promoters), one or more pol II promoters (eg, 1 , 2, 3, 4, 5 or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5 or more pol I promoters), or combinations thereof including. Examples of pol III promoters include, but are not limited to, U6 and H1 promoters. Examples of pol II promoters include, but are not limited to, the retroviral rous sarcoma virus (RSV) LTR promoter (optionally with a RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer) [e.g. et al, Cell, 41:521-530 (1985)], the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter and the EF1α promoter. The term "regulatory element" also includes enhancer elements such as WPRE; CMV enhancer; R-U5' segment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. USA., Vol.78(3), p.1527-31, 1981); SV40 enhancer; and intron sequence between exons 2 and 3 of rabbit β-globin is also included. Those skilled in the art will appreciate that the design of the expression vector may depend on factors such as the choice of host cell to be transformed, the level of expression desired, and the like. The vector can be introduced into a host cell such that transcripts, proteins or peptides encoded by the nucleic acids as described herein are produced as fusion proteins or peptides (e.g. clustered at regular intervals). short palindromic repeat (CRISPR) transcripts, proteins, enzymes, mutant forms thereof, fusion proteins thereof, etc.).
[0431] Advantageous vectors include lentiviruses and adeno-associated viruses, and such vector types can also be selected to target particular types of cells.
[0432] As used herein, a "crRNA" or "guide RNA" or "single guide RNA" or "sgRNA" or "one or more nucleic acid components" of a Type VI CRISPR-Cas locus effector protein The term includes any polynucleotide sequence having sufficient complementarity with a target nucleic acid sequence to hybridize with the target nucleic acid sequence and direct sequence-specific binding of the RNA targeting complex to the target RNA sequence. .
[0433] In certain embodiments, a CRISPR system as provided herein can utilize crRNA or similar polynucleotides comprising guide sequences, wherein the polynucleotides are RNA, DNA, or RNA and DNA. and / or the polynucleotide comprises one or more nucleotide analogues. This sequence may comprise any structure, including but not limited to those of native crRNA, such as bulge, hairpin or stem-loop structures. In certain embodiments, a polynucleotide comprising a guide sequence forms a duplex with a second polynucleotide sequence, which can be an RNA or DNA sequence.
[0434] In certain embodiments, the guide of the invention comprises non-naturally occurring nucleic acids and / or non-naturally occurring nucleotides and / or nucleotide analogs and / or chemical modifications. A non-naturally occurring nucleic acid can include, for example, a mixture of naturally occurring and non-naturally occurring nucleotides. Non-naturally occurring nucleotides and / or nucleotide analogs may be modified at the ribose, phosphate and / or base moieties. In some embodiments of the invention, the guide nucleic acid comprises ribonucleotides and non-ribonucleotides. In one such embodiment, the guide comprises one or more ribonucleotides and one or more deoxyribonucleotides. In certain embodiments of the invention, the guide is a nucleotide with a phosphorothioate linkage, a boranophosphate linkage, a locked nucleic acid (LNA) nucleotide containing a methylene bridge between the 2' and 4' carbons of the ribose ring, or a bridged nucleic acid (BNA), etc. , contains one or more non-naturally occurring nucleotides or nucleotide analogues. Other examples of modified nucleotides include 2'-O-methyl analogs, 2'-deoxy analogs, 2-thiouridine analogs, N6-methyladenosine analogs or 2'-fluoro analogs. Further examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine (Ψ), N 1 -methylpseudouridine (me 1 Ψ), 5-methoxyuridine (5moU), inosine, 7-methylguanosine. Examples of chemical modifications of the guide RNA include, without limitation, 2'-O-methyl (M), 2'-O-methyl 3' phosphorothioate (MS), S-constrained ethyl ( cEt) or 2'-O-methyl 3'thio PACE (MSP) incorporation. Such chemically modified guide RNAs can contain increased stability and increased activity when compared to unmodified guide RNAs, however on-target versus off-target specificity is unpredictable. (Hendel, 2015, Nat Biotechnol. 33(9): 985-9, doi: 10.1038 / nbt. 3290, published online 29 June 2015; Allerson et al., J. Med. Chem. 2005, 48: 901-904; Bramsen et al., Front.Genet., 2012, 3:154; Deng et al., PNAS, 2015, 112: 11870-11875; Sharma et al., MedChemComm., 2014, 5: 1454-1471; Li et al. ., Nature Biomedical Engineering, 2017, 1,0066 DOI:10.1038 / s41551-017-0066).
[0435] In some embodiments, the 5' and / or 3' ends of the guide RNA are modified with various functional moieties, including fluorescent dyes, polyethylene glycol, cholesterol, proteins or detection tags (Kelly et al., 2016, J. Biotech. 233:74-83). In certain embodiments, the guide comprises ribonucleotides in the region that binds the target DNA and one or more deoxyribonucleotides and / or nucleotide analogues in the region that binds Cas9, Cpf1 or C2c1. In certain embodiments of the invention, deoxyribonucleotides and / or nucleotide analogues are incorporated into engineered guide structures including, without limitation, 5' and / or 3' termini, stem-loop regions and seed regions. In certain embodiments, the modification is not in the 5'-handle of the stem-loop region. Chemical alterations at the 5'-handle of the stem-loop region of the guide can abolish its function (see Li, et al., Nature Biomedical Engineering, 2017, 1:0066). In certain embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 of the guides , 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or 75 nucleotides are chemically modified. In some embodiments, 3-5 nucleotides on either the 3' or 5' end of the guide are chemically modified. In some embodiments, only minor modifications are introduced into the seed region, such as 2'-F modifications. In some embodiments, a 2'-F modification is introduced at the 3' end of the guide. In certain embodiments, the 3-5 nucleotides at the 5' and / or 3' end of the guide are 2'-O-methyl (M), 2'-O-methyl-3'-phosphorothioate (MS), S-constrained It is chemically modified with ethyl (cEt) or 2'-O-methyl-3'-thioPACE (MSP). Such modifications can enhance genome editing efficiency (see Hendel et al., Nat. Biotechnol. (2015) 33(9):985-989). In certain embodiments, all phosphodiester bonds of the guide are replaced with phosphorothioate (PS) to enhance the level of gene disruption. In certain embodiments, more than 5 nucleotides at the 5' and / or 3' ends of the guide are chemically modified with 2'-O-Me, 2'-F or S-constrained ethyl (cEt). Such chemically modified guides can mediate enhanced levels of gene disruption (see Ragdarm et al., 0215, PNAS, E7110-E7111). In some embodiments of the invention, the guide is modified to include chemical moieties at its 3' and / or 5' ends. Such moieties include, but are not limited to, amines, azides, alkynes, thios, dibenzocyclooctynes (DBCO) or rhodamines. In certain embodiments, chemical moieties are conjugated to guides by linkers such as alkyl chains. In certain embodiments, modified chemical moieties of the guide can be used to attach the guide to another molecule such as DNA, RNA, protein or nanoparticle. Such chemically modified guides can be used to identify or enrich cells that have been gene-edited by the CRISPR system (see Lee et al., eLife, 2017, 6:e25312, DOI:10.7554). .
[0436] In one aspect of the invention, the guide comprises a modified crRNA against Cpf1 with a 5'-handle and a guide segment further comprising a seed region and a 3' end. In some embodiments, the modified guide is Acidaminococcus sp. BV3L6 Cpf1 (AsCpf1); Francisella tularensis subsp. Novicida U112 Cpf1 (FnCpf1); L.bacterium) MC2017 Cpf1 (Lb3Cpf1); Butyrivibrio proteoclasticus Cpf1 (BpCpf1); Parcubacteria bacterium GWC2011_GWC2_44_17 Cpf1 (PbCpf1); bacterium) GW2011_GWA_33_10 Cpf1 (PeCpf1) Leptospira inadai Cpf1 (LiCpf1); Smithella sp. SC_K08D17 Cpf1 (SsCpf1); L.bacterium MA2020 Cpf1 (Lb2Cpf1); Porphyromonas crevioricanis Cpf1 (PcCpf1); Porphyromonas macacae Cpf1 (PmCpf1); Candidatus Methanoplasma termitum Cpf1 (CMtCpf1); Eubacterium eligens Cpf1 (EeCpf1); Moraxella bovoculi 237 Cpf1 (MbCpf1); Prevotella disiens Cpf1 (PdCpf1); or L. bacterium ND2006 Cpf1 (LbCpf1) can.
[0437] In some embodiments the modification of the guide is a chemical modification, insertion, deletion or splitting. In some embodiments, chemical modifications include, but are not limited to, 2'-O-methyl (M) analogs, 2'-deoxy analogs, 2-thiouridine analogs, N6-methyladenosine analogs, 2 '-fluoro analogs, 2-aminopurine, 5-bromo-uridine, pseudouridine (Ψ), N1-methylpseudouridine (me1Ψ), 5-methoxyuridine (5moU), inosine, 7-methylguanosine, 2'- Incorporation of O-methyl-3'-phosphorothioate (MS), S-constrained ethyl (cEt), phosphorothioate (PS) or 2'-O-methyl-3'-thioPACE (MSP) is included. In some embodiments, the guide includes one or more phosphorothioate modifications. In certain embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 25 of the guides Nucleotides are chemically modified. In certain embodiments, one or more nucleotides of the seed region are chemically modified. In certain embodiments, one or more nucleotides at the 3' terminus are chemically modified. In certain embodiments, no nucleotides of the 5'-handle are chemically modified. In some embodiments, chemical modifications of the seed region are minor modifications, such as incorporation of 2'-fluoro analogues. In a specific embodiment, one nucleotide of the seed region is replaced with a 2'-fluoro analogue. In some embodiments, the 3' terminal 5 or 10 nucleotides are chemically modified. Such chemical modification of the 3' end of Cpf1 CrRNA improves gene cleavage efficiency (see Li, et al., Nature Biomedical Engineering, 2017, 1:0066). In a specific embodiment, the 3' terminal 5 nucleotides are replaced with 2'-fluoro analogs. In a specific embodiment, the 3' terminal 10 nucleotides are replaced with 2'-fluoro analogs. In a specific embodiment, the 3'-terminal 5 nucleotides are replaced with 2'-O-methyl (M) analogues.
[0438] In some embodiments, the 5'-handle loop of the guide is modified. In some embodiments, the 5'-handle loop of the guide is modified to have a deletion, insertion, split or chemical modification. In certain embodiments, the loop comprises 3, 4 or 5 nucleotides. In certain embodiments, the loop comprises a sequence of UCUU, UUUU, UAUU or UGUU.
[0439] In one aspect, the guide comprises moieties that are chemically linked or conjugated via non-phosphodiester bonds. In one embodiment, the guide comprises, in non-limiting examples, tracr sequences and tracr mate sequence portions or direct repeats and targeting sequences chemically linked or conjugated via non-nucleotide loops. In some embodiments, these moieties are joined via non-phosphodiester covalent linkages. Examples of covalent conjugates include, but are not limited to, carbamates, ethers, esters, amides, imines, amidines, aminothridines, hydrozones, disulfides, thioethers, thioesters, phosphorothioates, phospho rhodithioates, sulfonamides, sulfonates, sulphones, sulfoxides, ureas, thioureas, hydrazides, oximes, triazoles, photolabile linkages, Diels-Alder cycloaddition pairs or ring-closing metathesis pairs, etc. Chemical moieties selected from the group consisting of C-C bond forming groups and Michael reactive pairs are included.
[0440] In some embodiments, the guide moieties are first synthesized using standard phosphoramidite synthesis protocols (Herdewijn, P., ed., Methods in Molecular Biology Col 288, “Oligonucleotide Synthesis: Methods and Applications”). , Humana Press, New Jersey (2012)). In some embodiments, the non-targeting guide moiety can be functionalized using standard protocols known in the art to contain functional groups suitable for ligation (Hermanson, G.T., Bioconjugate Techniques, Academic Press (2013). )). Examples of functional groups include, but are not limited to, hydroxyl, amine, carboxylic acid, carboxylic acid halide, carboxylic acid active ester, aldehyde, carbonyl, chlorocarbonyl, imidazolylcarbonyl, hydrozide, semicarbazide, thiosemicarbazide, thiol, maleimide, haloalkyl , sphonyl, ant, propargyl, diene, alkyne and azide. When the non-targeting portion of the guide is functionalized, it can form a covalent chemical bond or linkage between the two oligonucleotides. Examples of chemical bonds include, but are not limited to, carbamates, ethers, esters, amides, imines, amidines, aminothridines, hydrozones, disulfides, thioethers, thioesters, phosphorothioates, phosphorodithio C-C bonds such as ates, sulfonamides, sulfonates, sulphones, sulfoxides, ureas, thioureas, hydrazides, oximes, triazoles, photolabile linkages, Diels-Alder cycloaddition pairs or ring-closing metathesis pairs Included are those based on forming groups and Michael reaction pairs.
[0441] In some embodiments, one or more portions of the guide can be chemically synthesized. In some embodiments, chemical synthesis is performed using 2'-acetoxyethyl orthoester (2'-ACE) (Scaringe et al., J.Am.Chem.Soc. (1998) 120:11820-11821; Scaringe, Methods). Enzymol. (2000) 317:3-18) or 2'-thionocarbamate (2'-TC) chemistry (Dellinger et al., J.Am.Chem.Soc. (2011) 133:11540-11546; Hendel et al. al., Nat. Biotechnol. (2015) 33:985-989) is used.
[0442] In some embodiments, guide moieties are covalently linked using a variety of bioconjugation reactions, loops, crosslinks and sugar modifications, internucleotide phosphodiester linkages, non-nucleotidic linkages through purine and pyrimidine residues. can be concatenated. Sletten et al., Angew. Chem. Int. Ed. (2009) 48:6974-6998; Manoharan, M. Curr. Opin. Chem. Biol. 2008) 18:305-19; Watts, et al., Drug. Discov. Today (2008) 13:842-55; Shukla, et al., ChemMedChem (2010) 5:328-49.
[0443] In some embodiments, guide moieties can be covalently linked using click chemistry. In some embodiments, guide moieties can be covalently linked using a triazole linker. In some embodiments, guide moieties can be covalently linked using Huisgen 1,3-dipolar cycloaddition reactions involving alkynes and azides that yield highly stable triazole linkers (He et al., ChemBioChem (2015) 17:1809-1812; WO2016 / 186745). In some embodiments, the guide moiety can be covalently linked by ligation of the 5'-hexyne and 3'-azido moieties. In some embodiments, either or both of the 5'-hexyne guide moiety and the 3'-azido guide moiety can be protected with a 2'-acetoxyethyl orthoester (2'-ACE) group, which is It can then be removed using the Dharmacon protocol (Scaringe et al., J. Am. Chem. Soc. (1998) 120:11820-11821; Scaringe, Methods Enzymol. (2000) 317:3-18 ).
[0444] In some embodiments, guide moieties include linkers (e.g., non-nucleotide loops) that include moieties such as spacers, attachments, bioconjugates, chromophores, reporter groups, dye-labeled RNA, and non-naturally occurring nucleotide analogs. can be covalently linked via More specifically, suitable spacers for the purposes of the present invention include, but are not limited to, polyethers (such as polyethylene glycols, polyhydric alcohols, polypropylene glycol or mixtures of ethylene and propylene glycols), polyamines Groups (eg, spennine, spermidine and its polymeric derivatives), polyesters (eg, poly(ethyl acrylate)), polyphosphodiesters, alkylenes and combinations thereof. Suitable attachments include, but are not limited to, any moiety that can be added to the linker to impart additional properties to the linker, such as a fluorescent label. Suitable bioconjugates include, but are not limited to, peptides, glycosides, lipids, cholesterol, phospholipids, diacylglycerols and dialkylglycerols, fatty acids, carbohydrates, enzyme substrates, steroids, biotin, digoxigenin, carbohydrates, polysaccharides. mentioned. Suitable chromophores, reporter groups and dye-labeled RNAs include, but are not limited to, fluorescent dyes such as fluorescein and rhodamine, chemiluminescent, electrochemiluminescent and bioluminescent marker compounds. The design of exemplary linkers that conjugate two RNA components is also described in WO2004 / 015075.
[0445] A linker (eg, a non-nucleotide loop) can be of any length. In some embodiments, the linker has a length equal to about 0-16 nucleotides. In some embodiments, the linker has a length equal to about 0-8 nucleotides. In some embodiments, the linker has a length equal to about 0-4 nucleotides. In some embodiments, the linker has a length equal to about 2 nucleotides. Exemplary linker designs are also described in WO2011 / 008730.
[0446] In some embodiments, the degree of complementarity is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5% when optimally aligned using a suitable alignment algorithm. , greater than or equal to 99%. Optimal alignment can be determined using any algorithm suitable for aligning sequences, non-limiting examples of which include Smith-Waterman algorithm, Needleman-Wunsch algorithm, algorithms based on Burroughs-Wheeler transformation. (e.g. Burrows Wheeler Aligners), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (soap.genomics.org.cn) (available at maq.sourceforge.net) and Maq (available at maq.sourceforge.net). The ability of a guide sequence (within an RNA targeting guide RNA or crRNA) to direct sequence-specific binding of a nucleic acid targeting complex to a target nucleic acid sequence can be assessed by any suitable assay. For example, sufficient components of the RNA-targeting CRISPR Cas13b system to form a nucleic acid targeting complex are added to a host cell having the corresponding target nucleic acid sequence, including the guide sequence to be tested. can be provided, such as by transfection of a vector encoding a , followed by assessment of preferential targeting (eg, cleavage) within the target nucleic acid sequence, such as by Surveyor assays as described herein. Similarly, cleavage of a target nucleic acid sequence provides a target nucleic acid sequence, a component of a nucleic acid targeting complex that includes the guide sequence to be tested, and a control guide sequence that differs from the test guide sequence, and their test g...
Claims
**Claim 1** i) A Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A, optionally containing one or more mutations and having at least 90% sequence identity with the Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A; and ii) crRNA A non-naturally occurring or engineered composition comprising The crRNA comprises a) a guide sequence capable of hybridizing to a target RNA sequence and b) a direct repeat sequence, and can form a CRISPR complex comprising the Cas13b effector protein complexed with the guide sequence capable of hybridizing to the target RNA sequence, Optionally, it contains an accessory protein that enhances Cas13b effector protein activity, preferably, the accessory protein that enhances Cas13b effector protein activity is the csx28 protein, or it contains an accessory protein that suppresses Cas13b effector protein activity, preferably, the accessory protein that suppresses Cas13b effector protein activity is the csx27 protein, A non-naturally occurring or engineered composition. **Claim 2** The Cas13b effector protein is associated with one or more functional domains, optionally, the functional domain cleaves the target RNA sequence or the functional domain modifies the translation of the target RNA sequence, the non-naturally occurring or engineered composition according to claim 1. **Claim 3** The Cas13b effector protein is associated with one or more functional domains, and the effector protein contains one or more mutations within the HEPN domain, whereby the complex can deliver an epigenetic modifier or a translational activation or repression signal, the composition according to claim 1. **Claim 4** A Cas13b vector system for providing the composition according to claim 1, A first regulatory element operably linked to a nucleotide sequence encoding a Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A, optionally including one or more mutations, having at least 90% sequence identity with the Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A, and capable of forming a CRISPR complex comprising the Cas13b effector protein complexed with the guide sequence capable of hybridizing to the target RNA sequence; and A second regulatory element operably linked to a nucleotide sequence encoding crRNA, Comprising one or more vectors, Optionally, the nucleotide sequence encoding the Cas13b effector protein is codon-optimized for expression in eukaryotic cells, Cas13b vector system. **Claim 5** The Cas13b vector system according to claim 4, further comprising a regulatory element operably linked to a nucleotide sequence of an accessory protein. **Claim 6** The one or more vectors include viral vectors, and optionally the one or more vectors include one or more retroviral, lentiviral, adenoviral, adeno-associated or herpes simplex viral vectors, the Cas13b vector system according to claim 4. **Claim 7** i) A Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A, optionally including one or more mutations, having at least 90% sequence identity with the Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A; and ii) crRNA A non-naturally occurring or engineered composition comprising, The crRNA comprises: a) a guide sequence having the ability to hybridize to a target RNA sequence in a cell; and b) a direct repeat sequence. The Cas13b effector protein can form a complex with the crRNA. The guide sequence can direct sequence-specific binding to the target RNA sequence, and a CRISPR complex can be formed that includes the Cas13b effector protein complexed with the guide sequence that can hybridize to the target RNA sequence. A delivery system configured to deliver the Cas13b effector protein and one or more nucleic acid components of the non-naturally occurring or engineered composition. **Claim 8** The delivery system according to claim 7, comprising one or more vectors or one or more polynucleotide molecules, wherein the one or more vectors or polynucleotide molecules comprise one or more polynucleotide molecules encoding the Cas13b effector protein and one or more nucleic acid components of the non-naturally occurring or engineered composition. **Claim 9** The delivery system according to claim 7, comprising a delivery medium comprising liposomes, particles, exosomes, microvesicles, gene guns, or one or more viral vectors. **Claim 10** A non-naturally occurring or engineered composition according to any one of claims 1 to 3, a vector system according to any one of claims 4 to 6, or a delivery system according to any one of claims 7 to 9 for use in a therapeutic treatment method. **Claim 11** An in vitro or ex vivo method of modifying the expression of a target gene of interest, comprising contacting the target RNA with one or more non-naturally occurring or engineered compositions. The one or more non-naturally occurring or engineered compositions comprise: i) a Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A, optionally comprising one or more mutations and having at least 90% sequence identity with the Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A; and ii) crRNA and The crRNA comprises: a) a guide sequence capable of hybridizing to a target RNA sequence in a cell; and b) a direct repeat sequence. The Cas13b effector protein can form a complex with the crRNA. The guide sequence can direct sequence-specific binding to the target RNA sequence in the cell. Thereby, a CRISPR complex is formed that includes the Cas13b effector protein complexed with the guide sequence capable of hybridizing to the target RNA sequence, whereby the expression of the target locus of interest is modified. Optionally, the method further comprises contacting the target RNA with an accessory protein that enhances Cas13b effector protein activity, preferably the accessory protein that enhances Cas13b effector protein activity is the csx28 protein, or the method further comprises contacting the target RNA with an accessory protein that suppresses Cas13b effector protein activity, preferably the accessory protein that suppresses Cas13b effector protein activity is the csx27 protein. The method is not a method for modifying the genetic identity of the human germline. Method. Claim 12 The modifying the expression of the target gene includes cleaving the target RNA or modifying the splicing, transport, localization, translation, turnover of the target RNA, according to the in vitro or ex vivo method of claim 11. Claim 13 An isolated nucleic acid encoding a Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A, optionally comprising one or more mutations, having at least 90% sequence identity with the Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A, and capable of forming a CRISPR complex comprising the Cas13b effector protein complexed with a guide sequence capable of hybridizing to a target RNA sequence, the isolated nucleic acid being DNA and further comprising a sequence encoding crRNA. Claim 14 A Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A, optionally containing one or more mutations, having at least 90% sequence identity with the Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A, and capable of forming a CRISPR complex complexed with a guide sequence having the ability to hybridize to a target RNA sequence, or an isolated eukaryotic cell comprising a nucleic acid encoding said Cas13b effector protein. **Claim 15** i) An mRNA encoding a Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A, capable of forming a CRISPR complex comprising the Cas13b effector protein complexed with the guide sequence capable of hybridizing to a target RNA sequence, optionally containing one or more mutations, and having at least 90% sequence identity with the Cas13b effector protein derived from Bacteroidetes bacterium GWA2#31#9 of Table 1A; and ii) crRNA, wherein said crRNA comprises a) a guide sequence having the ability to hybridize to a target RNA sequence and b) a direct repeat sequence. A non-naturally occurring or engineered composition comprising the same. **Claim 16** The non-naturally occurring or engineered composition according to any one of claims 1 to 3, the Cas13b vector system according to any one of claims 4 to 6, the delivery system according to any one of claims 7 to 9, the in vitro or ex vivo method according to claim 11 or 12, or the non-naturally occurring or engineered composition according to claim 15, wherein the target RNA is a pathogen-derived target, preferably viral RNA.