CRISPR-Cas13 crRNA array
A tandem array of CRISPR RNAs with unique guide sequences and DR mutations addresses the synthesis challenges of crRNA arrays, enabling effective viral RNA targeting and degradation, thereby treating viral infections.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- UNIVERSITY OF ROCHESTER
- Filing Date
- 2021-03-26
- Publication Date
- 2026-05-01
AI Technical Summary
The production of crRNA tandem arrays is hindered by the size and repeatability of direct repeat (DR) sequences, which complicates vector size requirements and synthesis.
A tandem array comprising at least two CRISPR RNAs (crRNAs) with unique guide sequences and DR sequences, including nucleotide mutations, specifically designed to target coronavirus or influenza virus sequences, is developed, along with a Cas protein for targeted RNA degradation.
The solution enables efficient reduction of multiple target RNAs and provides a method for treating viral infections by administering the crRNA array and Cas protein, effectively targeting and degrading viral RNA sequences.
Smart Images

Figure 0007854190000033 
Figure 0007854190000034 
Figure 0007854190000035
Abstract
Description
Background Art
[0001] Cross - reference to related applications This application claims priority to U.S. Provisional Application No. 63 / 000,757, filed Mar. 27, 2020, which is hereby incorporated by reference in its entirety.
[0002] Generally, mammalian guide RNA expression cassettes are produced by cloning annealed oligonucleotides containing a guide sequence into a cassette consisting of a mammalian Pol III promoter, direct repeats, and a terminator of six or more Ts. Usually, multiple guide RNAs are expressed by adding a Pol III promoter cassette, which can significantly increase the complexity and size of the vector. The production of crRNA tandem arrays significantly relaxes the vector size requirements, but the nucleotide synthesis of long arrays is hindered by the size and repeatability of the DR sequences.
Summary of the Invention
[0003] In one embodiment, the present disclosure provides a tandem array comprising at least two CRISPR RNAs (crRNAs), wherein each crRNA comprises a guide sequence and a direct repeat (DR) sequence.
[0004] In one embodiment, each guide sequence in the tandem array is different. In one embodiment, each DR sequence in the tandem array is different. In one embodiment, each DR sequence comprises a nucleotide mutation within the loop region of the DR sequence.
[0005] In one embodiment, each DR sequence comprises a T17 or T18 nucleotide mutation. In one embodiment, each crRNA comprises a DR sequence independently selected from SEQ ID NOs: 265 - 274. In one embodiment, each DR sequence is present 3' to the guide sequence.
[0006] In one embodiment, each guide sequence is independently substantially complementary to a coronavirus genome mRNA sequence or a coronavirus subgenome mRNA sequence. In one embodiment, each guide sequence is substantially complementary to a coronavirus leader sequence, a coronavirus S sequence, a coronavirus E sequence, a coronavirus M sequence, a coronavirus N sequence, or a coronavirus S2M sequence. In one embodiment, each guide sequence is independently substantially complementary to a sequence that is at least 80% homologous to a sequence selected from SEQ ID NOs. 168-174, 176-181, 186, and 187. In one embodiment, each guide sequence includes a sequence that is at least 80% homologous to a sequence selected from SEQ ID NOs. 189-224. In one embodiment, the tandem array includes a sequence that is at least 80% homologous to SEQ ID NO. 275.
[0007] In one embodiment, each guide sequence is independently substantially complementary to an influenza virus genomic RNA sequence or an influenza virus subgenomic RNA sequence. In one embodiment, each guide sequence is substantially complementary to an influenza virus PB2 sequence, an influenza virus PB1 sequence, an influenza virus PA sequence, an influenza virus NP sequence, or an influenza virus M sequence. In one embodiment, each guide sequence is independently substantially complementary to a sequence that is at least 80% homologous to a sequence selected from SEQ ID NOs. 225-244. In one embodiment, each guide sequence includes a sequence that is at least 80% homologous to a sequence selected from SEQ ID NOs. 245-249. In one embodiment, the tandem array includes a sequence that is at least 80% homologous to a sequence selected from SEQ ID NOs. 276-279.
[0008] In one embodiment, the Disclosure provides a composition comprising a tandem array of the Disclosure. In one embodiment, the composition further comprises a Cas protein. In one embodiment, the Cas protein is Cas13. In one embodiment, the Cas protein comprises a sequence that is at least 80% identical to a sequence selected from SEQ ID NOs: 1 to 47. In one embodiment, the Cas protein further comprises a localization signal or a transport signal. In one embodiment, the Cas protein comprises a NES, wherein the NES comprises a sequence that is at least 80% identical to SEQ ID NOs: 58 to 59. In one embodiment, the Cas protein comprises a nuclear localization signal (NLS), wherein the NLS comprises a sequence that is at least 80% identical to SEQ ID NOs: 50 to 57 and 323 to 935. In one embodiment, the Cas protein comprises a localization signal, wherein the localization signal comprises a sequence that is at least 80% identical to SEQ ID NOs: 60 to 66.
[0009] In one embodiment, the disclosure provides a method for constructing a tandem array comprising at least two CRISPR RNAs (crRNAs), wherein each crRNA comprises a guide sequence and a direct repeat (DR) sequence. In one embodiment, the method comprises ligating at least two crRNA sequences, wherein each crRNA sequence comprises a unique DR sequence, and the ligation generates a tandem array.
[0010] In one embodiment, the Disclosure provides a method for reducing the number of two or more unique target RNAs in a subject. In one embodiment, the Method comprises administering the tandem array of the Disclosure and a Cas protein or nucleic acid encoding a Cas protein to the subject, or administering the composition of the Disclosure. In one embodiment, the tandem array comprises at least two unique crRNA sequences substantially complementary to the target RNA sequence.
[0011] In one embodiment, the present disclosure provides a method for treating a viral infection. In one embodiment, the method comprises administering to: (a) a tandem array comprising at least two CRISPR RNAs (crRNAs), each crRNA comprising a guide sequence and a direct repeat (DR) sequence substantially complementary to a viral RNA sequence; and (b) a Cas protein or a nucleic acid encoding a Cas protein.
[0012] In one embodiment, the Cas protein is Cas13. In one embodiment, the Cas protein contains a sequence that is at least 80% identical to a sequence selected from SEQ ID NOs: 1 to 47. In one embodiment, the Cas protein further contains a localization signal or a transport signal. In one embodiment, the Cas protein contains a NES, wherein the NES contains a sequence that is at least 80% identical to SEQ ID NOs: 58 to 59. In one embodiment, the Cas protein contains a nuclear localization signal (NLS), wherein the NLS contains a sequence that is at least 80% identical to SEQ ID NOs: 50 to 57 and 323 to 935. In one embodiment, the Cas protein contains a localization signal, wherein the localization signal contains a sequence that is at least 80% identical to SEQ ID NOs: 60 to 66. In one embodiment, the Cas protein contains a sequence that is at least 80% identical to SEQ ID NOs: 68 to 100.
[0013] In one embodiment, the nucleic acid encoding the Cas protein contains a sequence that is at least 80% identical to sequence numbers 132-133. In another embodiment, the nucleic acid encoding the Cas protein contains a sequence that is at least 80% identical to sequence numbers 147-166.
[0014] In one embodiment, the viral infection is a coronavirus infection, where each crRNA independently includes a guide sequence substantially complementary to a coronavirus genome mRNA sequence or a coronavirus subgenome mRNA sequence.
[0015] In one embodiment, each crRNA independently contains a guide sequence substantially complementary to a coronavirus leader sequence, a coronavirus S sequence, a coronavirus E sequence, a coronavirus M sequence, an N sequence, or a coronavirus S2M sequence. In one embodiment, each crRNA independently contains a guide sequence substantially complementary to a sequence at least 80% homologous to a sequence selected from SEQ ID NOs. 168-174, 176-181, 186, and 187. In one embodiment, each crRNA independently contains a guide sequence at least 80% homologous to a sequence selected from SEQ ID NOs. 189-224. In one embodiment, the tandem array contains a sequence at least 80% homologous to SEQ ID NO. 275.
[0016] In one embodiment, the viral infection is an influenza infection, where each crRNA independently includes a guide sequence substantially complementary to an influenza virus genomic RNA sequence or an influenza virus subgenomic RNA sequence. In one embodiment, each crRNA independently includes a guide sequence substantially complementary to an influenza virus PB2 sequence, an influenza virus PB1 sequence, an influenza virus PA sequence, an influenza virus NP sequence, or an influenza virus M sequence. In one embodiment, each crRNA independently includes a guide sequence substantially complementary to a sequence at least 80% homologous to a sequence selected from SEQ ID NOs. 225-244. In one embodiment, each crRNA independently includes a guide sequence at least 80% homologous to a sequence selected from SEQ ID NOs. 245-264. In one embodiment, the tandem array includes a sequence at least 80% homologous to a sequence selected from SEQ ID NOs. 276-279.
[0017] In one embodiment, the present disclosure provides a delivery system comprising a packaging plasmid, a transdermal plasmid, and an envelope plasmid, wherein the packaging plasmid comprises a nucleic acid sequence encoding a gag-pol polyprotein, the transdermal plasmid comprises a nucleic acid sequence encoding a tandem array containing at least two CRISPR RNAs (crRNAs), each crRNA comprising a nucleic acid sequence including a guide sequence and a direct repeat (DR) sequence, and a nucleic acid sequence encoding a Cas protein, and the envelope plasmid comprises a nucleic acid sequence encoding an envelope protein.
[0018] In one embodiment, the Cas protein includes a sequence that is at least 80% identical to a sequence selected from SEQ ID NOs: 1-47. In one embodiment, the Cas protein further includes a localization signal or a transport signal. In one embodiment, the localization signal or transport signal includes a sequence that is 80% identical to a sequence selected from SEQ ID NOs: 50-66 and 323-935. In one embodiment, the envelope protein is a coronavirus spike glycoprotein. In one embodiment, the envelope protein includes a sequence that is at least 80% identical to a sequence selected from SEQ ID NOs: 101-129.
[0019] In one embodiment, each crRNA independently contains a guide sequence substantially complementary to a coronavirus genome mRNA sequence or a coronavirus subgenome mRNA sequence. In one embodiment, each crRNA independently contains a guide sequence substantially complementary to a coronavirus leader sequence, coronavirus S sequence, coronavirus E sequence, coronavirus M sequence, N sequence, or coronavirus S2M sequence. In one embodiment, each crRNA independently contains a guide sequence substantially complementary to a sequence at least 80% homologous to a sequence selected from SEQ ID NOs. 168-174, 176-181, 186, and 187. In one embodiment, the crRNA independently contains a guide sequence at least 80% homologous to a sequence selected from SEQ ID NOs. 189-224. In one embodiment, the tandem array contains a sequence at least 80% homologous to SEQ ID NO. 275.
[0020] In one embodiment, each crRNA independently contains a guide sequence substantially complementary to an influenza virus genomic RNA sequence or an influenza virus subgenomic RNA sequence. In one embodiment, each crRNA independently contains a guide sequence substantially complementary to an influenza virus PB2 sequence, an influenza virus PB1 sequence, an influenza virus PA sequence, an influenza virus NP sequence, or an influenza virus M sequence. In one embodiment, each crRNA independently contains a guide sequence substantially complementary to a sequence at least 80% homologous to a sequence selected from SEQ ID NOs. 225-244. In one embodiment, each crRNA independently contains a guide sequence at least 80% homologous to a sequence selected from SEQ ID NOs. 245-264. In one embodiment, the tandem array contains a sequence at least 80% homologous to a sequence selected from SEQ ID NOs. 276-279.
[0021] The following detailed description of various embodiments of the present invention will be better understood when read in conjunction with the accompanying drawings. Exemplary embodiments are shown in the drawings to illustrate the present invention. However, it should be understood that the present invention is not limited to the exact configurations and means of the embodiments shown in the drawings. [Brief explanation of the drawing]
[0022] [Figure 1] A schematic diagram of the coronavirus genome mRNA and subgenome mRNA is shown. [Figure 2] A schematic diagram of the eraserR platform is shown. [Figure 3] A schematic diagram of delivery using a pseudotyped integrase-deficient lentiviral vector is shown. [Figure 4A] Figure 4(A-C) shows schematic diagrams of guide RNA testing, lentiviral generation, and cell targeting. Figure 4A shows a schematic diagram of the design of luciferase reporter constructs encoding 5' and 3' CoV target sequences. [Figure 4B]Figures 4 (A - C) show schematic diagrams of guide RNA testing, lentivirus production, and cell targeting. Figure 4B shows a schematic diagram demonstrating that a lentiviral construct encoding CRISPR - Cas13 components can be packaged into non - integrating lentiviral particles pseudotyped with a spike glycoprotein derived from the SARS - CoV - 2 coronavirus, which provides specificity for entry into cells expressing the viral envelope protein, e.g., the ACE2 receptor. This enables specific targeting to cell types "targeted by the coronavirus". [Figure 4C] Figures 4 (A - C) show schematic diagrams of guide RNA testing, lentivirus production, and cell targeting. Figure 4C shows a schematic diagram demonstrating that transient expression of CRISPR - Cas13 components enables rapid target degradation of CoV genomic viral mRNA and sub - genomic viral mRNA through post - transduction processing and formation of non - integrating lentiviral episomes. [Figure 5] Shows conservation and target sites of the SARS - CoV - 2 leader sequence. [Figure 6] Shows tiled SARS - CoV - 2 leader crRNA. [Figure 7A] Figures 7 (A - C) show verification of CRISPR - Cas13 guide RNA targeting the SARS - CoV - 2 leader sequence. Figure 7A shows a schematic diagram of a luciferase reporter containing the SARS - 2 - CoV leader sequence and the target site of crRNA. [Figure 7B] Figures 7 (A - C) show verification of CRISPR - Cas13 guide RNA targeting the SARS - CoV - 2 leader sequence. Figure 7B shows the sequence alignment of tiled crRNAs targeting the SARS - CoV - 2 leader sequence. The transcription regulatory sequence (TRS) is highlighted in yellow. [Figure 7C]Figures 7 (A - C) show the validation of CRISPR - Cas13 guide RNAs targeting the SARS - CoV - 2 leader sequence. Figure 7C shows a cell - based luciferase assay demonstrating potent knockdown by crRNAs targeting the SARS - CoV - 2 leader sequence (crRNAs A - G) or the sequence encoding luciferase (Luc) of CoV leader Luc reporter activity in cells, compared to non - targeting crRNAs. [Figure 8A] Figures 8 (A - C) show the validation of CRISPR - Cas13 guide RNAs targeting the SARS - CoV - 2 stem - loop - like 2 (S2M) sequence. Figure 8A shows a schematic diagram of a luciferase reporter containing the SARS - 2 - CoV S2M sequence and the target sites of crRNAs. [Figure 8B] Figures 8 (A - C) show the validation of CRISPR - Cas13 guide RNAs targeting the SARS - CoV - 2 stem - loop - like 2 (SQM) sequence. Figure 8B shows the sequence alignment of tiled crRNAs targeting the SARS - CoV - 2 S2M sequence. [Figure 8C] Figures 8 (A - C) show the validation of CRISPR - Cas13 guide RNAs targeting the SARS - CoV - 2 stem - loop - like 2 (S2M) sequence. Figure 8C shows a cell - based luciferase assay demonstrating potent knockdown by crRNAs targeting the SARS - CoV - 2 S2M sequence (crRNAs A - F) or the sequence encoding luciferase (Luc) of CoV S2M Luc reporter activity in cells, compared to non - targeting crRNAs. [Figure 9-1]Figures A-C illustrate the directed one-step assembly of a CRISPR-Cas13 crRNA array. Figure A is a schematic diagram showing the genomic structure of a bacterial CRISPR-Cas13 locus, typically consisting of a single Cas13 protein and a CRISPR array containing multiple spacer sequences and direct repeat (DR) sequences. Figure B is a schematic diagram showing that each functional CRISPR guide RNA is processed to include spacers and direct repeats. The spacer sequences are antisense to the target sequence and provide target specificity, while the DR sequences function as handles for binding to the Cas13 protein. Figure C is a schematic diagram showing that a typical mammalian crRNA expression cassette is constructed by annealing and ligating oligonucleotides containing the desired spacer sequences. Figure D is a schematic diagram showing that ordered arrays of multiple guide RNAs can be efficiently constructed by utilizing permissible nucleotide substitutions within the loop region of the DR. Figure E shows possible permissible nucleotide substitutions within the loop region of the PspCas13b DR that can be used for array assembly. [Figure 9-2]Figures D-E illustrate the directed one-step assembly of a CRISPR-Cas13 crRNA array. Figure A is a schematic diagram showing the genomic structure of a bacterial CRISPR-Cas13 locus, typically consisting of a CRISPR array containing a single Cas13 protein and multiple spacer and direct repeat (DR) sequences. Figure B is a schematic diagram showing that each functional CRISPR guide RNA is processed to include spacers and direct repeats. The spacer sequences are antisense to the target sequence and provide target specificity, while the DR sequences function as handles for binding to the Cas13 protein. Figure C is a schematic diagram showing that a typical mammalian crRNA expression cassette is constructed by annealing and ligating oligonucleotides containing the desired spacer sequences. Figure D is a schematic diagram showing that ordered arrays of multiple guide RNAs can be efficiently constructed by utilizing permissible nucleotide substitutions within the loop region of the DR. Figure E shows possible permissible nucleotide substitutions within the loop region of the PspCas13b DR that can be used for array assembly. [Figure 10] This paper demonstrates the identification and validation of non-essential loop residues in the Cas13b direct repeat (DR). A shows all possible mutations at positions T17 and T18 of the PspCas13b direct repeat. B shows a schematic diagram indicating the luciferase reporter and crRNA target sites. C shows experimental results demonstrating CRISPR-Cas13b-mediated knockdown of luciferase activity using two independent guide RNAs containing distinct DR loop mutations. [Figure 11A] Figure 11(A-C) shows targeted knockdown of a SARS-CoV-2 luciferase reporter using guide RNA sequences. Figure 11A is a schematic diagram showing a lentiviral gene transplasm encoding a CRISPR-Cas13 expression cassette that encodes either a single guide RNA or an array of triplicate guide RNAs. [Figure 11B]Figure 11(A-C) shows targeted knockdown of a SARS-CoV-2 luciferase reporter using a guide RNA sequence. Figure 11B is a schematic diagram of a luciferase reporter containing multiple SARS-CoV-2 virus sequences within the 5' and 3' UTR. [Figure 11C] Figure 11(A-C) shows targeted knockdown of the SARS-CoV-2 luciferase reporter by guide RNA sequences. Figure 11C shows experimental results demonstrating relative luciferase activity knockdown compared to negative control untargeted crRNA (NC) by expression of RNA-targeting components of CRISPR-Cas13 driven by a single (LDR-D) or triplicate guide RNA (LDR-D / NB / S2M-D) targeting the SARS-CoV-2 luciferase reporter. [Figure 12] A schematic diagram of a CRISPR-Cas13 expression cassette encoding a tripty of guide RNAs that can be packaged into an AAV viral vector is shown. [Figure 13] A schematic diagram of the influenza virus is shown. A shows a schematic diagram of influenza virus RNA (vRNA). Influenza is an enclosed negative-strand RNA virus composed of eight vRNA segments. A shows a schematic diagram of the influenza virus particle. All eight vRNAs are packaged within an outer membrane virus that utilizes the HA and NA of viral proteins for binding and fusion to the host cell. [Figure 14A] Figure 14(A and B) is a schematic diagram of the packaging and delivery of components for CRISPR-Cas13 RNA editing targeting influenza. Figure 14A is a schematic diagram showing that components for CRISPR-Cas13 editing, including a CRISPR guide RNA sequence and Cas13 protein, can be packaged into a viral gene therapy vector, such as an integrase-deficient lentiviral vector. Pseudotyped lentiviral vectors with influenza NA and HA envelope proteins is one method for delivery to host cells targeted by the influenza virus. [Figure 14B] Figure 14(A and B) is a schematic diagram of the packaging and delivery of components for CRISPR-Cas13 RNA editing targeting influenza. Figure 14A is a schematic diagram showing that the expression of CRISPR-Cas13 components leads to targeted degradation of vRNA or viral mRNA during viral vector fusion and delivery. Strong nuclear localization of the Cas13 protein may be required for targeting of vRNA. [Figure 15] This document presents experimental results demonstrating the pseudotyping of lentiviral vectors using the SARS-CoV spike envelope protein. A schematic diagram demonstrates that N and C-terminal modifications (4LV) are required for lentiviral pseudotyping using CoV spike proteins derived from SARS-CoV-1 and SARS-CoV-2. B shows experimental results demonstrating that the wild-type (WT) CoV spike protein is unsuitable for lentiviral pseudotyping for transduction of HEK293T cells or HEK293T cells expressing human ACE2 (ACE2-HEK293T). Expression of the human ACE2 receptor in HEK293T cells is necessary and sufficient for transduction with the 4LV pseudotyped lentiviral vector. Apart from ACE2 expression, the VSV-G envelope enables lentiviral pseudotyping for broad transduction of numerous cell types in vitro. [Figure 16] This paper presents experimental results demonstrating the activity of Cas13b crRNA targeting highly conserved positive and negative RNA sequences in influenza A. A shows experimental results with guide RNA targeting conserved positive RNA sequences in influenza A segments 1, 2, 3, 5, and 7. B shows experimental results with conserved negative RNA sequences in influenza A segments 1, 2, 3, 5, and 7. All crRNAs targeting influenza A showed potent knockdown effects compared to untargeted (NT) crRNAs of luciferase reporters possessing target sequences specific to the corresponding influenza A segments. [Modes for carrying out the invention]
[0023] In some embodiments, the disclosure provides a tandem array of crRNA sequences. In one embodiment, the tandem array includes one or more, two or more, three or more, four or more, five or more, six or more, seven or more, or eight or more unique crRNA sequences.
[0024] In one embodiment, the crRNA sequence of the tandem array includes a guide sequence and direct repeat (DR) sequences. In one embodiment, each DR sequence is unique. In one embodiment, the tandem array can be formed by a single ligation step.
[0025] definition Unless otherwise defined, technical and scientific terms used herein have the same meanings as those generally understood by those skilled in the art to which this invention pertains.
[0026] In general, the nomenclature and experimental procedures used herein in cell culture, molecular genetics, organic chemistry, nucleic acid chemistry, and hybridization are well known and commonly used in the art.
[0027] Standard techniques are used for nucleic acid and peptide synthesis. These techniques and procedures are generally carried out in accordance with conventional methods in the art and various general references provided throughout this specification (e.g., Sambrook and Russell, 2012, Molecular Cloning, A Laboratory Approach, Cold Spring Harbor Press, Cold Spring Harbor, NY, and Ausubel et al., 2012, Current Protocols in Molecular Biology, John Wiley & Sons, NY).
[0028] The nomenclature used herein and the experimental procedures used in analytical chemistry and organic synthesis described below are well known and commonly used in the art. Standard techniques or their modifications are used in chemical synthesis and chemical analysis.
[0029] In the context of this invention (particularly in the context of the claims), the terms “a,” “an,” and “the,” as well as similar terms, are to be interpreted as encompassing both singular and plural forms unless otherwise specifically stated herein or clearly negated by the context.
[0030] When referring to measurable values such as quantity or duration, the term “about” as used herein is intended to include variations of ±20%, ±10%, ±5%, ±1%, or ±0.1% from the specified value, where such variations are appropriate for performing the methods of this disclosure.
[0031] "Antisense" specifically refers to the nucleic acid sequence of the non-coding strand of a double-stranded DNA molecule that codes for a protein, or a sequence substantially homologous to such a non-coding strand. As defined herein, an antisense sequence is complementary to the sequence of a double-stranded DNA molecule that codes for a protein. An antisense sequence does not need to be complementary only to the coding portion of the coding strand of the DNA molecule. An antisense sequence may be complementary to a regulatory sequence defined in the coding strand of a protein-coding DNA molecule, which controls the expression of the coding sequence.
[0032] "Disease" is a state of health in an animal in which the animal is unable to maintain homeostasis and, if the disease does not improve, the animal's health continues to deteriorate.
[0033] In contrast, a “disorder” in animals is a state of health in which the animal can maintain homeostasis, but the animal’s health is not as good as it would be without the disorder. If left untreated, a disorder does not necessarily lead to a further deterioration of the animal’s health.
[0034] A disease or disorder is considered “mitigated” if the severity of its signs or symptoms, the frequency with which a patient experiences such signs or symptoms, or both are reduced.
[0035] "Code" refers to the inherent characteristic of a particular nucleotide sequence in a polynucleotide, such as a gene, cDNA, or mRNA, that in a biological process functions as a template for the synthesis of either a predetermined nucleotide sequence (i.e., rRNA, tRNA, and mRNA) or a predetermined amino acid sequence, and other polymers and macromolecules having the biological properties they provide. Thus, a gene codes for a protein if the transcription and translation of the mRNA corresponding to that gene produces a protein in a cell or other biological system. Both the coding strand, which is usually shown in the sequence listing, and the non-coding strand, which is used as a template for the transcription of the gene or cDNA, can be said to code for a protein or other product of that gene or cDNA if the nucleotide sequence is identical to the mRNA sequence.
[0036] The terms “patient,” “subject,” and “individual” are used interchangeably herein and refer to any animal or cell suitable for the methods described herein, whether in vitro or in vivo. In one embodiment, subjects include vertebrates and invertebrates. Invertebrates include, but are not limited to, Drosophila melanogaster and Caenorhabditis elegans. Vertebrates include, but are not limited to, primates, rodents, domesticated animals, or game animals. Primates include, but are not limited to, chimpanzees, crab-eating macaques, spider monkeys, and macaques (e.g., rhesus macaques). Rodents include, but are not limited to, mice, rats, marmots, ferrets, rabbits, and hamsters. Examples of domesticated and game animals include, but are not limited to, cattle, horses, pigs, deer, bison, buffalo, feline species (e.g., domestic cats), canid species (e.g., dogs, foxes, wolves), bird species (e.g., chickens, emus, ostriches), and fish (e.g., zebrafish, trout, catfish, and salmon). In some embodiments, the subject is a mammal, such as a primate, such as a human. In certain non-limiting embodiments, the patient, subject, or individual is a human.
[0037] As used herein with respect to antibodies, the term "specifically binding" means an antibody that recognizes a specific antigen but substantially does not recognize or bind to other molecules in the sample. For example, an antibody that specifically binds to an antigen from one species may also bind to that antigen from one or more species. However, such cross-reactivity itself does not change the classification of the antibody as a specific antibody. In another example, an antibody that specifically binds to an antigen may also bind to different allele forms of that antigen. However, such cross-reactivity itself does not change the classification of the antibody as a specific antibody.
[0038] In some cases, the terms “specific binding” or “specifically binding” may be used to mean that, in relation to the interaction between an antibody, protein, or peptide and a second chemical species, the interaction depends on the presence of a specific structure of that chemical species (e.g., an antigenic determinant or epitope). For example, an antibody recognizes a specific protein structure and binds to it, rather than broadly recognizing and binding to a protein. If an antibody is specific to epitope “A”, the presence of labeled “A” and molecules containing epitope A (or free, unlabeled A) in the reaction product containing the antibody reduces the amount of labeled A that binds to the antibody.
[0039] The "coding region" of a gene consists of nucleotide residues in the coding chain and nucleotides in the non-coding chain of the gene, which are homologous or complementary to the coding region of the mRNA molecule produced by the transcription of the gene.
[0040] Furthermore, the "coding region" of an mRNA molecule consists of nucleotide residues of that mRNA molecule that either fit into the anticodon region of the transfer RNA molecule or encode a stop codon during the translation of that mRNA molecule. Therefore, the coding region may include nucleotide residues containing codons of amino acid residues that are not present in the mature protein encoded by the mRNA molecule (for example, amino acid residues in the protein transport signal sequence).
[0041] As used herein to refer to nucleic acids, “complementary” refers to the broad concept of sequence complementarity between regions of two nucleic acid chains, or between two regions of the same nucleic acid chain. It is known that an adenine residue in a first nucleic acid region can form a specific hydrogen bond ("base pairing") with a residue in a second nucleic acid region that is antiparallel to the first region, if the residue is thymine or uracil. Similarly, it is known that a cytosine residue in a first nucleic acid chain can form a base pair with a residue in a second nucleic acid chain that is antiparallel to the first chain, if the residue is guanine. When two regions are arranged antiparallel, if at least one nucleotide residue in the first region can form a base pair with a residue in the second region, then the first region of the nucleic acid is complementary to the second region of the same or a different nucleic acid. In one embodiment, the first region includes a first portion, and the second region includes a second portion, and so the first and second portions are arranged antiparallel, at least about 50%, at least about 75%, at least about 90%, or at least about 95% of the nucleotide residues of the first portion can form base pairs with the nucleotide residues of the second portion. In one embodiment, all of the nucleotide residues of the first portion can form base pairs with the nucleotide residues of the second portion.
[0042] As used herein, the term "DNA" is defined as deoxyribonucleic acid.
[0043] As used herein, the term “expression” is defined as the transcription and / or translation of a particular nucleotide sequence, facilitated by its promoter.
[0044] As used herein, the term “expression vector” refers to a vector containing nucleic acid sequences that encode at least a portion of a gene product that can be transcribed. In some cases, the RNA molecule is then translated into a protein, polypeptide, or peptide. In other cases, for example, in the production of antisense molecules, siRNA, ribozymes, etc., these sequences are not translated. Expression vectors may contain various regulatory sequences, which are nucleic acid sequences necessary for the transcription and possibly translation of the operationally linked coding sequence in a particular host organism. In addition to regulatory sequences that control transcription and translation, vectors and expression vectors may also contain nucleic acid sequences that perform other functions.
[0045] As used herein, the term “wild type” is a term understood by those skilled in the art and means a typical type of organism, strain, gene, or characteristic found in nature, distinct from a variant or mutant.
[0046] The term "homology" refers to the degree of complementarity. Partial homology or complete homology (i.e., identity) can exist. Homologousity is often measured using sequence analysis software (e.g., Sequence Analysis Software Package from Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705). Such software compares similar sequences by assigning a degree of homology to various substitutions, deletions, insertions, and other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine.
[0047] The term "nucleic acid" means any nucleic acid, whether composed of deoxyribonucleosides or ribonucleosides, and composed of phosphodiester bonds or modified bonds, such as phosphotriesters, phosphoramidates, siloxanes, carbonates, carboxymethyl esters, acetamidates, carbamates, thioethers, cross-linked phosphoramidates, cross-linked methylenephosphonates, phosphorothioates, methylphosphonates, phosphorodithioates, cross-linked phosphorothioates, or sulfone bonds, or combinations of such bonds. The term nucleic acid also specifically includes nucleic acids composed of bases other than the five biologically occurring bases (adenine, guanine, thymine, cytosine, and uracil). The term "nucleic acid" usually refers to large polynucleotides.
[0048] Conventional notation for representing polynucleotide sequences is used herein. That is, the left end of a single-stranded polynucleotide sequence is called the 5' end, and the leftward direction of a double-stranded polynucleotide sequence is called the 5' direction.
[0049] The direction in which nucleotides are added to a newly synthesized RNA transcript from 5' to 3' is called the transcription direction. The DNA strand having the same sequence as the mRNA is called the "coding strand," the DNA strand sequence located 5' to the reference point on the DNA is called the "upstream sequence," and the DNA strand sequence located 3' to the reference point on the DNA is called the "downstream sequence."
[0050] In the context of this invention, the following abbreviations for commonly existing nucleic acid bases are used: "A" refers to adenosine, "C" to cytosine, "G" to guanosine, "T" to thymidine, and "U" to uridine.
[0051] As used herein, the terms “peptide,” “polypeptide,” and “protein” are interchangeable and refer to compounds composed of amino acid residues covalently linked by peptide bonds. A protein or peptide must contain at least two amino acids, and there is no limit to the maximum number of amino acids that can constitute a protein or peptide sequence. Polypeptides include any peptide or protein containing two or more amino acids linked to each other by peptide bonds. As used herein, this term refers to both short chains, commonly referred to in the art as, for example, peptides, oligopeptides, and oligomers, and longer chains, commonly referred to in the art as proteins, of which there are many types. Examples of “polypeptides” include, among others, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, polypeptide variants, modified polypeptides, derivatives, analogs, and fusion proteins. Polypeptides include native peptides, recombinant peptides, synthetic peptides, or combinations thereof.
[0052] As used herein, the term "RNA" is defined as ribonucleic acid.
[0053] Where used herein, a “variant” is a nucleic acid sequence or peptide sequence that differs in sequence from a reference nucleic acid sequence or reference peptide sequence, but retains the fundamental biological properties of the reference molecule. Sequence changes in nucleic acid variants do not necessarily alter the amino acid sequence of the peptide encoded by the reference nucleic acid; they may result in amino acid substitutions, additions, deletions, fusions, and shortenings. Sequence changes in peptide variants are usually limited or conserved, so the sequences of the reference peptide and the variant are generally very similar and identical in many regions. Variants and reference peptides may differ in amino acid sequence due to one or more substitutions, additions, or deletions in any combination. Nucleic acid or peptide variants may be naturally occurring, such as allelic variants, or they may be variants whose natural occurrence is not known. Non-natural variants of nucleic acids and peptides may be produced by mutagenesis or direct synthesis.
[0054] A "vector" is a composition of substances containing an isolated nucleic acid that can be used to deliver the isolated nucleic acid into a cell. Many vectors are known in the art, but are not limited to linear polynucleotides, polynucleotides associated with ionic or amphiphilic compounds, plasmids, and viruses. Therefore, the term "vector" encompasses autonomously replicating plasmids or viruses. The term should also be interpreted to include non-plasmid and non-viral compounds that facilitate the introduction of nucleic acids into cells, such as polylysine compounds and liposomes. Examples of viral vectors include, but are not limited to, adenovirus vectors, adeno-associated virus vectors, and retroviral vectors.
[0055] As used herein, the terms “guide sequence,” “spacer sequence,” “crRNA,” “guide RNA,” “single guide RNA,” or “gRNA” refer to a polynucleotide comprising any polynucleotide sequence that hybridizes with a target nucleic acid sequence and has sufficient complementarity with the target nucleic acid sequence to induce sequence-specific binding of a complex targeting the RNA, which includes the guide sequence and a CRISPR effector protein, to the target nucleic acid sequence. In some exemplary embodiments, the degree of complementarity is about 50% or greater than about 50%, about 60% or greater than about 60%, about 75% or greater than about 75%, about 80% or greater than about 80%, about 85% or greater than about 85%, about 90% or greater than about 90%, about 95% or greater than about 95%, about 97.5% or greater than about 97.5%, about 99% or greater than about 99%, or more, when optimally aligned using a preferred alignment algorithm. The optimal alignment can be determined using any suitable algorithm for aligning sequences, and non-limiting examples of these include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, Calif), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). The ability of a guide sequence (of a nucleic acid-targeting guide RNA) to induce sequence-specific binding of a nucleic acid-targeting complex to a target nucleic acid sequence can be evaluated by any suitable assay.For example, components of a nucleic acid-targeting CRISPR system sufficient to form a nucleic acid-targeting complex, including the guide sequence to be tested, may be provided to a host cell having the corresponding target nucleic acid sequence by translocation using a vector encoding the components of the nucleic acid-targeting complex, and subsequently, selective targeting (e.g., cleavage) into the target nucleic acid sequence is evaluated by a Surveyor assay as described herein. Similarly, components of a nucleic acid-targeting complex, including the target nucleic acid sequence, the guide sequence to be tested, and a control guide sequence different from the guide sequence to be tested, may be provided, and cleavage of the target nucleic acid sequence in vitro can be evaluated by comparing the binding or cleavage rate within the target sequence between the reaction with the guide sequence to be tested and the reaction with the control guide sequence. Other assays are available and will be conceivable by those skilled in the art. The guide sequence and the nucleic acid-targeting guide thereof may be selected to target any target nucleic acid sequence. The target sequence may be DNA. The target sequence may be any RNA sequence. In some embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of messenger RNA (mRNA), premRNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), small interfering RNA (siRNA), nuclear small RNA (snRNA), nucleolar small RNA (snoRNA), double-stranded RNA (dsRNA), non-coding RNA (ncRNA), long non-coding RNA (lncRNA), and cytoplasmic small RNA (scRNA). In some preferred embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of mRNA, premRNA, and rRNA. In some preferred embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of ncRNA and lncRNA. In some more preferred embodiments, the target sequence may be a sequence within an mRNA molecule or a premRNA molecule.
[0056] Scope: Throughout this disclosure, various aspects of the invention may be presented in scope form. It should be understood that scope form is merely for convenience and brevity and should not be interpreted as an inflexible limitation on the scope of this disclosure. Accordingly, scope descriptions should be considered to specifically disclose all possible subranges and individual numbers within that range. For example, a scope description such as 1 to 6 should be considered to specifically disclose subranges, e.g., 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numbers within that range, e.g., 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the width of the range.
[0057] protein In one embodiment, the disclosure is based on the development of a novel editing protein that results in targeted RNA cleavage. In some embodiments, the protein includes a localization signal. In one embodiment, the localization signal localizes the protein to a site where target RNA is present. In one embodiment, the protein includes a nuclear localization signal (NLS) for targeting RNA in the nucleus. In one embodiment, the protein includes a nuclear export signal (NES) for targeting RNA in the cytoplasm. In one embodiment, the fusion protein includes a purification and / or detection tag.
[0058] In one embodiment, the disclosure is based on the development of a novel editing protein that results in targeted RNA cleavage and is effectively delivered. In some embodiments, the protein includes a localization signal. In one embodiment, the localization signal localizes the protein to a site where the target RNA is present. In one embodiment, the protein includes a purification and / or detection tag.
[0059] In one embodiment, the disclosure is based on the development of a novel editing protein that can regulate the cleavage and / or polyadenylation of nuclear RNA and is effectively delivered to the nucleus. In some embodiments, the protein comprises a cleavage and / or polyadenylation protein. In some embodiments, the protein comprises a nuclear localization signal. In one embodiment, the protein comprises a purification and / or detection tag.
[0060] Edited proteins In one embodiment, the editing protein may include, but is not limited to, CRISPR-related (Cas) proteins, zinc finger nuclease (ZFN) proteins, and proteins having a DNA or RNA binding domain.
[0061] Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cm Examples include r4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx15, Csf1, Csf2, Csf3, Csf4, SpCas9, StCas9, NmCas9, SaCas9, CjCas9, CjCas9, AsCpf1, LbCpf1, FnCpf1, VRER SpCas9, VQR SpCas9, xCas9 3.7, their homologs, their orthologues, or their variants. In some embodiments, the Cas protein has DNA or RNA cleavage activity. In some embodiments, the Cas protein induces cleavage of one or both strands of a nucleic acid molecule at a target sequence site, e.g., within the target sequence and / or within the complement of the target sequence. In some embodiments, the Cas protein induces a break in one or both strands within a range of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 base pairs, or more, from the first or last nucleotide of the target sequence. In one embodiment, the Cas protein is Cas9, Cas13, or Cpf1. In one embodiment, the Cas protein lacks catalytic activity (dCas).
[0062] In one embodiment, the Cas protein has RNA-binding activity. In one embodiment, the Cas protein is Cas13. In one embodiment, the Cas protein is PspCas13b, a shortened form of PspCas13b, AdmCas13d, AspCas13b, AspCas13c, BmaCas13a, BzoCas13b, CamCas13a, CcaCas13b, Cga2Cas13a, CgaCas13a, EbaCas13a, EreCas13a, EsCas13d, FbrCas13b, FnbCas13c, FndCas13c, FnfCas13c, FnsCas13c, FpeCas13c, FulCas13c, Hh eCas13a, LbfCas13a, LbmCas13a, LbnCas13a, LbuCas13a, LseCas13a, LshCas13a, LspCas13a, Lwa2cas13a, LwaCas13a, LweCas13a, PauCas13 b, PbuCas13b, PgiCas13b, PguCas13b, Pin2Cas13b, Pin3Cas13b, PinCas13b, Pprcas13a, PsaCas13b, PsmCas13b, RaCas13d, RanCas13b, RcdC as13a, RcrCas13a, RcsCas13a, RfxCas13d, UrCas13d, dPspCas13b, PspCas13b_A133H, PspCas13b_A1058H, truncated form of dPspCas13b, dAdmCas13d, dA spCas13b, dAspCas13c, dBmaCas13a, dBzoCas13b, dCamCas13a, dCcaCas13b, dCga2Cas13a, dCgaCas13a, dEbaCas13a, dEreCas13a, dEsCas13 d, dFbrCas13b, dFnbCas13c, dFndCas13c, dFnfCas13c, dFnsCas13c, dFpeCas13c, dFulCas13c, dHheCas13a, dLbfCas13a, dLbmCas13a, dLbnC as13a, dLbuCas13a, dLseCas13a, dLshCas13a, dLspCas13a, dLwa2cas13a, dLwaCas13a, dLweCas13a, dPauCas13b, dPbuCas13b, dPgiCas13b,These are dPguCas13b, dPin2Cas13b, dPin3Cas13b, dPinCas13b, dPprCas13a, dPsaCas13b, dPsmCas13b, dRaCas13d, dRanCas13b, dRcdCas13a, dRcrCas13a, dRcsCas13a, dRfxCas13d, or dUrCas13d. Additional Cas proteins are known in the art (e.g., incorporated herein by reference: Konermann et al., Cell, 2018, 173:665-676 e14; Yan et al., Mol Cell, 2018, 7:327-339 e5; Cox, DBT, et al., Science, 2017, 358:1019-1027; Abudayyeh et al., Nature, 2017, 550:280-284; Gootenberg et al., Science, 2017, 356:438-442; and East-Seletsky et al., Mol Cell, 2017, 66:373-383 e3).
[0063] In one embodiment, the Cas protein contains a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 1 to 49. In one embodiment, the Cas protein contains a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 1 to 49. In one embodiment, the Cas protein contains a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, or at least 99% identical to one of sequence numbers 1 to 47.
[0064] Localized signals In some embodiments, the protein may contain localization signals, such as nuclear localization signals (NLS), nuclear export signals (NES), or other localization signals for localizing to organelles such as mitochondria, or to the cytoplasm. In one embodiment, the localization signal localizes the protein to a site where the target RNA is present.
[0065] Nuclear localization signals In one embodiment, the protein comprises an NLS. In one embodiment, the NLS is a retrotransposon NLS. In one embodiment, the NLS is derived from Ty1, yeast GAL4, SKI3, L29, or histone H2B protein, polyomavirus large T protein, VP1 or VP2 capsid protein, SV40 VP1 or VP2 capsid protein, adenovirus Ela or DBP protein, influenza virus NS1 protein, hepatitis virus core antigen, or mammalian lamin, c-myc, max, c-myb, p53, c-erbA, jun, Tax, steroid receptor, or Mx protein, nucleoplasmin (NPM2), nucleophosmin (NPM1), or Simianvirus 40 ("SV40") T antigen. In one embodiment, the NLS is Ty1 or Ty1-derived NLS, Ty2 or Ty2-derived NLS, or MAK11 or MAK11-derived NLS. In one embodiment, Ty1 NLS contains the amino acid sequence of SEQ ID NO: 50. In one embodiment, Ty2 NLS contains the amino acid sequence of SEQ ID NO: 51. In one embodiment, MAK11 NLS contains the amino acid sequence of SEQ ID NO: 52. In one embodiment, the NLS contains a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of SEQ ID NOs: 50-57 and 323-935. In one embodiment, the NLS includes one sequence from sequence numbers 50-57 and 323-935.
[0066] In one embodiment, the NLS is a Ty1-like NLS. For example, in one embodiment, the Ty1-like NLS contains a KKRX motif. In one embodiment, the Ty1-like NLS contains a KKRX motif at its N-terminus. In one embodiment, the Ty1-like NLS contains a KKR motif. In one embodiment, the Ty1-like NLS contains a KKR motif at its C-terminus. In one embodiment, the Ty1-like NLS contains both a KKRX motif and a KKR motif. In one embodiment, the Ty1-like NLS contains a KKRX motif at its N-terminus and a KKR motif at its C-terminus. In one embodiment, the Ty1-like NLS contains at least 20 amino acids. In one embodiment, the Ty1-like NLS contains 20 to 40 amino acids. In one embodiment, the Ty1-like NLS contains a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 323 to 935. In one embodiment, the NLS includes one sequence from sequence numbers 323 to 935, wherein the sequence includes one or more insertions, deletions, or substitutions of one, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more. In one embodiment, the Ty1-like NLS includes one sequence from sequence numbers 323 to 935.
[0067] In one embodiment, the NLS comprises two copies of the same NLS. For example, in one embodiment, the NLS comprises a polymer of an NLS derived from a first Ty1 and an NLS derived from a second Ty1.
[0068] Nuclear export signals In one embodiment, the protein contains a nuclear export signal (NES). In one embodiment, the NES is bound to the N-terminus of the Cas protein. In one embodiment, the NES localizes the protein to the cytoplasm to target cytoplasmic RNA. In one embodiment, the NES contains an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 58 or 59.
[0069] Organelle localization signals In one embodiment, the protein includes a localization signal that localizes the protein to an organelle. In one embodiment, the localization signal localizes the protein to the nucleolus, ribosome, vesicle, rough endoplasmic reticulum, Golgi apparatus, cytoskeleton, smooth endoplasmic reticulum, mitochondria, vacuole, cytosol, lysosome, or centrioles. Numerous localization signals are known in the art.
[0070] In one embodiment, the protein includes a localization signal that localizes the protein to an organelle or extracellular space. In one embodiment, the localization signal localizes the protein to the nucleolus, ribosome, vesicle, rough endoplasmic reticulum, Golgi apparatus, cytoskeleton, smooth endoplasmic reticulum, mitochondria, vacuole, cytosol, lysosome, or centrioles.
[0071] Numerous localization signals are known in the art. Examples of localization signals include, but are not limited to, 1x mitochondrial targeting sequences, 4x mitochondrial targeting sequences, secretory signaling sequences (IL-2), myristylation sequences, calcechestrin reader sequences, KDEL retention sequences, and peroxisome targeting sequences.
[0072] In one embodiment, the localization signal includes sequences identical to sequences 60-66 by at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the localization signal includes sequences 60-66.
[0073] Purification and / or detection of tags In some embodiments, the protein may contain a purification and / or detection tag. In one embodiment, the tag is located at the N-terminus of the protein. In one embodiment, the tag is a 3xFLAG tag. In one embodiment, the tag contains an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 67. In one embodiment, the tag contains the amino acid sequence of SEQ ID NO: 67.
[0074] Cleavage and / or polyadenylated proteins In one embodiment, the fusion protein comprises an edited protein and a cleaved and / or polyadenylated protein that are effectively delivered to the nucleus. In one embodiment, the cleaved and / or polyadenylated protein is an RNA-binding protein of the human 3' end processing mechanism. In one embodiment, the cleaved and / or polyadenylated protein is CPSF30, WDR33, or NUDT21. In one embodiment, the cleaved and / or polyadenylated protein is NUDT21. In one embodiment, the cleaved and / or polyadenylated protein is NUDT21, a NUDT21 variant, a NUDT21 dimer, a NUDT21 fusion protein, or any combination thereof. In one embodiment, the cleaved and / or polyadenylated protein is human NUDT21, helminth NUDT21, fly NUDT21, zebrafish NUDT21, NUDT21_R63S, NUDT21_F103A, or a tandem dimer of NUDT21.
[0075] In one embodiment, the cleaved and / or polyadenylated protein contains an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of SEQ ID NOs: 298-307.
[0076] Fusion protein A fusion protein containing Cas protein and localization signals. In one embodiment, the protein of the Disclosure is effectively delivered to the nucleus, organelles, cytoplasm, or extracellular space, enabling targeted RNA cleavage. In one embodiment, the protein contains an amino acid sequence that is 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of SEQ ID NOs.
[0077] A fusion protein comprising Cas protein and cleaved and / or polyadenylated proteins. In one embodiment, the protein of the Disclosure is effectively delivered to the nucleus, enabling RNA targeted cleavage and / or polyadenylation. In one embodiment, the protein contains an amino acid sequence that is 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of SEQ ID NOs.
[0078] Proteins, peptides, and fusion proteins The proteins of this disclosure can be prepared using chemical methods. For example, the proteins can be synthesized by solid-phase methods (Roberge JY et al (1995) Science 269:202-204), cleaved from resin, and purified by preparative high-performance liquid chromatography. Automated synthesis can be achieved, for example, using an ABI 431A Peptide Synthesizer (Perkin Elmer) according to the instructions for use provided by the manufacturer.
[0079] The proteins of this disclosure can be prepared using recombinant protein expression. The recombinant expression vectors of this disclosure contain the nucleic acids of the present invention in a form suitable for expression of said nucleic acids in a host cell, meaning that the recombinant expression vector contains one or more regulatory sequences operably ligated to the nucleic acid sequence to be expressed, selected based on the host cell used for expression. In a recombinant expression vector, “operably ligated” means that the nucleotide sequence of interest is ligated to the regulatory sequences in a manner that enables the expression of the nucleotide sequence (for example, in an in vitro transcription / translation system, or in the host cell if the vector is introduced into a host cell).
[0080] The term “regulatory sequence” is intended to encompass promoters, enhancers, and other expression regulatory elements (e.g., polyadenylation signals). Such regulatory sequences are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, Calif. (1990). Regulatory sequences include those that induce constitutive expression of nucleotide sequences in a number of host cell types, and those that induce expression of nucleotide sequences only in specific host cells (e.g., tissue-specific regulatory sequences). Those skilled in the art will understand that the design of expression vectors may depend on factors such as the selection of host cells to be transformed and the desired level of protein expression. Expression vectors of the present invention can be introduced into host cells to produce proteins or peptides containing fusion proteins or peptides encoded by nucleic acids described herein.
[0081] The recombinant expression vectors of the present invention may be designed for the generation of variant proteins in prokaryotic or eukaryotic cells. For example, the proteins of the present invention may be expressed in bacterial cells such as Escherichia coli, insect cells (using baculovirus expression vectors), yeast cells, or mammalian cells. Suitable host cells are further described in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, Calif. (1990). Alternatively, the recombinant expression vectors may be transcribed and translated in vitro, for example, using a T7 promoter regulatory sequence and T7 polymerase.
[0082] Protein expression in prokaryotes is almost always carried out in Escherichia coli using vectors containing constitutive or inducible promoters that induce the expression of fusion or non-fusion proteins. Fusion vectors add numerous amino acids to the amino-terminus or C-terminus of the recombinant protein to the protein encoded therein. Such fusion vectors typically serve three purposes: (i) to increase the expression of the recombinant protein; (ii) to increase the solubility of the recombinant protein; and (iii) to aid in the purification of the recombinant protein by acting as a ligand in affinity purification. In fusion expression vectors, proteolytic cleavage sites are often introduced at the junction of the fusion site and the recombinant protein to allow for the separation of the recombinant protein from the fusion site after the purification of the fusion protein. Examples of such enzymes and their homologous recognition sequences include factor Xa, thrombin, prescission, TEV, and enterokinase. Representative fusion expression vectors include pGEX (Pharmacia Biotech Inc; Smith and Johnson, 1988. Gene 67:31-40), pMAL (New England Biolabs, Beverly, Mass.), and pRITS (Pharmacia, Piscataway, NJ), which fuse glutathione S-transferase (GST), maltose E-binding protein, or protein A to the target recombinant protein, respectively.
[0083] Suitable examples of inducible non-fusion E. coli expression vectors include pTrc (Amrann et al., (1988) Gene 69:301-315) and pET11d (Studier et al., Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, Calif. (1990) 60-89), however, this is not entirely accurate, as pET11a-d have an N-terminal T7 tag.
[0084] One strategy to maximize recombinant protein expression in E. coli is to express the protein in a host bacterium whose ability to cleave the recombinant protein by proteolysis is impaired. See, for example, Gottesman, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, Calif. (1990) 119-128. Another strategy is to modify the nucleic acid sequence of the nucleic acid inserted into the expression vector so that individual codons for each amino acid are preferentially utilized in E. coli (see, for example, Wada, et al., 1992. Nucl. Acids Res. 20: 2111-2118). Such modification of the nucleic acid sequence in this invention can be carried out by standard DNA synthesis techniques. Another strategy to address codon bias is to use BL21 codon-plus strains (Invitrogen) or Rosetta strains (Novagen), which have extra copies of the rare tRNA gene in E. coli.
[0085] In another embodiment, the expression vector encoding the protein of this disclosure is a yeast expression vector. Examples of vectors for expression in the yeast Saccharomyces cerevisiae include pYepSec1 (Baldari, et al., 1987. EMBO J.6:229-234), pMFa (Kurjan and Herskowitz, 1982. Cell 30:933-943), pJRY88 (Schultz et al., 1987. Gene 54:113-123), pYES2 (Invitrogen Corporation, San Diego, Calif.), and picZ (InVitrogen Corp, San Diego, Calif.).
[0086] Alternatively, the polypeptides of the present invention can be generated in insect cells using baculovirus expression vectors. Examples of baculovirus vectors available for protein expression in cultured insect cells (e.g., SF9 cells) include the pAc series (Smith, et al., 1983. Mol. Cell. Biol. 3:2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170:31-39).
[0087] In yet another embodiment, the nucleic acids of this disclosure are expressed in mammalian cells using mammalian expression vectors. Mammalian cell lines available in the art for the expression of heterologous polypeptides include, but are not limited to, Chinese hamster ovary (CHO) cells, HeLa cells, baby hamster kidney cells, NSO mouse melanoma cells, YB2 / 0 rat myeloma cells, human embryonic kidney cells, human embryonic retinal cells, and numerous other cells. Examples of mammalian expression vectors include pCDM8 (Seed, 1987. Nature 329:840), and pMT2PC (Kaufman, et al., 1987. EMBO J.6:187-195), pIRESpuro (Clontech), pUB6 (Invitrogen), pCEP4 (Invitrogen), pREP4 (Invitrogen), and pcDNA3 (Invitrogen). When used in mammalian cells, the regulatory function of the expression vector is often provided by viral regulatory elements. For example, commonly used promoters are derived from polyomavirus, adenovirus 2, cytomegalovirus, Rous sarcoma virus, and Simianvirus 40. For other expression systems suitable for both prokaryotic and eukaryotic cells, see, for example, Chapters 16 and 17 of Sambrook, et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
[0088] In another embodiment, a mammalian recombinant expression vector can preferentially induce nucleic acid expression in a specific cell type (for example, a tissue-specific regulatory element is used to express nucleic acid). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promoters include albumin promoters (liver-specific; Pinkert, et al., 1987. Genes Dev. 1:268-277), lymphoid-specific promoters (Calame and Eaton, 1988. Adv. Immunol. 43:235-275), particularly T cell receptor promoters (Winoto and Baltimore, 1989. EMBO J. 8:729-733) and immunoglobulin promoters (Banerji, et al., 1983. Cell 33:729-740; Queen and Baltimore, 1983. Cell 33:741-748), and neuronal-specific promoters (e.g., neuronal filament promoter; Byrne and Ruddle, 1989. Proc. Natl. Acad. Sci. USA) Examples include the 86:5473-5477), pancreas-specific promoter (Edlund, et al., 1985. Science 230:912-916), and mammary gland-specific promoter (e.g., whey promoter; U.S. Patent No. 4,873,316 and European Patent Application Publication No. 264,166). Promoter regulated during development, such as the mouse hox promoter (Kessel and Gruss, 1990. Science 249:374-379) and the α-fetoprotein promoter (Campes and Tilghman, 1989. Genes Dev. 3:537-546), are also included.
[0089] The present invention should be interpreted to also include any form of protein having substantial homology to the proteins disclosed herein. In one embodiment, a protein that is "substantially homologous" is one that is about 50%, about 70%, about 80%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% homologous to the amino acid sequence of the fusion protein disclosed herein.
[0090] Alternatively, proteins can be produced by recombinant methods or by cleavage from longer polypeptides. The protein composition can be determined by amino acid analysis or sequencing.
[0091] Variants of the protein according to the present invention may be (i) variants in which one or more amino acid residues are substituted with conserved or non-conserved amino acid residues, wherein such substituted amino acid residues are encoded by the genetic code or not; (ii) variants containing one or more modified amino acid residues, for example, residues modified by the attachment of substituents; (iii) variants in which the peptide is an alternative splicing variant of the protein of the present invention; (iv) peptide fragments; and / or (v) variants in which the protein is fused to another peptide, for example, a leader sequence or secretion sequence, or a sequence used for purification (e.g., a His tag) or a sequence used for detection (e.g., an Sv5 epitope tag). Fragments include peptides produced by proteolytic cleavage (including multi-site proteolysis) of the original sequence. Variants may be post-translationally modified or chemically modified. Such variants are considered to be within the scope of the art as taught herein.
[0092] As is known in the art, the “similarity” between two fusion proteins is determined by comparing the amino acid sequence and its conserved amino acid substitutions of one polypeptide with the sequence of the other polypeptide. A variant is defined as containing a peptide sequence different from the original sequence. In one embodiment, a variant is one in which less than 40% of residues per segment of interest differ from the original sequence, less than 25% of residues per segment of interest differ from the original sequence, less than 10% of residues per segment of interest differ from the original protein sequence, or a very small number of residues per segment of interest differ from the original protein sequence, while at the same time being sufficiently homologous to the original sequence to maintain the function of the original sequence and / or its ability to stimulate differentiation of stem cells into osteoblasts. The present invention includes an amino acid sequence that is at least 60%, 65%, 70%, 72%, 74%, 76%, 78%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% similar to or identical to the original amino acid sequence. The degree of identity between two peptides is determined using computer algorithms and methods widely known to those skilled in the art. Identity between two amino acid sequences can be determined using the BLASTP algorithm [BLAST Manual, Altschul, S., et al., NCBI NLM NIH Bethesda, Md.20894, Altschul, S., et al., J.Mol.Biol.215:403-410 (1990)].
[0093] The proteins of this disclosure may be post-translation modified. For example, post-translation modifications within the scope of the present invention include signal peptide cleavage, glycosylation, acetylation, isoprenylation, proteolysis, myristoylation, protein folding, and proteolytic processing. Some modifications or processing events require the introduction of additional biological mechanisms. For example, processing events such as signal peptide cleavage and core glycosylation can be examined by adding canine microsomal membrane or Xenopus egg extract (U.S. Patent No. 6,103,489) to a standard translation reaction.
[0094] The proteins disclosed herein may contain unnatural amino acids formed by post-translational modification or by introducing unnatural amino acids during translation. Various methods are available for introducing unnatural amino acids during protein translation.
[0095] The proteins disclosed herein can be phosphorylated using conventional methods, such as those described in Reedijk et al. (The EMBO Journal 11(4):1365, 1992).
[0096] Cyclic derivatives of the fusion proteins of the present invention are also part of the present invention. Cyclization can allow a protein to adopt a conformational structure favorable for association with other molecules. Cyclization can be achieved using techniques known in the art. For example, a disulfide bond can be formed between two appropriately spaced components having free sulfhydryl groups, or an amide bond can be formed between an amino group of one component and a carboxyl group of another. Cyclization may also be achieved using azobenzene-containing amino acids, as described in Ulysse, L., et al., J. Am. Chem. Soc. 1995, 117, 8466-8467. The components forming the bond can be amino acid side chains, non-amino acid components, or a combination of the two. In embodiments of the present invention, the cyclic peptide may contain a β-turn at an appropriate position. The β-turn can be introduced into the peptide of the present invention by adding the amino acid Pro-Gly at an appropriate position.
[0097] In some cases, it is desirable to create cyclic proteins that are more flexible than cyclic peptides containing the peptide bonds described above. More flexible peptides can be prepared by introducing cysteine at the left and right positions of the peptide and forming disulfide crosslinks between the two cysteine. The two cysteine are positioned so as not to deform the β-sheet and turn. This peptide is more flexible as a result of the shorter length of the disulfide bond and the fewer hydrogen bonds in the β-sheet portion. The relative flexibility of a cyclic peptide can be determined by molecular dynamics simulations.
[0098] The disclosure also relates to a peptide comprising a fusion protein containing Cas13 and an RNase protein, wherein the fusion protein itself is fused to or incorporated into a targeting domain capable of inducing a target protein and / or a chimeric protein to a desired cellular component, cell type, or tissue. The chimeric protein may also contain additional amino acid sequences or domains. The chimeric protein is recombinant in that its various components originate from different sources and are therefore not found together in nature (i.e., heterogeneous).
[0099] In one embodiment, the targeting domain may be a transmembrane domain, a membrane-bound domain, or a sequence that induces the protein to associate with, for example, a vesicle or nucleus. In one embodiment, the targeting domain can target the peptide to a specific cell type or tissue. For example, the targeting domain may be a cell surface ligand or an antibody against a cell surface antigen of the target tissue. The targeting domain can target the peptide of the present invention to cellular components.
[0100] The peptides of the present invention can be synthesized by prior art. For example, peptides or chimeric proteins can be synthesized by chemical synthesis using solid-phase peptide synthesis. These methods employ either solid-phase or liquid-phase synthesis (see, for example, J. M. Stewart, and J. D. Young, Solid Phase Peptide Synthesis, 2nd Ed., Pierce Chemical Co., Rockford Ill. (1984) and G. Barany and R. B. Merrifield, The Peptides: Analysis Synthesis, Biology editors E. Gross and J. Meienhofer Vol. 2 Academic Press, New York, 1980, pp. 3-254 for solid-phase synthesis, and M. Bodansky, Principles of Peptide Synthesis, Springer-Verlag, Berlin 1984 and E. Gross and J. Meienhofer, Eds., The Peptides: Analysis, Synthesis, Biology, supr, Vol 1). As an example, the peptide of the present invention can be synthesized using a 9-fluorenylmethoxycarbonyl (Fmoc) solid-phase chemistry method, in which phosphothreonine is directly incorporated as an N-fluorenylmethoxycarbonyl-O-benzyl-L-phosphothreonine derivative.
[0101] N-terminal or C-terminal fusion proteins comprising the peptide or chimeric protein of the present invention conjugated with other molecules can be prepared by fusion of the N-terminus or C-terminus of the peptide or chimeric protein with the sequence of a selected protein or selected marker having a desired biological function using recombination techniques. The resulting fusion protein contains the protein fused to the selected protein or marker protein described herein. Examples of proteins that can be used to prepare fusion proteins include immunoglobulins, glutathione-S-transferase (GST), hemagglutinin (HA), and the truncated form myc.
[0102] The peptides of the present invention can be developed using biological expression systems. These systems enable the generation of large libraries of random peptide sequences and the screening of these libraries for peptide sequences that bind to specific proteins. Libraries can be constructed by cloning synthetic DNA encoding random peptide sequences into appropriate expression vectors (see Christian et al 1992, J.Mol.Biol.227:711; Devlin et al, 1990 Science 249:404; Cwirla et al 1990, Proc.Natl.Acad, Sci.USA, 87:6378). Libraries can also be constructed by the simultaneous synthesis of duplicate peptides (see U.S. Patent No. 4,708,871).
[0103] The peptides and chimeric proteins of the present invention can be converted into pharmaceutical salts by reacting them with inorganic acids, such as hydrochloric acid, sulfuric acid, hydrobromic acid, phosphoric acid, or organic acids, such as formic acid, acetic acid, propionic acid, glycolic acid, lactic acid, pyruvic acid, oxalic acid, succinic acid, malic acid, tartaric acid, citric acid, benzoic acid, salicylic acid, benzenesulfonic acid, and toluenesulfonic acid.
[0104] nucleic acid In one embodiment, the disclosure relates to a novel nucleic acid molecule encoding an editing protein that results in targeted RNA cleavage. In some embodiments, the nucleic acid molecule includes a nucleic acid sequence encoding a localization signal. In one embodiment, the localization signal localizes the protein to a site where target RNA is present. In one embodiment, the nucleic acid molecule includes a nucleic acid sequence encoding a nuclear localization signal (NLS) for targeting RNA in the nucleus. In one embodiment, the nucleic acid molecule includes a nucleic acid sequence encoding a nuclear export signal (NES) for targeting RNA in the cytoplasm. Other localization signals (known in the art) may be used for targeting RNA in organelles such as mitochondria. In other embodiments, the nucleic acid molecule does not include a nucleic acid sequence encoding a localization signal for targeting RNA in the cytoplasm. In one embodiment, the nucleic acid molecule includes a nucleic acid sequence encoding a purification and / or detection tag.
[0105] This disclosure also provides targeted nucleic acids, including CRISPR RNA (crRNA), for targeting the proteins of this disclosure to target RNA.
[0106] In one embodiment, the disclosure provides a novel nucleic acid molecule encoding an editing protein that results in RNA targeted cleavage. In some embodiments, the nucleic acid molecule includes a nucleic acid sequence encoding a localization signal. In one embodiment, the localization signal localizes the protein to a site where the target RNA is present. Thus, the disclosure provides a nucleic acid molecule encoding a protein for RNA targeted cleavage that can be localized.
[0107] Edited proteins In one embodiment, the nucleic acid molecule includes a sequence nucleic acid encoding an edited protein. In one embodiment, the edited protein may include, but is not limited to, CRISPR-related (Cas) proteins, zinc finger nuclease (ZFN) proteins, and proteins having a DNA or RNA binding domain.
[0108] Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cm Examples include r4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx15, Csf1, Csf2, Csf3, Csf4, SpCas9, StCas9, NmCas9, SaCas9, CjCas9, CjCas9, AsCpf1, LbCpf1, FnCpf1, VRER SpCas9, VQR SpCas9, xCas9 3.7, their homologs, their orthologues, or their variants. In some embodiments, the Cas protein has DNA or RNA cleavage activity. In some embodiments, the Cas protein induces cleavage of one or both strands of a nucleic acid molecule at a target sequence site, e.g., within the target sequence and / or within the complement of the target sequence. In some embodiments, the Cas protein induces a break in one or both strands within a range of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 base pairs, or more, from the first or last nucleotide of the target sequence. In one embodiment, the Cas protein is Cas9, Cas13, or Cpf1. In one embodiment, the Cas protein lacks catalytic activity (dCas).
[0109] In one embodiment, the Cas protein has RNA-binding activity. In one embodiment, the Cas protein is Cas13. In one embodiment, the Cas protein is PspCas13b, a shortened form of PspCas13b, AdmCas13d, AspCas13b, AspCas13c, BmaCas13a, BzoCas13b, CamCas13a, CcaCas13b, Cga2Cas13a, CgaCas13a, EbaCas13a, EreCas13a, EsCas13d, FbrCas13b, FnbCas13c, FndCas13c, FnfCas13c, FnsCas13c, FpeCas13c, FulCas13c, Hh eCas13a, LbfCas13a, LbmCas13a, LbnCas13a, LbuCas13a, LseCas13a, LshCas13a, LspCas13a, Lwa2cas13a, LwaCas13a, LweCas13a, PauCas13 b, PbuCas13b, PgiCas13b, PguCas13b, Pin2Cas13b, Pin3Cas13b, PinCas13b, Pprcas13a, PsaCas13b, PsmCas13b, RaCas13d, RanCas13b, RcdC as13a, RcrCas13a, RcsCas13a, RfxCas13d, UrCas13d, dPspCas13b, PspCas13b_A133H, PspCas13b_A1058H, truncated form of dPspCas13b, dAdmCas13d, dA spCas13b, dAspCas13c, dBmaCas13a, dBzoCas13b, dCamCas13a, dCcaCas13b, dCga2Cas13a, dCgaCas13a, dEbaCas13a, dEreCas13a, dEsCas13 d, dFbrCas13b, dFnbCas13c, dFndCas13c, dFnfCas13c, dFnsCas13c, dFpeCas13c, dFulCas13c, dHheCas13a, dLbfCas13a, dLbmCas13a, dLbnC as13a, dLbuCas13a, dLseCas13a, dLshCas13a, dLspCas13a, dLwa2cas13a, dLwaCas13a, dLweCas13a, dPauCas13b, dPbuCas13b, dPgiCas13b,These are dPguCas13b, dPin2Cas13b, dPin3Cas13b, dPinCas13b, dPprCas13a, dPsaCas13b, dPsmCas13b, dRaCas13d, dRanCas13b, dRcdCas13a, dRcrCas13a, dRcsCas13a, dRfxCas13d, or dUrCas13d. Additional Cas proteins are known in the art (e.g., incorporated herein by reference: Konermann et al., Cell, 2018, 173:665-676 e14; Yan et al., Mol Cell, 2018, 7:327-339 e5; Cox, DBT, et al., Science, 2017, 358:1019-1027; Abudayyeh et al., Nature, 2017, 550:280-284; Gootenberg et al., Science, 2017, 356:438-442; and East-Seletsky et al., Mol Cell, 2017, 66:373-383 e3).
[0110] In one embodiment, the nucleic acid sequence encoding the Cas protein includes a nucleic acid sequence encoding an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 1 to 49. In one embodiment, the nucleic acid sequence encoding the Cas protein includes a nucleic acid sequence encoding one of sequence numbers 1 to 49.
[0111] In one embodiment, the nucleic acid sequence encoding the Cas protein includes a nucleic acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 132 to 135. In one embodiment, the nucleic acid sequence encoding the Cas protein includes one of sequence numbers 132 or 133.
[0112] Localized signals In some embodiments, the nucleic acid molecule includes a nucleic acid sequence that encodes a localization signal, such as a nuclear localization signal (NLS), a nuclear export signal (NES), or other localization signals for localization to organelles such as the cytoplasm or mitochondria. In one embodiment, the localization signal localizes the protein to a site where a target RNA is present.
[0113] Nuclear localization signals In one embodiment, the nucleic acid molecule comprises a nucleic acid sequence encoding a nuclear localization signal (NLS). In one embodiment, the NLS is a retrotransposon NLS. In one embodiment, the NLS is derived from Ty1, yeast GAL4, SKI3, L29, or histone H2B protein, polyomavirus large T protein, VP1 or VP2 capsid protein, SV40 VP1 or VP2 capsid protein, adenovirus Ela or DBP protein, influenza virus NS1 protein, hepatitis virus core antigen, or mammalian lamin, c-myc, max, c-myb, p53, c-erbA, jun, Tax, steroid receptor, or Mx protein, nucleoplasmin (NPM2), nucleophosmin (NPM1), or Simianvirus 40 ("SV40") T antigen.
[0114] In one embodiment, the NLS is Ty1 or NLS derived from Ty1, Ty2 or NLS derived from Ty2, or MAK11 or NLS derived from MAK11. In one embodiment, the Ty1 NLS contains the amino acid sequence of SEQ ID NO: 50. In one embodiment, the Ty2 NLS contains the amino acid sequence of SEQ ID NO: 51. In one embodiment, the MAK11 NLS contains the amino acid sequence of SEQ ID NO: 52. In one embodiment, the nucleic acid sequence encoding the NLS includes a nucleic acid sequence encoding an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 50-57 and 323-935.
[0115] In one embodiment, the NLS is a Ty1-like NLS. For example, in one embodiment, the Ty1-like NLS contains a KKRX motif. In one embodiment, the Ty1-like NLS contains a KKRX motif at its N-terminus. In one embodiment, the Ty1-like NLS contains a KKR motif. In one embodiment, the Ty1-like NLS contains a KKR motif at its C-terminus. In one embodiment, the Ty1-like NLS contains both a KKRX motif and a KKR motif. In one embodiment, the Ty1-like NLS contains a KKRX motif at its N-terminus and a KKR motif at its C-terminus. In one embodiment, the Ty1-like NLS contains at least 20 amino acids. In one embodiment, the Ty1-like NLS contains 20 to 40 amino acids. In one embodiment, the nucleic acid sequence encoding the Ty1-like NLS includes a nucleic acid sequence encoding an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 323 to 935. In one embodiment, the nucleic acid sequence encoding the Ty1-like NLS includes a nucleic acid sequence encoding one amino acid sequence from sequence numbers 323 to 935, wherein the sequence includes one or more insertions, deletions, or substitutions of one, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more. In one embodiment, the nucleic acid sequence encoding the Ty1-like NLS includes a nucleic acid sequence encoding one amino acid sequence from sequence numbers 323 to 935.
[0116] In one embodiment, the nucleic acid sequence encoding the NLS encodes two copies of the same NLS. For example, in one embodiment, the nucleic acid sequence encodes a polymer of an NLS derived from a first Ty1 and an NLS derived from a second Ty1.
[0117] In one embodiment, the nucleic acid sequence encoding the NLS includes a nucleic acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to sequence number 136.
[0118] Nuclear export signals In one embodiment, the nucleic acid molecule includes a nucleic acid sequence encoding a nuclear export signal (NES). In one embodiment, the NES localizes a protein to the cytoplasm to target cytoplasmic RNA. In one embodiment, the nucleic acid sequence encoding the NES includes a sequence encoding an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 58 or 59.
[0119] In one embodiment, the nucleic acid sequence encoding the NES includes a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to sequence number 137 or 138.
[0120] Organelle localization signals In one embodiment, the nucleic acid molecule comprises a nucleic acid sequence that encodes a localization signal that localizes a protein to an organelle or extracellular space. In one embodiment, the localization signal localizes the protein to the nucleolus, ribosome, vesicle, rough endoplasmic reticulum, Golgi apparatus, cytoskeleton, smooth endoplasmic reticulum, mitochondria, vacuole, cytosol, lysosome, or centrioles. Numerous localization signals are known in the art.
[0121] Examples of localization signals include, but are not limited to, 1x mitochondrial targeting sequences, 4x mitochondrial targeting sequences, secretory signaling sequences (IL-2), myristylation sequences, calcechestrin reader sequences, KDEL retention sequences, and peroxisome targeting sequences.
[0122] In one embodiment, the nucleic acid molecule includes a nucleic acid sequence encoding a localization signal. In one embodiment, the localization signal localizes the protein to an organelle or extracellularly. In one embodiment, the nucleic acid sequence encoding the localization signal includes at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, and at least 89% of the sequence encoding an amino acid sequence. In one embodiment, the nucleic acid molecule includes a nucleic acid sequence encoding a localization signal. In one embodiment, the localization signal localizes the protein to an organelle or extracellularly. In one embodiment, the nucleic acid sequence encoding the localization signal includes a sequence encoding an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 60 to 66.
[0123] In one embodiment, the nucleic acid sequence encoding the localization signal includes a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 139 to 145.
[0124] Purification and / or detection of tags In one embodiment, the nucleic acid molecule includes a nucleic acid sequence encoding a purification and / or detection tag. In one embodiment, the tag is located at the N-terminus of the protein. In one embodiment, the tag is a 3xFLAG tag. In one embodiment, the nucleic acid sequence encoding the purification and / or detection tag encodes an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 67. In one embodiment, the nucleic acid sequence encoding the purification and / or detection tag encodes the amino acid sequence of SEQ ID NO: 67.
[0125] In one embodiment, the nucleic acid sequence encoding the purified and / or detection tag includes a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to sequence number 146.
[0126] Cleavage and / or polyadenylated proteins In one embodiment, the nucleic acid molecule comprises a nucleic acid sequence encoding a cleaved and / or polyadenylated protein. In one embodiment, the cleaved and / or polyadenylated protein is an RNA-binding protein of the human 3' end processing mechanism. In one embodiment, the cleaved and / or polyadenylated protein is CPSF30, WDR33, or NUDT21. In one embodiment, the cleaved and / or polyadenylated protein is NUDT21. In one embodiment, the cleaved and / or polyadenylated protein is NUDT21, a NUDT21 variant, a NUDT21 dimer, a NUDT21 fusion protein, or any combination thereof. In one embodiment, the cleaved and / or polyadenylated protein is human NUDT21, helminth NUDT21, fly NUDT21, zebrafish NUDT21, NUDT21_R63S, NUDT21_F103A, or a tandem dimer of NUDT21.
[0127] In one embodiment, the nucleic acid sequence encoding the cleaved and / or polyadenylated protein encodes an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of SEQ ID NOs: 298-307.
[0128] In one embodiment, the nucleic acid sequence encoding the cleaved and / or polyadenylated protein includes sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to sequences 311 to 319.
[0129] Fusion protein A fusion protein containing Cas protein and localization signals. In one embodiment, the nucleic acid molecule comprises a nucleic acid sequence encoding a protein of the present disclosure that is effectively delivered to the nucleus, organelles, cytoplasm, or extracellular space, enabling targeted RNA cleavage. In one embodiment, the nucleic acid sequence encoding the protein encodes an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 68-100.
[0130] In one embodiment, the nucleic acid sequence encoding the protein contains a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 147 to 166.
[0131] A fusion protein comprising Cas protein and cleaved and / or polyadenylated proteins. In one embodiment, the nucleic acid molecule comprises a nucleic acid sequence encoding a protein of the Disclosure that enables RNA targeted cleavage and / or polyadenylation. In one embodiment, the nucleic acid sequence encoding the protein encodes an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of SEQ ID NOs. In one embodiment, the nucleic acid sequence encoding the protein encodes an amino acid sequence of one of SEQ ID NOs.
[0132] In one embodiment, the nucleic acid sequence encoding the protein includes a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 320 to 322.
[0133] Targeting nucleic acids and CRISPR RNA (crRNA) In one embodiment, the disclosure provides a CRISPR RNA (crRNA) for targeting Cas to a target RNA. In one embodiment, the crRNA includes a guide sequence. In one embodiment, the crRNA includes a direct repeat (DR) sequence. In one embodiment, the crRNA includes a direct repeat sequence and a guide sequence fused to or ligated to a guide sequence or spacer sequence. In one embodiment, the direct repeat sequence may be located upstream (i.e., 5') of the guide sequence or spacer sequence. In other embodiments, the direct repeat sequence may be located downstream (i.e., 3') of the guide sequence or spacer sequence.
[0134] In some embodiments, the crRNA includes a stem loop. In one embodiment, the crRNA includes a single stem loop. In one embodiment, the direct repeat sequence forms a stem loop. In one embodiment, the direct repeat sequence forms a single stem loop.
[0135] In one embodiment, the length of the guide RNA spacer is 15 to 35 nucleotides. In one embodiment, the length of the guide RNA spacer is at least 15 nucleotides. In one embodiment, the length of the spacer is 15 to 17 nucleotides, e.g., 15, 16, or 17 nucleotides; 17 to 20 nucleotides, e.g., 17, 18, 19, or 20 nucleotides; 20 to 24 nucleotides, e.g., 20, 21, 22, 23, or 24 nucleotides; 23 to 25 nucleotides, e.g., 23, 24, or 25 nucleotides; 24 to 27 nucleotides, e.g., 24, 25, 26, or 27 nucleotides; 27 to 30 nucleotides, e.g., 27, 28, 29, or 30 nucleotides; 30 to 35 nucleotides, e.g., 30, 31, 32, 33, 34, or 35 nucleotides; or 35 nucleotides or more.
[0136] Generally, the guide sequence is any polynucleotide sequence that hybridizes with the target sequence and has sufficient complementarity to the target polynucleotide sequence to induce sequence-specific binding of the CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence is about 50% or greater than 50%, about 60% or greater than 60%, about 75% or greater than 75%, about 80% or greater than 80%, about 85% or greater than 85%, about 90% or greater than 90%, about 95% or greater than 95%, about 97.5% or greater than 97.5%, about 99% or greater than 99%, or higher, when optimally aligned using a suitable alignment algorithm. The optimal alignment can be determined using any suitable algorithm for aligning sequences, and these non-limiting examples include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).In some embodiments, the guide array is approximately 5 or greater than approximately 5, approximately 10 or greater than approximately 10, approximately 11 or greater than approximately 11, approximately 12 or greater than approximately 12, approximately 13 or greater than approximately 13, approximately 14 or greater than approximately 14, approximately 15 or greater than approximately 15, approximately 16 or greater than approximately 16, approximately 17 or greater than approximately 17, approximately 18 or greater than approximately 18, approximately 19 or greater than approximately 19, approximately 20 or greater than approximately 20, approximately 21 or greater than approximately 21, approximately 22 or The nucleotide length is approximately greater than 22, approximately 23 or greater than 23, approximately 24 or greater than 24, approximately 25 or greater than 25, approximately 26 or greater than 26, approximately 27 or greater than 27, approximately 28 or greater than 28, approximately 29 or greater than 29, approximately 30 or greater than 30, approximately 35 or greater than 35, approximately 40 or greater than 40, approximately 45 or greater than 45, approximately 50 or greater than 50, approximately 75 or greater than 75, or greater than that. In some embodiments, the guide sequence has a nucleotide length of less than approximately 75, less than 50, less than 45, less than 40, less than 35, less than 30, less than 25, less than 20, less than 15, less than 12, or less. Preferably, the guide sequence is 10 to 30 nucleotides long. The ability of the guide sequence to induce sequence-specific binding of the CRISPR complex to the target sequence can be evaluated by any suitable assay. For example, sufficient CRISPR system components to form a CRISPR complex, including the guide sequence to be tested, may be provided to host cells having the corresponding target sequence by transduction using a vector encoding the components of the CRISPR sequence, and subsequently, selective cleavage within the target sequence is evaluated by a Surveyor assay as described herein. Similarly, components of a CRISPR complex, including the target sequence, the guide sequence to be tested, and a control guide sequence different from the guide sequence to be tested, can be provided, and in vitro cleavage of the target polynucleotide sequence can be evaluated by comparing the binding or cleavage rate within the target sequence between the reaction with the guide sequence to be tested and the reaction with the control guide sequence. Other assays are available and will be conceivable by those skilled in the art.
[0137] In some embodiments of the CRISPR-Cas system, the degree of complementarity between the guide sequence and its corresponding target sequence can be approximately 50% or greater than 50%, approximately 60% or greater than 60%, approximately 75% or greater than 75%, approximately 80% or greater than 80%, approximately 85% or greater than 85%, approximately 90% or greater than 90%, approximately 95% or greater than 95%, approximately 97.5% or greater than 97.5%, approximately 99% or greater than 99%, or 100%, and the guide RNA or sgRNA can be approximately 5 or greater than 5, approximately 10 or greater than 10, approximately 11 or greater than 11, approximately 12 or greater than 12, approximately 13 or greater than 13, approximately 14 or greater than 14, approximately 15 or greater than 15, approximately 16 or greater than 16, approximately 17 or greater than 17, approximately 18 or greater than 18, approximately 19 or greater than 19 The nucleotide length may be approximately 20 or greater than approximately 20, approximately 21 or greater than approximately 21, approximately 22 or greater than approximately 22, approximately 23 or greater than approximately 23, approximately 24 or greater than approximately 24, approximately 25 or greater than approximately 25, approximately 26 or greater than approximately 26, approximately 27 or greater than approximately 27, approximately 28 or greater than approximately 28, approximately 29 or greater than approximately 29, approximately 30 or greater than approximately 30, approximately 35 or greater than approximately 35, approximately 40 or greater than approximately 40, approximately 45 or greater than approximately 45, approximately 50 or greater than approximately 50, approximately 75 or greater than approximately 75, or greater than or equal to the nucleotide length of the guide RNA or sgRNA, which may be less than approximately 75, less than 50, less than 45, less than 40, less than 35, less than 30, less than 25, less than 20, less than 15, less than 12, or less, and favorably, the tracrRNA is 30 or 50 nucleotides long. However, aspects of the present disclosure involve reducing off-target interactions, for example, reducing guide sequences that interact with target sequences having low complementarity. In fact, the examples show that the present disclosure includes mutations that result in a CRISPR-Cas system that can distinguish between a target sequence and off-target sequences having more than 80% to about 95% complementarity, for example, 83% to 84% or 88% to 89% or 94% to 95% complementarity (for example, distinguishing between a target having 18 nucleotides and an off-target of 18 nucleotides having 1, 2, or 3 mismatches).Therefore, in the context of this disclosure, the degree of complementarity between the guide sequence and its corresponding target sequence is greater than 94.5%, greater than 95%, greater than 95.5%, greater than 96%, greater than 96.5%, greater than 97%, greater than 97.5%, greater than 98%, greater than 98.5%, greater than 99%, greater than 99.5%, greater than 99.9%, or 100%. Off-target sequences are those with complementarity between their sequence and the guide that is less than 100%, less than 99.9%, less than 99.5%, less than 99%, less than 98.5%, less than 98%, less than 97%, less than 96.5%, less than 96%, less than 95.5%, less than 94.5%, less than 93%, less than 92%, less than 91%, less than 90%, less than 89%, or 8 The off-target is less than 8% or less than 87% or less than 86% or less than 85% or less than 84% or less than 83% or less than 82% or less than 81% or less than 80%, and favorably, the off-target is less than 8% or less than 80
[0138] In one embodiment, the crRNA includes a sequence substantially complementary to the viral RNA sequence. In one embodiment, the crRNA includes a sequence substantially complementary to the coronavirus genome mRNA sequence or the coronavirus subgenome mRNA sequence. For example, in one embodiment, the crRNA includes a sequence substantially complementary to the coronavirus leader sequence, S sequence, E sequence, M sequence, N sequence, or S2M sequence. In one embodiment, the crRNA includes a sequence substantially complementary to the coronavirus leader sequence, N sequence, or S2M sequence.
[0139] In one embodiment, the crRNA contains a sequence substantially complementary to a sequence or fragment selected from SEQ ID NOs. 167-187 that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous.
[0140] In one embodiment, the crRNA contains a sequence substantially complementary to a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to a sequence fragment selected from sequence numbers 167, 175, or 182-185.
[0141] In one embodiment, the crRNA contains a sequence substantially complementary to a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to a sequence selected from SEQ ID NOs.
[0142] In one embodiment, the crRNA contains a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to a sequence selected from SEQ ID NOs.
[0143] In one embodiment, the disclosure provides a crRNA having a sequence substantially complementary to an influenza virus sequence. In one embodiment, the crRNA includes a sequence substantially complementary to a genomic mRNA sequence or subgenomic mRNA sequence of the influenza virus. For example, in one embodiment, the crRNA includes a sequence substantially complementary to the PB2, PB1, PA, HA, NP, NA, M, or NS sequence of the influenza virus. In one embodiment, the crRNA includes a sequence substantially complementary to the PB2, PB1, PA, NP, or M sequence of the influenza virus.
[0144] In one embodiment, the crRNA contains a sequence substantially complementary to a sequence or fragment selected from SEQ ID NOs. 225-244 that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous.
[0145] In one embodiment, the crRNA contains a sequence substantially complementary to the viral RNA sequence. In one embodiment, the crRNA contains a sequence substantially complementary to the positive-strand viral RNA sequence. In one embodiment, the crRNA contains a sequence substantially complementary to the negative-strand viral RNA sequence. In one embodiment, the crRNA contains a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to a sequence selected from SEQ ID NOs. In one embodiment, the crRNA includes a sequence selected from sequence numbers 245-264.
[0146] In one embodiment, the crRNA includes a direct repeat (DR) sequence. In one embodiment, the DR sequence is located at 5' of a sequence substantially complementary to the target sequence. For example, in one embodiment, the DR sequence is located at 5' of a sequence substantially complementary to a coronavirus genome mRNA sequence or a coronavirus subgenome mRNA sequence. In one embodiment, the DR sequence is located at 5' of a sequence substantially complementary to an influenza virus genome RNA sequence or an influenza virus subgenome RNA sequence. In one embodiment, the DR sequence is located at 5' of a sequence substantially complementary to an elongated RNA repeat sequence. In one embodiment, the DR sequence enhances the targeting activity of Cas13 to the target sequence, the catalytic activity of Cas13, or both. For example, in one embodiment, the DR sequence includes a mutation. For example, in one embodiment, the DR sequence includes a T17C point mutation. In one embodiment, the DR sequence includes a T18C point mutation. In one embodiment, the DR sequence is located at 5' of a sequence that is at least 80% homologous to a sequence selected from sequence numbers 189-224 and 245-264.
[0147] In one embodiment, the DR sequence is located at 3' of a sequence substantially complementary to the target sequence. For example, in one embodiment, the DR sequence is located at 3' of a sequence substantially complementary to a coronavirus genome mRNA sequence or a coronavirus subgenome mRNA sequence. In one embodiment, the DR sequence is located at 3' of a sequence substantially complementary to an influenza virus genome mRNA sequence or an influenza virus subgenome mRNA sequence. In one embodiment, the DR sequence is located at 3' of a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to a sequence selected from sequence numbers 189-224 and 245-264.
[0148] In one embodiment, the selection of the 5' or 3' DR sequence depends on the Cas protein orthologue used. In one embodiment, the DR sequence includes a sequence selected from SEQ ID NOs. 265-274.
[0149] Tandem array In one embodiment, the present invention provides a tandem CRISPR RNA (crRNA) sequence. In one embodiment, the tandem crRNA array allows a single promoter to drive the expression of multiple crRNAs. In one embodiment, the tandem array includes one or more, two or more, three or more, four or more, five or more, six or more, seven or more, or eight or more crRNA sequences.
[0150] In one embodiment, each crRNA in a tandem crRNA array includes a direct repeat (DR) sequence and a spacer sequence. In one embodiment, the direct repeat sequence may be located upstream (i.e., 5') of the guide sequence or spacer sequence. In other embodiments, the direct repeat sequence may be located downstream (i.e., 3') of the guide sequence or spacer sequence.
[0151] In one embodiment, the DR sequence is specific to the associated Cas protein. For example, in one embodiment, the Cas protein is Cas13, and the direct repeat sequence contains one of the sequences from SEQ ID NOs. 265 to 274. In one embodiment, the direct repeat sequence contains a single mutation in a poly-T stretch. For example, in one embodiment, the direct repeat sequence contains a sequence selected from SEQ ID NOs. 268 to 274.
[0152] In one embodiment, each crRNA in a tandem crRNA array contains a different direct repeat sequence. For example, in one embodiment, nucleotide substitutions within the loop region of the direct repeat, along with multiple guide RNAs, lead to the efficient generation of an ordered array of crRNAs.
[0153] In one embodiment, the tandem array includes at least two or more crRNAs, where each crRNA contains a sequence substantially complementary to the target RNA. For example, in one embodiment, each crRNA contains different sequences substantially complementary to different sequences within a single target RNA. In one embodiment, each crRNA contains different sequences substantially complementary to different sequences within different target RNAs. In one embodiment, the different target RNAs are associated with a single disease, disorder, or infection. For example, in one embodiment, the different target RNAs are each viral RNA sequence of a virus.
[0154] For example, in one embodiment, each crRNA contains a different sequence substantially complementary to a different sequence within the coronavirus genomic RNA sequence and / or coronavirus subgenomic RNA sequence. In one embodiment, the crRNA contains a sequence substantially complementary to the coronavirus genomic mRNA sequence or coronavirus subgenomic mRNA sequence. For example, in one embodiment, the crRNA contains a sequence substantially complementary to the coronavirus leader sequence, S sequence, E sequence, M sequence, N sequence, or S2M sequence. In one embodiment, the crRNA contains a sequence substantially complementary to the coronavirus leader sequence, N sequence, or S2M sequence.
[0155] In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to the coronavirus genomic RNA sequence and / or coronavirus subgenomic RNA sequence. In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to a sequence or fragment selected from SEQ ID NOs. 167-188 that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to that sequence or fragment. In one embodiment, the tandem array includes at least two crRNAs, each containing a sequence or fragment of a sequence selected from sequence numbers 167-188 that is substantially complementary to that sequence.
[0156] In one embodiment, the tandem array includes at least two crRNAs that are substantially complementary to sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous. In one embodiment, the tandem array includes at least two crRNAs, each containing a sequence substantially complementary to a sequence selected from sequence numbers 168-174, 176-181, 186, and 187.
[0157] In one embodiment, the tandem array includes at least two crRNAs containing sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to sequences selected from SEQ ID NOs.
[0158] In one embodiment, the tandem array contains sequences homologous to sequence number 275 by at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the tandem array contains the sequence of sequence number 275.
[0159] In one embodiment, the tandem array comprises at least two crRNAs, each independently containing a sequence substantially complementary to an influenza virus sequence. In one embodiment, the tandem array comprises at least two crRNAs, each containing a sequence substantially complementary to a genomic mRNA sequence or subgenomic mRNA sequence of the influenza virus. In one embodiment, the tandem array comprises at least two crRNAs, each containing a sequence substantially complementary to a PB2 sequence, PB1 sequence, PA sequence, HA sequence, NP sequence, NA sequence, M sequence, or NS sequence of the influenza virus.
[0160] In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to sequences or fragments of sequences selected from SEQ ID NOs. 225-244, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous.
[0161] In one embodiment, the tandem array includes at least two crRNAs, each containing a sequence that targets a different positive-strand vRNA segment 1, 2, 3, 5, or 7. In one embodiment, the tandem array includes at least two crRNAs, each containing a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to a sequence selected from SEQ ID NOs. In one embodiment, the tandem array includes at least two crRNAs containing sequences selected from SEQ ID NOs. 245-264. In one embodiment, the tandem array includes at least two crRNAs containing sequences homologous to at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the tandem array includes at least two crRNAs containing sequences selected from SEQ ID NOs. 255-259.
[0162] In one embodiment, the tandem array includes at least two crRNAs, each containing a sequence that targets a different minus-chain vRNA segment 1, 2, 3, 5, or 7. In one embodiment, the tandem array includes at least two crRNAs, each containing a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to a sequence selected from SEQ ID NOs. In one embodiment, the tandem array includes at least two crRNAs containing sequences selected from SEQ ID NOs. 245-249. In one embodiment, the tandem array includes at least two crRNAs containing sequences homologous to sequences selected from SEQ ID NOs. 260-264 by at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the tandem array includes at least two crRNAs containing sequences selected from SEQ ID NOs. 260-264.
[0163] In one embodiment, the tandem array includes at least two crRNAs containing sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to sequences selected from sequence numbers 250 to 254.
[0164] In one embodiment, the tandem array includes sequences homologous to at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the tandem array includes sequences 276–279.
[0165] In one embodiment, the tandem array contains at least two crRNAs, each independently containing a sequence substantially complementary to an mRNA sequence within a cellular pathway. In some embodiments, the cellular pathway results in dysregulation of cellular processes, such as uncontrolled proliferation and enhanced cell survival. For example, in one embodiment, each crRNA independently contains a sequence substantially complementary to an mRNA sequence within an intracellular RAS, JAK-STAT, PI3K / AKT ErbB, p53-mediated apoptosis, GSK3, Hippo, Wnt, estrogen, insulin, mTOR, NF-κB, Notch, TGF-β, Toll-like receptor, VEGF, AMPK, or MAPK signaling pathway.
[0166] In one embodiment, the tandem array contains at least two crRNAs, each independently containing sequences substantially complementary to mRNA sequences for use in biofuel production in microorganisms or plants, production of CAR-T cells in biotechnology, silencing stress response pathways, silencing inflammatory pathways in immune cells, and targeting multiple different types or subtypes of infectious viruses.
[0167] In one embodiment, the crRNA includes a direct repeat (DR) sequence. In one embodiment, the DR sequence is located at 5' of a sequence substantially complementary to the target sequence. For example, in one embodiment, the DR sequence is located at 5' of a sequence substantially complementary to a coronavirus genome mRNA sequence or a coronavirus subgenome mRNA sequence. In one embodiment, the DR sequence enhances the targeting activity of Cas13 to the target sequence, the Cas13 catalytic activity, or both. For example, in one embodiment, the DR sequence includes a mutation. For example, in one embodiment, the DR sequence includes a T17C point mutation. In one embodiment, the DR sequence includes a T18C point mutation. In one embodiment, the DR sequence includes a T17A point mutation. In one embodiment, the DR sequence includes a T18A point mutation. In one embodiment, the DR sequence includes a T17G point mutation. In one embodiment, the DR sequence includes a T18G point mutation. In one embodiment, the DR sequence is located at 5' of a sequence at least 80% homologous to a sequence selected from sequence numbers 189-224. In one embodiment, the DR sequence is located at 3' of a sequence substantially complementary to the target sequence. For example, in one embodiment, the DR sequence is located at 3' of a sequence substantially complementary to a coronavirus genome mRNA sequence or a coronavirus subgenome mRNA sequence. In one embodiment, the DR sequence is located at 3' of a sequence that is at least 80% homologous to a sequence selected from SEQ ID NOs. 189-224. In one embodiment, the selection of the 5' or 3' DR sequence depends on the Cas protein orthologue used. In one embodiment, the DR sequence includes a sequence selected from SEQ ID NOs. 265-274.
[0168] In one embodiment, the length of the guide RNA spacer is 15 to 35 nucleotides. In one embodiment, the length of the guide RNA spacer is at least 15 nucleotides. In one embodiment, the length of the spacer is 15 to 17 nucleotides, e.g., 15, 16, or 17 nucleotides; 17 to 20 nucleotides, e.g., 17, 18, 19, or 20 nucleotides; 20 to 24 nucleotides, e.g., 20, 21, 22, 23, or 24 nucleotides; 23 to 25 nucleotides, e.g., 23, 24, or 25 nucleotides; 24 to 27 nucleotides, e.g., 24, 25, 26, or 27 nucleotides; 27 to 30 nucleotides, e.g., 27, 28, 29, or 30 nucleotides; 30 to 35 nucleotides, e.g., 30, 31, 32, 33, 34, or 35 nucleotides; or 35 nucleotides or more.
[0169] Generally, the term “guide sequence,” used interchangeably with the term “spacer sequence,” is any polynucleotide sequence that hybridizes with the target sequence and has sufficient complementarity to the target polynucleotide sequence to induce sequence-specific binding of the CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence is about 50% or greater than 50%, about 60% or greater than 60%, about 75% or greater than 75%, about 80% or greater than 80%, about 85% or greater than 85%, about 90% or greater than 90%, about 95% or greater than 95%, about 97.5% or greater than 97.5%, about 99% or greater than 99%, or higher, when optimally aligned using a preferred alignment algorithm. The optimal alignment can be determined using any suitable algorithm for aligning sequences, and these non-limiting examples include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).In some embodiments, the guide array is approximately 5 or greater than approximately 5, approximately 10 or greater than approximately 10, approximately 11 or greater than approximately 11, approximately 12 or greater than approximately 12, approximately 13 or greater than approximately 13, approximately 14 or greater than approximately 14, approximately 15 or greater than approximately 15, approximately 16 or greater than approximately 16, approximately 17 or greater than approximately 17, approximately 18 or greater than approximately 18, approximately 19 or greater than approximately 19, approximately 20 or greater than approximately 20, approximately 21 or greater than approximately 21, approximately 22 or The nucleotide length is approximately greater than 22, approximately 23 or greater than 23, approximately 24 or greater than 24, approximately 25 or greater than 25, approximately 26 or greater than 26, approximately 27 or greater than 27, approximately 28 or greater than 28, approximately 29 or greater than 29, approximately 30 or greater than 30, approximately 35 or greater than 35, approximately 40 or greater than 40, approximately 45 or greater than 45, approximately 50 or greater than 50, approximately 75 or greater than 75, or greater than that. In some embodiments, the guide sequence has a nucleotide length of less than approximately 75, less than 50, less than 45, less than 40, less than 35, less than 30, less than 25, less than 20, less than 15, less than 12, or less. Preferably, the guide sequence is 10 to 30 nucleotides long. The ability of the guide sequence to induce sequence-specific binding of the CRISPR complex to the target sequence can be evaluated by any suitable assay. For example, sufficient CRISPR system components to form a CRISPR complex, including the guide sequence to be tested, may be provided to host cells having the corresponding target sequence by transduction using a vector encoding the components of the CRISPR sequence, and subsequently, selective cleavage within the target sequence is evaluated by a Surveyor assay as described herein. Similarly, components of a CRISPR complex, including the target sequence, the guide sequence to be tested, and a control guide sequence different from the guide sequence to be tested, can be provided, and in vitro cleavage of the target polynucleotide sequence can be evaluated by comparing the binding or cleavage rate within the target sequence between the reaction with the guide sequence to be tested and the reaction with the control guide sequence. Other assays are available and will be conceivable by those skilled in the art.
[0170] In some embodiments of the CRISPR-Cas system, the degree of complementarity between the guide sequence and its corresponding target sequence can be approximately 50% or greater than 50%, approximately 60% or greater than 60%, approximately 75% or greater than 75%, approximately 80% or greater than 80%, approximately 85% or greater than 85%, approximately 90% or greater than 90%, approximately 95% or greater than 95%, approximately 97.5% or greater than 97.5%, approximately 99% or greater than 99%, or 100%, and the guide RNA or sgRNA can be approximately 5 or greater than 5, approximately 10 or greater than 10, approximately 11 or greater than 11, approximately 12 or greater than 12, approximately 13 or greater than 13, approximately 14 or greater than 14, approximately 15 or greater than 15, approximately 16 or greater than 16, approximately 17 or greater than 17, approximately 18 or greater than 18, approximately 19 or greater than 19 The nucleotide length may be approximately 20 or greater than approximately 20, approximately 21 or greater than approximately 21, approximately 22 or greater than approximately 22, approximately 23 or greater than approximately 23, approximately 24 or greater than approximately 24, approximately 25 or greater than approximately 25, approximately 26 or greater than approximately 26, approximately 27 or greater than approximately 27, approximately 28 or greater than approximately 28, approximately 29 or greater than approximately 29, approximately 30 or greater than approximately 30, approximately 35 or greater than approximately 35, approximately 40 or greater than approximately 40, approximately 45 or greater than approximately 45, approximately 50 or greater than approximately 50, approximately 75 or greater than approximately 75, or greater than or equal to the nucleotide length of the guide RNA or sgRNA. However, aspects of the present invention involve reducing off-target interactions, for example, reducing guide sequences that interact with target sequences having low complementarity. In fact, examples show that the present invention includes mutations that result in a CRISPR-Cas system that can distinguish between a target sequence and off-target sequences having more than 80% to about 95% complementarity, for example, 83% to 84% or 88% to 89% or 94% to 95% complementarity (for example, distinguishing between a target having 18 nucleotides and an off-target of 18 nucleotides having 1, 2, or 3 mismatches).Therefore, in the context of the present invention, the degree of complementarity between the guide sequence and its corresponding target sequence is greater than 94.5%, greater than 95%, greater than 95.5%, greater than 96%, greater than 96.5%, greater than 97%, greater than 97.5%, greater than 98%, greater than 98.5%, greater than 99%, greater than 99.5%, greater than 99.9%, or 100%. Off-target sequences are those with complementarity between their sequence and the guide that is less than 100%, less than 99.9%, less than 99.5%, less than 99%, less than 98.5%, less than 98%, less than 97%, less than 96.5%, less than 96%, less than 95.5%, less than 94.5%, less than 93%, less than 92%, less than 91%, less than 90%, less than 89%, or 8 The off-target is less than 8% or less than 87% or less than 86% or less than 85% or less than 84% or less than 83% or less than 82% or less than 81% or less than 80%, and favorably, the off-target is less than 8% or less than 80
[0171] nucleic acid The isolated nucleic acid sequences of this disclosure may be obtained using any of the numerous recombination methods known in the art, for example, by standard techniques, by screening libraries derived from cells expressing the gene, by isolating the gene from vectors known to contain the gene, or by directly isolating the gene from cells and tissues containing the gene. Alternatively, the gene of interest may be generated by synthesis rather than cloning.
[0172] The isolated nucleic acids may include any type of nucleic acid, but are not limited to DNA and RNA. For example, in one embodiment, the composition includes an isolated DNA molecule, for example, an isolated cDNA molecule encoding the protein of the disclosure. In one embodiment, the composition includes an isolated RNA molecule encoding the protein of the disclosure or a functional fragment thereof.
[0173] The nucleic acid molecules of the present invention may be modified to improve their stability in serum or growth media for cell culture. Modifications may be added to enhance the stability, functionality, and / or specificity of the nucleic acid molecules of the present invention and to minimize immunostimulatory effects. For example, to enhance stability, 3' residues may be stabilized against degradation. For example, they may be selected to consist of purine nucleotides, particularly adenosine or guanosine nucleotides. Alternatively, substitution of pyrimidine nucleotides with modified analogs, such as substitution of uridine with 2'-deoxythymidine, is acceptable and does not affect the function of the molecule.
[0174] In one embodiment of the present invention, a nucleic acid molecule may contain at least one modified nucleotide analog. For example, the terminus may be stabilized by incorporating a modified nucleotide analog. Non-limiting examples of nucleotide analogs include ribonucleotides with modified sugars and / or backbones (i.e., modifications to the phosphate-sugar backbone). For example, the phosphodiester bond in native RNA can be modified to include at least one of nitrogen or sulfur heteroatoms. In exemplary backbone-modified ribonucleotides, the phosphate ester group attached to an adjacent ribonucleotide is replaced by a modifying group, such as a phosphorothioate group. In exemplary sugar-modified ribonucleotides, the 2'OH group is replaced by a group selected from H, OR, R, halo, SH, SR, NH2, NHR, NR2, or ON, where R is a C1-C6 alkyl, alkenyl, or alkynyl group, and halo is F, Cl, Br, or I.
[0175] Other examples of modification include ribonucleotides with modified nucleic acid bases, i.e., ribonucleotides containing at least one non-natural nucleic acid base in place of a naturally occurring nucleic acid base. The bases may be modified to inhibit the activity of adenosine deaminase. Exemplary modified nucleic acid bases include, but are not limited to, uridine and / or cytidine modified at position 5, e.g., 5-(2-amino)propyluridine, 5-bromouridine; adenosine and / or guanosine modified at position 8, e.g., 8-bromoguanosine; deazanucleotides, e.g., 7-deaza-adenosine; and O-alkylated and N-alkylated nucleotides, e.g., N6-methyladenosine, which are preferred. Note that the above modifications may be combined.
[0176] In some cases, the nucleic acid molecule may include at least one of the following chemical modifications: 2'-H, 2'-O-methyl, or 2'-OH modification of one or more nucleotides. In certain embodiments, the nucleic acid molecule of the present invention may have enhanced resistance to nucleases. For increased nuclease resistance, the nucleic acid molecule may include, for example, a 2'-modified ribose unit and / or a phosphorothioate bond. For example, the 2'-hydroxyl group (OH) may be modified or replaced with a number of different "oxy" or "deoxy" substituents. For increased nuclease resistance, the nucleic acid molecule of the present invention may include 2'-O-methyl, 2'-fluorine, 2'-O-methoxyethyl, 2'-O-aminopropyl, 2'-amino, and / or a phosphorothioate bond. Inclusion of locked nucleic acids (LNA), ethylene nucleic acids (ENA), such as 2'-4'-ethylene-bridged nucleic acids, and specific nucleic acid base modifications, such as 2-amino-A modification, 2-thio (e.g., 2-thio-U) modification, and G-clamp modification, can also increase binding affinity to the target.
[0177] In one embodiment, the nucleic acid molecule comprises 2'-modified nucleotides, such as 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-methyl, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), or 2'-ON-methylacetamide (2'-O-NMA). In one embodiment, the nucleic acid molecule comprises at least one 2'-O-methyl modified nucleotide, and in some embodiments, all nucleotides of the nucleic acid molecule contain 2'-O-methyl modification.
[0178] In certain embodiments, the nucleic acid molecule of the present invention has one or more of the following characteristics.
[0179] The nucleic acid agents described herein include RNA and DNA that are otherwise unmodified, as well as RNA and DNA that have been modified, for example, to improve efficacy, and polymers of nucleoside substitutes. Unmodified RNA refers to molecules in which the components of nucleic acid, namely sugars, bases, and phosphate groups, are the same as, or essentially the same as, those that occur naturally or in the human body. In the art, RNA that is rare or unusual but occurs naturally is considered modified RNA. See, for example, Limbach et al. (Nucleic Acids Res., 1994, 22:2183-2196). Such rare or unusual RNA, often referred to as modified RNA, is typically the result of post-transcriptional modification and falls within the scope of the term unmodified RNA as used herein. As used herein, modified RNA refers to molecules in which one or more of the components of nucleic acid, namely sugars, bases, and phosphate groups, differ from those that occur naturally or in the human body. These are called "modified RNAs," but naturally, due to the modifications, they also include molecules that are not strictly RNA. Nucleoside substitutes are molecules in which the ribose phosphate skeleton is replaced by a non-ribose phosphate structure, for example, an uncharged mimic of the ribose phosphate skeleton, which allows the bases to be presented in the correct spatial relationships so that hybridization is substantially similar to that seen in the ribose phosphate skeleton.
[0180] The nucleic acid modification of the present invention may be present in one or more of the following: a phosphate group, a sugar group, a skeleton, the N-terminus, the C-terminus, or a nucleic acid base.
[0181] The present invention also includes vectors into which the isolated nucleic acids of the present invention are inserted. A wealth of suitable vectors useful in the present invention exist in the art.
[0182] In short, the expression of native or synthetic nucleic acids encoding the proteins of the Disclosure is typically achieved by operably ligating the nucleic acid encoding the protein or a portion thereof to a promoter and incorporating the construct into an expression vector. The vector used is suitable for incorporation due to replication and, optionally, in eukaryotic cells. Typical vectors contain transcriptional and translational terminators, start sequences, and promoters useful for regulating the expression of the desired nucleic acid sequence.
[0183] The vectors of the present invention may also be used in nucleic acid immunoassays and gene therapies using standard gene delivery protocols. Methods for gene delivery are known in the art. See, for example, U.S. Patents 5,399,346, 5,580,859 and 5,589,466 (all of which are incorporated herein by reference). In another embodiment, the present invention provides gene therapy vectors.
[0184] The isolated nucleic acids of the present invention can be cloned into a number of types of vectors. For example, nucleic acids can be cloned into vectors that include, but are not limited to, plasmids, phagemids, phage derivatives, animal viruses, and cosmids. Vectors of particular interest include expression vectors, replication vectors, probe generation vectors, and sequencing vectors.
[0185] Furthermore, vectors can be supplied to cells in the form of viral vectors. Viral vector technology is well known in the art and is described, for example, in Sambrook et al. (2012, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York), as well as in other virology and molecular biology manuals. Useful viruses as vectors include, but are not limited to, retroviruses, adenoviruses, adeno-associated viruses, herpesviruses, and lentiviruses. Generally, a suitable vector contains a replication origin, promoter sequence, useful restriction endonuclease site, and one or more selection markers that function in at least one organism (e.g., WO01 / 96584; WO01 / 29058; and U.S. Patent No. 6,326,193).
[0186] Delivery system and delivery method In one embodiment, the disclosure relates to the development of novel lentiviral packaging and delivery systems. Lentiviral particles deliver viral enzymes as proteins. In this mode, the short lifespan of lentiviral enzymes limits the possibility of off-target editing due to long-term expression throughout the cell's lifespan. Accordingly, in one embodiment, the disclosure provides a novel delivery system for delivering genes or genetic material.
[0187] Incorporating editing components or conventional CRISPR-Cas editing components as proteins into lentiviral particles is advantageous, given that their required activity is needed only for a short period. Accordingly, in one embodiment, the present disclosure provides a lentiviral delivery system, as well as a method for delivering the composition of the present invention using the lentiviral delivery system, a method for editing genetic material, and a method for delivering nucleic acids.
[0188] In one embodiment, the delivery system comprises (1) a packaging plasmid, (2) a transdermal plasmid, and (3) an envelope plasmid. In one embodiment, the delivery system comprises (1) a packaging plasmid, (2) an envelope plasmid, and (3) a VPR plasmid. In one embodiment, the packaging plasmid comprises a nucleic acid sequence encoding a gag-pol polyprotein. In one embodiment, the gag-pol polyprotein comprises a catalytically dead integrase. In one embodiment, the gag-pol polyprotein comprises mutations selected from D116N, D116A, D116E, D64V, D64E, and D64A.
[0189] In one embodiment, the transplasmid comprises a nucleic acid sequence encoding the crRNA sequence and the Cas protein of the Disclosure. For example, in one embodiment, the transplasmid comprises a nucleic acid sequence encoding the protein of the Disclosure, which includes a crRNA sequence and the Cas protein. In one embodiment, the transplasmid comprises a nucleic acid sequence encoding the protein of the Disclosure, which includes a crRNA sequence and the Cas protein and a localization signal. In one embodiment, the transplasmid comprises a nucleic acid sequence encoding the protein of the Disclosure, which includes a crRNA sequence and the Cas protein and the NLS, NES, or other localization signal.
[0190] For example, in one embodiment, the introduced plasmid includes a nucleic acid sequence encoding the tandem crRNA array of the Disclosure. In one embodiment, the tandem array includes one or more, two or more, three or more, four or more, five or more, six or more, seven or more, or eight or more crRNA sequences. In one embodiment, each crRNA in the tandem crRNA array includes a direct repeat (DR) sequence and a spacer sequence. In one embodiment, the direct repeat sequence may be located upstream (i.e., 5') of the guide sequence or spacer sequence. In other embodiments, the direct repeat sequence may be located downstream (i.e., 3') of the guide sequence or spacer sequence.
[0191] In one embodiment, each crRNA in a tandem crRNA array contains a different direct repeat sequence. For example, in one embodiment, nucleotide substitutions within the loop region of the direct repeat, along with multiple guide RNAs, lead to the efficient generation of an ordered array of crRNAs.
[0192] In one embodiment, the tandem array includes at least two or more crRNAs, where each crRNA contains a sequence substantially complementary to the target RNA. For example, in one embodiment, each crRNA contains different sequences substantially complementary to different sequences within a single target RNA. In one embodiment, each crRNA contains different sequences substantially complementary to different sequences within different target RNAs. In one embodiment, the different target RNAs are associated with a single disease, disorder, or infection. For example, in one embodiment, the different target RNAs are each viral RNA sequence of a virus.
[0193] In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to the coronavirus genomic RNA sequence and / or coronavirus subgenomic RNA sequence. In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to a sequence or fragment selected from SEQ ID NOs. 167-188 that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to that sequence or fragment. In one embodiment, the tandem array includes at least two crRNAs, each containing a sequence or fragment of a sequence selected from sequence numbers 167-188 that is substantially complementary to that sequence.
[0194] In one embodiment, the tandem array includes at least two crRNAs containing sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to sequences selected from sequence numbers 189 to 22.
[0195] In one embodiment, the tandem array contains sequences homologous to sequence number 275 by at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the tandem array contains the sequence of sequence number 275.
[0196] In one embodiment, the tandem array comprises at least two crRNAs, each independently containing a sequence substantially complementary to an influenza virus sequence. In one embodiment, the tandem array comprises at least two crRNAs, each containing a sequence substantially complementary to a genomic mRNA sequence or subgenomic mRNA sequence of the influenza virus. In one embodiment, the tandem array comprises at least two crRNAs, each containing a sequence substantially complementary to a PB2 sequence, PB1 sequence, PA sequence, HA sequence, NP sequence, NA sequence, M sequence, or NS sequence of the influenza virus.
[0197] In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to sequences or fragments of sequences selected from SEQ ID NOs. 225-244, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous.
[0198] In one embodiment, the tandem array includes at least two crRNAs containing sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to sequences selected from sequence numbers 245 to 249.
[0199] In one embodiment, the tandem array includes sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous. In one embodiment, the tandem array includes sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, or at least 99% homologous.
[0200] In one embodiment, the tandem array contains at least two crRNAs, each independently containing a sequence substantially complementary to an mRNA sequence within a cellular pathway. In some embodiments, the cellular pathway results in dysregulation of cellular processes, such as enhanced uncontrolled proliferation and cell survival. For example, in one embodiment, each crRNA independently contains a sequence substantially complementary to an mRNA sequence within the RAS, JAK-STAT, PI3K / AKT, or MAPK cellular pathway.
[0201] In one embodiment, the nucleic acid sequence encoding the Cas protein includes a sequence encoding an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to one of SEQ ID NOs. In one embodiment, the nucleic acid sequence encoding the Cas protein includes a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to one of sequence numbers 132-133 or 14-166.
[0202] In one embodiment, the introduced plasmid contains the sequence of sequence number 283.
[0203] In one embodiment, the envelope plasmid includes a nucleic acid sequence encoding an envelope protein. In one embodiment, the envelope protein may be selected based on a desired cell type. In one embodiment, the envelope plasmid includes a nucleic acid sequence encoding an HIV envelope protein. In one embodiment, the envelope plasmid includes a nucleic acid sequence encoding a vesicular stomatitis virus g protein (VSV-g) envelope protein. In one embodiment, the envelope plasmid contains a nucleic acid sequence encoding an amino acid sequence homologous to SEQ ID NO: 130 by at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the envelope plasmid contains a nucleic acid sequence encoding the amino acid sequence of SEQ ID NO: 130.
[0204] In one embodiment, the envelope plasmid includes a nucleic acid sequence encoding a coronavirus spike protein or a protein derived from the coronavirus spike protein. For example, in one embodiment, the envelope plasmid includes a nucleic acid sequence encoding an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to one of SEQ ID NOs.
[0205] In one embodiment, coronavirus-derived viral envelope proteins are not efficient at pseudotyping lentiviral vectors. Therefore, in one embodiment, the disclosure also provides novel coronavirus envelope proteins for use in pseudotyping lentiviral vectors. In one embodiment, the coronavirus envelope protein comprises an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to one of SEQ ID NOs. In one embodiment, the coronavirus envelope protein contains one amino acid sequence from sequence numbers 101 to 129.
[0206] In one embodiment, the VPR plasmid comprises a nucleic acid sequence encoding a fusion protein containing VPR and the Cas protein of the Disclosure.
[0207] In one embodiment, a packaging plasmid, a transdermal plasmid, and an envelope plasmid are introduced into a cell. In one embodiment, the cell transcribes and translates the nucleic acid sequence encoding the gag-pol protein encoded by the packaging plasmid to produce a gag-pol polyprotein. In one embodiment, the cell transcribes and translates the nucleic acid sequence encoding the envelope protein of the envelope plasmid to produce an envelope protein. In one embodiment, the cell transcribes the nucleic acid sequence encoding the crRNA sequence or crRNA array of the transdermal plasmid to produce a crRNA or crRNA array. In one embodiment, the cell transcribes and translates the nucleic acid sequence encoding the Cas protein of the transdermal plasmid to produce a Cas protein or a Cas fusion protein.
[0208] In one embodiment, the transcribed introduce plasmid and gag-pol protein are packaged into a lentiviral vector. In one embodiment, the lentiviral vector is collected from cell culture medium. In one embodiment, the viral particles are transduced into target cells, where the transcribed crRNA and Cas protein are cleaved and translated, thereby producing Cas protein and crRNA. The crRNA binds to the Cas protein, inducing it to an RNA having a sequence substantially complementary to the crRNA sequence.
[0209] In one embodiment, a packaging plasmid, a transdermal plasmid, and an envelope plasmid are introduced into a cell. In one embodiment, the cell transcribes and translates the nucleic acid sequence encoding the gag-pol protein encoded by the packaging plasmid to produce a gag-pol polyprotein. In one embodiment, the cell transcribes and translates the nucleic acid sequence encoding the envelope protein of the envelope plasmid to produce an envelope protein. In one embodiment, the cell transcribes the nucleic acid sequence encoding the gene to produce a gene. In one embodiment, the cell transcribes and translates the nucleic acid sequence encoding the gene of the transdermal plasmid to produce a protein.
[0210] In one embodiment, the transcribed introduce plasmid and gag-pol protein are packaged into a lentiviral vector. In one embodiment, the lentiviral vector is collected from cell culture medium. In one embodiment, the viral particles are transduced into target cells, where the transcribed gene is delivered to the cells and inserted into their genome.
[0211] In one embodiment, the transcribed introduce plasmid and gag-pol protein are packaged into a lentiviral vector. In one embodiment, the lentiviral vector is collected from the cell culture medium. In one embodiment, the viral particles are transduced into target cells, where the transcribed and translated genes are delivered to the cells.
[0212] In one embodiment, the gene or protein is delivered to a respiratory cell type, angiocyte type, renal cell type, or cardiovascular cell type. Thus, in one embodiment, the envelope protein is derived from a coronavirus. In one embodiment, the coronavirus envelope protein contains an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to one of SEQ ID NOs. In one embodiment, the coronavirus envelope protein contains one amino acid sequence from sequence numbers 101 to 129.
[0213] Furthermore, numerous additional virus-based systems have been developed for gene delivery into mammalian cells. For example, retroviruses provide a convenient platform for gene delivery systems. Selected genes can be inserted into vectors using techniques known in the art and packaged into retroviral particles. Recombinant viruses can then be isolated and delivered to target cells either in vivo or ex vivo. Numerous retroviral systems are known in the art. In some embodiments, adenovirus vectors are used. Numerous adenovirus vectors are known in the art. In one embodiment, lentiviral vectors are used.
[0214] For example, retrovirus-derived vectors, such as lentiviruses, are suitable tools for achieving long-term gene transfer because they enable the long-term, stable integration of the introduced gene and its transmission in daughter cells. Lentiviral vectors have an additional advantage over vectors derived from oncoretroviruses, such as mouse leukemia virus, in that they can transduce non-proliferating cells, such as hepatocytes. They also have the additional advantage of low immunogenicity.
[0215] In one embodiment, the composition comprises a vector derived from adeno-associated virus (AAV). The term “AAV vector” means a vector derived from adeno-associated virus serotypes, including, but not limited to, AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, and AAV-9. AAV vectors have become a powerful gene delivery tool for treating a variety of disorders. AAV vectors possess several characteristics that make them ideally suited for gene therapy, including non-pathogenicity, minimal immunogenicity, and the ability to stably and efficiently transduce cells that have completed cell division. The expression of specific genes contained within an AAV vector can be specifically targeted to one or more types of cells by selecting an appropriate combination of AAV serotype, promoter, and delivery method.
[0216] For example, in one embodiment, the AAV vector includes a crRNA having a sequence substantially complementary to the coronavirus genome mRNA sequence or the coronavirus subgenome mRNA sequence. In one embodiment, the AAV vector includes a crRNA array containing two or more crRNAs having sequences substantially complementary to the coronavirus genome mRNA sequence or the coronavirus subgenome mRNA sequence. In one embodiment, the AAV vector includes a sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to SEQ ID NO: 284. In one embodiment, the introduced plasmid contains the sequence of sequence number 284.
[0217] In one embodiment, the AAV vector comprises a crRNA having a sequence substantially complementary to the influenza virus genomic RNA sequence or the influenza virus subgenomic RNA sequence. In one embodiment, the transplasm comprises a crRNA array containing two or more crRNAs having sequences substantially complementary to the influenza virus genomic RNA sequence or the influenza virus subgenomic RNA sequence.
[0218] AAV vectors may have deletions of one or more AAV wild-type genes, preferably all or part of the rep and / or cap genes, but may retain functional adjacent ITR sequences. Despite high homology, different serotypes have different tissue orientations. The receptor for AAV1 is unknown, but AAV1 is known to transduce skeletal muscle and cardiac muscle more efficiently than AAV2. Most studies have been performed using pseudotype vectors in which vector DNA flanked by AAV2 ITRs is packaged in a capsid of an alternative serotype, so it is clear that the biological differences are related to the capsid rather than the genome. Recent evidence shows that DNA expression cassettes packaged in an AAV1 capsid are at least 1 log10 more efficient in transduction of cardiomyocytes than those packaged in an AAV2 capsid. In one embodiment, the viral delivery system is an adeno-associated virus delivery system. Adeno-associated viruses may be serotype 1 (AAV1), serotype 2 (AAV2), serotype 3 (AAV3), serotype 4 (AAV4), serotype 5 (AAV5), serotype 6 (AAV6), serotype 7 (AAV7), serotype 8 (AAV8), or serotype 9 (AAV9).
[0219] Desired AAV fragments for assembly into vectors include cap proteins containing vp1, vp2, vp3, and the hypervariable region, rep proteins containing rep78, rep68, rep52, and rep40, and sequences encoding these proteins. These fragments can be readily used in various vector systems and host cells. Such fragments can be used alone, in combination with sequences or fragments of other AAV serotypes, or in combination with elements derived from other AAV or non-AAV viral sequences. As used herein, artificial AAV serotypes include, but are not limited to, AAVs having non-natural capsid proteins. Such artificial capsids can be generated by any preferred technique using selected AAV sequences (e.g., fragments of the vp1 capsid protein) in combination with heterologous sequences that may be obtained from selected different AAV serotypes, non-adjacent portions of the same AAV serotype, non-AAV viral sources, or non-viral sources. Artificial AAV serotypes may include, but are not limited to, chimeric AAV capsids, recombinant AAV capsids, or "humanized" AAV capsids. Therefore, exemplary AAVs or artificial AAVs suitable for the expression of one or more proteins include, in particular, AAV2 / 8 (see U.S. Patent No. 7,282,199), AAV2 / 5 (available from the National Institutes of Health), AAV2 / 9 (International Patent Publication WO2005 / 033321), AAV2 / 6 (U.S. Patent No. 6,156,303), and AAVrh8 (International Patent Publication WO2003 / 042397).
[0220] In certain embodiments, the vector also includes conventional regulatory elements operably ligated to the transgene in a manner that enables transcription, translation, and / or expression of the transgene in cells transfected with the plasmid vector produced by the present invention or in cells infected with the virus produced by the present invention. As used herein, “operably ligated” sequences include both expression regulatory sequences adjacent to the gene of interest and expression regulatory sequences that act trans or at a distance to control the gene of interest. Expression regulatory sequences include appropriate transcription start sequences, transcription termination sequences, promoter sequences, and enhancer sequences; efficient RNA processing signals such as splicing signals and polyadenylation (poly-A) signals; sequences that stabilize cytoplasmic mRNA; sequences that enhance translation efficiency (i.e., Kozak consensus sequences); sequences that enhance protein stability; and, if necessary, sequences that enhance the secretion of the encoded product. Numerous expression regulatory sequences, including innate promoters, constitutive promoters, inducible promoters, and / or tissue-specific promoters, are known and available in the art.
[0221] Additional promoter elements, such as enhancers, regulate the frequency of transcription initiation. Typically, these are located 30–110 bp upstream of the initiation site, but in recent years, many promoters have been shown to also contain functional elements downstream of the initiation site. Often, the spacing between promoter elements is flexible, and as a result, promoter function is maintained even if the elements are inverted or moved relative to each other. In thymidine kinase (TK) promoters, the spacing between promoter elements can be as large as 50 bp without causing a decrease in activity. In some promoters, individual elements appear to be able to activate transcription cooperatively or independently.
[0222] One example of a suitable promoter is the cytomegalovirus (CMV) initial promoter sequence. This promoter sequence is a potent constitutive promoter sequence that can activate high levels of expression of any functionally linked polynucleotide sequence. Another example of a suitable promoter is elongation growth factor-1α (EF-1α). However, other constitutive promoter sequences may also be used, including, but not limited to, the Simian virus 40 (SV40) initial promoter, mouse mammary cancer virus (MMTV), human immunodeficiency virus (HIV) terminal repeat sequence (LTR) promoter, MoMuLV promoter, avian leukemia virus promoter, Epstein-Barr virus initial promoter, Roussarcoma virus promoter, and human gene promoters such as actin promoters, myosin promoters, hemoglobin promoters, and creatine kinase promoters. Furthermore, the present invention should not be limited to the use of constitutive promoters. Inducible promoters are also intended as part of the present invention. The use of inductive promoters provides a molecular switch that can turn on the expression of a functionally linked polynucleotide sequence when such expression is desired, or turn it off when expression is not desired. Examples of inductive promoters include, but are not limited to, metallothionein promoters, glucocorticoid promoters, progesterone promoters, and tetracycline promoters.
[0223] Enhancer sequences found on a vector also regulate the expression of genes contained within the vector. Typically, enhancers bind to protein factors to enhance gene transcription. Enhancers may be located upstream or downstream of the gene they regulate. Enhancers may also be tissue-specific to enhance transcription in a particular cell or tissue type. In one embodiment, the vector of the present invention includes one or more enhancers to promote the transcription of genes present within the vector.
[0224] To evaluate the expression of the fusion protein of the present invention, the expression vector introduced into cells may also contain either a selection marker gene or a reporter gene, or both, to facilitate the identification and selection of expressing cells from a cell population to be transfected or infected with the viral vector. In other embodiments, the selection marker may be contained in a separate DNA fragment and used in a co-transfection procedure. Both the selection marker gene and the reporter gene may be adjacent to appropriate regulatory sequences to enable expression in host cells. Useful selection markers include, for example, antibiotic resistance genes such as neo.
[0225] Reporter genes are used to identify potentially transfected cells and to evaluate the function of regulatory sequences. Generally, a reporter gene is a gene that is not present in or expressed by the recipient organism or tissue, and is a polypeptide whose expression is revealed by several readily detectable characteristics, such as enzymatic activity. Reporter gene expression is tested at a suitable time after the DNA has been introduced into the recipient cells. Suitable reporter genes may include genes encoding luciferase, β-galactosidase, chloramphenicol acetyltransferase, secreted alkaline phosphatase, or green fluorescent protein (e.g., Ui-Tei et al., 2000 FEBS Letters 479:79-82). Suitable expression systems are well known and can be prepared using known techniques or are commercially available. Generally, a construct with a minimal 5' facile region exhibiting the highest expression level of the reporter gene is identified as a promoter. Such a promoter region can be ligated to the reporter gene and used to evaluate the action of the agent for its ability to regulate promoter-driven transcription.
[0226] Methods for introducing and expressing genes in cells are known in the art. In relation to expression vectors, vectors can be readily introduced into host cells, such as mammalian cells, bacterial cells, yeast cells, or insect cells, by any method in the art. For example, expression vectors can be introduced into host cells by physical, chemical, or biological means.
[0227] Physical methods for introducing polynucleotides into host cells include calcium phosphate precipitation, lipofection, particulate gun techniques, microinjection, and electroporation. Methods for preparing cells containing vectors and / or exogenous nucleic acids are well known in the art. See, for example, Sambrook et al. (2012, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York). An exemplary method for introducing polynucleotides into host cells is calcium phosphate transfection.
[0228] Biological methods for introducing polynucleotides of interest into host cells include the use of DNA and RNA vectors. Viral vectors, particularly retroviral vectors, are the most widely used method for inserting genes into mammalian cells, such as human cells. Other viral vectors may be derived from lentiviruses, poxviruses, herpes simplex virus type 1, adenoviruses, and adeno-associated viruses. See, for example, U.S. Patents 5,350,674 and 5,585,362.
[0229] Chemical means for introducing polynucleotides into host cells include colloidal dispersions, such as polymer complexes, nanocapsules, microspheres, beads, and lipid-based systems, which include oil-in-water emulsions, micelles, mixed micelles, and liposomes. An exemplary colloidal system for use as a delivery medium in vitro and in vivo is liposomes (e.g., artificial membrane vesicles).
[0230] When nonviral delivery systems are used, exemplary delivery media are liposomes. Lipid formulations are intended for the (in vitro, ex vivo, or in vivo) delivery of nucleic acids to host cells. In another embodiment, nucleic acids may be associated with lipids. Lipid-associated nucleic acids may be encapsulated within the aqueous interior of liposomes, dispersed within the lipid bilayer of liposomes, attached to liposomes via linking molecules associated with both liposomes and oligonucleotides, contained within liposomes, complexed with liposomes, dispersed in lipid-containing solutions, mixed with lipids, combined with lipids, contained in lipids as suspensions, contained in micelles, complexed with them, or otherwise associated with lipids. Compositions of associated lipids, lipids / DNA, or lipids / expression vectors are not limited to any particular structure in solution. For example, they may exist in bilayer structures, as micelles, or as "disintegrated" structures. They may also simply be scattered in solution, or possibly forming aggregates of non-uniform size or shape. Lipids are fatty substances that can be naturally occurring or synthetic lipids. Examples of lipids include naturally occurring lipid droplets in the cytoplasm, as well as a class of compounds including long-chain aliphatic hydrocarbons and their derivatives, such as fatty acids, alcohols, amines, amino alcohols, and aldehydes.
[0231] Suitable lipids for use can be obtained from private suppliers. For example, dimyristylphosphatidylcholine ("DMPC") can be obtained from Sigma, St. Louis, Mo., dicetyl phosphatidyl ("DCP") can be obtained from K&K Laboratories (Plainview, NY), cholesterol ("Choi") can be obtained from Calbiochem-Behring, and dimyristylphosphatidylglycerol ("DMPG") and other lipids can be obtained from Avanti Polar Lipids, Inc. (Birmingham, AL). Storage solutions of lipids in chloroform or chloroform / methanol can be stored at approximately -20°C. Chloroform is used as the sole solvent because it evaporates more readily than methanol. "Liposomes" is a general term encompassing various monolayer and multilayer lipid media formed by the formation of closed lipid bilayers or aggregates. Liposomes may be characterized by having a vesicular structure with a phospholipid bilayer membrane and an internal aqueous medium. Multilayer liposomes have multiple lipid layers separated by an aqueous medium. They spontaneously form when phospholipids are suspended in an excess aqueous solution. The lipid components undergo self-rearrangement, followed by the formation of a closed structure that encapsulates water and dissolved solutes between the lipid bilayers (Ghosh et al., 1991 Glycobiology 5:505-10). However, compositions having structures different from the usual vesicular structure in solution are also included. For example, lipids may take on a micelle structure or simply exist as heterogeneous aggregates of lipid molecules. Lipofectamine-nucleic acid complexes are also intended.
[0232] Regardless of the method used to introduce exogenous nucleic acids into host cells, various assays may be performed to confirm the presence of recombinant DNA sequences in host cells. Such assays include, for example, “molecular biological” assays well known to those skilled in the art, such as Southern blotting and Northern blotting, RT-PCR, and PCR; and “biochemical” assays, such as those by immunological means (ELISA and Western blotting), or the assays described herein for identifying active ingredients that fall within the scope of the present invention, for the detection of the presence or absence of specific peptides.
[0233] system In one embodiment, the present invention provides a system for reducing the number of one or more RNA transcripts in a subject. In one embodiment, the system comprises one or more vectors, each containing a nucleic acid sequence encoding a protein, wherein the protein comprises a CRISPR-related (Cas) protein and optionally a localization sequence, such as an NLS, NES, or organelle localization signal; and nucleic acid sequences encoding a tandem crRNA array.
[0234] In one embodiment, the tandem array contains one or more, two or more, three or more, four or more, five or more, six or more, seven or more, or eight or more crRNA sequences. In one embodiment, each crRNA contains a different sequence substantially complementary to a different sequence within a single target RNA. In one embodiment, each crRNA contains a different sequence substantially complementary to a different sequence within different target RNAs. In one embodiment, the different target RNAs are associated with a single disease, disorder, or infection. For example, in one embodiment, the different target RNAs are each viral RNA sequence of a virus.
[0235] In one embodiment, the nucleic acid sequence encoding Cas and the nucleic acid sequence encoding crRNA are present in the same vector. In another embodiment, the nucleic acid sequence encoding protein and the nucleic acid sequence encoding crRNA are present in different vectors.
[0236] In one embodiment, the nucleic acid sequence encoding a protein is (1) a nucleic acid that encodes an amino acid sequence identical to one of SEQ ID NOs. 1 to 47 by at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. The nucleic acid sequence comprises (1) a nucleic acid sequence encoding one amino acid from sequence numbers 1 to 47; and optionally (2) a nucleic acid sequence encoding one amino acid from sequence numbers 50 to 66 and 323 to 935, with respect to at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the nucleic acid sequence encoding a protein comprises (1) a nucleic acid sequence encoding one amino acid from sequence numbers 1 to 47; and optionally (2) a nucleic acid sequence encoding one amino acid from sequence numbers 50 to 66 and 323 to 935.In one embodiment, the nucleic acid sequence encoding a protein includes a nucleic acid sequence encoding an amino acid that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 68 to 100.
[0237] In one embodiment, the tandem array contains one or more, two or more, three or more, four or more, five or more, six or more, seven or more, or eight or more crRNA sequences. In one embodiment, each crRNA in the tandem crRNA array contains a direct repeat (DR) sequence and a spacer sequence. In one embodiment, the direct repeat sequence may be located upstream (i.e., 5') of the guide sequence or spacer sequence. In other embodiments, the direct repeat sequence may be located downstream (i.e., 3') of the guide sequence or spacer sequence.
[0238] In one embodiment, each crRNA in a tandem crRNA array contains a different direct repeat sequence. For example, in one embodiment, nucleotide substitutions within the loop region of the direct repeat, along with multiple guide RNAs, lead to the efficient generation of an ordered array of crRNAs.
[0239] In one embodiment, the tandem array includes at least two or more crRNAs, where each crRNA contains a sequence substantially complementary to the target RNA. For example, in one embodiment, each crRNA contains different sequences substantially complementary to different sequences within a single target RNA. In one embodiment, each crRNA contains different sequences substantially complementary to different sequences within different target RNAs. In one embodiment, the different target RNAs are associated with a single disease, disorder, or infection. For example, in one embodiment, the different target RNAs are each viral RNA sequence of a virus.
[0240] In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to the coronavirus genomic RNA sequence and / or coronavirus subgenomic RNA sequence. In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to a sequence or fragment selected from SEQ ID NOs. 167-188 that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to that sequence or fragment. In one embodiment, the tandem array includes at least two crRNAs, each containing a sequence or fragment of a sequence selected from sequence numbers 167-188 that is substantially complementary to that sequence.
[0241] In one embodiment, the tandem array includes at least two crRNAs containing sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to sequences selected from SEQ ID NOs.
[0242] In one embodiment, the tandem array contains sequences homologous to sequence number 275 by at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the tandem array contains the sequence of sequence number 275.
[0243] In one embodiment, the tandem array comprises at least two crRNAs, each independently containing a sequence substantially complementary to an influenza virus sequence. In one embodiment, the tandem array comprises at least two crRNAs, each containing a sequence substantially complementary to a genomic mRNA sequence or subgenomic mRNA sequence of the influenza virus. In one embodiment, the tandem array comprises at least two crRNAs, each containing a sequence substantially complementary to a PB2 sequence, PB1 sequence, PA sequence, HA sequence, NP sequence, NA sequence, M sequence, or NS sequence of the influenza virus.
[0244] In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to sequences or fragments of sequences selected from SEQ ID NOs. 225-244, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous.
[0245] In one embodiment, the tandem array includes at least two crRNAs containing sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to sequences selected from SEQ ID NOs.
[0246] In one embodiment, the tandem array includes sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous. In one embodiment, the tandem array includes sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, or at least 99% homologous.
[0247] In one embodiment, the tandem array contains at least two crRNAs, each independently containing a sequence substantially complementary to an mRNA sequence within a cellular pathway. In some embodiments, the cellular pathway results in dysregulation of cellular processes, such as enhanced uncontrolled proliferation and cell survival. For example, in one embodiment, each crRNA independently contains a sequence substantially complementary to an mRNA sequence within the RAS, JAK-STAT, PI3K / AKT, or MAPK cellular pathway.
[0248] Compositions and Formulations In one embodiment, the present invention provides a composition for reducing the number of RNA transcripts in a subject. In one embodiment, the composition comprises a fusion protein, wherein the fusion protein comprises a CRISPR-related (Cas) protein and optionally a localization sequence, such as an NLS, NES, or organelle localization signal. In one embodiment, the composition comprises a tandem crRNA array. In one embodiment, the tandem crRNA array comprises two or more crRNAs, each substantially hybridizing to a target RNA sequence of one or more RNA transcripts.
[0249] In one embodiment, the composition is (1) an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of SEQ ID NOs: 1 to 47; and optionally (2) A protein comprising an amino acid sequence identical to at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the composition comprises (1) an amino acid from one of the SEQ ID NOs: 1 to 47; and optionally (2) a protein comprising one amino acid from the SEQ ID NOs: 50 to 66 and 323 to 935.
[0250] In one embodiment, the composition comprises a protein having an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to one of sequence numbers 68 to 100. In one embodiment, the nucleic acid sequence encoding the protein comprises a protein having an amino acid sequence that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, or at least 99% identical to one of sequence numbers 68 to 100.
[0251] In one embodiment, the composition comprises a tandem crRNA array, wherein the array comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, or eight or more crRNA sequences. In one embodiment, each crRNA in the tandem crRNA array comprises a direct repeat (DR) sequence and a spacer sequence. In one embodiment, the direct repeat sequence may be located upstream (i.e., 5') of the guide sequence or spacer sequence. In other embodiments, the direct repeat sequence may be located downstream (i.e., 3') of the guide sequence or spacer sequence.
[0252] In one embodiment, each crRNA in a tandem crRNA array contains a different direct repeat sequence. For example, in one embodiment, nucleotide substitutions within the loop region of the direct repeat, along with multiple guide RNAs, lead to the efficient generation of an ordered array of crRNAs.
[0253] In one embodiment, the composition comprises a tandem array, where the tandem array comprises at least two or more crRNAs, each crRNA containing a sequence substantially complementary to a target RNA. For example, in one embodiment, each crRNA contains different sequences substantially complementary to different sequences within a single target RNA. In one embodiment, each crRNA contains different sequences substantially complementary to different sequences within different target RNAs. In one embodiment, the different target RNAs are associated with a single disease, disorder, or infection. For example, in one embodiment, the different target RNAs are each viral RNA sequence of a virus.
[0254] In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to the coronavirus genomic RNA sequence and / or coronavirus subgenomic RNA sequence. In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to a sequence or fragment selected from SEQ ID NOs. 167-188 that is at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to that sequence or fragment. In one embodiment, the tandem array includes at least two crRNAs, each containing a sequence or fragment of a sequence selected from sequence numbers 167-188 that is substantially complementary to that sequence.
[0255] In one embodiment, the tandem array includes at least two crRNAs containing sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to sequences selected from SEQ ID NOs.
[0256] In one embodiment, the tandem array contains sequences homologous to sequence number 275 by at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In one embodiment, the tandem array contains the sequence of sequence number 275.
[0257] In one embodiment, the tandem array comprises at least two crRNAs, each independently containing a sequence substantially complementary to an influenza virus sequence. In one embodiment, the tandem array comprises at least two crRNAs, each containing a sequence substantially complementary to a genomic mRNA sequence or subgenomic mRNA sequence of the influenza virus. In one embodiment, the tandem array comprises at least two crRNAs, each containing a sequence substantially complementary to a PB2 sequence, PB1 sequence, PA sequence, HA sequence, NP sequence, NA sequence, M sequence, or NS sequence of the influenza virus.
[0258] In one embodiment, the tandem array includes at least two crRNAs containing sequences substantially complementary to sequences or fragments of sequences selected from SEQ ID NOs. 225-244, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous.
[0259] In one embodiment, the tandem array includes at least two crRNAs containing sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous to sequences selected from SEQ ID NOs.
[0260] In one embodiment, the tandem array includes sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homologous. In one embodiment, the tandem array includes sequences that are at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, or at least 99% homologous.
[0261] In one embodiment, the tandem array contains at least two crRNAs, each independently containing a sequence substantially complementary to an mRNA sequence within a cellular pathway. In some embodiments, the cellular pathway results in dysregulation of cellular processes, such as enhanced uncontrolled proliferation and cell survival. For example, in one embodiment, each crRNA independently contains a sequence substantially complementary to an mRNA sequence within the RAS, JAK-STAT, PI3K / AKT, or MAPK cellular pathway.
[0262] This disclosure also includes the use of the pharmaceutical compositions of this disclosure to carry out the methods of this disclosure. Such a pharmaceutical composition may consist of at least one modulating (e.g., inhibitory or activating) composition of the present invention or a salt thereof in a form suitable for administration to a subject, or the pharmaceutical composition may consist of at least one modulating (e.g., inhibitory or activating) composition of the present invention or a salt thereof, and one or more pharmaceutically acceptable carriers, one or more additional components, or several combinations thereof. The compounds of the present invention may exist in the pharmaceutical composition in the form of physiologically acceptable salts, for example, in combination with physiologically acceptable cations or anions, as is well known in the art.
[0263] In one embodiment, a pharmaceutical composition useful for carrying out the method of the present invention may be administered to deliver a dose of 1 ng / kg / day to 100 mg / kg / day. In another embodiment, a pharmaceutical composition useful for carrying out the method of the present invention may be administered to deliver a dose of 1 ng / kg / day to 500 mg / kg / day.
[0264] The relative amounts of the active ingredient, pharmaceutically acceptable carrier, and any additional ingredients in the pharmaceutical composition of the present invention will vary depending on the characteristics, size, and condition of the target being treated, as well as the route through which the composition is administered. As an example, the composition may contain 0.1% to 100% (w / w) of the active ingredient.
[0265] Pharmaceutical compositions useful in the methods of the present invention can be suitably developed for oral, rectal, vaginal, parenteral, topical, intrapulmonary, intranasal, oral, ophthalmic, or other routes of administration. Compositions useful in the methods of the present invention can be administered directly to the skin or any other tissue of a mammal. Other intended formulations include liposomal formulations, reencapsulated red blood cells containing the active ingredient, and immunological formulations. The route of administration(s) will be readily apparent to those skilled in the art and will depend on several factors, including the type and severity of the disease being treated, and the type and age of the animal or human subject being treated.
[0266] Formulations of the pharmaceutical compositions described herein may be prepared by any method known or to be developed in the field of pharmacology. Generally, such preparation methods include mixing the active ingredient with a carrier or one or more other auxiliary ingredients, and then, if necessary or desirable, forming or packaging the product into desired single or multi-dose units.
[0267] As used herein, “unit dose” refers to a specific amount of a pharmaceutical composition containing a predetermined amount of the active ingredient. The amount of the active ingredient is generally equal to the dose of the active ingredient that would be administered to the subject, or a convenient fraction of such a dose, for example, half or one-third of such a dose. The unit dosage form may be for once-daily doses or for multiple-daily doses (e.g., about 1 to 4 times or more per day). When multiple-daily doses are used, the unit dosage form may be the same or different for each dose.
[0268] In one embodiment, the composition of the present invention is formulated using one or more pharmaceutically acceptable excipients or carriers. In one embodiment, the pharmaceutical composition of the present invention comprises a therapeutically effective amount of a compound or conjugate of the present invention and a pharmaceutically acceptable carrier. Pharmaceutically acceptable carriers that are useful include, but are not limited to, glycerol, water, saline, ethanol, and other pharmaceutically acceptable salt solutions, such as salts of phosphates and organic acids. Examples of these and other pharmaceutically acceptable carriers are described in Remington’s Pharmaceutical Sciences (1991, Mack Publication Co., New Jersey).
[0269] The carrier can be, for example, a solvent or dispersion medium containing water, ethanol, polyols (such as glycerol, propylene glycol, and liquid polyethylene glycol, etc.), suitable mixtures thereof, and vegetable oils. Appropriate fluidity can be maintained, for example, by using coatings such as lecithin, by maintaining the required particle size in the case of dispersions, and by using surfactants. Prevention of the activity of microorganisms can be achieved by various antibacterial and antifungal agents, such as parabens, chlorobutanol, phenol, ascorbic acid, thimerosal, etc. In many cases, isotonic agents, such as sugars, sodium chloride, or polyhydric alcohols, such as mannitol and sorbitol, are included in the composition. Prolongation of the absorption of injectable compositions can be caused by incorporating agents that delay absorption, such as aluminum monostearate or gelatin, into the composition. In one embodiment, the pharmaceutically acceptable carrier is not only DMSO.
[0270] The formulation can be used in a mixture with conventional excipients, i.e., pharmaceutically acceptable organic or inorganic carrier substances suitable for oral, intravaginal, parenteral, nasal, intravenous, subcutaneous, enteral, or any other suitable mode of administration known to those skilled in the art. The pharmaceutical formulation can be sterilized and, if necessary, can be mixed with adjuvants such as lubricants, preservatives, stabilizers, wetting agents, emulsifying agents, salts for affecting osmotic buffers, colorants, flavoring agents, and / or fragrances. They may, if desired, be combined with other active agents, such as other analgesics.
[0271] As used herein, "additional components" include, but are not limited to, one or more of the following: excipients; surfactants; dispersants; inert diluents; granulating and disintegrating agents; binders; lubricants; sweetening agents; flavoring agents; colorants; preservatives; physiologically degradable compositions such as gelatin; aqueous vehicles and solvents; oily vehicles and solvents; suspending agents; dispersing or wetting agents; emulsifying agents, viscous agents; buffering agents; salts; thickening agents; fillers; emulsifying agents; antioxidants; antibiotics; antifungal agents; stabilizers; and pharmaceutically acceptable polymeric or hydrophobic materials. Other "additional components" that can be included in the pharmaceutical compositions of the present invention are known in the art and are described, for example, in Genaro, ed. (1985, Remington’s Pharmaceutical Sciences, Mack Publishing Co., Easton, PA), which is incorporated herein by reference.
[0272] The compositions of the present invention can contain about 0.005% to 2.0% preservative, based on the total weight of the composition. Preservatives are used to prevent spoilage when exposed to contaminants in the environment. Examples of preservatives useful according to the present invention include, but are not limited to, those selected from the group consisting of benzyl alcohol, sorbic acid, parabens, imidurea, and combinations thereof. An exemplary preservative is a combination of about 0.5% to 2.0% benzyl alcohol and 0.05% to 0.5% sorbic acid.
[0273] In one embodiment, the composition contains antioxidants and chelating agents that inhibit the degradation of the compound. Exemplary antioxidants for some compounds include BHT, BHA, α-tocopherol, and ascorbic acid in amounts ranging from about 0.01% to 0.3% of the total weight of the composition, and BHT in amounts ranging from 0.03% to 0.1% of the weight. In one embodiment, the chelating agent is present in an amount ranging from 0.01% to 0.5% of the weight based on the total weight of the composition. Exemplary chelating agents include EDTA salts (e.g., disodium EDTA) and citric acid in amounts ranging from about 0.01% to 0.20% of the weight. In some embodiments, the chelating agent is in the range of 0.02% to 0.10% of the total weight of the composition. The chelating agent is useful for chelating metal ions in the composition that may be detrimental to the shelf life of the formulation. BHT and disodium edetate are exemplary antioxidants and chelating agents in some compounds, respectively, but other suitable and equivalent antioxidants and chelating agents known to those skilled in the art may be substituted for them.
[0274] Liquid suspensions can be prepared using conventional methods for achieving suspension of active ingredients in aqueous or oily vehicles. Examples of aqueous vehicles include water and isotonic saline. Examples of oily vehicles include almond oil, oily esters, ethyl alcohol, vegetable oils such as peanut oil, olive oil, sesame oil, or coconut oil, fractionated vegetable oils, and mineral oils such as liquid paraffin. Liquid suspensions may further contain one or more additional components, but are not limited to suspending agents, dispersants or wetting agents, emulsifiers, lubricants, preservatives, buffers, salts, flavoring agents, coloring agents, and sweeteners. Oily suspensions may further contain thickeners. Known suspending agents, but are not limited to sorbitol syrup, hydrogenated edible fats, sodium alginate, polyvinylpyrrolidone, tragacanth gum, acacia gum, and cellulose derivatives such as sodium carboxymethylcellulose, methylcellulose, and hydroxypropyl methylcellulose. Known dispersants or wetting agents include, but are not limited to, naturally occurring phospholipids such as lecithin, alkylene oxides, partial esters derived from fatty acids, long-chain aliphatic alcohols, fatty acids and hexitol, or condensates of partial esters derived from fatty acids and hexitol anhydrides (e.g., polyoxyethylene stearate, heptadecaethyleneoxycetanol, polyoxyethylene sorbitol monooleate, and polyoxyethylene sorbitan monooleate, respectively). Known emulsifiers include, but are not limited to, lecithin and acacia. Known preservatives include, but are not limited to, methyl parahydroxybenzoate, ethyl parahydroxybenzoate, or n-propyl parahydroxybenzoate, ascorbic acid, and sorbic acid. Known sweeteners include, for example, glycerol, propylene glycol, sorbitol, sucrose, and saccharin. Known thickeners in oily suspensions include, for example, beeswax, solid paraffin, and cetyl alcohol.
[0275] Liquid solutions of active ingredients in aqueous or oily solvents can be prepared in substantially the same manner as liquid suspensions, the main difference being that the active ingredient is dissolved rather than suspended in the solvent. As used herein, “oily” liquids contain carbon-containing liquid molecules and exhibit lower polarity than water. Liquid solutions of the pharmaceutical compositions of the present invention may contain each of the components described with respect to liquid suspensions, but it will be understood that the suspending agent does not necessarily aid in the dissolution of the active ingredient in the solvent. Examples of aqueous solvents include water and isotonic saline. Examples of oily solvents include almond oil, oily esters, ethyl alcohol, vegetable oils such as peanut oil, olive oil, sesame oil, or coconut oil, fractionated vegetable oils, and mineral oils such as liquid paraffin.
[0276] Powder and granular formulations of the pharmaceutical preparations of the present invention can be prepared using known methods. Such formulations can be administered directly to a subject or used, for example, to form tablets, to fill capsules, or to prepare aqueous or oily suspensions or solutions by adding an aqueous or oily vehicle thereto. Each of these formulations may further contain one or more of the following: dispersants or wetting agents, suspending agents, and preservatives. Additional excipients, such as fillers and sweeteners, flavorings, or colorants, may also be included in these formulations.
[0277] The pharmaceutical compositions of the present invention may also be prepared, packaged, or sold in the form of oil-in-water emulsions or water-in-oil emulsions. The oil phase may be a vegetable oil such as olive oil or peanut oil, a mineral oil such as liquid paraffin, or a combination thereof. Such compositions may further contain one or more emulsifiers, for example, naturally occurring gums such as gum arabic or tragacanth gum, naturally occurring phosphatides such as soy or lecithin phosphatides, esters or partial esters derived from a combination of fatty acids and hexitol anhydrides, for example, sorbitan monooleate, and condensates of such partial esters with ethylene oxide, for example, polyoxyethylene sorbitan monooleate. These emulsions may also contain additional components, for example, sweeteners or flavorings.
[0278] Methods for impregnating or coating materials with chemical compositions are known in the art and are not limited thereto, but include methods for depositing or bonding chemical compositions onto a surface, methods for incorporating chemical compositions into the structure of a material during its synthesis (i.e., synthesis using physiologically biodegradable materials, etc.), and methods for absorbing an aqueous or oily solution or suspension into an absorbent material, followed by drying or not drying.
[0279] The administration plan may affect what constitutes the effective dose. The therapeutic agent may be administered to the subject either before or after the diagnosis of the disease. Furthermore, several divided doses and time-staggered doses may be administered daily or sequentially, or the dose may be administered by continuous infusion or bolus injection. In addition, the dose of the therapeutic agent may be increased or decreased accordingly if there is an urgency in the treatment or prevention situation.
[0280] The administration of the compositions of the present invention to subjects, including mammals such as humans, may be carried out using known procedures in doses and durations effective for preventing or treating a disease. The effective amount of the therapeutic compound required to achieve a therapeutic effect may vary depending on factors such as the activity of the particular compound used; the time of administration; the elimination rate of the compound; the duration of treatment; other drugs, compounds, or materials used in combination with the compound; the state of the disease or disorder; the age, sex, weight, condition, health status, and medical history of the subject being treated, as well as similar factors well known in the medical field. The administration plan may be adjusted to obtain an optimal therapeutic response. For example, several divided doses may be administered daily, or the dose may be reduced accordingly if there is an urgency in the treatment situation. A non-limiting example of the effective dose range of the therapeutic compound of the present invention is about 1 to 5,000 mg / kg body weight per day. Those skilled in the art will be able to study the relevant factors and make decisions regarding the effective amount of the therapeutic compound without excessive experimentation.
[0281] The compound may be administered to the subject several times a day, or less frequently, for example, once a day, once a week, once every two weeks, once a month, or even less frequently, for example, once every few months or even once a year or less. It will be understood that the amount of the compound administered per day may, in non-limiting examples, be daily, every other day, every two days, every three days, every four days, or every five days. For example, in an every-other-day administration, a dose of 5 mg / day may be started on Monday, a first subsequent dose of 5 mg / day may be administered on Wednesday, a second subsequent dose of 5 mg / day may be administered on Friday, and so on. The frequency of administration will be readily apparent to those skilled in the art and will depend on several factors, such as, but are not limited to, the type and severity of the disease being treated, the type and age of the animal, etc.
[0282] The actual dose level of the active ingredient in the pharmaceutical composition of the present invention may vary for a particular subject, composition, and mode of administration, in order to obtain an amount of the active ingredient that is effective in achieving a desired therapeutic response without being toxic to the subject.
[0283] A physician or veterinarian with ordinary skill in the art can easily determine and prescribe the effective amount of the required pharmaceutical composition. For example, a physician or veterinarian may start the dose of the compound of the present invention used in the pharmaceutical composition at a level lower than the level required to achieve the desired therapeutic effect and gradually increase the dose until the desired effect is achieved.
[0284] In certain embodiments, it is particularly advantageous to formulate the compound in unit dosage forms for ease of administration and uniformity of dosage. As used herein, a unit dosage form refers to a physically separated unit suitable as a unit dose for the subject to be treated, each unit containing a predetermined amount of the therapeutic compound calculated to produce the desired therapeutic effect in conjunction with the required excipients. The unit dosage forms of the present invention are determined and directly depend on (a) the inherent characteristics of the therapeutic compound and the specific therapeutic effect to be achieved, and (b) the inherent limitations of the technique for compounding such therapeutic compounds to treat a disease in the subject.
[0285] In one embodiment, the composition of the present invention is administered to a subject in a dosage range of 1 to 5 times or more per day. In another embodiment, the composition of the present invention is administered to a subject in a dosage range that includes, but is not limited to, once daily, once every two days, once every three days to once a week, and once every two weeks. It will be readily apparent to those skilled in the art that the administration frequency of various combinations of compositions of the present invention will vary from subject to subject depending on a number of factors, including, but is not limited to, age, disease or disorder to be treated, sex, overall health condition, and other factors. Accordingly, the present invention should not be construed as being limited to any particular dosing schedule, and the exact dosage and composition to be administered to any subject will be determined by the attending physician, taking into account all other factors relating to the subject.
[0286] The compounds of the present invention for administration are available in doses of approximately 1 mg to 10,000 mg, 20 mg to 9,500 mg, 40 mg to 9,000 mg, 75 mg to 8,500 mg, 150 mg to 7,500 mg, 200 mg to 7,000 mg, 3050 mg to 6,000 mg, 500 mg to 5,000 mg, 750 mg to 4,000 mg, and 1 mg to 3,000 mg. The range may be 0 mg, approximately 10 mg to approximately 2,500 mg, approximately 20 mg to approximately 2,000 mg, approximately 25 mg to approximately 1,500 mg, approximately 50 mg to approximately 1,000 mg, approximately 75 mg to approximately 900 mg, approximately 100 mg to approximately 800 mg, approximately 250 mg to approximately 750 mg, approximately 300 mg to approximately 600 mg, approximately 400 mg to approximately 500 mg, and all or part of any increment between these ranges.
[0287] In some embodiments, the dose of the compound of the present invention is about 1 mg to about 2,500 mg. In some embodiments, the dose of the compound of the present invention used in the compositions described herein is less than about 10,000 mg, or less than about 8,000 mg, or less than about 6,000 mg, or less than about 5,000 mg, or less than about 3,000 mg, or less than about 2,000 mg, or less than about 1,000 mg, or less than about 500 mg, or less than about 200 mg, or less than about 50 mg. Similarly, in some embodiments, the dose of the second compound described herein (i.e., a drug used to treat the same disease as treated by the composition of the present invention or a different disease) is less than about 1,000 mg, less than about 800 mg, or less than about 600 mg, or less than about 500 mg, or less than about 400 mg, or less than about 300 mg, or less than about 200 mg, or less than about 100 mg, or less than about 50 mg, or less than about 40 mg, or less than about 30 mg, or less than about 25 mg, or less than about 20 mg, or less than about 15 mg, or less than about 10 mg, or less than about 5 mg, or less than about 2 mg, or less than about 1 mg, or less than about 0.5 mg, and all or part of any increment thereof.
[0288] In one embodiment, the present invention relates to a packaged pharmaceutical composition comprising a container for containing a therapeutically effective amount of the compound or conjugate of the present invention, either alone or in combination with a second pharmaceutical product, and instructions for use of the compound or conjugate for treating, preventing, or reducing one or more symptoms of a disease in a subject.
[0289] The term “container” includes any container for housing a pharmaceutical composition. For example, in one embodiment, the container is packaging for housing a pharmaceutical composition. In other embodiments, the container is not packaging for housing a pharmaceutical composition; that is, the container is a box or vial or other container for housing a packaged or unpackaged pharmaceutical composition and instructions for use of the pharmaceutical composition. Furthermore, packaging techniques are well known in the art. Instructions for use of a pharmaceutical composition may be included in the packaging for housing the pharmaceutical composition, and it should be understood that the instructions for use form a strong functional relationship with the packaged product. However, it should be understood that the instructions for use may include information about the compound’s intended function, such as the treatment or prevention of a disease in a subject, or the ability of the compound to deliver an imaging agent or diagnostic agent to a subject.
[0290] Routes of administration of any of the compositions of the present invention include oral, nasal, parenteral, sublingual, transdermal, transmucosal (e.g., sublingual, tongue, (trans)oral, and nasal (intra)), intravesical, intraduodenal, intragastric, rectal, intraperitoneal, subcutaneous, intramuscular, intradermal, intra-arterial, intravenous, or administration.
[0291] Suitable compositions and dosage forms include, for example, tablets, capsules, caplets, pills, gel capsules, lozenges, dispersants, suspensions, solutions, syrups, granules, beads, transdermal patches, gels, powders, pellets, magmas, licks, creams, pastes, ointments, lotions, discs, suppositories, liquid sprays for nasal or oral administration, dry powders or aerosolized formulations for inhalation, and compositions and formulations for intravesical administration. It should be understood that the formulations and compositions that may be useful in the present invention are not limited to the specific formulations and compositions described herein.
[0292] Method for producing a tandem array The present invention provides a method for producing a tandem array of the present invention. For example, in one embodiment, the tandem array can be produced in a single step. In one embodiment, the production of the tandem array enables the transfer of the entire array length.
[0293] In one embodiment, the method includes ligating at least two crRNA sequences, where each crRNA sequence includes a unique direct repeat sequence (DR), and the tandem array is generated by the ligation. In one embodiment, each DR sequence includes a mutation unique to the poly-T stretch of SEQ ID NO: 267. For example, in one embodiment, each DR sequence includes a unique mutation at T17 or T18 of SEQ ID NO: 267. In one embodiment, one crRNA sequence includes a wild-type DR sequence, and each additional crRNA sequence includes a DR sequence having a mutation unique to the poly-T stretch of SEQ ID NO: 267. In one embodiment, each DR sequence is independently selected from SEQ ID NOs: 268 - 274.
[0294] Method for reducing RNA and treatment method In one aspect, the present disclosure provides a method for reducing the number of one or more RNA transcripts in a subject. In one embodiment, the method reduces the number of two or more RNA transcripts in the subject. In one embodiment, two or more RNA transcripts in the subject are related. For example, in one embodiment, the two or more RNA transcripts are each viral transcript of the same virus. In one embodiment, the two or more RNA transcripts are each involved in the same cellular pathway.
[0295] In one embodiment, RNA is localized in the cytoplasm. In one embodiment, RNA is localized in the nucleus. In one embodiment, RNA is localized in organelles. For example, in one embodiment, the method reduces RNA localized in the nucleolus, ribosomes, vesicles, rough endoplasmic reticulum, Golgi apparatus, cytoskeleton, smooth endoplasmic reticulum, mitochondria, vacuoles, cytosol, lysosomes, or centrioles. In one embodiment, the method reduces RNA associated with the cell membrane. In one embodiment, the method reduces extracellular RNA.
[0296] In one embodiment, the method comprises administering to (1) a nucleic acid molecule encoding a fusion protein of the Disclosure comprising a Cas protein and optionally a localization sequence, e.g., NLS, NES, or organelle localization signal; or a fusion protein of the Disclosure comprising a Cas protein and optionally a localization sequence, e.g., NLS, NES, or organelle localization signal; and (2) a nucleic acid molecule encoding a crRNA array comprising two or more crRNAs.
[0297] In some embodiments, cytoplasmic RNA. In such embodiments, the method comprises administering to (1) a protein of the disclosure comprising a Cas protein and NES or a nucleic acid molecule encoding a protein of the disclosure comprising a Cas protein and NES; and (2) a nucleic acid molecule encoding a crRNA array comprising two or more crRNAs.
[0298] In some embodiments, nuclear RNA. In such embodiments, the method comprises administering to (1) a protein of the disclosure comprising a Cas protein and NLS or a nucleic acid molecule encoding a protein of the disclosure comprising a Cas protein and NLS; and (2) a nucleic acid molecule encoding a crRNA array comprising two or more crRNAs.
[0299] In one embodiment, the subject is a cell. In one embodiment, the cell is a prokaryotic cell or a eukaryotic cell. In one embodiment, the cell is a eukaryotic cell. In one embodiment, the cell is a plant cell, an animal cell, or a fungal cell. In one embodiment, the cell is a plant cell. In one embodiment, the cell is an animal cell. In one embodiment, the cell is a yeast cell.
[0300] In one embodiment, the subject is a mammal. For example, in one embodiment, the subject is a human, a non-human primate, a dog, a cat, a horse, a cattle, a goat, a sheep, a rabbit, a pig, a rat, or a mouse. In one embodiment, the subject is a non-mammalian subject. For example, in one embodiment, the subject is a zebrafish, a fruit fly, or a roundworm.
[0301] In one embodiment, the amount of viral RNA is reduced in vitro. In another embodiment, the amount of viral RNA is reduced in vivo.
[0302] In one embodiment, the present invention provides a method for cleaving one or more target viral RNAs in a subject. In one embodiment, the method comprises administering to a subject (1) a nucleic acid molecule encoding the protein of the disclosure comprising a Cas protein and optionally a localization sequence, e.g., NLS, NES, or organelle localization signal; or the protein of the disclosure comprising a Cas protein and optionally a localization sequence, e.g., NLS, NES, or organelle localization signal; and (2) a nucleic acid molecule encoding a crRNA array comprising two or more crRNAs, where each crRNA independently comprises a sequence substantially complementary to the RNA sequence of the target viral RNA.
[0303] Treatment method and method of use The present invention provides methods for treating diseases or disorders in a subject, for reducing their symptoms, and / or for reducing the risk of developing them. For example, in one embodiment, the method of the present invention treats diseases or disorders in mammals, for reducing their symptoms, and / or for reducing the risk of developing them. In one embodiment, the method of the present invention treats diseases or disorders in plants, for reducing their symptoms, and / or for reducing the risk of developing them. In one embodiment, the method of the present invention treats diseases or disorders in yeast, for reducing their symptoms, and / or for reducing the risk of developing them.
[0304] In one embodiment, the subject is a cell. In one embodiment, the cell is a prokaryotic cell or a eukaryotic cell. In one embodiment, the cell is a eukaryotic cell. In one embodiment, the cell is a plant cell, an animal cell, or a fungal cell. In one embodiment, the cell is a plant cell. In one embodiment, the cell is an animal cell. In one embodiment, the cell is a yeast cell.
[0305] In one embodiment, the subject is a mammal. For example, in one embodiment, the subject is a human, a non-human primate, a dog, a cat, a horse, a cattle, a goat, a sheep, a rabbit, a pig, a rat, or a mouse. In one embodiment, the subject is a non-mammalian subject. For example, in one embodiment, the subject is a zebrafish, a fruit fly, or a roundworm.
[0306] In one embodiment, the disease or disorder is caused by the activation of a cellular pathway. In some embodiments, the cellular pathway results in dysregulation of cellular processes, such as uncontrolled proliferation and enhanced cell survival. For example, in one embodiment, each crRNA independently contains a sequence substantially complementary to the mRNA sequence in the RAS, JAK-STAT, PI3K / AKT, or MAPK cellular pathway.
[0307] In one embodiment, the method comprises administering (1) a protein of the Disclosure or a nucleic acid molecule encoding a protein of the Disclosure, and (2) a nucleic acid molecule encoding a crRNA array comprising two or more crRNAs, wherein each crRNA independently comprises a sequence substantially complementary to the nucleotide sequence of the RNA transcript of the pathway. In one embodiment, the crRNA array comprises two or more crRNAs, wherein each crRNA independently comprises a sequence substantially complementary to the nucleotide sequence of the mRNA sequence in the RAS, JAK-STAT, PI3K / AKT, or MAPK cellular pathway. In one embodiment, the Cas protein cleaves the RNA transcript(s).
[0308] In one embodiment, the diseases or disorders include, but are not limited to, cancer, heart disease, atherosclerosis, cardiac fibrosis, cardiac arrhythmia, hypertension, respiratory disease, stroke, Alzheimer's disease, diabetes, nephritis, and liver disease.
[0309] In one embodiment, the disease or disorder is a viral infection. In one embodiment, the disease or disorder is caused by a viral infection. In one embodiment, the viral infection is an RNA virus infection. For example, in one embodiment, the viral infection is a positive-strand ssRNA virus infection, a negative-strand ssRNA virus infection, a dsRNA virus infection, or an ssRNA-RT virus infection. In one embodiment, the viral infection is a DNA virus infection. For example, in one embodiment, the viral infection is a dsDNA virus infection, an ssDNA virus infection, or a dsDNA-RT virus infection. Accordingly, in one embodiment, the disease or disorder can be treated, reduced, or the risk reduced by a component that interferes with or reduces viral RNA transcripts. Accordingly, in one embodiment, the disease or disorder can be treated, reduced, or the risk reduced by a component that interferes with or reduces viral mRNA transcripts or prevents or reduces the translation of viral proteins.
[0310] In one embodiment, the method comprises administering (1) a protein of the Disclosure or a nucleic acid molecule encoding a protein of the Disclosure, and (2) a nucleic acid molecule encoding a crRNA array containing two or more crRNAs, wherein each crRNA independently contains a sequence substantially complementary to a nucleotide sequence complementary to a viral RNA transcript. In one embodiment, the Cas protein cleaves the viral RNA transcript.
[0311] In one embodiment, the virus is an RNA virus. In one embodiment, the virus produces RNA during its life cycle. In one embodiment, the virus is a human virus, a plant virus, or an animal virus. Examples of viruses include, but are not limited to, viruses from the families Adenoviridae, Alphaflexiviridae, Anelloviridae, Arenavirus, Arteriviridae, Asfarviridae, Astroviridae, Benyviridae, Betaflexiviridae, Birnaviridae, Bornaviridae, Bromoviridae, Caliciviridae, Caulimoviridae, Circoviridae, Closteroviridae, Coronaviridae, Filoviridae, Flaviviridae, Geminiviridae, Hantaviridae, and Hepadnaviridae. Examples of viruses include those belonging to the families Hepeviridae, Herpesviridae, Kitaviridae, Luteoviridae, Nairoviridae, Nanoviridae, Nimaviridae, Orthomyxoviridae, Paramyxoviridae, Phenuiviridae, Picornaviridae, Polyomaviridae, Pospiviridae, Potyviridae, Poxviridae, Reoviridae, Retroviridae, Retrovirus, Rhabdoviridae, Seboviridae, Togaviridae, Tombusviridae, Tospoviridae, Tymoviridae, and Virgaviridae. For example, examples of viruses, though not limited to them, include African swine fever virus, avian hepatitis E virus, avian infectious laryngotracheitis virus, chicken nephritis virus, bamboo mosaic virus, banana bunchy top virus, wheat mosaic virus, barley yellow wilt virus, potato leaf curl virus, Borna disease virus, brom mosaic virus, wheat,Cauliflower mosaic virus, Chikungunya virus, Eastern equine encephalitis virus, Citrus leprosis virus, Citrus sudden death associated virus, Citrus tristeza virus, Coconut cadang-cadang viroid, Curlytop virus, African cassava mosaic Numerous examples of viruses that harm crops, including viruses such as cytomegalovirus, Epstein-Barr virus, dengue virus, yellow fever virus, West Nile virus, Zika virus, Ebola virus, Marburg virus, equine arteritis virus, porcine reproductive and respiratory syndrome virus, equine infectious anemia virus, foot-and-mouth disease virus, enterovirus, rhinovirus, hepatitis B virus, hepatitis E virus, HIV, HIV-1, HIV-2, infectious bursal disease virus (poultry), infectious pancreatic necrosis virus (salmon), canine infectious hepatitis virus, poultry aviadenovirus, influenza virus, Lassa fever virus, lymphocytic choriomeningitis virus, monkeypox virus, Nairobi sheep disease virus, Newcastle disease virus (poultry), Norwalk virus, potato Y virus, and Porcine circovirus. 2. Examples include beak feather disease virus (poultry), potato M virus, rabies virus, respiratory and digestive adenoviruses, respiratory syncytial virus, rice stripe necrosis virus, Rift Valley fever virus, rotavirus, SARS-CoV-2, MERS virus, sheeppox virus, Lumpy Skin disease virus, Sin Nombre virus, Andes virus, SV40, tobacco ring spot virus, tomato bushy stunt virus, tomato spotted wilt virus, Torkutenovirus, Venezuelan encephalitis virus, vesicular stomatitis Indiana virus, viral hemorrhagic septicemic virus (trout), and white spot disease virus (shrimp).
[0312] In one embodiment, exemplary viruses include, but are not limited to, primate T lymphotropic virus 1, primate T lymphotropic virus 2, primate T lymphotropic virus 3, human immunodeficiency virus 1, human immunodeficiency virus 2, monkey foam virus, human picovirnavirus, Colorado tick fever virus, Changinora virus, Great Island virus, Lebombo virus, Orungo virus, rotavirus A, rotavirus B, rotavirus C, Banna virus, Borna disease virus, Lake Victoria Marburg virus, Reston Ebola virus, Sudan Ebola virus, Thai Forest Ebola virus, Zaire virus, human parainfluenza virus 2, human parainfluenza virus 4, mumps virus, Newcastle disease virus, human parainfluenza virus 1, human parainfluenza virus 3, Hendra virus, Nipah virus, measles virus, human respiratory syncytial virus, human metapneumovirus, Chandipla virus, Isfahan virus, pirivirus, vesicular virus Stomatitis Aragoas virus, vesicular stomatitis Indiana virus, vesicular stomatitis New Jersey virus, Australian bat lyssavirus, Dubenhaji virus, European bat lyssavirus 1, European bat lyssavirus 2, Mocola virus, rabies virus, Guanalitovirus, Junin virus, Lassa fever virus, lymphocytic choriomeningitis virus, Machupo virus, Pichinde virus, Sabia virus, Whitewater Arroyo virus, Bunyanbella virus, Bwamba virus, California encephalitis virus, Calapal virus, Katu virus, Guama virus, Guarroa virus, Kairi Viruses, Maritsuba virus, Olivoca virus, Sacral virus, Schuni virus, Tacaiuma virus, Weomiya virus, Andes virus, Bayo virus, Black Creek Canal virus, Dobraba-Bergledo virus, Hunter virus, Laguna Negra virus, New York virus, Pumara virus, Seoul virus, Sin Nombre virus, Crimean-Congo hemorrhagic fever virus, Djugbe virus, Candiru virus, Punta Torovirus, Rift Valley fever virus, sandfly fever Napoli virus, influenza A virus, influenza B virus, influenza C virus, Dori virus, Togoto virus, hepatitis delta virus, human coronavirus 229E, human coronavirus NL63, human coronavirus HKU1, human coronavirus OC43, SARS coronavirus, human torovirus, human enterovirus A, human enterovirus B, human enterovirus C, human enterovirus D, human rhinovirus A, human rhinovirus B, human rhinovirus C, encephalomyocarditis virus, tylovirus, equine rhinitis A virus, foot-and-mouth disease virus, hepatitis A virus, human parechovirus, Yungan virus, Aichi virus, human astrovirus, human astrovirus 2, human astrovirus 3, human astrovirus 4, human astrovirus 5, human astrovirus 6, human astrovirus 7, human astrovirus 8, Norwalk virus, Sapporo virus, Aroa Examples of viruses include vanzi virus, dengue virus, Ilheus virus, Japanese encephalitis virus, Kokobella virus, Kasanur Forest disease virus, jumping disease virus, Malay Valley encephalitis virus, Untaya virus, Omsk hemorrhagic fever virus, Poissant virus, Rio Bravo virus, St. Louis encephalitis virus, tick-borne encephalitis virus, Usutu virus, Wesselsbron virus, West Nile virus, yellow fever virus, Zika virus, hepatitis C virus, hepatitis E virus, Burma Forest virus, Chikungunya virus, Eastern equine encephalitis virus, Everglades virus, Geta virus, Mayaro virus, Mukambo virus, Onyonnyon virus, Pixna virus, Ross River virus, Semlik Forest virus, Sindbis virus, Venezuelan horse encephalitis virus, Western equine encephalitis virus, Wataroa virus, and rubella virus. In one embodiment, exemplary viruses include, but are not limited to, frog herpesvirus 1, frog herpesvirus 2, frog herpesvirus 3, eel herpesvirus 1, carp herpesvirus 1, carp herpesvirus 2, carp herpesvirus 3, sturgeon herpesvirus 2, American catfish herpesvirus 1, American catfish herpesvirus 2, salmonid fish herpesvirus 1, salmonid fish herpesvirus 2, salmonid fish herpesvirus 3, trialphaherpesvirus 1, parrot alphaherpesvirus Duck alphaherpesvirus 1, pigeon alphaherpesvirus 1, avian alphaherpesvirus 2, avian alphaherpesvirus 3, turkey alphaherpesvirus 1, penguin alphaherpesvirus 1, sea turtle alphaherpesvirus 5, tortoise alphaherpesvirus 3, spider monkey alphaherpesvirus 1, bovine alphaherpesvirus 2, cercomare alphaherpesvirus 2, human alphaherpesvirus 1, human alphaherpesvirus 2, rabbit alphaherpesvirus 4, Macaque alphaherpesvirus 1, Kangaroo alphaherpesvirus 1, Kangaroo alphaherpesvirus 2, Chimpanzee alphaherpesvirus 3, Baboon alphaherpesvirus 2, Fruit bat alphaherpesvirus 1, Squirrel monkey alphaherpesvirus 1, Bovine alphaherpesvirus 1, Bovine alphaherpesvirus 5, Water buffalo alphaherpesvirus 1, Dog alphaherpesvirus 1, Goat alphaherpesvirus 1, Cercopithecus alphaherpesvirus 9, Deer alphaherpesvirus 1, C Cavalier alpha herpesvirus 2, Equine alpha herpesvirus 1, Equine alpha herpesvirus 3, Equine alpha herpesvirus 4, Equine alpha herpesvirus 8, Equine alpha herpesvirus 9, Feline alpha herpesvirus 1, Human alpha herpesvirus 3, Narwhal alpha herpesvirus 1, Seal alpha herpesvirus 1, Pig alpha herpesvirus 1, Sea turtle alpha herpesvirus 6, Occult monkey beta herpesvirus 1, Capuchin monkey beta herpesvirus 1, Cercopithecus beta herpesvirus 5Human betaherpesvirus 5, Macaque betaherpesvirus 3, Macaque betaherpesvirus 8, Mandrilline betaherpesvirus 1, Chimpanzee betaherpesvirus 2, Baboon betaherpesvirus 3, Baboon betaherpesvirus 4, Squirrel monkey herpesvirus 4, Mouse betaherpesvirus 1, Mouse betaherpesvirus 2, Mouse betaherpesvirus 8, Elephant betaherpesvirus 1, Elephant betaherpesvirus 4, Elephant betaherpesvirus 5, Human betaherpesvirus 7, Human betaherpesvirus 6A, Human betaherpesvirus 6B, Macaque betaherpesvirus 9, Mouse betaherpesvirus Rupesvirus 3, porcine betaherpesvirus 2, guinea pig betaherpesvirus 2, tree shrew betaherpesvirus 1, sycamore gammaherpesvirus 3, cercomoni gammaherpesvirus 14, gorilla gammaherpesvirus 1, human gammaherpesvirus 4, macaque gammaherpesvirus 4, macaque gammaherpesvirus 10, chimpanzee gammaherpesvirus 1, baboon gammaherpesvirus 1, drosophila gammaherpesvirus 2, bovine gammaherpesvirus 1 , Bovine gamma herpesvirus 2, Bovine gamma herpesvirus 6, Goat gamma herpesvirus 2, Bluebuck gamma herpesvirus 1, Sheep gamma herpesvirus 2, Pig gamma herpesvirus 3, Pig gamma herpesvirus 4, Pig gamma herpesvirus 5, Horse gamma herpesvirus 2, Horse gamma herpesvirus 5, Cat gamma herpesvirus 1, Weasel gamma herpesvirus 1, Seal gamma herpesvirus 3, Baby bat gamma herpesvirus 1, C Mammal gamma herpesvirus 2, spider monkey gamma herpesvirus 3, bovine gamma herpesvirus 4, pygmy rat gamma herpesvirus 2, human gamma herpesvirus 8, macaque gamma herpesvirus 5, macaque gamma herpesvirus 8, macaque gamma herpesvirus 11, macaque gamma herpesvirus 12, rat gamma herpesvirus 4, rat gamma herpesvirus 7, squirrel monkey gamma herpesvirus 2, horse gamma herpesvirus 7, seal gamma herpesvirus 2,Saguinine gammaherpesvirus 1 イアアルペウス2 イゥゥルウイSalmonella 1, Salmonella virus SKML39, Shigella virus AG3, Dickeya virus Limestone, Dickeya virus RC2014, Escherichia virus CBA120, Escherichia virus PhaxI, Salmonella virus 38, Salmonella virus Det7, Salmonella virus GG32, Salmonella virus PM10, Salmonella virus SFP10, Salmonella virus SH19, Salmonella virus SJ3, Escherichia virus KWBSE43-6, Klebsiella virus 0507KN21, Klebsiella virus KpS110, Klebsiella virus May, Klebsiella virus Menlow, Serratia virus IME250, Erwinia virus Ea2809, Serratia virus MAM1, Acinetobacter virus Acibel007, Acinetobacter virus AB3, Acinetobacter virus AbKT21III、Acinetobacter virus Abp1、Acinetobacter virus Aci07、Acinetobacter virus Aci08、Acinetobacter virus AS11、Acinetobacter virus AS12、Acinetobacter virus Fri1,Acinetobacter virus IME200,Acinetobacter virus PD6A3, Acinetobacter virus PDAB9, Acinetobacter virus phiAB1, Acinetobacter virus SH-Ab 15519, Acinetobacter virus SWHAb1, Acinetobacter virus SWHAb3, Acinetobacter virus WCHABP5.Acinetobacter virus B1、Acinetobacter virus B2、Acinetobacter virus B5、Acinetobacter virus D2、Acinetobacter virus P1、Acinetobacter virus P2、Acinetobacter virus phiAB6、Acinetobacter virus Petty、Vibrio virus Vc1、Vibrio virus A318、Vibrio virus AS51、Vibrio virus Vp670、Marinomonas virus CB5A、Marinomonas virus CPP1m、Vibrio virus VEN、Pseudomonas virus Achelous、Pseudomonas virus Alpheus、Pseudomonas virus Nerthus、Pseudomonas virus Njord、Pseudomonas virus uligo、Pseudomonas virus C171、Pectobacterium virus PP16、Pectobacterium virus PPWS1、Pectobacterium virus PPWS2、Pectobacterium virus CB5、Pectobacterium virus Clickz、Pectobacterium virus fM1、Pectobacterium virus Gaspode、Pectobacterium virus Khlen、Pectobacterium virus Koot、Pectobacterium virus Lelidair、Pectobacterium virus Nobby、Pectobacterium virus Peat1、Pectobacterium virus Phoria、Pectobacterium virus PP90、Pectobacterium virus Zenivior、Dickeya virus BF25-12、Pseudomonas virus NV3、Pseudomonas virus 130-113、Pseudomonas virus 15pyo、Pseudomonas virus Ab05、Pseudomonas virus ABTNL、Pseudomonas virus DL62、Pseudomonas virus kF77、Pseudomonas virus LKD16、Pseudomonas virus LUZ19、Pseudomonas virus MPK6、Pseudomonas virus MPK7、Pseudomonas virus NFS、Pseudomonas virus PAXYB1、Pseudomonas virus phiKMV、Pseudomonas virus PT2、Pseudomonas virus PT5、Pseudomonas virus RLP、Pseudomonas virus LKA1、Pseudomonas virus f2、Aeromonas virus 25AhydR2PP、Aeromonas virus AS7、Aeromonas virus ZPAH7、Yersinia virus ISAO8、Aeromonas virus Ahp1、Aeromonas virus CF7、Cronobacter virus DevCD23823、Cronobacter virus GAP227、Salmonella virus Spp16、Yersinia virus R8-01、Yersinia virus fHeYen301、Yersinia virus Phi80-18、Pectobacterium virus Arno160、Pectobacterium virus PP2、Proteus virus PM85、Proteus virus PM93、Proteus virus PM116、Proteus virus Pm5460、Pectobacterium virus PP1、Erwinia virus Era103、Erwinia virus S2、Lelliottia virus phD2B、Citrobacter CrRp3、Escherichia virus LL11、Escherichia virus AAPEc6、Escherichia virus ACGC91、Escherichia virus B、Escherichia virus C、Escherichia virus K、Escherichia virus K1-5、Escherichia virus K1E、Escherichia virus mutPK1A2、Escherichia virus VEc3、Escherichia virus UAB78、Salmonella virus BP12B、Sa、 lmonella virus SP6, Burkholderia virus BpAMP1, Ralstonia virus RSPI1, Ralstonia virus RSB1, Ralstonia virus RsoP1IDN, Burkholderia virus JG068, Ralstonia virus RSJ2, Ralstonia virus RSJ5, Ralstonia virus RSPII1, Shigella virus Buco, Escherichia virus Minorna, Leprosy virus AltoGao, Leprosy virus BO1E, Leprosy virus F19, Leprosy virus K244, Leprosy virus Kp2, Leprosy virus KP34, Leprosy virus KPRio2015, Leprosy virus KpS2, Leprosy virus KpV41, Leprosy virus KpV48, Leprosy virus KpV71, Leprosy virus KpV74, Leprosy virus KpV475, Leprosy virus KPV811, Klebsiella virus myPSH1235, Klebsiella virus SU503, Klebsiella virus SU552A, Shigella virus SFN6B, Enterobacter virus KDA1, Proteus virus PM16, Proteus virus PM75, Dickeya virus Dagda、Dickeya virus Katbat、Dickeya virus Luksen、Dickeya virus Mysterion、Yersinia virus AP10、Erwinia virus FE44、Escherichia virus 285P、Escherichia virus BA14、Escherichia virus P483、Escherichia virus P694, Escherichia virus S523, Kluyvera virusKvp1、Pectobacterium virus PP74、Salmonella virus BP12A、Salmonella virus BSP161、Shigella virus A7、Yersinia virus Berlin、Yersinia virus PYPS50、Yersinia virus Yepe2、Yersinia virus Yepf、Citrobacter virus CR8、Vibrio virus ICP3、Vibrio virus N4、Vibrio virus VP4、Enterobacter virus Eap1、Erwinia virus L1、Escherichia virus SRT7、Pseudomonas virus 17A、Pseudomonas virus gh1、Pseudomonas virus Henninger、Pseudomonas virus KNP、Pseudomonas virus Pf1ERZ2017、Pseudomonas virus PhiPSA2、Pseudomonas virus PhiPsa17、Pseudomonas virus PPPL1、Pseudomonas virus shl2、Pseudomonas virus WRT、Yersinia virus fPS9、Yersinia virus fPS53、Yersinia virus fPS59、Yersinia virus fPS54ocr、Pectobacterium virus Jarilo、Citrobacter virus CR44b、Citrobacter virus SH3、Citrobacter virus SH4、Cronobacter virus Dev2、Cronobacter virus GW1、Enterobacter virus EcpYZU01、Escherichia virus EcoDS1、Escherichia virus F、Escherichia virus GA2A、Escherichia virus IMM002、Escherichia virus K1F、Escherichia virus LM33P1、Escherichia virus PE3-1、Escherichia virusRo45lw、Escherichia virus ST31、Escherichia virus Vec13、Escherichia virus YZ1、Escherichia virus ZG49、Shigella virus SFPH2、Morganella virus MmP1、Morganella virus MP2、Dickeya virus JA10、Dickeya virus Ninurta、Pectobacterium virus PP47、Pectobacterium virus PP81、Pectobacterium virus PPWS4、Pseudomonas virus PPpW4、Pseudomonas virus 22PfluR64PP、Pseudomonas virus IBBPF7A、Pseudomonas virus Pf10、Pseudomonas virus PFP1、Pseudomonas virus PhiS1、Pseudomonas virus UNOSLW1、Pseudomonas virus PspYZU08、Escherichia virus K30、Klebsiella virus 2044-307w、Klebsiella virus BIS33、Klebsiella virus Henu1、Klebsiella virus IL33、Klebsiella virus IME205、Klebsiella virus IME321、Klebsiella virus K5、Klebsiella virus K11、Klebsiella virus K5-2、Klebsiella virus K5-4、Klebsiella virus KN1-1、Klebsiella virus KN3-1、Klebsiella virus KN4-1、Klebsiella virus Kp1、Klebsiella virus KP32、Klebsiella virus KP32i192、Klebsiella virus KP32i194、Klebsiella virus KP32i195、Klebsiella virus KP32i196、Klebsiella virus kpssk3、Klebsiella virusKpV289, Leprosy virus KpV763, Leprosy virus KpV766, Leprosy virus KpV767, Leprosy virus Pharr, Leprosy virus PRA33, Leprosy virus SHKp152234, Leprosy virus SHKp152410, Citrobacter virus CFP1, Citrobacter virus SH1, Citrobacter virus SH2, Enterobacter virus E2, Enterobacter virus E3, Enterobacter virus KPN3, Enterobacter virus T7M, Escherichia virus ECA2, Escherichia virus LL2, Escherichia virus T3, Escherichia virus T3Luria, Leclercia virus 10164-302, Salmonella virus SG-JL2, Serratia virus 2050H2, Serratia virus SM9-3Y, Yersinia virus AP5、Yersinia virus YeF10、Yersinia virus YeO3-12、Enterobacterial virus IME390、Escherichia virus 13a、Escherichia virus 64795ec1、Escherichia virus C5、Escherichia virus CICC80001、Escherichia virus Ebrios, Escherichia virus EG1, Escherichia virus HZ2R8, Escherichia virus HZP2, Escherichia virus N30, Escherichia virus NCA, Escherichia virus T7, Salmonella virus 3A8767, Salmonella virus Vi06, Stenotrophomonas virus IME15, Yersinia virus YpPY, Yersinia virusYpsPG、Pseudomonas virus Phi15、Pectobacterium virus DUPPII、Synechococcus virus SCBP42、Aquamicrobium virus P14、Ashivirus S45C4、Agrobacterium virus Atuph02、Agrobacterium virus Atuph03、Ralstonia virus Ap1、Ayaqvirus S45C18、Prochlorococcus virus SS120-1、Pseudomonas virus Andromeda、Pseudomonas virus Bf7、Escherichia virus J8-65、Escherichia virus Lidtsur、Prochlorococcus virus NATL1A7、Chosvirus KM23C739、Rhizobium virus RHEph02、Rhizobium virus RHEph08、Rhizobium virus RHEph09、Vibrio virus Cyclit、Escherichia virus PGT2、Escherichia virus PhiKT、Alteromonas virus H4-4、Foussvirus S46C10、Fussvirus S30C28、Escherichia virus ECBP5、Pectobacterium virus PP99、Ralstonia virus DURPI、Ralstonia virus RsoP1EGY、Synechococcus STIP37、Jalkavirus S08C159、Ralstonia virus RSB3、Kawavirus SWcelC56、Synechococcus virus SRIP1、Providencia virus PS3、Curvibacter virus P26059B、Ralstonia virus RSB2、Synechococcus virus SCBP2、Krakvirus S39C11、Podovirus Lau218、Pantoea virus LIMElight、Prochlorococcus virus PGSP1、Synechococcus virusSCBP3、Caulobacter virus Lullwater、Vibrio virus KF1、Vibrio virus KF2、Vibrio virus OWB、Vibrio virus VP93、Pseudomonas virus VSW3、Nohivirus S31C1、Oinezvirus S37C6、Rhi zobium virus RHEph01、Pagavirus S05C849、Mesorhizobium virus Lo5R7ANS、Pedosvirus S28C3、Pekhitvirus S04C24、Pelagibacter virus HTVC019P、Pelagivirus S35C6、Caulobacter virus Percy、Delftia virus IMEDE1、Podivirus S05C243、Pseudomonas virus PollyC、Synechococcus virus SCBP4、Powvirus S08C41、Xanthomonas virus f20、Xanthomonas virus f30、Xanthomonas virus XAJ24、Xanthomonas virus Xc10、Xylella virus Prado、Synechococcus virus SB28、Sphingomonas virus Scott、Synechococcus virus SRIP2、Ralstonia virus ITL1、Sieqvirus S42C7、Ralstonia virus RPSC1、Stopalavirus S38C3、Pelagibacter virus HTVC011P、Stupnyavirus KM16C193、Prochlorococcus virus 951510a、Prochlorococcus virus NATL2A133、Prochlorococcus virus PSSP10、Vibrio virus JSF7、Prochlorococcus virus PSSP7、Synechococcus virus P60、Prochlorococcus virus PSSP3、Synechococcus virus PSSP2、Synechococcus virus Syn5、Votkovvirus S28C10、Pantoea virus LIMEzero、Pasteurella virus PHB01、Pasteurella virus PHB02、Escherichia virus GJ1、Escherichia virus ST32、Erwinia virus Faunus、Erwiniavirus Y2、Aeromonas virus pAh6C、Pectobacterium virus PM1、Pectobacterium virus PP101、Shewanella virus Spp001、Shewanella virus SppYZU05、Vibrio virus Ceto、Vibrio virus Thalassa、Vibrio virus JSF10、Vibrio virus JSF12、Vibrio virus phi3、Vibrio virus pVp1、Escherichia virus EPS7、Escherichia virus mar003J3、Escherichia virus saus132、Salmonella virus 123、Salmonella virus 329、Salmonella virus 118970sal2、Salmonella virus LVR16A、Salmonella virus S113、Salmonella virus S114、Salmonella virus S116、Salmonella virus S124、Salmonella virus S126、Salmonella virus S132、Salmonella virus S133、Salmonella virus S147、Salmonella virus Seafire、Salmonella virus SH9、Salmonella virus STG2、Salmonella virus Stitch、Salmonella virus Sw2、Yersinia virus phiR201、Escherichia virus AKFV33、Escherichia virus BF23、Escherichia virus chee24、Escherichia virus DT5712、Escherichia virus DT57C、Escherichia virus FFH1、Escherichia virus Gostya9、Escherichia virus H8、Escherichia virus mar004NP2、Escherichia virus OSYSP、Escherichia virusphiAPCEc03, Escherichia virus phiLLS, Escherichia virus slur09, Escherichia virus T5, Salmonella virus NR01, Salmonella virus S131, Salmonella virus Shivani, Salmonella virus SP01, Salmonella virus SP3, Salmonella virus SPC35, Shigella virus SHSML45, Shigella virus SSP1, Pectobacterium virus DUPPV, Pectobacterium virus My1, Proteus virus PM135, Proteus virus Stubb, Vibrio virus PG07, Vibrio virus VspSw1, Aeromonas virus AhSzq1, Aeromonas virus AhSzw1, Klebsiella virus IME260, Klebsiella virus Sugarland, Escherichia virus IME542, Escherichia virus ACGM12, Escherichia virus EC3a, Escherichia virus DTL, Escherichia virus IME253, Escherichia virus Rtp, Shigella virus Sf12, Escherichia virus phiEB49, Escherichia virus AHP42, Escherichia virus AHS24, Escherichia virus AKS96, Escherichia virus C119, Escherichia virus E41c, Escherichia virus Eb49, Escherichia virus Jk06, Escherichia virus KP26, Escherichia virus phiJLA23, Escherichia virus Rogue1, Shigella virus Sd1, Shigella virus pSf1、Citrobacter virus DK2017、Citrobacter virusSazh, Citrobacter virus Stevie, Escherichia virus LL5, Escherichia virus TLS, Salmonella virus 36, Salmonella virus PHB07, Salmonella virus phSE2, Salmonella virus SP126, Salmonella virus YSP2, Escherichia virus 95, Escherichia virus mar001J1, Escherichia virus mar002J2, Escherichia virus SECphi27, Escherichia virus swan01, Escherichia virus IME347, Escherichia virus SRT8, Escherichia virus ADB2, Escherichia virus BIFF, Escherichia virus IME18, Escherichia virus JMPW1, Escherichia virus JMPW2, Escherichia virus SH2, Escherichia virus T1, Shigella virus 008, Shigella virus ISF001, Shigella virus PSf2, Shigella virus Sfin1, Shigella virus SH6, Shigella virus Shfl1, Shigella virus ISF002, Cronobacter virus Esp2949-1, Enterobacter virus EcL1, Cronobacter virus PhiCS01, Escherichia virus ESCO41, Pantoea virus AAS23, Escherichia virus NBD2, Enterobacter virus F20, Klebsiella virus 1513, Klebsiella virus GHK3, Klebsiella virus KLPN1, Klebsiella virus KOX1, Klebsiella virus KP36, Klebsiella virus KpCol1, Klebsiella virusKpKT21phi1、Klebsiella virus KPN N141、Klebsiella virus KpV522、Klebsiella virus MezzoGao、Klebsiella virus NJR15、Klebsiella virus NJS1、Klebsiella virus NJS2、Klebsiella virus PKP126、Klebsiella virus Sushi、Klebsiella virus TAH8、Klebsiella virus TSK1、Bacillus virus Agate、Bacillus virus Bobb、Bacillus virus Bp8pC、Bacillus virus Bastille、Bacillus virus CAM003、Bacillus virus Evoli、Bacillus virus HoodyT、Bacillus virus AvesoBmore、Bacillus virus B4、Bacillus virus Bigbertha、Bacillus virus Riley、Bacillus virus Spock、Bacillus virus Troll、Bacillus virus Bc431、Bacillus virus Bcp1、Bacillus virus BCP82、Bacillus virus BM15、Bacillus virus Deepblue、Bacillus virus JBP901、Bacillus virus Grass、Bacillus virus NIT1、Bacillus virus SPG24、Bacillus virus BCP78、Bacillus virus TsarBomba、Bacillus virus BPS13、Bacillus virus BPS10C、Bacillus virus Hakuna、Bacillus virus Megatron、Bacillus virus WPh、Bacillus virus Mater、Bacillus virus Moonbeam、Bacillus virus SIOphi、Enterococcus virus ECP3、Enterococcus virus EF24C、Enterococcusvirus EFLK1、Enterococcus virus EFDG1、Enterococcus virus EFP01、Enterococcus virus EfV12、Listeria virus A511、Listeria virus AG20、Listeria virus List36 , Listeria virus LMSP25, Listeria virus LMTA34, Listeria virus LMTA148, Listeria virus LP048, Listeria virus LP064, Listeria virus LP083-2, Listeria virus P100, Listeria virus WIL1, Bacillus virus Camphawk, Bacillus virus SPO1, Bacillus virus CP51, Bacillus virus JL, Bacillus virus Shanette, Staphylococcus virus BS1, Staphylococcus virus BS2, Lactobacillus virus Bacchae, Lactobacillus virus Bromius, Lactobacillus virus Iacchus, Lactobacillus virus Lpa804, Lactobacillus virus Semele, Staphylococcus virus G1, Staphylococcus virus G15, Staphylococcus virus JD7, Staphylococcus virus K, Staphylococcus virus MCE2014, Staphylococcus virus P108, Staphylococcus virus Rodi, Staphylococcus virus S253, Staphylococcus virus S25-4, Staphylococcus virus SA12, Staphylococcus virus Sb1, Staphylococcus virus SscM1、Staphylococcus virus IPLAC1C、Staphylococcus virus SEP1、Staphylococcus virus Remus、Staphylococcus virus SA11、Staphylococcus virus Stau2、Staphylococcus virus Twort、Brochothrix virus A9、Lactobacillus virus Lb338-1、Lactobacillusvirus LP65、Campylobacter virus CP21、Campylobacter virus CP220、Campylobacter virus CPt10、Campylobacter virus IBB35、Campylobacter virus CP81、Campylobacter virus CP30A、Campylobacter virus CPX、Campylobacter virus Los1、Campylobacter virus NCTC12673、Escherichia virus Alf5、Escherichia virus AYO145A、Escherichia virus EC6、Escherichia virus HY02、Escherichia virus JH2、Escherichia virus TP1、Escherichia virus VpaE1、Escherichia virus wV8、Salmonella virus BPS15Q2、Salmonella virus BPS17L1、Salmonella virus BPS17W1、Salmonella virus FelixO1、Salmonella virus Mushroom、Salmonella virus Si3、Salmonella virus SP116、Salmonella virus UAB87、Erwinia virus Ea214、Erwinia virus M7、Citrobacter virus Moogle、Citrobacter virus Mordin、Shigella virus Sf13、Shigella virus Sf14、Shigella virus Sf17、Escherichia virus SUSP1、Escherichia virus SUSP2、Ralstonia virus RSA1、Ralstonia virus RSY1、Mannheimia virus 1127AP1、Mannheimia virus PHL101、Aeromonas virus phiO18P、Vibrio virus Canoe、Pseudoalteromonas virus C5a、Pseudomonas virusDobby、Pseudomonas virus phiCTX、Erwinia virus EtG、Escherichia virus 186、Salmonella virus PsP3、Salmonella virus SEN1、Erwinia virus ENT90、Klebsiella virus 4LV2017、Salmonella virus Fels2、Salmonella virus RE2010、Salmonella virus SEN8、Salmonella virus SopEphi、Haemophilus virus HP1、Haemophilus virus HP2、Vibrio virus Kappa、Pasteurella virus F108、Burkholderia virus KS14、Burkholderia virus AP3、Burkholderia virus KS5、Vibrio virus K139、Burkholderia virus ST79、Escherichia virus fiAA91ss、Escherichia virus P2、Escherichia virus pro147、Escherichia virus pro483、Escherichia virus Wphi、Yersinia virus L413C、Pseudomonas virus phi3、Salinivibrio virus SMHB1、Klebsiella virus 3LV2017、Salmonella virus SEN4、Cronobacter virus ESSI2、Stenotrophomonas virus Smp131、Salmonella virus FSLSP004、Burkholderia virus KL3、Burkholderia virus phi52237、Burkholderia virus phiE122、Burkholderia virus phiE202、Vibrio virus PV94、Escherichia virus P88、Escherichia virus Bp7、Escherichia virus IME08、Escherichia virus JS10、Escherichia virusJS98, Escherichia virus MX01, Escherichia virus QL01, Escherichia virus VR5, Escherichia virus WG01, Escherichia virus VR7, Escherichia virus VR20, Escherichia virus VR25, Escherichia virus VR26, Shigella virus SP18, Salmonella virus Melville, Salmonella virus S16, Salmonella virus STML198, Salmonella virus STP4a, Klebsiella virus JD18, Klebsiella virus PKO111, Enterobacter virus PG7, Escherichia virus CC31, Escherichia virus ECD7, Escherichia virus GEC3S, Escherichia virus JSE, Escherichia virus phi1, Escherichia virus RB49, Citrobacter virus CF1, Citrobacter virus Merlin, Citrobacter virus Moon、Escherichia virus APCEc01, Escherichia virus HP3, Escherichia virus HX01, Escherichia virus JS09, Escherichia virus O157tp3, Escherichia virus O157tp6, Escherichia virus PhAPEC2, Escherichia virus RB69, Escherichia virus ST0, Shigella virus SHSML521, Shigella virus UTAM, Vibrio virus KVP40, Vibrio virus nt1, Vibrio virus ValKK3, Enterobacter virus Eap3, Klebsiella virus KP15, Klebsiella virus KP27、Klebsiella virus Matisse、KlebsiellaMiro virus、Klebsiella virus PMBT1、Escherichia virus AR1、Escherichia virus C40、Escherichia virus CF2、Escherichia virus E112、Escherichia virus ECML134、Escherichia virus HY01、Escherichia virus HY03, Escherichia virus Ime09, Escherichia virus RB3, Escherichia virus RB14, Escherichia virus slur03, Escherichia virus slur04, Escherichia virus T4, Shigella virus Pss1, Shigella virus Sf21, Shigella virus Sf22, Shigella virus Sf24, Shigella virus SHBML501, Shigella virus Shfl2, Yersinia virus D1, Yersinia virus PST, Acinetobacter virus 133, Aeromonas virus 65, Aeromonas virus Aeh1, Escherichia virus RB16, Escherichia virus RB32, Escherichia virus RB43, Pseudomonas virus 42, Escherichia virus Av05, Cronobacter virus CR3, Cronobacter virus CR8, Cronobacter virus CR9, Cronobacter virus PBES02, Pectobacterium virus phiTE, Cronobacter virus GAP31, Escherichia virus 4MG, Salmonella virus PVPSE1, Salmonella virus SSE121, Escherichia virus APECc02, Escherichia virus FFH2, Escherichia virus FV3 Escherichia virus JES2013 Escherichia virusMurica, Escherichia virus slur16, Escherichia virus V5, Escherichia virus V18, Brevibacillus virus Abouo, Brevibacillus virus D avies、Synechococcus virus SMbCM100、Erwinia virus Deimos、Erwinia virus Desertfox、Erwinia virus Ea35-70、Erwinia virus RAY、Erwinia virus Simmy50、Erwinia virus SpecialG、Synechococcus virus SShM2、Klebsiella virus K64-1、Klebsiella virus RaK2、Dickeya virus AD1、Erwinia virus Alexandra、Lactobacillus virus LBR48、Synechococcus virus SCAM1、Synechococcus virus SCBWM1、Vibrio virus Aphrodite1、Escherichia virus 121Q、Escherichia virus PBECO4、Synechococcus virus AC2014fSyn7803C8、Synechococcus virus ACG2014f、Synechococcus virus ACG2014fSyn7803US26、Synechococcus virus STIM5、Pseudomonas virus PaBG、Rheinheimera virus Barba18A、Rheinheimera virus Barba19A、Rheinheimera virus Barba21A、Rheinheimera virus Barba5S、Rheinheimera virus Barba8S、Burkholderia virus BcepMu、Burkholderia virus phiE255、Synechococcus virus Bellamy、Gordonia virus GMA6、Aeromonas virus 44RR2、Mycobacterium virus Alice、Mycobacterium virus Bxz1、Mycobacterium virus Dandelion、Mycobacterium virus HyRo、Mycobacterium virus I3、Mycobacterium virusLukilu、Mycobacterium virus Nappy、Mycobacterium virus Sebata、Faecalibacterium virus Brigit、Prochlorococcus virus Syn33、Synechococcus virus SRIM12-01、Synechococcus virus SRIM12-06、Synechococcus virus SRIM12-08、Salmonella virus SEN34、Acidovorax virus ACP17、Xanthomonas virus Carpasina、Xanthomonas virus XcP1、Pseudomonas virus pf16、Synechococcus virus SCAM3、Ralstonia virus RSF1、Ralstonia virus RSL2、Synechococcus virus SWAM2、Erwinia virus Derbicus、Pseudomonas virus EL、Sinorhizobium virus M7、Sinorhizobium virus M12、Sinorhizobium virus N3、Serratia virus BF、Yersinia virus Yen9-04、Faecalibacterium virus Epona、Erwinia virus Asesino、Erwinia virus EaH2、Prochlorococcus virus MED4-213、Prochlorococcus virus PHM1、Prochlorococcus virus PHM2、Flavobacterium virus FCL2、Flavobacterium virus FCV1、Pseudomonas virus KIL2、Pseudomonas virus KIL4、Edwardsiella virus GF2、Escherichia virus Goslar、Halomonas virus HAP1、Vibrio virus VP882、Lactobacillus virus Lb、Erwinia virus EaH1、Iodobacter virus PLPE、Delftia virusPhiW14, Klebsiella virus JD001, Klebsiella virus KpV52, Klebsiella virus KpV80, Escherichia virus CVM10, Escherichia virus ECOO78, Escherichia virus ep3, Brevibacillus virus Jimmer, Brevibacillus Osiris virus、Synechococcus virus SCAM9、Rhizobium virus RHEph4、Faecalibacterium virus Lagaffe、Synechococcus virus SP4、Synechococcus virus Syn30、Prochlorococcus virus PTIM40、Synechococcus virus SSKS1, Salmonella virus ZCSE2, Clostridium virus phiC2, Clostridium virus phiCD27, Clostridium virus phiCD119, Erwinia virus Machina, Arthrobacter virus BarretLemon, Arthrobacter virus Beans, Arthrobacter virus Brent, Arthrobacter virus Jawnski、Arthrobacter virus Martha、Arthrobacter virus Piccoletto、Arthrobacter virus Shade、Arthrobacter virus Sonny、Synechococcus virus SCAM7、Acinetobacter virus ME3、Ralstonia virus RSL1、Cronobacter virus GAP32.Pectobacterium virus CBB, Faecalibacterium virus Mushu, Escherichia virus Mu, Shigella virus SfMu, Halobacterium virus phiH, Burkholderia virus Bcep1, Burkholderia virus Bcep43, Burkholderiavirus Bcep781、Burkholderia virus BcepNY3、Xanthomonas virus OP2、Synechococcus virus SMbCM6、Pseudomonas virus Ab03、Pseudomonas virus G1、Pseudomonas virus KPP10、Pseudomonas virus PAKP3、Pseudomonas virus PS24、Synechococcus virus SRIM8、Synechococcus virus SRIM50、Synechococcus virus ACG2014bSyn7803C61、Synechococcus virus ACG2014bSyn9311C4、Synechococcus virus SRIM2、Synechococcus virus SPM2、Pseudomonas virus Noxifer、Acinetobacter virus AB1、Acinetobacter virus AB2、Acinetobacter virus AbC62、Acinetobacter virus AbP2、Acinetobacter virus AP22、Acinetobacter virus LZ35、Acinetobacter virus WCHABP1、Acinetobacter virus WCHABP12、Pseudomonas virus Psa374、Pseudomonas virus VCM、Pseudomonas virus CAb1、Pseudomonas virus CAb02、Pseudomonas virus JG004、Pseudomonas virus MAG1、Pseudomonas virus PA10、Pseudomonas virus PAKP1、Pseudomonas virus PAKP2、Pseudomonas virus PAKP4、Pseudomonas virus PaP1、Pseudomonas virus phiMK、Pseudomonas virus Zigelbrucke、Prochlorococcus virus PSSM7、Burkholderia virus BcepF1、Pseudomonasvirus 141、Pseudomonas virus Ab28、Pseudomonas virus CEBDP1、Pseudomonas virus DL60、Pseudomonas virus DL68、Pseudomonas virus E215、Pseudomonas virus E217、Pseudomonas virus F8、Pseudomonas virus JG024、Pseudomonas virus KPP12、Pseudomonas virus KTN6、Pseudomonas virus LBL3、Pseudomonas virus LMA2、Pseudomonas virus NH4、Pseudomonas virus PA5、Pseudomonas virus PB1、Pseudomonas virus PS44、Pseudomonas virus SN、Pectobacterium virus PEAT2、Edwardsiella virus pEtSU、Bordetella virus PHB04、Escherichia phage ESCO13、Escherichia virus ESCO5、Escherichia virus phAPEC8、Escherichia virus Schickermooser、Klebsiella virus ZCKP1、Pseudomonas virus PA7、Pseudomonas virus phiKZ、Pseudomonas virus SL2、Pseudomonas virus PMW、Agrobacterium virus Atuph07、Synechococcus virus Syn19、Aeromonas virus 56、Aeromonas virus 43、Escherichia virus P1、Escherichia virus RCS47、Salmonella virus SJ46、Pseudoalteromonas virus J2-1、Arthrobacter virus ArV1、Arthrobacter virus Colucci、Arthrobacter virus Trina、Ralstonia virus RP12、Erwinia virusRisingsun、Salmonella virus BP63、Acinetobacter virus Aci05、Acinetobacter virus Aci01-1、Acinetobacter virus Aci02-2、Prochlorococcus virus PSSM 2、Dickeya virus JA11、Dickeya virus JA29、Erwinia virus Y3、Agrobacterium virus 7-7-1、Salmonella virus SPN3US、Bacillus virus Shbh1、Bacillus virus 1、Geobacillus virus GBSV1、Pseudomonas virus tabernarius、Synechococcus virus ST4、Faecalibacterium virus Taranis、Synechococcus virus SIOM18、Yersinia virus R1RT、Yersinia virus TG1、Synechococcus virus STIM4、Synechococcus virus SSM1、Bacillus virus SP15、Vibrio virus pTD1、Vibrio virus VP4B、Tetrasphaera virus TJE1、Faecalibacterium virus Toutatis、Aeromonas virus 25、Aeromonas virus Aes12、Aeromonas virus Aes508、Aeromonas virus AS4、Aeromonas virus Asgz、Stenotrophomonas virus IME13、Prochlorococcus virus Syn1、Synechococcus virus SRIM44、Vibrio virus MAR、Vibrio virus VHML、Vibrio virus VP585、Escherichia virus ECML4、Salmonella virus Marshall、Salmonella virus Maynard、Salmonella virus SJ2、Salmonella virus STML131、Salmonella virus ViI、Erwinia virus Wellington、Escherichia virus ECML-117、Escherichia virus FEC19、Escherichia virus WFC、Escherichia virus WFH、Serratiavirus CHI14、Edwardsiella virus MSW3、Edwardsiella virus PEi21、Erwinia virus Yoloswag、Bacillus virus G、Bacillus virus PBS1、Microcystis virus Ma-LMM01、Streptococcus virus Cp1、Streptococcus virus Cp7、Lactococcus virus WP2、Bacillus virus B103、Bacillus virus GA1、Bacillus virus phi29、Kurthia virus 6、Actinomyces virus Av1、Mycoplasma virus P1、Staphylococcus virus Andhra、Staphylococcus virus St134、Staphylococcus virus 66、Staphylococcus virus 44AHJD、Staphylococcus virus BP39、Staphylococcus virus CSA13、Staphylococcus virus GRCS、Staphylococcus virus Pabna、Staphylococcus virus phiAGO13、Staphylococcus virus PSa3、Staphylococcus virus S24-1、Staphylococcus virus SAP2、Staphylococcus virus SCH1、Staphylococcus virus SLPW、Shigella virus 7502Stx、Shigella virus POCJ13、Escherichia virus 191、Escherichia virus PA2、Escherichia virus TL2011、Shigella virus VASD、Escherichia virus 24B、Escherichia virus 933W、Escherichia virus Min27、Escherichia virus PA28、Escherichia virus Stx2 II、Dinoroseobacter virusDFL12、Pseudomonas virus Bjorn、Pseudomonas virus Ab22、Pseudomonas virus CHU、Pseudomonas virus LUZ24、Pseudomonas virus PAA2、Pseudomonas virus PaP3、Pseudomonas virus PaP4、Pseudomonas virus TL、Vibrio virus VC8、Vibrio virus VP2、Vibrio virus VP5、Escherichia virus N4、Flavobacterium virus Fpv1、Flavobacterium virus Fpv4、Streptococcus virus C1、Escherichia virus APEC5、Escherichia virus APEC7、Escherichia virus Bp4、Escherichia virus EC1UPM、Escherichia virus ECBP1、Escherichia virus G7C、Escherichia virus IME11、Shigella virus Sb1、Escherichia virus C1302、Pseudomonas virus F116、Pseudomonas virus H66、Escherichia virus Pollock、Salmonella virus FSL SP-058、Salmonella virus FSL SP-076、Arthrobacter virus Adat、Arthrobacter virus Jasmine、Erwinia virus Ea9-2、Erwinia virus Frozen、Achromobacter virus Axp3、Achromobacter virus JWAlpha、Edwardsiella virus KF1、Burkholderia virus KL4、Pseudomonas virus KPP25、Pseudomonas virus R18、Pseudomonas virus tf、Escherichia virus 172-1、Escherichia virus ECB2、Escherichia virusNJ01、Escherichia virus phiEco32、Escherichia virus Septima11、Escherichia virus SU10、Escherichia virus HK620、Salmonella virus BTP1、Salmonella virus P22、Salmonella virus SE1Spa、Salmonella virus ST64T、Shigella virus Sf6、Burkholderia virus Bcep22、Burkholderia virus Bcepil02、Burkholderia virus Bcepmigl、Burkholderia virus DC1、Cellulophaga virus Cba41、Cellulophaga virus Cba172、Pseudomonas virus Ab09、Pseudomonas virus LIT1、Pseudomonas virus PA26、Pseudomonas virus KPP21、Pseudomonas virus LUZ7、Vibrio virus 48B1、Vibrio virus 51A6、Vibrio virus 51A7、Vibrio virus 52B1、Myxococcus virus Mx8、Bacillus virus Page、Bacillus virus Palmer、Bacillus virus Pascal、Bacillus virus Pony、Bacillus virus Pookie、Brucella virus Pr、Brucella virus Tb、Bordetella virus BPP1、Burkholderia virus BcepC6B、Helicobacter virus 1961P、Helicobacter virus KHP30、Helicobacter virus KHP40、Pseudomonas virus phCDa、Escherichia virus Skarpretter、Escherichia virus Sortsne、Klebsiella virus IME279、Escherichia virus phiV10、Salmonella virusEpsilon15、Salmonella virus SPN1S、Pseudomonas virus NV1、Pseudomonas virus UFVP2、Escherichia virus PTXU04、Hamiltonella virus APSE1、Lactococcus virus KSY1、Phormidium virus WMP3、Phormidium virus WMP4、Pseudomonas virus 119X、Roseobacter virus SIO1、Vibrio virus VpV262、Streptomyces virus ELB20、Streptomyces virus R4、Streptomyces virus Amela、Streptomyces virus phiCAM、Streptomyces virus Aaronocolus、Streptomyces virus Caliburn、Streptomyces virus Danzina、Streptomyces virus Hydra、Streptomyces virus Izzy、Streptomyces virus Lannister、Streptomyces virus Lika、Streptomyces virus Sujidade、Streptomyces virus Zemlya、Streptomyces virus phiHau3、Mycobacterium virus Acadian、Mycobacterium virus Baee、Mycobacterium virus Reprobate、Mycobacterium virus Adawi、Mycobacterium virus Bane1、Mycobacterium virus BrownCNA、Mycobacterium virus Chrisnmich、Mycobacterium virus Cooper、Mycobacterium virus JAMaL、Mycobacterium virus Nigel、Mycobacterium virus Stinger、Mycobacterium virus Vincenzo、Mycobacterium virusZemanar、Mycobacterium virus Apizium、Mycobacterium virus Manad、Mycobacterium virus Oline、Mycobacterium virus Osmaximus、Mycobacterium virus Pg1、Mycobacterium virus Soto、Mycobac terium virus Suffolk, Mycobacterium virus Athena, Mycobacterium virus Bernardo, Mycobacterium virus Gadjet, Mycobacterium virus Pipefish, Mycobacterium virus Godines, Mycobacterium virus Rosebush, Mycobacterium virus TA17a, Mycobacterium virus Babsiella, Mycobacterium virus Brujita, Mycobacterium virus Hawkeye, Mycobacterium virus Plot, Caulobacter virus CcrBL9, Caulobacter virus CcrSC, Caulobacter virus CcrColossus, Caulobacter virus CcrPW, Caulobacter virus CcrBL10, Caulobacter virus CcrRogue, Caulobacter virus phiCbK, Caulobacter virus Swift, Salmonella virus SP31, Salmonella virus AG11, Salmonella virus Ent1, Salmonella virus f18SE, Salmonella virus Jersey, Salmonella virus L13, Salmonella virus LSPA1, Salmonella virus SE2, Salmonella virus SETP3, Salmonella virus SETP7, Salmonella virus SETP13, Salmonella virus SP101, Salmonella virus SS3e, Salmonella virus wksl3、Escherichia virus K1G、Escherichia virus K1H、Escherichia virus K1ind1、Escherichia virus K1ind2、Escherichia virus Golestan、Raoultella virus RP180、Gordoniavirus Asapag、Gordonia virus BENtherdunthat、Gordonia virus Getalong、Gordonia virus Kenna、Gordonia virus Horus、Gordonia virus Phistory、Leuconostoc virus Lmd1、Leuconostoc virus LN03、Leuconostoc virus LN04、Leuconostoc virus LN12、Leuconostoc virus LN6B、Leuconostoc virus P793、Leuconostoc virus 1A4、Leuconostoc virus Ln8、Leuconostoc virus Ln9、Leuconostoc virus LN25、Leuconostoc virus LN34、Leuconostoc virus LNTR3、Mycobacterium virus Bongo、Mycobacterium virus Rey、Mycobacterium virus Butters、Mycobacterium virus Michelle、Mycobacterium virus Charlie、Mycobacterium virus Pipsqueaks、Mycobacterium virus Xeno、Mycobacterium virus Panchino、Mycobacterium virus Phrann、Mycobacterium virus Redi、Mycobacterium virus Skinnyp、Gordonia virus BaxterFox、Gordonia virus Yeezy、Gordonia virus Kita、Gordonia virus Nymphadora、Gordonia virus Zirinka、Mycobacterium virus Bignuz、Mycobacterium virus Brusacoram、Mycobacterium virus Donovan、Mycobacterium virus Fishburne、Mycobacterium virus Jebeks、Mycobacterium virusMalithi、Mycobacterium virus Phayonce、Lactobacillus virus B2、Lactobacillus virus Lenus、Lactobacillus virus Nyseid、Lactobacillus virus SAC12、Lactobacillus virus Ldl1、Lactobacillus virus ViSo2018a、Lactobacillus virus Maenad、Lactobacillus virus P1、Lactobacillus virus Satyr、Streptomyces virus AbbeyMikolon、Pseudomonas virus Ab18、Pseudomonas virus Ab19、Pseudomonas virus PaMx11、Burkholderia virus AH2、Arthrobacter virus Amigo、Arthrobacter virus Molivia、Propionibacterium virus Anatole、Propionibacterium virus B3、Arthrobacter virus Andrew、Bacillus virus Andromeda、Bacillus virus Blastoid、Bacillus virus Curly、Bacillus virus Eoghan、Bacillus virus Finn、Bacillus virus Glittering、Bacillus virus Riggi、Bacillus virus Taylor、Microbacterium virus Appa、Gordonia virus Apricot、Microbacterium virus Armstrong、Gordonia virus Attis、Streptomyces virus Attoomi、Streptomyces virus Austintatious、Streptomyces virus Ididsumtinwong、Streptomyces virus PapayaSalad、Gordonia virus Bantam、Mycobacterium virusBarnyard、Mycobacterium virus Konstantine、Mycobacterium virus Predator、Pseudomonas virus B3、Pseudomonas virus JBD67、Pseudomonas virus JD18、Pseudomonas virus PM105、Mycobacterium virus Bernal13、Gordonia virus BetterKatz、Streptomyces virus Bing、Staphylococcus virus 13、Staphylococcus virus 77、Staphylococcus virus 108PVL、Gordonia virus Bowser、Arthrobacter virus Bridgette、Arthrobacter virus Constance、Arthrobacter virus Eileen、Arthrobacter virus Judy、Arthrobacter virus Peas、Gordonia virus Britbrat、Mycobacterium virus Bron、Mycobacterium virus Faith1、Mycobacterium virus JoeDirt、Mycobacterium virus Rumpelstiltskin、Streptococcus virus 858、Streptococcus virus 2972、Streptococcus virus ALQ132、Streptococcus virus O1205、Streptococcus virus Sfi11、Pseudomonas virus D3112、Pseudomonas virus DMS3、Pseudomonas virus FHA0480、Pseudomonas virus LPB1、Pseudomonas virus MP22、Pseudomonas virus MP29、Pseudomonas virus MP38、Pseudomonas virus PA1KOR、Cellulophaga virus ST、Bacillus virus 250、Bacillus virusIEBH、Lactococcus virus bIL67、Lactococcus virus c2、Corynebacterium virus C3PO、Corynebacterium virus Darwin、Corynebacterium virus Zion、Lactobacillus virus c5、Lactobacillus virus Ld3、Lactobacillus virus Ld17、Lactobacillus virus Ld25A、Lactobacillus virus LLKu、Lactobacillus virus phiLdb、Mycobacterium virus Che9c、Mycobacterium virus Sbash、Mycobacterium virus Ardmore、Mycobacterium virus Avani、Mycobacterium virus Boomer、Mycobacterium virus Che8、Mycobacterium virus Che9d、Mycobacterium virus DeadP、Mycobacterium virus Dlane、Mycobacterium virus Dorothy、Mycobacterium virus DotProduct、Mycobacterium virus Drago、Mycobacterium virus Fruitloop、Mycobacterium virus GUmbie、Mycobacterium virus Ibhubesi、Mycobacterium virus Llij、Mycobacterium virus Mozy、Mycobacterium virus Mutaforma13、Mycobacterium virus Pacc40、Mycobacterium virus PMC、Mycobacterium virus Ramsey、Mycobacterium virus RockyHorror、Mycobacterium virus SG4、Mycobacterium virus Shauna1、Mycobacterium virus Shilan、Mycobacterium virusSpartacus、Mycobacterium virus Taj、Mycobacterium virus Tweety、Mycobacterium virus Wee、Mycobacterium virus Yoshi、Salmonella virus Chi、Salmonella virus FSLSP030、Salmonella virus FSLSP088、Sal monella virus iEPS5, Salmonella virus SPN19, Corynebacterium virus P1201, Clavibacter virus CMP1, Clavibacter virus CN1A, Lactobacillus virus ATCC8014, Lactobacillus virus phiJL1, Pediococcus virus cIP1, Arthrobacter virus Coral, Arthrobacter virus Kepler, Mycobacterium virus Corndog, Mycobacterium virus Firecracker, Rhodobacter virus RcCronus, Gordonia virus DareDevil, Arthrobacter virus Decurro, Stenotrophomonas virus DLP5, Gordonia virus Demosthenes, Gordonia virus Katyusha, Gordonia virus Kvothe, Pseudomonas virus D3, Pseudomonas virus PMG1, Escherichia virus EK99P1, Escherichia virus HK578, Escherichia virus JL1, Escherichia virus SSL2009a, Escherichia virus YD2008s, Shigella virus EP23, Sodalis virus SO1, Microbacterium virus Dismas, Propionibacterium virus B22, Propionibacterium virus Doucette, Propionibacterium virus E6, Propionibacterium virus G4, Microbacterium virus Eden, Enterococcus virus AL2, Enterococcus virus AL3, Enterococcus virus AUEF3, Enterococcus virus EcZZ2, Enterococcus virus EF3, Enterococcus virusEF4, Enterococcus virus EfaCPT1, Enterococcus virus IME196, Enterococcus virus LY0322, Enterococcus virus phiSHEF2, Enterococcus virus phiSHEF4, Enterococcus virus phiSHEF5, Enterococcus virus PMBT2, Enterococcus virus SANTOR1, Edwardsiella virus eiAU, Xanthomonas virus PhiL7, Microbacterium virus Eleri, Gordonia virus Cozz, Gordonia virus Emalyn, Gordonia virus GTE2, Gordonia virus Troje, Gordonia virus Eyre, Gordonia virus Fairfaxidumvirus, Microbacterium virus ISF9, Erwinia virus Eho49, Erwinia virus Eho59, Staphylococcus virus 2638A, Staphylococcus virus QT1, Colwellia virus 9A, Mycobacterium virus Alma, Mycobacterium virus Arturo, Mycobacterium virus Astro, Mycobacterium virus Backyardigan, Mycobacterium virus Benedict, Mycobacterium virus Bethlehem, Mycobacterium virus Billknuckles, Mycobacterium virus BPBiebs31, Mycobacterium virus Bruns、Mycobacterium virus Bxb1、Mycobacterium virus Bxz2、Mycobacterium virus Che12、Mycobacterium virus Cuco、Mycobacterium virus D29、Mycobacterium virus Doom、Mycobacterium virusEricb、Mycobacterium virus Euphoria、Mycobacterium virus George、Mycobacterium virus Gladiator、Mycobacterium virus Goose、Mycobacterium virus Hammer、Mycobacterium virus Heldan、Mycobacterium virus Jasper、Mycobacterium virus JC27、Mycobacterium virus Jeffabunny、Mycobacterium virus JHC117、Mycobacterium virus KBG、Mycobacterium virus Kssjeb、Mycobacterium virus Kugel、Mycobacterium virus L5、Mycobacterium virus Lesedi、Mycobacterium virus LHTSCC、Mycobacterium virus lockley、Mycobacterium virus Marcell、Mycobacterium virus Microwolf、Mycobacterium virus Mrgordo、Mycobacterium virus Museum、Mycobacterium virus Nepal、Mycobacterium virus Packman、Mycobacterium virus Peaches、Mycobacterium virus Perseus、Mycobacterium virus Pukovnik、Mycobacterium virus Rebeuca、Mycobacterium virus Redrock、Mycobacterium virus Ridgecb、Mycobacterium virus Rockstar、Mycobacterium virus Saintus、Mycobacterium virus Skipole、Mycobacterium virus Solon、Mycobacterium virus Switzer、Mycobacterium virus SWU1、Mycobacterium virusTiger, Mycobacterium virus Timshel, Mycobacterium virus Trixie, Mycobacterium virus Turbido, Mycobacterium virus Twister, Mycobacterium virus U2, Mycobacterium virus Violet, Mycobacterium virus Wonder, Mycobacterium virus Gaia, Arthrobacter virus Abidatro, Arthrobacter virus Galaxy, Gordonia virus GAL1, Gordonia virus GMA3, Gordonia virus Gsput1, Gordonia virus GMA7, Gordonia virus GTE7, Gordonia virus Ghobes, Mycobacterium virus Giles, Microbacterium virus OneinaGillian, Gordonia virus GodonK, Microbacterium virus Goodman, Arthrobacter virus Captnmurica, Arthrobacter virus Gordon, Gordonia virus GordTnk2, Proteus virus Isfahan, Gordonia virus Jumbo, Gordonia virus Gustav, Gordonia virus Mahdia, Paenibacillus virus Harrison, Gordonia virus Hedwig, Cellulophaga virus Cba121, Cellulophaga virus Cba171, Cellulophaga virus Cba181, Escherichia virus HK022, Escherichia virus HK75、Escherichia virus HK97、Escherichia virus HK106、Escherichia virus HK446、Escherichia virus HK542、Escherichia virus HK544、Escherichia virusHK633、Escherichia virus mEp234、Escherichia virus mEpX1、Escherichia virus mEpX2、Streptomyces virus Hiyaa、Salinibacter virus M1EM1、Salinibacter virus M8CR30-2、Listeria virus LP26、Listeria virus LP37、Listeria virus LP110、Listeria virus LP114、Listeria virus P70、Corynebacterium virus phi673、Corynebacterium virus phi674、Microbacterium virus Hamlet、Microbacterium virus Ilzat、Polaribacter virus P12002L、Polaribacter virus P12002S、Nonlabens virus P12024L、Nonlabens virus P12024S、Gordonia virus Jace、Brevibacillus virus Jenst、Corynebacterium virus Juicebox、Salinibacter virus M31CR41-2、Salinibacter virus SRUTV1、Arthrobacter virus Kellezzio、Arthrobacter virus Kitkat、Burkholderia virus KL1、Xanthomonas virus CP1、Microbacterium virus Golden、Microbacterium virus Koji、Arthrobacter virus Bennie、Arthrobacter virus DrRobert、Arthrobacter virus Glenn、Arthrobacter virus HunterDalle、Arthrobacter virus Joann、Arthrobacter virus Korra、Arthrobacter virus Preamble、Arthrobacter virus Pumancara、Arthrobacter virusWayne、Mycobacterium virus 244、Mycobacterium virus Bask21、Mycobacterium virus CJW1、Mycobacterium virus Eureka、Mycobacterium virus Kostya、Mycobacterium virus P orky、Mycobacterium virus Pumpkin、Mycobacterium virus Sirduracell、Mycobacterium virus Toto、Microbacterium virus Krampus、Salinibacter virus M8CC19、Salinibacter virus M8CRM1、Sphingobium virus Lacusarx、Escherichia virus DE3、Escherichia virus HK629、Escherichia virus HK630、Escherichia virus Lambda、Pseudomonas virus Lana、Arthrobacter virus Laroye、Eggerthella virus PMBT5、Arthrobacter virus Liebe、Mycobacterium virus Halo、Mycobacterium virus Liefie、Acinetobacter virus IMEAB3、Acinetobacter virus Loki、Streptomyces virus phiBT1、Streptomyces virus phiC31、Brevibacterium virus LuckyBarnes、Gordonia virus Lucky10、Faecalibacterium virus Lugh、Bacillus virus BMBtp2、Bacillus virus TP21、Bacillus virus Mgbh1、Arthrobacter virus Maja、Arthrobacter virus DrManhattan、Mycobacterium virus Ff47、Mycobacterium virus Muddy、Vibrio virus MAR10、Vibrio virus SSP002、Mycobacterium virus Marvin、Mycobacterium virus Mosmoris、Pseudomonas virus PMBT3、Microbacterium virus MementoMori、Microbacterium virus Fireman、Microbacteriumvirus Metamorphoo, Microbacterium virus RobsFeet, Microbacterium virus Min1, Streptococcus virus 7201, Streptococcus virus DT1, Streptococcus virus phiAbc2, Streptococcus virus Sfi19, Streptococcus virus Sfi21, Gordonia virus Birksandsocks, Gordonia virus Flakey, Gordonia virus Monty, Gordonia virus Stevefrench, Arthrobacter virus Circum, Arthrobacter virus Mudcat, Escherichia virus EC2, Salmonella virus Lumpael, Dinoroseobacter virus D5C, Burkholderia virus BcepNazgul, Microbacterium virus Neferthena, Pseudomonas virus nickie, Pseudomonas virus NP1, Pseudomonas virus PaMx25, Escherichia virus 9g, Escherichia virus JenK1, Escherichia virus JenP1, Escherichia virus JenP2, Salmonella virus SE1Kor, Salmonella virus 9NA, Salmonella virus SP069, Gordonia virus Nyceirae, Faecalibacterium virus Oengus, Mycobacterium virus Baka, Mycobacterium virus Courthouse, Mycobacterium virus Littlee, Mycobacterium virus Omega, Mycobacterium virus Optimus, Mycobacterium virus Thibault, Gordonia virus BrutonGaster, Gordonia virus OneUp, GordoniaOrchid virus, Thermus virus P23-45, Thermus virus P74-26, Propionibacterium virus ATCC29399BC, Propionibacterium virus ATCC29399BT, Propionibacterium virus Attacne, Propionibacterium virus Keiki, Propionibacterium virus Kubed, Propionibacterium virus Lauchelly, Propionibacterium virus MrAK, Propionibacterium virus Ouroboros, Propionibacterium virus P91, Propionibacterium virus P105, Propionibacterium virus P144, Propionibacterium virus P1001, Propionibacterium virus P1.1, Propionibacterium virus P100A, Propionibacterium virus P100D, Propionibacterium virus P101A, Propionibacterium virus P104A, Propionibacterium virus PA6, Propionibacterium virus Pacnes201215, Propionibacterium virus PAD20, Propionibacterium virus PAS50, Propionibacterium virus PHL009M11, Propionibacterium virus PHL025M00, Propionibacterium virus PHL037M02、Propionibacterium virus PHL041M10、Propionibacterium virus PHL060L00、Propionibacterium virus PHL067M01、Propionibacterium virus PHL070N00、Propionibacterium virus PHL071N05、Propionibacteriumvirus PHL082M03、Propionibacterium virus PHL092M00、Propionibacterium virus PHL095N00、Propionibacterium virus PHL111M01、Propionibacterium virus PHL112N00、Propionibacterium virus PHL113M01、Propionibacterium virus PHL114L00、Propionibacterium virus PHL116M00、Propionibacterium virus PHL117M00、Propionibacterium virus PHL117M01、Propionibacterium virus PHL132N00、Propionibacterium virus PHL141N00、Propionibacterium virus PHL151M00、Propionibacterium virus PHL151N00、Propionibacterium virus PHL152M00、Propionibacterium virus PHL163M00、Propionibacterium virus PHL171M01、Propionibacterium virus PHL179M00、Propionibacterium virus PHL194M00、Propionibacterium virus PHL199M00、Propionibacterium virus PHL301M00、Propionibacterium virus PHL308M00、Propionibacterium virus Pirate、Propionibacterium virus Procrass1、Propionibacterium virus SKKY、Propionibacterium virus Solid、Propionibacterium virus Stormborn、Propionibacterium virus Wizzo、Pseudomonas virus PaMx28、Pseudomonas virus PaMx74、Mycobacterium virusPapyrus、Mycobacterium virus Send513、Mycobacterium virus Patience、Mycobacterium virus PBI1、Rhodococcus virus Pepy6、Rhodococcus virus Poco6、Staphylococcus virus 11、Staphylococcus virus 29、Staphylococcus virus 37、Staphylococcus virus 53、Staphylococcus virus 55、Staphylococcus virus 69、Staphylococcus virus 71、Staphylococcus virus 80、Staphylococcus virus 85、Staphylococcus virus 88、Staphylococcus virus 92、Staphylococcus virus 96、Staphylococcus virus 187、Staphylococcus virus 52a、Staphylococcus virus 80alpha、Staphylococcus virus CNPH82、Staphylococcus virus EW、Staphylococcus virus IPLA5、Staphylococcus virus IPLA7、Staphylococcus virus IPLA88、Staphylococcus virus PH15、Staphylococcus virus phiETA、Staphylococcus virus phiETA2、Staphylococcus virus phiETA3、Staphylococcus virus phiMR11、Staphylococcus virus phiMR25、Staphylococcus virus phiNM1、Staphylococcus virus phiNM2、Staphylococcus virus phiNM4、Staphylococcus virus SAP26、Staphylococcus virus X2、Enterococcus virus FL1、Enterococcus virusFL2、Enterococcus virus FL3、Streptomyces virus Picard、Microbacterium virus Pikmin、Corynebacterium virus Poushou、Providencia virus PR1、Listeria virus LP302、 Listeria virus PSA、Psimunavirus psiM2、Propionibacterium virus PFR1、Microbacterium phage KaiHaiDragon、Microbacterium phage Paschalis、Microbacterium phage Quhwah、Streptomyces virus Darolandstone、Streptomyces virus Raleigh、Escherichia virus N15、Rhodococcus virus RER2、Rhizobium virus P106B、Streptomyces virus Drgrey、Streptomyces virus Rima、Microbacterium virus Hendrix、Gordonia virus Fryberger、Gordonia virus Ronaldo、Aeromonas virus pIS4A、Streptomyces virus Rowa、Gordonia virus Ruthy、Streptomyces virus Jay2Jay、Streptomyces virus Mildred21、Streptomyces virus NootNoot、Streptomyces virus Paradiddles、Streptomyces virus Peebs、Streptomyces virus Samisti12、Pseudomonas virus SM1、Corynebacterium virus SamW、Xylella virus Salvo、Xylella virus Sano、Caulobacter virus Sansa、Enterococcus virus BC611、Enterococcus virus IMEEF1、Enterococcus virus SAP6、Enterococcus virus VD13、Streptococcus virus SPQS1、Salmonella virus Sasha、Corynebacterium virus BFK20、Geobacillus virus Tp84、Streptomyces virus Scap1、Gordonia virusSchnabeltier、Microbacterium virus Schubert、Pseudomonas virus 73、Pseudomonas virus Ab26、Pseudomonas virus Kakheti25、Escherichia virus Cajan、Escherichia virus Seurat、Caulobacter virus Seuss、Staphylococcus virus SEP9、Staphylococcus virus Sextaec、Paenibacillus virus Diva、Paenibacillus virus Hb10c2、Paenibacillus virus Rani、Paenibacillus virus Shelly、Paenibacillus virus Sitara、Paenibacillus virus Willow、Lactococcus virus 712、Lactococcus virus ASCC191、Lactococcus virus ASCC273、Lactococcus virus ASCC281、Lactococcus virus ASCC465、Lactococcus virus ASCC532、Lactococcus virus Bibb29、Lactococcus virus bIL170、Lactococcus virus CB13、Lactococcus virus CB14、Lactococcus virus CB19、Lactococcus virus CB20、Lactococcus virus jj50、Lactococcus virus P2、Lactococcus virus P008、Lactococcus virus sk1、Lactococcus virus Sl4、Bacillus virus Slash、Bacillus virus Stahl、Bacillus virus Staley、Bacillus virus Stills、Gordonia virus Bachita、Gordonia virus ClubL、Gordonia virus Smoothie、Arthrobacter virus Sonali、Gordonia virusSoups、Gordonia virus Strosahl、Gordonia virus Wait、Gordonia virus Sour、Bacillus virus SPbeta、Microbacterium virus Hyperion、Microbacterium virus Squash、Burkholderia virus phi6442、Burkholderia virus phi1026b、Burkholderia virus phiE125、Achromobacter virus 83-24、Achromobacter virus JWX、Arthrobacter virus Tank、Gordonia virus Suzy、Gordonia virus Terapin、Streptomyces virus TG1、Mycobacterium virus Anaya、Mycobacterium virus Angelica、Mycobacterium virus CrimD、Mycobacterium virus Fionnbharth、Mycobacterium virus JAWS、Mycobacterium virus Larva、Mycobacterium virus MacnCheese、Mycobacterium virus Pixie、Mycobacterium virus TM4、Tsukamurella virus TIN2、Tsukamurella virus TIN3、Tsukamurella virus TIN4、Rhodobacter virus RcSpartan、Rhodobacter virus RcTitan、Mycobacterium virus Tortellini、Staphylococcus virus 47、Staphylococcus virus 3a、Staphylococcus virus 42e、Staphylococcus virus IPLA35、Staphylococcus virus phi12、Staphylococcus virus phiSLT、Mycobacterium virus 32HC、Rhodococcus virus Trina、Gordonia virusTrine、Paenibacillus virus Tripp、Flavobacterium virus 1H、Flavobacterium virus 23T、Flavobacterium virus 2A、Flavobacterium virus 6H、Streptomyces virus Lilbooboo、Streptomyces virus Vash、Paenibacillus virus Vegas、Gordonia virus Vendetta、Paracoccus virus Shpa、Pantoea virus Vid5、Acinetobacter virus B1251、Acinetobacter virus R3177、Gordonia virus Brandonk123、Gordonia virus Lennon、Gordonia virus Vivi2、Bordetella virus CN1、Bordetella virus CN2、Bordetella virus FP1、Bordetella virus MW2、Bacillus virus Wbeta、Rhodococcus virus Weasel、Mycobacterium virus Wildcat、Gordonia virus Billnye、Gordonia virus Twister6、Gordonia virus Wizard、Gordonia virus Hotorobo、Gordonia virus Woes、Streptomyces virus TP1604、Streptomyces virus YDN12、Roseobacter virus RDJL1、Roseobacter virus RDJL2、Xanthomonas virus OP1、Xanthomonas virus Xop411、Xanthomonas virus Xp10、Arthrobacter virus Yang、Alphaproteobacteria virus phiJl001、Pseudomonas virus LKO4、Pseudomonas virus M6、Pseudomonas virus MP1412、Pseudomonas virus PAE1、Pseudomonasvirus Yua、Gordonia virus Yvonnetastic、Microbacterium virus Zeta1847、Rhodococcus virus RGL3、Paenibacillus virus Lily、Vibrio virus CTXphi、Propionibacterium virus B5、Vibrio virus KSF1、Xanthomonas virus Cf1c、Vibrio virus fs1、Vibrio virus VGJ、Ralstonia virus RS551、Ralstonia virus RS603、Ralstonia virus RSM1、Ralstonia virus RSM3、Escherichia virus If1、Escherichia virus M13、Escherichia virus I22、Salmonella virus IKe、Ralstonia virus PE226、Pseudomonas virus Pf1、Stenotrophomonas virus PSH1、Ralstonia virus RSS1、Vibrio virus fs2、Vibrio virus VFJ、Stenotrophomonas virus SMA6、Stenotrophomonas virus SMA9、Stenotrophomonas virus SMA7、Pseudomonas virus Pf3、Thermus virus OH3、Vibrio virus VfO3K6、Vibrio virus VCY、Vibrio virus Vf33、Xanthomonas virus Xf109、Acholeplasma virus L51、Spiroplasma virus SVTS2、Spiroplasma virus C74、Spiroplasma virus R8A2B、Spiroplasma virus SkV1CR23x、Escherichia virus alpha3、Escherichia virus ID21、Escherichia virus ID32、Escherichia virus ID62、Escherichia virus NC28、Escherichia virusNC29, Escherichia virus NC35, Escherichia virus phiK, Escherichia virus St1, Escherichia virus WA45, Escherichia virus G4, Escherichia v irus ID52, Escherichia virus Talmos, Escherichia virus phiX174, Bdellovibrio virus MAC1, Bdellovibrio virus MH2K, Chlamydia virus Chp1, Chlamydia virus Chp2, Chlamydia virus CPAR39, Chlamydia virus CPG1, Spiroplasma virus SpV4, Bombyx mori bidensovirus, Acerodon celebensis polyomavirus 1, Artibeus planirostris polyomavirus 2, Artibeus planirostris polyomavirus 3, Ateles paniscus polyomavirus 1, Cardioderma cor polyomavirus 1, Carollia perspicillata polyomavirus 1. Chlorocebus pygerythrus polyomavirus 1. Chlorocebus pygerythrus polyomavirus 3. Dobsonia moluccensis polyomavirus 1. Eidolon helvum polyomavirus 1. Gorilla gorilla polyomavirus 1. 5. 5. 8. 8. 5. 8. 8. 5. 8. 8. 8. 8 Levi 9, Lilac 13. Macaca fascicularis polyomavirus 1, Mesocricetus auratus polyomavirus 1, Miniopterus schreibersii polyomavirus 1, Miniopterus schreibersii polyomavirus 2, Molossus molossus polyomavirus 1, Mus musculus polyomavirus 1, Otomops martienseni polyomavirus 1, Otomops martienseni polyomavirus 2、Pan troglodytes polyomavirus 1、Pan troglodytes polyomavirus 2、Pan troglodytes polyomavirus 3、Pan troglodytes polyomavirus 4、Pan troglodytes polyomavirus 5、Pan troglodytes polyomavirus 6、Pan troglodytes polyomavirus 7、Papio cynocephalus polyomavirus 1、Piliocolobus badius polyomavirus 1、Piliocolobus rufomitratus polyomavirus 1、Pongo abelii polyomavirus 1、Pongo pygmaeus polyomavirus 1、Procyon lotor polyomavirus 1、Pteropus vampyrus polyomavirus 1、Rattus norvegicus polyomavirus 1、Sorex araneus polyomavirus 1、Sorex coronatus polyomavirus 1、Sorex minutus polyomavirus 1、Sturnira lilium polyomavirus 1、Tupaia belangeri polyomavirus 1、Acerodon celebensis polyomavirus 2、Artibeus planirostris polyomavirus 1、Canis familiaris polyomavirus 1、Cebus albifrons polyomavirus 1、Cercopithecus erythrotis polyomavirus 1、Chlorocebus pygerythrus polyomavirus 2、Desmodus rotundus polyomavirus 1、Dobsonia moluccensis polyomavirus 2、Dobsonia moluccensis polyomavirus 3、Enhydra lutris polyomavirus 1、Equus caballus polyomavirus 1, human polyomavirus 1, human polyomavirus 2, human polyomavirus 3, human polyomavirus 4, Leptonychotes weddellii polyomavirus 1, Loxodonta africana polyomavirus 1, Macaca mulatta polyomavirus 1, Mastomys natalensis polyomavirus 1, Meles meles polyomavirus 1, Microtus arvalis polyomavirus 1, Miniopterus africanus polyomavirus 1, Mus musculus polyomavirus 2, Mus musculus polyomavirus 3, Myodes glareolus polyomavirus 1, Myotis lucifugus polyomavirus 1, Pan troglodytes polyomavirus 8, Papio cynocephalus polyomavirus 2, Pteronotus davyi polyomavirus 1, Pteronotus parnellii polyomavirus 1. Rattus norvegicus polyomavirus 2. Rousettus aegyptiacus polyomavirus 1. Saimiri boliviensis polyomavirus 1. Saimiri sciureus polyomavirus 1. Vicugna pacos polyomavirus 1. Zalophus californianus polyomavirus 1. Human polyomavirus 6. Human polyomavirus 7. Human polyomavirus 10. Human polyomavirus 11. Anser anser polyomavirus 1. Aves polyomavirus 1. Corvus monedula polyomavirus 1. Cracticus torquatus polyomavirus 1. Erythrura gouldiae polyomavirus 1,Lonchura maja polyomavirus 1, Pygoscelis adeliae polyomavirus 1, Pyrrhula pyrrhula polyomavirus 1, Serinus canaria polyomavirus 1, Ailuropoda melanoleuca polyomavirus 1, Bos taurus polyomavirus 1, Centropristis striata polyomavirus 1, Delphinus delphis polyomavirus 1, Procyon lotor polyomavirus 2, Rhynchobatus djiddensis polyomavirus 1, Sparus aurata polyomavirus 1, Trematomus bernacchii polyomavirus 1, Trematomus pennellii polyomavirus 1. Alpha papillomavirus 1, Alpha papillomavirus 2, Alpha papillomavirus 3, Alpha papillomavirus 4, Alpha papillomavirus 5, Alpha papillomavirus 6, Alpha papillomavirus 7, Alpha papillomavirus 8, Alpha papillomavirus 9, Alpha papillomavirus 10, Alpha papillomavirus 11, Alpha papillomavirus 12, Alpha papillomavirus 13, Alpha papillomavirus 14, Beta papillomavirus 1, Beta papillomavirus 2, Beta papillomavirus 3, Beta papillomavirus 4, Beta papillomavirus -mavirus 5, beta papillomavirus 6, ky papillomavirus 1, ky papillomavirus 2, ky papillomavirus 3, delta papillomavirus 1, delta papillomavirus 2, delta papillomavirus 3, delta papillomavirus 4, delta papillomavirus 5, delta papillomavirus 6, delta papillomavirus 7, diok...
Claims
1. A tandem array comprising at least two types of CRISPR RNA (crRNA), wherein each crRNA comprises a guide sequence and a direct repeat (DR) sequence, each DR sequence in the tandem array is different, and each DR sequence is a PspCas13b direct repeat sequence containing only at least one T17 or T18 nucleotide mutation of SEQ ID NO:
267.
2. The tandem array according to claim 1, wherein each guide array in the tandem array is different.
3. The tandem array according to claim 1 or 2, wherein each crRNA comprises a DR sequence independently selected from sequence numbers 268 to 273.
4. A tandem array according to any one of claims 1 to 3, wherein the DR sequence is located at 3' of the guide sequence.
5. The tandem array according to any one of claims 1 to 4, wherein each guide sequence is independently complementary to a coronavirus genome mRNA sequence or a coronavirus subgenome mRNA sequence.
6. The tandem array according to claim 5, wherein each guide sequence is complementary to a coronavirus leader sequence, a coronavirus S sequence, a coronavirus E sequence, a coronavirus M sequence, an N sequence, or a coronavirus S2M sequence.
7. The tandem array according to claim 5 or 6, wherein each guide sequence is independently complementary to a sequence that is at least 95% identical to a sequence selected from sequence numbers 168-174, 176-181, 186, and 187.
8. The tandem array according to any one of claims 5 to 7, wherein each guide sequence includes a sequence that is at least 95% identical to a sequence selected from sequence numbers 189 to 208 and 216 to 224.
9. The tandem array according to any one of claims 5 to 8, wherein the tandem array includes an array that is at least 95% identical to sequence number 275.
10. The tandem array according to any one of claims 1 to 4, wherein each guide sequence is independently complementary to an influenza virus genomic RNA sequence or an influenza virus subgenomic RNA sequence.
11. The tandem array according to claim 10, wherein each guide sequence is complementary to the influenza virus PB2 sequence, influenza virus PB1 sequence, influenza virus PA sequence, influenza virus NP sequence, or influenza virus M sequence.
12. The tandem array according to claim 10 or 11, wherein each guide sequence is independently complementary to a sequence that is at least 95% identical to a sequence selected from sequence numbers 225 to 244.
13. The tandem array according to any one of claims 10 to 12, wherein each guide sequence includes a sequence that is at least 95% identical to a sequence selected from sequence numbers 245 to 249.
14. The tandem array according to any one of claims 5 to 8, wherein the tandem array includes sequences that are at least 95% identical to sequences selected from sequence numbers 276 and 277.
15. A composition comprising a tandem array according to any one of claims 1 to 14.
16. The composition according to claim 15, wherein the composition further comprises a Cas protein, the Cas protein being Cas13 that targets a PspCas13b direct repeat sequence.
17. The composition according to claim 16, wherein the Cas protein comprises a sequence that is at least 95% identical to a sequence selected from sequence numbers 1 and 2.
18. The composition according to claim 16 or 17, wherein the Cas protein further comprises a localization signal or a transport signal.
19. The composition according to claim 18, wherein the Cas protein comprises NES, and the NES comprises a sequence that is at least 95% identical to sequence numbers 58-59.
20. The composition according to claim 18, wherein the Cas protein comprises a nuclear localization signal (NLS), and the NLS comprises a sequence that is at least 95% identical to sequence numbers 50-57 and 323-935.
21. The composition according to claim 18, wherein the Cas protein contains a localization signal, and the localization signal contains a sequence that is at least 95% identical to sequence numbers 60-66.
22. A pharmaceutical composition for reducing the number of two or more unique target RNAs in a subject, comprising a tandem array and a Cas protein or nucleic acid encoding a Cas protein according to any one of claims 1 to 14, or a composition according to any one of claims 15 to 21.
Citation Information
Patent Citations
Crispr system based antiviral therapy
WO2019010422A1