Compositions and methods of nucleic acid targeting nucleic acid

By introducing mutations into the P-domain of the target nucleic acid, the binding specificity and dissociation constant of the target nucleic acid are optimized, solving the problems of specificity and dissociation constant of nucleic acid targeting in existing technologies, and realizing efficient genome engineering modification.

CN111454951BActive Publication Date: 2026-06-02CARIBOU BIOSCIENCES INC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CARIBOU BIOSCIENCES INC
Filing Date
2014-03-12
Publication Date
2026-06-02

Smart Images

  • Figure BDA0002443731720000141
    Figure BDA0002443731720000141
  • Figure BDA0002443731720000151
    Figure BDA0002443731720000151
  • Figure BDA0002443731720000391
    Figure BDA0002443731720000391
Patent Text Reader

Abstract

The present disclosure provides compositions and methods of use of nucleic acids and complexes thereof targeted to nucleic acids. Genomic engineering can refer to the alteration of a genome by deletion, insertion, mutation, or replacement of a particular nucleic acid sequence. The alteration can be gene or position specific. Genomic engineering can utilize nucleases to cleave nucleic acids, thereby generating a site for alteration. Engineering of non-genomic nucleic acids is also contemplated.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application is a divisional application of Chinese Patent Application No. 201480028062.1.

[0003] This application claims U.S. Provisional Application No. 61 / 818,386 (Agent No. 44287-710.101), filed May 1, 2013; U.S. Provisional Application No. 61 / 902,723 (Agent No. 44287-710.103), filed November 11, 2013; U.S. Provisional Application No. 61 / 818,382 (Agent No. 44287-712.101), filed May 1, 2013; U.S. Provisional Application No. 61 / 859,661 (Agent No. 44287-712.102), filed July 29, 2013; and U.S. Provisional Application No. 26, 2013. 61 / 858,767 [Agent's Case No. 44287-713.102], U.S. Provisional Application No. 61 / 822,002 filed May 10, 2013 [Agent's Case No. 44287-717.101], U.S. Provisional Application No. 61 / 832,690 filed June 7, 2013 [Agent's Case No. 44287-719.101], U.S. Provisional Application No. 61 / 906,211 filed November 19, 2013 [Agent's Case No. 44287-719.102], U.S. Provisional Application No. 61 / 900,311 filed November 5, 2013 [Agent's Case No. 44287-713.102] 19.103], U.S. Provisional Application No. 61 / 845,714 filed July 12, 2013 [Agent No. 44287-721.101], U.S. Provisional Application No. 61 / 883,804 filed September 27, 2013 [Agent No. 44287-721.102], U.S. Provisional Application No. 61 / 781,598 filed March 14, 2013 [Agent No. 44287-722.101], U.S. Provisional Application No. 61 / 899,712 filed November 4, 2013 [Agent No. 44287-727.101], U.S. Provisional Application No. 61 / 899,712 filed August 14, 2013 Please refer to U.S. Provisional Application No. 61 / 865,743 [Agent's Case No. 44287-733.101], U.S. Provisional Application No. 61 / 907,777 [Agent's Case No. 44287-734.101] filed November 22, 2013, U.S. Provisional Application No. 61 / 903,232 [Agent's Case No. 44287-751.101] filed November 12, 2013, U.S. Provisional Application No. 61 / 906,335 [Agent's Case No. 44287-752.101] filed November 19, 2013, and U.S. Provisional Application No. 61 / 907,216 [Agent's Case No. 44287-733.101] filed November 21, 2013.The rights and interests of [44287-753.101], the entire contents of which are hereby incorporated herein by reference.

[0004] sequence list

[0005] This application contains a sequence list that has been electronically submitted in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy created on March 10, 2014, is named 44287-722-601_SeqList and is 7,828,964 bytes in size. Background of the Invention

[0007] Genome engineering refers to altering the genome by deleting, inserting, mutating, or substituting specific nucleic acid sequences. This alteration can be gene- or location-specific. Genome engineering utilizes nucleases to cleave nucleic acids, thereby creating sites for modification. The engineering of non-genomic nucleic acids is also envisioned. Proteins containing nuclease domains can bind to and cleave target nucleic acids by forming complexes with them. In one instance, this cleavage can introduce double-strand breaks in the target nucleic acid. Nucleic acids can be repaired, for example, through endogenous non-homologous end joining (NHEJ) mechanisms. In another instance, a segment of nucleic acid can be inserted. N-modifications of target nucleic acids and site-directed peptides can introduce novel functions for genome engineering. Invention Overview

[0009] On one hand, this disclosure provides an engineered nucleic acid targeting a nucleic acid, comprising: a mutation in the P-domain of the nucleic acid targeting the nucleic acid. In some embodiments, the P-domain begins downstream of the last nucleotide pair of the double strand between the CRISPR repeat and the tracrRNA sequence of the nucleic acid targeting the nucleic acid. In some embodiments, the engineered nucleic acid targeting the nucleic acid further comprises an adapter sequence. In some embodiments, the adapter sequence links the CRISPR repeat and the tracrRNA sequence. In some embodiments, the engineered nucleic acid targeting the nucleic acid is an isolated engineered nucleic acid targeting the nucleic acid. In some embodiments, the engineered nucleic acid targeting the nucleic acid is a recombinant engineered nucleic acid targeting the nucleic acid. In some embodiments, the engineered nucleic acid targeting the nucleic acid is adapted to hybridize with a target nucleic acid. In some embodiments, the P-domain comprises two adjacent nucleotides. In some embodiments, the P-domain comprises three adjacent nucleotides. In some embodiments, the P-domain comprises four adjacent nucleotides. In some embodiments, the P-domain comprises 5 adjacent nucleotides. In some embodiments, the P-domain comprises 6 or more adjacent nucleotides. In some embodiments, the P-domain begins 1 nucleotide downstream of the last pair of nucleotides in the duplex. In some embodiments, the P-domain begins 2 nucleotides downstream of the last pair of nucleotides in the duplex. In some embodiments, the P-domain begins 3 nucleotides downstream of the last pair of nucleotides in the duplex. In some embodiments, the P-domain begins 4 nucleotides downstream of the last pair of nucleotides in the duplex. In some embodiments, the P-domain begins 5 nucleotides downstream of the last pair of nucleotides in the duplex. In some embodiments, the P-domain begins 6 or more nucleotides downstream of the last pair of nucleotides in the duplex. In some embodiments, the mutation comprises one or more mutations. In some embodiments, the one or more mutations are adjacent to each other. In some embodiments, the one or more mutations are spaced apart from each other. In some embodiments, the mutation is adapted to hybridize the engineered nucleic acid-targeted nucleic acid with different protospace radjacent motifs. In some embodiments, the different pre-interstitial region sequence adjacent motifs contain at least 4 nucleotides. In some embodiments, the different pre-interstitial region sequence adjacent motifs contain at least 5 nucleotides. In some embodiments, the different pre-interstitial region sequence adjacent motifs contain at least 6 nucleotides. In some embodiments, the different pre-interstitial region sequence adjacent motifs contain at least 7 or more nucleotides.In some embodiments, the different pre-interstitial region sequence adjacent motifs comprise two non-adjacent regions. In some embodiments, the different pre-interstitial region sequence adjacent motifs comprise three non-adjacent regions. In some embodiments, the mutation is adapted to cause the engineered nucleic acid-targeting nucleic acid to bind to the target nucleic acid with a lower dissociation constant than the unengineered nucleic acid-targeting nucleic acid. In some embodiments, the mutation is adapted to cause the engineered nucleic acid-targeting nucleic acid to bind to the target nucleic acid with higher specificity than the unengineered nucleic acid-targeting nucleic acid. In some embodiments, the mutation is adapted to reduce the binding of the engineered nucleic acid-targeting nucleic acid to non-specific sequences in the target nucleic acid compared to the unengineered nucleic acid-targeting nucleic acid. In some embodiments, the engineered nucleic acid targeting nucleic acid further comprises two hairpin structures, wherein one of the two hairpin structures comprises a double strand between a polynucleotide comprising at least 50% identity with CRISPR RNA in six adjacent nucleotides and a polynucleotide comprising at least 50% identity with tracrRNA in six adjacent nucleotides; and wherein one of the two hairpin structures is the 3' of the first hairpin structure, and wherein the second hairpin structure comprises an engineered P-domain. In some embodiments, the second hairpin structure is adapted to unwind upon contact between the nucleic acid and the target nucleic acid. In some embodiments, the P-domain is adapted to: hybridize with a first polynucleotide comprising a region of the engineered nucleic acid targeting nucleic acid, hybridize with a second polynucleotide comprising the target nucleic acid, and specifically hybridize with the first or second polynucleotide. In some embodiments, the first polynucleotide comprises at least 50% identity with tracrRNA in six adjacent nucleotides. In some embodiments, the first polynucleotide is downstream of a duplex between a polynucleotide containing at least 50% identity with a CRISPR repeat fragment in six adjacent nucleotides and a polynucleotide containing at least 50% identity with a tracrRNA sequence in six adjacent nucleotides. In some embodiments, the second polynucleotide contains a motif adjacent to a pre-interstitial region sequence. In some embodiments, the engineered nucleic acid-targeting nucleic acid is adapted to bind to a site-directed polypeptide. In some embodiments, the mutation comprises inserting one or more nucleotides into the P-domain. In some embodiments, the mutation comprises deleting one or more nucleotides from the P-domain. In some embodiments, the mutation comprises a mutation of one or more nucleotides. In some embodiments, the mutation is constructed to allow the nucleic acid-targeting nucleic acid to hybridize with different pre-interstitial region sequence adjacent motifs.In some embodiments, the different pre-interstitial sequence neighbor motifs comprise pre-interstitial sequence neighbor motifs selected from: 5'-NGGNG-3', 5'-NNAAAAW-3', 5'-NNNNGATT-3', 5'-GNNNCNNA-3', and 5'-NNNACA-3', or any combination thereof. In some embodiments, the mutation is configured to cause the engineered nucleic acid-targeted nucleic acid to bind with a lower dissociation constant than that of an unengineered nucleic acid-targeted nucleic acid. In some embodiments, the mutation is configured to cause the engineered nucleic acid-targeted nucleic acid to bind with higher specificity than that of an unengineered nucleic acid-targeted nucleic acid. In some embodiments, the mutation is configured to reduce the binding of the engineered nucleic acid-targeted nucleic acid to non-specific sequences in the target nucleic acid compared to an unengineered nucleic acid-targeted nucleic acid.

[0010] On one hand, this disclosure provides a method for modifying a target nucleic acid, comprising contacting the target nucleic acid with an engineered nucleic acid-targeted nucleic acid, said engineered nucleic acid-targeted nucleic acid comprising: a mutation in the P-domain of said nucleic acid-targeted nucleic acid, and modifying the target nucleic acid. In some embodiments, the method further comprises inserting a donor polynucleotide into the target nucleic acid. In some embodiments, the modification comprises cleaving the target nucleic acid. In some embodiments, the modification comprises modifying the transcription of the target nucleic acid.

[0011] On the one hand, this disclosure provides a vector containing a polynucleotide sequence encoding an engineered nucleic acid-targeting nucleic acid, wherein the engineered nucleic acid-targeting nucleic acid contains a mutation in the P-domain of the nucleic acid-targeting nucleic acid.

[0012] On one hand, this disclosure provides a kit comprising: an engineered nucleic acid targeting a nucleic acid, comprising: a mutation in the P-domain of the nucleic acid targeting a nucleic acid; and a buffer. In some embodiments, the kit further comprises a site-directed peptide. In some embodiments, the kit further comprises a donor polynucleotide. In some embodiments, the kit further comprises instructions for use.

[0013] On one hand, this disclosure provides an engineered nucleic acid targeting a nucleic acid, comprising: a mutation in a protrusion region of the nucleic acid targeting the nucleic acid. In some embodiments, the protrusion is located within a double-stranded body between a CRISPR repeat fragment and a tracrRNA sequence of the nucleic acid targeting the nucleic acid. In some embodiments, the engineered nucleic acid targeting the nucleic acid further comprises an adapter sequence. In some embodiments, the adapter sequence links the CRISPR repeat fragment and the tracrRNA sequence. In some embodiments, the engineered nucleic acid targeting the nucleic acid is an isolated engineered nucleic acid targeting the nucleic acid. In some embodiments, the engineered nucleic acid targeting the nucleic acid is a recombinant engineered nucleic acid targeting the nucleic acid. In some embodiments, the protrusion comprises at least one unpaired nucleotide on the CRISPR repeat fragment and one unpaired nucleotide on the tracrRNA sequence. In some embodiments, the protrusion comprises at least one unpaired nucleotide on the CRISPR repeat fragment and at least two unpaired nucleotides on the tracrRNA sequence. In some embodiments, the protrusion comprises at least one unpaired nucleotide on the CRISPR repeat and at least three unpaired nucleotides on the tracrRNA sequence. In some embodiments, the protrusion comprises at least one unpaired nucleotide on the CRISPR repeat and at least four unpaired nucleotides on the tracrRNA sequence. In some embodiments, the protrusion comprises at least one unpaired nucleotide on the CRISPR repeat and at least five unpaired nucleotides on the tracrRNA sequence. In some embodiments, the protrusion comprises at least two unpaired nucleotides on the CRISPR repeat and one unpaired nucleotide on the tracrRNA sequence. In some embodiments, the protrusion comprises at least three unpaired nucleotides on the CRISPR repeat and at least two unpaired nucleotides on the tracrRNA sequence. In some embodiments, the protrusion comprises at least four unpaired nucleotides on the CRISPR repeat and at least three unpaired nucleotides on the tracrRNA sequence. In some embodiments, the protrusion comprises at least five unpaired nucleotides on a CRISPR repeat and at least four unpaired nucleotides on a tracrRNA sequence. In some embodiments, the protrusion comprises at least one nucleotide on a CRISPR repeat adapted to form a wobble pair with at least one nucleotide on the tracrRNA sequence. In some embodiments, the mutation comprises one or more mutations. In some embodiments, the one or more mutations are adjacent to each other. In some embodiments, the one or more mutations are spaced apart from each other. In some embodiments, the mutation is adapted to allow the engineered nucleic acid-targeted nucleic acid to bind to different site-directed peptides.In some embodiments, the distinct site-directed peptides are homologs of Cas9. In some embodiments, the distinct site-directed peptides are mutant forms of Cas9. In some embodiments, the distinct site-directed peptides contain 10% amino acid sequence identity with Cas9 in a nuclease domain selected from: RuvC nuclease domains and HNH nuclease domains or any combination thereof. In some embodiments, the mutation is adapted to allow the engineered nucleic acid-targeted nucleic acid to hybridize with adjacent motifs of the distinct pre-intercalation region. In some embodiments, the mutation is adapted to allow the engineered nucleic acid-targeted nucleic acid to bind to the site-directed peptide with a lower dissociation constant than the unengineered nucleic acid-targeted nucleic acid. In some embodiments, the mutation is adapted to allow the engineered nucleic acid-targeted nucleic acid to bind to the site-directed peptide with higher specificity than the unengineered nucleic acid-targeted nucleic acid. In some embodiments, the mutation is adapted to allow the engineered nucleic acid-targeted nucleic acid to guide the site-directed peptide to cleave the target nucleic acid with higher specificity than the unengineered nucleic acid-targeted nucleic acid. In some embodiments, the mutation is adapted to reduce the binding of the engineered nucleic acid-targeting nucleic acid to nonspecific sequences in the target nucleic acid compared to an unengineered nucleic acid-targeting nucleic acid. In some embodiments, the engineered nucleic acid-targeting nucleic acid is adapted to hybridize with the target nucleic acid. In some embodiments, the mutation comprises inserting one or more nucleotides into the protrusion. In some embodiments, the mutation comprises deleting one or more nucleotides from the protrusion. In some embodiments, the mutation comprises a mutation of one or more nucleotides. In some embodiments, the mutation is configured to allow the engineered nucleic acid-targeting nucleic acid to hybridize with adjacent motifs of different pre-interstitial regions compared to an unengineered nucleic acid-targeting nucleic acid. In some embodiments, the mutation is configured to allow the engineered nucleic acid-targeting nucleic acid to bind to a site-directed peptide with a lower dissociation constant than an unengineered nucleic acid-targeting nucleic acid. In some embodiments, the mutation is configured to allow the engineered nucleic acid-targeting nucleic acid to bind to a site-directed peptide with higher specificity than an unengineered nucleic acid-targeting nucleic acid. In some implementations, the mutation is constructed to reduce the binding of the engineered nucleic acid-targeted nucleic acid to nonspecific sequences in the target nucleic acid compared to an unengineered nucleic acid-targeted nucleic acid.

[0014] On one hand, this disclosure provides a method for modifying a target nucleic acid, comprising: contacting the target nucleic acid with an engineered nucleic acid-targeted nucleic acid, said engineered nucleic acid-targeted nucleic acid comprising: a mutation in a protrusion region of the nucleic acid-targeted nucleic acid; and modifying the target nucleic acid. In some embodiments, the method further comprises inserting a donor polynucleotide into the target nucleic acid. In some embodiments, the modification comprises cleaving the target nucleic acid. In some embodiments, the modification comprises modifying the transcription of the target nucleic acid.

[0015] On the one hand, this disclosure provides a vector comprising a polynucleotide sequence encoding an engineered nucleic acid-targeted nucleic acid, wherein the engineered nucleic acid-targeted nucleic acid comprises: a mutation in a protrusion region of the nucleic acid-targeted nucleic acid; and modification of the target nucleic acid.

[0016] On one hand, this disclosure provides a kit comprising: an engineered nucleic acid targeting a nucleic acid, comprising: a mutation in a protrusion region of the nucleic acid targeting the nucleic acid; and modification of the target nucleic acid; and a buffer. In some embodiments, the kit further comprises a site-directed peptide. In some embodiments, the kit further comprises a donor polynucleotide. In some embodiments, the kit further comprises instructions for use.

[0017] On one hand, this disclosure provides a method for manufacturing donor polynucleotide-labeled cells, comprising: lysing target nucleic acids in cells using a complex comprising a site-directed polypeptide and a nucleic acid targeting a nucleic acid; inserting a donor polynucleotide into the lysed target nucleic acid; proliferating cells carrying the donor polynucleotide; and determining the source of the donor polynucleotide-labeled cells. In some embodiments, the method is performed in vivo. In some embodiments, the method is performed in vitro. In some embodiments, the method is performed in situ. In some embodiments, the proliferation produces a cell population. In some embodiments, the proliferation produces a cell line. In some embodiments, the method further comprises determining the nucleic acid sequence of the nucleic acid in the cells. In some embodiments, the nucleic acid sequence determines the source of the cells. In some embodiments, the determination comprises determining the genotype of the cells. In some embodiments, the proliferation comprises differentiating the cells. In some embodiments, the proliferation comprises dedifferentiating the cells. In some embodiments, the proliferation comprises differentiating the cells and then dedifferentiating the cells. In some embodiments, the proliferation comprises passage the cells. In some embodiments, the proliferation comprises inducing cell division. In some embodiments, the proliferation comprises inducing the cells to enter the cell cycle. In some embodiments, the propagation includes the formation of transfer cells. In some embodiments, the propagation includes cells that differentiate into pluripotent cell differentiation components. In some embodiments, the cells are differentiated cells. In some embodiments, the cells are dedifferentiated cells. In some embodiments, the cells are stem cells. In some embodiments, the cells are pluripotent stem cells. In some embodiments, the cells are eukaryotic cell lines. In some embodiments, the cells are primary cell lines. In some embodiments, the cells are patient-derived cell lines. In some embodiments, the method further includes transplanting the cells into an organism. In some embodiments, the organism is a human. In some embodiments, the organism is a mammal. In some embodiments, the organism is selected from: humans, dogs, rats, mice, chickens, fish, cats, plants, and primates. In some embodiments, the method further includes selecting the cells. In some embodiments, the donor polynucleotide is inserted into a target nucleic acid expressed in one cell state. In some embodiments, the donor polynucleotide is inserted into a target nucleic acid expressed in multiple cell types. In some embodiments, the donor polynucleotide is inserted into a target nucleic acid expressed in a pluripotent state. In some embodiments, the donor polynucleotide is inserted into a target nucleic acid expressed in a differentiated state.

[0018] On one hand, this disclosure provides a method for producing a cloned and expanded cell line, comprising: introducing a complex comprising a site-directed polypeptide and a nucleic acid targeting a nucleic acid into a cell, contacting the complex with the target nucleic acid, lysing the target nucleic acid, wherein the lysis is performed by the complex to produce lysed target nucleic acid, inserting a donor polynucleotide into the lysed target nucleic acid, and proliferating the cell, wherein the proliferation produces a cloned and expanded cell line. In some embodiments, the cells are selected from: HeLa cells, Chinese hamster ovary cells, 293-T cells, pheochromocytoma, neuroblastoma fibroblasts, rhabdomyosarcoma, dorsal root ganglion cells, NSO cells, CV-I (ATCC CCL 70), COS-I (ATCC CRL 1650), COS-7 (ATCC CRL 1651), CHO-K1 (ATCC CCL 61), 3T3 (ATCC CCL 92), NIH / 3T3 (ATCC CRL 1658), HeLa (ATCC CCL 2), C 1271 (ATCC CRL 1616), BS-CI (ATCC CCL 26), MRC-5 (ATCC CCL 171), L-cells, HEK-293 (ATCC CRL 1573) and PC 12 (ATCC CRL-1721), HEK293T (ATCC CRL 1573), and PC 12 (ATCC CRL 1721). CRL-11268), RBL (ATCC CRL-1378), SH-SY5Y (ATCC CRL-2266), MDCK (ATCC CCL-34), SJ-RH30 (ATCC CRL-2061), HepG2 (ATCC HB-8065), ND7 / 23 (ECACC 92090903), CHO (ECACC 85050302), Vera (ATCC CCL 81), Caco-2 (ATCC HTB 37), K562 (ATCC CCL 243), Jurkat (ATCC TIB-152), Per.Có, Huvec (ATCC human primary PCS 100-010, mouse CRL 2514, CRL 2515, CRL 2516), HuH-7D12 (ECACC 01042712), 293 (ATCC CRL The cells may be 10852, A549 (ATCC CCL 185), IMR-90 (ATCC CCL 186), MCF-7 (ATC HTB-22), U-2 OS (ATCC HTB-96), and T84 (ATCC CCL 248), or any combination thereof. In some embodiments, the cells are stem cells. In some embodiments, the cells are differentiated cells. In some embodiments, the cells are pluripotent cells.

[0019] On one hand, this disclosure provides a method for multiplex cell type analysis, comprising: cleaving at least one target nucleic acid in two or more cell types using a complex comprising a site-directed polypeptide and a nucleic acid targeting a nucleic acid to produce two cleaved target nucleic acids; inserting different donor polynucleotides into each cleaved target nucleic acid; and analyzing the two or more cell types. In some embodiments, the analysis comprises simultaneously analyzing the two or more cell types. In some embodiments, the analysis comprises determining the sequence of the target nucleic acid. In some embodiments, the analysis comprises comparing the two or more cell types. In some embodiments, the analysis comprises determining the genotype of the two or more cell types. In some embodiments, the cells are differentiated cells. In some embodiments, the cells are dedifferentiated cells. In some embodiments, the cells are stem cells. In some embodiments, the cells are pluripotent stem cells. In some embodiments, the cells are eukaryotic cell lines. In some embodiments, the cells are primary cell lines. In some embodiments, the cells are patient-derived cell lines. In some embodiments, multiple donor polynucleotides are inserted into multiple cleaved target nucleic acids in the cells.

[0020] On one hand, this disclosure provides a composition comprising: an engineered nucleic acid targeting a nucleic acid including a 3' hybridizing extension, and a donor polynucleotide, wherein the donor polynucleotide hybridizes to the 3' hybridizing extension. In some embodiments, the 3' hybridizing extension is adapted to hybridize with at least five nucleotides from the 3' of the donor polynucleotide. In some embodiments, the 3' hybridizing extension is adapted to hybridize with at least five nucleotides from the 5' of the donor polynucleotide. In some embodiments, the 3' hybridizing extension is adapted to hybridize with at least five adjacent nucleotides of the donor polynucleotide. In some embodiments, the 3' hybridizing extension is adapted to hybridize with the entire donor polynucleotide. In some embodiments, the 3' hybridizing extension comprises a reverse transcription template. In some embodiments, the reverse transcription template is adapted to be reverse transcribed by a reverse transcriptase. In some embodiments, the composition further comprises a reverse-transcribed DNA polynucleotide. In some embodiments, the reverse-transcribed DNA polynucleotide is adapted to hybridize with the reverse transcription template. In some embodiments, the donor polynucleotide is DNA. In some embodiments, the 3' hybridizing extension is RNA. In some embodiments, the engineered nucleic acid targeting nucleic acid is an isolated engineered nucleic acid targeting nucleic acid. In some embodiments, the engineered nucleic acid targeting nucleic acid is a recombinant engineered nucleic acid targeting nucleic acid.

[0021] On one hand, this disclosure provides a method for introducing a donor polynucleotide into a target nucleic acid, comprising: contacting the target nucleic acid with a composition comprising: an engineered nucleic acid targeting a nucleic acid containing a 3' hybridization overhang, and a donor polynucleotide, wherein the donor polynucleotide hybridizes with the 3' hybridization overhang. In some embodiments, the method further comprises cleaving the target nucleic acid to produce cleaved target nucleic acid. In some embodiments, the cleavage is performed via a site-directed peptide. In some embodiments, the method further comprises inserting the donor polynucleotide into the cleaved target nucleic acid.

[0022] On one hand, this disclosure provides a composition comprising: an effector protein and a nucleic acid, wherein the nucleic acid contains at least 50% sequence identity with crRNA in 6 adjacent nucleotides and at least 50% sequence identity with tracrRNA in 6 adjacent nucleotides; and a non-natural sequence, wherein the nucleic acid is adapted to bind to the effector protein. In some embodiments, the composition further comprises a polypeptide containing at least 10% amino acid sequence identity with the nuclease domain of Cas9, wherein the nucleic acid binds to the polypeptide. In some embodiments, the polypeptide contains at least 60% amino acid sequence identity with the nuclease domain of Cas9. In some embodiments, the polypeptide is Cas9. In some embodiments, the nucleic acid further comprises an adapter sequence, wherein the adapter sequence is linked to a sequence containing at least 50% sequence identity with crRNA in 6 adjacent nucleotides and a sequence containing at least 50% sequence identity with tracrRNA in 6 adjacent nucleotides. In some embodiments, the non-natural sequence is located at a nucleic acid position selected from the 5' end and the 3' end or any combination thereof. In some embodiments, the nucleic acid comprises two nucleic acid molecules. In some embodiments, the nucleic acid comprises a single, continuous nucleic acid molecule. In some embodiments, the non-natural sequence comprises a CRISPR RNA-binding protein binding sequence. In some embodiments, the non-natural sequence comprises a binding sequence selected from Cas5 RNA-binding sequences, Cas6 RNA-binding sequences, and Csy4 RNA-binding sequences, or any combination thereof. In some embodiments, the effector protein comprises a CRISPR RNA-binding protein. In some embodiments, the effector protein contains at least 15% amino acid sequence identity with a protein selected from Cas5, Cas6, and Csy4, or any combination thereof. In some embodiments, the RNA-binding domain of the effector protein contains at least 15% amino acid sequence identity with the RNA-binding domain of a protein selected from Cas5, Cas6, and Csy4, or any combination thereof. In some embodiments, the effector protein is selected from Cas5, Cas6, and Csy4, or any combination thereof. In some embodiments, the effector protein further comprises one or more non-natural sequences. In some embodiments, the non-natural sequence provides enzymatic activity to the effector protein.In some embodiments, the enzyme activity is selected from: methyltransferase activity, demethylase activity, acetylation activity, deacetylation activity, ubiquitination activity, deubiquitination activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity or glycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristylation activity, remodeling activity, protease activity, oxidoreductase activity, transferase activity, hydrolase activity, lyase activity, isomerase activity, synthase activity, synthase activity, and demyristylation activity, or any combination thereof. In some embodiments, the nucleic acid is RNA. In some embodiments, the effector protein comprises a fusion protein containing an RNA-binding protein and a DNA-binding protein. In some embodiments, the composition further comprises a donor polynucleotide. In some embodiments, the donor polynucleotide binds directly to the DNA-binding protein, and wherein the RNA-binding protein binds to the nucleic acid targeting the nucleic acid. In some embodiments, the 5' end of the donor polynucleotide binds to the DNA-binding protein. In some embodiments, the 3' end of the donor polynucleotide binds to the DNA-binding protein. In some embodiments, at least five nucleotides of the donor polynucleotide bind to the DNA-binding protein. In some embodiments, the nucleic acid is an isolated nucleic acid. In some embodiments, the nucleic acid is a recombinant nucleic acid.

[0023] On one hand, this disclosure provides a method for introducing a donor polynucleotide into a target nucleic acid, comprising: contacting the target nucleic acid with a complex comprising a site-directed polypeptide and a composition comprising: an effector protein and a nucleic acid, wherein the nucleic acid comprises at least 50% sequence identity with crRNA in 6 adjacent nucleotides, at least 50% sequence identity with tracrRNA in 6 adjacent nucleotides, and a non-natural sequence, wherein the nucleic acid is adapted to bind to the effector protein. In some embodiments, the method further comprises cleaving the target nucleic acid. In some embodiments, the cleavage is performed by the site-directed polypeptide. In some embodiments, the method further comprises inserting the donor polynucleotide into the target nucleic acid.

[0024] On one hand, this disclosure provides a method for regulating a target nucleic acid, comprising: contacting the target nucleic acid with one or more complexes, each complex comprising a site-directed polypeptide and a composition comprising: an effector protein and a nucleic acid, wherein the nucleic acid comprises at least 50% sequence identity with crRNA in six adjacent nucleotides, at least 50% sequence identity with tracrRNA in six adjacent nucleotides, and a non-natural sequence, wherein the nucleic acid is adapted to bind to the effector protein, and regulating the target nucleic acid. In some embodiments, the site-directed polypeptide comprises at least 10% amino acid sequence identity with the nuclease domain of Cas9. In some embodiments, the regulation is performed by the effector protein. In some embodiments, the regulation comprises activities selected from: methyltransferase activity, demethylase activity, acetylation activity, deacetylation activity, ubiquitination activity, deubiquitination activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity or glycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristylation activity, remodeling activity, protease activity, oxidoreductase activity, transferase activity, hydrolase activity, lyase activity, isomerase activity, synthase activity, synthase activity, and demyristylation activity, or any combination thereof. In some embodiments, the effector protein comprises one or more effector proteins.

[0025] On one hand, this disclosure provides a method for detecting whether two complexes are adjacent to each other, comprising: contacting a first target nucleic acid with a first complex, wherein the first complex comprises a first site-directed polypeptide, a first modified nucleic acid targeting a nucleic acid, and a first effector protein, wherein the effector protein is adapted to bind to the modified nucleic acid targeting a nucleic acid, and wherein the first effector protein comprises a non-natural sequence containing a first portion of the separation system; and contacting a second target nucleic acid with a second complex, wherein the second complex comprises a second site-directed polypeptide, a second modified nucleic acid targeting a nucleic acid, and a second effector protein, wherein the effector protein is adapted to bind to the modified nucleic acid targeting a nucleic acid, and wherein the second effector protein comprises a non-natural sequence containing a second portion of the separation system. In some embodiments, the first target nucleic acid and the second target nucleic acid are on the same polynucleotide polymer. In some embodiments, the separation system comprises two or more protein fragments that are inactive individually but generate an active protein complex when forming a complex. In some embodiments, the method further comprises detecting an interaction between the first portion and the second portion. In some embodiments, the detection indicates that the first and second complexes are adjacent to each other. In some embodiments, the site-directed polypeptide is adapted to be unable to cleave the target nucleic acid. In some embodiments, the detection includes determining the occurrence of a genetic movement event. In some embodiments, the genetic movement event includes a translocation. In some embodiments, prior to the genetic movement event, the two parts of the separation system do not interact. In some embodiments, after the genetic movement event, the two parts of the separation system interact. In some embodiments, the genetic movement event is a translocation between the BCR and Abl genes. In some embodiments, the interaction activates the separation system. In some embodiments, the interaction indicates that the target nucleic acids bound by the complex are close together. In some embodiments, the separation system is selected from: a GFP separation system, a ubiquitin separation system, a transcription factor separation system, and an affinity marker separation system, or any combination thereof. In some embodiments, the separation system includes a GFP separation system. In some embodiments, the detection indicates a genotype. In some embodiments, the method further includes determining a course of treatment for the disease based on the genotype. In some embodiments, the method further includes treating the disease. In some embodiments, the treatment includes administration. In some embodiments, the treatment includes administering a complex comprising a nucleic acid targeting a nucleic acid and a site-specific polypeptide, wherein the complex can modify genetic factors involved in the disease. In some embodiments, the modification is selected from: adding a nucleic acid sequence to the genetic factor, replacing a nucleic acid sequence in the genetic factor, and deleting a nucleic acid sequence from the genetic factor, or any combination thereof. In some embodiments, the method further includes: communicating the genotype from a caregiver to a patient.In some embodiments, the communication includes communication from a storage memory system to a remote computer. In some embodiments, the detection diagnoses a disease. In some embodiments, the method further includes: communicating the diagnosis from a caregiver to a patient. In some embodiments, the detection indicates the presence of a single nucleotide polymorphism (SNP). In some embodiments, the method further includes: communicating the occurrence of a genetic movement event from a caregiver to a patient. In some embodiments, the communication includes communication from a storage memory system to a remote computer. In some embodiments, the site-directed peptide contains at least 20% amino acid sequence identity with Cas9. In some embodiments, the site-directed peptide contains at least 60% amino acid sequence identity with Cas9. In some embodiments, the site-directed peptide contains at least 60% amino acid sequence identity in the nuclease domain of Cas9. In some embodiments, the site-directed peptide is Cas9. In some embodiments, the modified nucleic acid-targeted nucleic acid contains a non-natural sequence. In some embodiments, the non-natural sequence is located at a position selected from the following on the modified nucleic acid-targeted nucleic acid: the 5' end and the 3' end, or any combination thereof. In some embodiments, the modified nucleic acid targeting nucleic acid comprises two nucleic acid molecules. In some embodiments, the nucleic acid comprises a single continuous nucleic acid molecule comprising: a first portion containing at least 50% identity with a CRISPR repeat fragment in six adjacent nucleotides and a second portion containing at least 50% identity with a tracrRNA sequence in six adjacent nucleotides. In some embodiments, the first and second portions are linked by a linker. In some embodiments, the non-natural sequence comprises a CRISPR RNA-binding protein binding sequence. In some embodiments, the non-natural sequence comprises a binding sequence selected from: Cas5 RNA-binding sequences, Cas6 RNA-binding sequences, and Csy4 RNA-binding sequences, or any combination thereof. In some embodiments, the modified nucleic acid targeting nucleic acid is adapted to bind to an effector protein. In some embodiments, the effector protein is a CRISPR RNA-binding protein. In some embodiments, the effector protein contains at least 15% amino acid sequence identity with a protein selected from: Cas5, Cas6, and Csy4, or any combination thereof. In some embodiments, the RNA-binding domain of the effector protein contains at least 15% amino acid sequence identity with the RNA-binding domain of a protein selected from Cas5, Cas6, and Csy4, or any combination thereof. In some embodiments, the effector protein is selected from Cas5, Cas6, and Csy4, or any combination thereof. In some embodiments, the nucleic acid targeting the nucleic acid is RNA. In some embodiments, the target nucleic acid is DNA. In some embodiments, the interaction involves the formation of an affinity marker.In some embodiments, the detection comprises capturing the affinity marker. In some embodiments, the method further comprises sequencing the nucleic acid bound to the first and second complexes. In some embodiments, the method further comprises cleaving the nucleic acid prior to the capture. In some embodiments, the interaction forms an activation system. In some embodiments, the method further comprises altering the transcription of the first or second target nucleic acid, wherein the alteration is performed via the activation system. In some embodiments, the second target nucleic acid is not linked to the first target nucleic acid. In some embodiments, trans-transcription of the altered second target nucleic acid is performed. In some embodiments, cis-transcription of the altered first target nucleic acid is performed. In some embodiments, the first or second target nucleic acid is selected from: endogenous nucleic acids and exogenous nucleic acids, or any combination thereof. In some embodiments, the alteration comprises enhancing the transcription of the first or second target nucleic acid; in some embodiments, the first or second target nucleic acid comprises a polynucleotide encoding one or more genes causing cell death. In some embodiments, the first or second target nucleic acid comprises a polynucleotide encoding a cell lysis-inducing peptide. In some embodiments, the first or second target nucleic acid comprises a polynucleotide encoding an immune cell recruitment antigen. In some embodiments, the first or second target nucleic acid comprises a polynucleotide encoding one or more genes involved in apoptosis. In some embodiments, the one or more genes involved in apoptosis comprise caspase. In some embodiments, the one or more genes involved in apoptosis comprise cytokines. In some embodiments, the one or more genes involved in apoptosis are selected from: tumor necrosis factor (TNF), TNF receptor 1 (TNFR1), TNF receptor 2 (TNFR2), Fas receptor, FasL, caspase-8, caspase-10, caspase-3, caspase-9, caspase-3, caspase-6, caspase-7, Bcl-2, and apoptosis-inducing factor (AIF), or any combination thereof. In some embodiments, the first or second target nucleic acid comprises a polynucleotide encoding one or more nucleic acids that target nucleic acids. In some embodiments, the one or more nucleic acid-targeting nucleic acids target multiple target nucleic acids. In some embodiments, the detection comprises generating genetic data. In some embodiments, the method further comprises transmitting the genetic data from a storage memory system to a remote computer. In some embodiments, the genetic data indicates a genotype. In some embodiments, the genetic data indicates the occurrence of genetic movement events. In some embodiments, the genetic data indicates the spatial location of genes.

[0026] On one hand, this disclosure provides a kit comprising: a site-directed peptide, a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence; an effector protein adapted to bind to the non-natural sequence; and a buffer. In some embodiments, the kit further comprises instructions for use.

[0027] On one hand, this disclosure provides a vector comprising a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence. In some embodiments, the polynucleotide sequence may be operatively linked to a promoter. In some embodiments, the promoter is an inducible promoter.

[0028] On one hand, this disclosure provides a vector comprising: a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a sequence constructed to bind to an effector protein and a site-directed polypeptide. In some embodiments, the polynucleotide sequence may be operatively linked to a promoter. In some embodiments, the promoter is an inducible promoter.

[0029] On one hand, this disclosure provides a vector comprising: a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence, a site-directed polypeptide, and an effector protein. In some embodiments, the polynucleotide sequence may be operatively linked to a promoter. In some embodiments, the promoter is an inducible promoter.

[0030] On one hand, this disclosure provides a genetically modified cell comprising a composition comprising: an effector protein, and a nucleic acid, wherein the nucleic acid comprises at least 50% sequence identity with crRNA in 6 adjacent nucleotides, at least 50% sequence identity with tracrRNA in 6 adjacent nucleotides, and a non-natural sequence, wherein the nucleic acid is adapted to bind to the effector protein.

[0031] On one hand, this disclosure provides a gene-modified cell comprising a vector containing a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence.

[0032] On one hand, this disclosure provides a gene-modified cell comprising a vector comprising: a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a sequence configured to bind to an effector protein and a site-directed polypeptide.

[0033] On one hand, this disclosure provides a gene-modified cell comprising a vector, the vector comprising: a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence, a site-directed polypeptide, and an effector protein.

[0034] On one hand, this disclosure provides a kit comprising: a vector containing a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence; and a buffer. In some embodiments, the kit further includes instructions for use.

[0035] On one hand, this disclosure provides a kit comprising: a vector containing a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a sequence constructed to bind to an effector protein and a site-directed polypeptide, and a buffer. In some embodiments, the kit further comprises instructions for use.

[0036] On one hand, this disclosure provides a kit comprising: a vector containing a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid includes a non-natural sequence, a site-directed peptide, and an effector protein; and a buffer. In some embodiments, the kit further includes instructions for use.

[0037] On one hand, this disclosure provides a composition comprising: a multiplex genetic target, wherein the multiplex genetic target comprises one or more nucleic acid modules, wherein the nucleic acid modules comprise a non-natural sequence, and wherein the nucleic acid modules are configured to bind to a polypeptide comprising at least 10% amino acid sequence identity with a nuclease domain of Cas9, and wherein the nucleic acid modules are configured to hybridize with a target nucleic acid. In some embodiments, the nucleic acid module comprises a first sequence comprising at least 50% sequence identity with crRNA in six adjacent nucleotides and a second sequence comprising at least 50% sequence identity with tracrRNA in six adjacent nucleotides. In some embodiments, the composition further comprises a linker sequence connecting the first and second sequences. In some embodiments, the one or more nucleic acid modules hybridize with one or more target nucleic acids. In some embodiments, the one or more nucleic acid modules are distinguished by at least one nucleotide in a spacer region of the one or more nucleic acid modules. In some embodiments, the one or more nucleic acid modules are RNA. In some embodiments, the multiplex genetic target is RNA. In some embodiments, the non-natural sequence comprises a ribozyme. In some embodiments, the non-natural sequence comprises a ribonuclease-binding sequence. In some embodiments, the ribonuclease-binding sequence is located at the 5' end of the nucleic acid module. In some embodiments, the ribonuclease-binding sequence is located at the 3' end of the nucleic acid module. In some embodiments, the ribonuclease-binding sequence is suitable for binding by a CRISPR ribonuclease. In some embodiments, the ribonuclease-binding sequence is suitable for binding by a ribonuclease containing a RAMP domain. In some embodiments, the ribonuclease-binding sequence is suitable for binding by a ribonuclease selected from: Cas5 superfamily member ribonucleases and Cas6 superfamily member ribonucleases, or any combination thereof. In some embodiments, the ribonuclease-binding sequence is suitable for binding by a ribonuclease containing at least 15% amino acid sequence identity with a protein selected from: Csy4, Cas5, and Cas6. In some embodiments, the ribonuclease-binding sequence is adapted to be bound by a ribonuclease containing at least 15% amino acid sequence identity with a nuclease domain selected from Csy4, Cas5, and Cas6. In some embodiments, the ribonuclease-binding sequence comprises a hairpin structure. In some embodiments, the hairpin structure comprises at least four consecutive nucleotides in a stem-loop configuration. In some embodiments, the ribonuclease-binding sequence contains at least 60% identity with sequences selected from:

[0038]

[0039]

[0040] and Or any combination thereof. In some embodiments, the one or more nucleic acid modules are adapted to be bound by different ribonucleases. In some embodiments, the multiplex genetic target is an isolated multiplex genetic target. In some embodiments, the multiplex genetic target is a recombinant multiplex genetic target.

[0041] On one hand, this disclosure provides a vector comprising a polynucleotide sequence encoding a multiple genetic target, wherein the multiple genetic target comprises one or more nucleic acid modules, wherein the nucleic acid modules comprise non-natural sequences, and wherein the nucleic acid modules are configured to bind to a polypeptide comprising at least 10% amino acid sequence identity with a nuclease domain of Cas9, and wherein the nucleic acid modules are configured to hybridize with a target nucleic acid. In some embodiments, the polynucleotide sequence may be operatively linked to a promoter. In some embodiments, the promoter is an inducible promoter.

[0042] On one hand, this disclosure provides a genetically modified cell comprising multiple genetic targets, wherein the multiple genetic targets comprise one or more nucleic acid modules, wherein the nucleic acid modules comprise non-natural sequences, and wherein the nucleic acid modules are configured to bind to a polypeptide comprising at least 10% amino acid sequence identity with a nuclease domain of Cas9, and wherein the nucleic acid modules are configured to hybridize with a target nucleic acid.

[0043] On one hand, this disclosure provides a genetically modified cell comprising a vector containing a multinucleotide sequence encoding a multiple genetic target, wherein the multiple genetic target comprises one or more nucleic acid modules, wherein the nucleic acid modules comprise non-natural sequences, and wherein the nucleic acid modules are configured to bind to a polypeptide comprising at least 10% amino acid sequence identity with a nuclease domain of Cas9, and wherein the nucleic acid modules are configured to hybridize with a target nucleic acid.

[0044] On one hand, this disclosure provides a kit comprising a multiplex genetic target, wherein the multiplex genetic target comprises one or more nucleic acid modules, wherein the nucleic acid modules comprise non-natural sequences, and wherein the nucleic acid modules are configured to bind to a polypeptide comprising at least 10% amino acid sequence identity with a nuclease domain of Cas9, and wherein the nucleic acid modules are configured to hybridize with a target nucleic acid, and a buffer. In some embodiments, the kit further comprises instructions for use.

[0045] On one hand, this disclosure provides a kit comprising: a vector containing a multinucleotide sequence encoding a multiple genetic target, wherein the multiple genetic target comprises one or more nucleic acid modules, wherein the nucleic acid modules comprise non-natural sequences, and wherein the nucleic acid modules are configured to bind to a polypeptide comprising at least 10% amino acid sequence identity with a nuclease domain of Cas9, and wherein the nucleic acid modules are configured to hybridize with a target nucleic acid, and a buffer. In some embodiments, the kit further comprises instructions for use.

[0046] On one hand, this disclosure provides a method for generating nucleic acids, wherein the nucleic acids are bound to a polypeptide containing at least 10% amino acid sequence identity with a nuclease domain of Cas9 and hybridized with a target nucleic acid. The method includes: introducing a multiplex genetic target agent, wherein the multiplex genetic target agent comprises one or more nucleic acid modules, wherein the nucleic acid modules contain non-natural sequences, and wherein the nucleic acid modules are configured to bind to a polypeptide containing at least 10% amino acid sequence identity with a nuclease domain of Cas9, and wherein the nucleic acid modules are configured to hybridize with a target nucleic acid into a host cell; processing the multiplex genetic target agent into the one or more nucleic acid modules; and contacting the processed one or more nucleic acid modules with one or more target nucleic acids in the cell. In some embodiments, the method further includes cleaving the target nucleic acid. In some embodiments, the method further includes modifying the target nucleic acid. In some embodiments, the modification comprises altering the transcription of the target nucleic acid. In some embodiments, the modification comprises inserting a donor polynucleotide into the target nucleic acid.

[0047] On one hand, this disclosure provides a modified site-directed polypeptide comprising: a first nuclease domain, a second nuclease domain, and an inserted nuclease domain. In some embodiments, the site-directed polypeptide contains at least 15% identity with the nuclease domain of Cas9. In some embodiments, the first nuclease domain comprises a nuclease domain selected from: an HNH domain and a RuvC domain, or any combination thereof. In some embodiments, the second nuclease domain comprises a nuclease domain selected from: an HNH domain and a RuvC domain, or any combination thereof. In some embodiments, the inserted nuclease domain comprises an HNH domain. In some embodiments, the inserted nuclease domain comprises a RuvC domain. In some embodiments, the inserted nuclease domain is the N-terminus of the first nuclease domain. In some embodiments, the inserted nuclease domain is the N-terminus of the second nuclease domain. In some embodiments, the inserted nuclease domain is the C-terminus of the first nuclease domain. In some embodiments, the inserted nuclease domain is the C-terminus of the second nuclease domain. In some embodiments, the inserted nuclease domain is tandem with a first nuclease domain. In some embodiments, the inserted nuclease domain is tandem with a second nuclease domain. In some embodiments, the inserted nuclease domain is adapted to cleave the target nucleic acid at a site different from the first or second nuclease domain. In some embodiments, the inserted nuclease domain is adapted to cleave the RNA in a DNA-RNA hybrid. In some embodiments, the inserted nuclease domain is adapted to cleave the DNA in a DNA-RNA hybrid. In some embodiments, the inserted nuclease domain is adapted to improve the binding specificity of the modified site-directed peptide to the target nucleic acid. In some embodiments, the inserted nuclease domain is adapted to improve the binding strength of the modified site-directed peptide to the target nucleic acid.

[0048] On the one hand, this disclosure provides a vector containing a polynucleotide sequence encoding a site-directed polypeptide, wherein the modified site-directed polypeptide comprises: a first nuclease domain, a second nuclease domain, and an inserted nuclease domain.

[0049] On one hand, this disclosure provides a kit comprising: a site-directed polypeptide modified with a first nuclease domain, a second nuclease domain, and an inserted nuclease domain, and a buffer. In some embodiments, the kit further comprises instructions for use.

[0050] In one aspect, this disclosure provides a composition comprising: a modified site-directed polypeptide, wherein the polypeptide is modified to be suitable for targeting a motif adjacent to a second pre-interstitial region sequence compared to a wild-type site-directed polypeptide. In some embodiments, the site-directed polypeptide is modified by a modification method selected from: amino acid addition, amino acid substitution, amino acid replacement, and amino acid deletion, or any combination thereof. In some embodiments, the modified site-directed polypeptide comprises a non-natural sequence. In some embodiments, the modified site-directed polypeptide is adapted to target a motif adjacent to a second pre-interstitial region sequence with higher specificity than a wild-type site-directed polypeptide. In some embodiments, the modified site-directed polypeptide is adapted to target a motif adjacent to a second pre-interstitial region sequence with a lower dissociation constant than a wild-type site-directed polypeptide. In some embodiments, the modified site-directed polypeptide is adapted to target a motif adjacent to a second pre-interstitial region sequence with a higher dissociation constant than a wild-type site-directed polypeptide. In some embodiments, the second anterior interstitial sequence neighbor motif comprises an anterior interstitial sequence neighbor motif selected from: 5'-NGGNG-3', 5'-NNAAAAW-3', 5'-NNNNGATT-3', 5'-GNNNCNNA-3', and 5'-NNNACA-3', or any combination thereof.

[0051] On the one hand, this disclosure provides a vector comprising a multinucleotide sequence encoding a modified site-directed polypeptide, wherein the polypeptide is modified to make it suitable for targeting motifs adjacent to a second pre-interstitial region sequence compared to a wild-type site-directed polypeptide.

[0052] On one hand, this disclosure provides a kit comprising: a modified site-targeting peptide, wherein the peptide is modified to be suitable for targeting motifs adjacent to a second pre-interstitial region sequence compared to a wild-type site-targeting peptide, and a buffer. In some embodiments, the kit further comprises instructions for use.

[0053] On one hand, this disclosure provides a composition comprising: a modified site-directed polypeptide, wherein the polypeptide is modified to be suitable for targeting a second nucleic acid-targeting nucleic acid compared to a wild-type site-directed polypeptide. In some embodiments, the site-directed polypeptide is modified by a modification method selected from: amino acid addition, amino acid substitution, amino acid replacement, and amino acid deletion, or any combination thereof. In some embodiments, the modified site-directed polypeptide comprises a non-natural sequence. In some embodiments, the modified site-directed polypeptide is adapted to target the second nucleic acid-targeting nucleic acid with higher specificity than the wild-type site-directed polypeptide. In some embodiments, the modified site-directed polypeptide is adapted to target the second nucleic acid-targeting nucleic acid with a lower dissociation constant than the wild-type site-directed polypeptide. In some embodiments, the modified site-directed polypeptide is adapted to target the second nucleic acid-targeting nucleic acid with a higher dissociation constant than the wild-type site-directed polypeptide. In some embodiments, the site-directed polypeptide targets the tracrRNA portion of the second nucleic acid-targeting nucleic acid.

[0054] On the one hand, this disclosure provides a vector comprising a polynucleotide sequence encoding a modified site-directed polypeptide, wherein the polypeptide is modified to make it suitable for targeting a second nucleic acid target compared to a wild-type site-directed polypeptide.

[0055] On one hand, this disclosure provides a kit comprising: a modified site-targeting peptide, wherein the peptide is modified to be suitable for targeting a second nucleic acid target compared to a wild-type site-targeting peptide, and a buffer. In some embodiments, the kit further comprises instructions for use.

[0056] On one hand, this disclosure provides a composition comprising: a modified site-directed polypeptide comprising, compared to SEQ ID:8, a modified portion in the bridging helix. In some embodiments, the composition is configured to cleave target nucleic acids.

[0057] On one hand, this disclosure provides a composition comprising: a modified site-directed polypeptide comprising, compared to SEQ ID:8, a modified high-basic patch. In some embodiments, the composition is configured to cleave target nucleic acids.

[0058] On one hand, this disclosure provides a composition comprising a modified site-directed polypeptide having a modified polymerase-like domain compared to SEQ ID:8. In some embodiments, the composition is configured to cleave target nucleic acids.

[0059] On one hand, this disclosure provides a composition comprising: a modified site-directed polypeptide, wherein the modified site-directed polypeptide, compared to SEQ ID:8, comprises modifications in a bridging helix, a high-basic patch, a nuclease domain, and a polymerase domain, or any combination thereof.

[0060] On the one hand, this disclosure provides a vector comprising a polynucleotide sequence encoding a site-directed polypeptide, wherein the site-directed polypeptide comprises, compared to SEQ ID:8, modifications in a bridging helix, a high-basic patch, a nuclease domain, and a polymerase domain, or any combination thereof.

[0061] On one hand, this disclosure provides a kit comprising: a modified site-directed polypeptide, said modified site-directed polypeptide comprising, compared to SEQ ID:8, modifications in a bridging helix, a high-basic patch, a nuclease domain, and a polymerase domain, or any combination thereof, and a buffer. In some embodiments, the kit further comprises instructions for use. In some embodiments, the kit further comprises a nucleic acid targeting a nucleic acid.

[0062] On one hand, this disclosure provides a gene-modified cell comprising a modified site-directed polypeptide, wherein the modified site-directed polypeptide comprises, compared to SEQ ID:8, modifications in a bridging helix, a high-basic patch, a nuclease domain, and a polymerase domain, or any combination thereof.

[0063] On one hand, this disclosure provides a genome engineering method comprising: contacting a target nucleic acid with a complex, wherein the complex comprises a modified site-directed polypeptide, the modified site-directed polypeptide comprising modifications in a bridging helix, a high-basic patch, a nuclease domain, and a polymerase domain, or any combination thereof, compared to SEQ ID:8, and a nucleic acid targeting the nucleic acid, and modifying the target nucleic acid. In some embodiments, the contact comprises contacting the complex with a motif adjacent to a pre-intercalation region sequence in the target nucleic acid. In some embodiments, the contact comprises contacting the complex with a target nucleic acid sequence longer than the unmodified site-directed polypeptide. In some embodiments, the modification comprises cleaving the target nucleic acid. In some embodiments, the target nucleic acid comprises RNA. In some embodiments, the target nucleic acid comprises DNA. In some embodiments, the modification comprises an RNA strand cleaving hybrid RNA and DNA. In some embodiments, the modification comprises a DNA strand cleaving hybrid RNA and DNA. In some embodiments, the modification comprises inserting a donor polynucleotide, a portion of a donor polynucleotide, a copy of a donor polynucleotide, or a portion of a copy of a donor polynucleotide, or any combination thereof, into the target nucleic acid. In some embodiments, the modification comprises modifying the transcriptional activity of the target nucleic acid. In some embodiments, the modification comprises deleting one or more nucleotides of the target nucleic acid.

[0064] On one hand, this disclosure provides a composition comprising: a modified site-directed polypeptide comprising a modified nuclease domain compared to SEQ ID:8. In some embodiments, the composition is configured to cleave target nucleic acids. In some embodiments, the modified nuclease domain comprises a RuvC domain nuclease domain. In some embodiments, the modified nuclease domain comprises an HNH nuclease domain. In some embodiments, the modified nuclease domain comprises a copy of the HNH nuclease domain. In some embodiments, the modified nuclease domain is adapted to enhance the specificity of the amino acid sequence for target nucleic acids compared to an unmodified site-directed polypeptide. In some embodiments, the modified nuclease domain is adapted to enhance the specificity of the amino acid sequence for nucleic acids targeting nucleic acids compared to an unmodified site-directed polypeptide. In some embodiments, the modified nuclease domain comprises a modification selected from: amino acid addition, amino acid substitution, amino acid replacement, and amino acid deletion, or any combination thereof. In some embodiments, the modified nuclease domain comprises an inserted non-natural sequence. In some embodiments, the non-natural sequence provides enzymatic activity to the modified site-directed polypeptide. In some embodiments, the enzyme activity is selected from: nuclease activity, methyltransferase activity, acetyltransferase activity, demethylase activity, deamination activity, superoxide dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristylation activity, remodeling activity, protease activity, oxidoreductase activity, transferase activity, hydrolase activity, lyase activity, isomerase activity, synthase activity, synthase activity, and demyristylation activity, or any combination thereof. In some embodiments, the enzyme activity is adapted to regulate the transcription of the target nucleic acid. In some embodiments, the modified nuclease domain is adapted to allow the amino acid sequence to bind to a motif adjacent to the interstitial region sequence, which is different from the interstitial region sequence adjacent to the unmodified site-directed polypeptide. In some embodiments, the modified nuclease domain is adapted to allow the amino acid sequence to bind to a nucleic acid targeting a nucleic acid, which is different from the nucleic acid targeting a nucleic acid that the unmodified site-directed polypeptide is adapted to bind to. In some embodiments, the modified site-directed polypeptide is adapted to bind to a target nucleic acid sequence longer than that of the unmodified site-directed polypeptide. In some embodiments, the modified site-directed polypeptide is adapted to cleave double-stranded DNA.In some embodiments, the modified site-directed peptide is adapted to cleave the RNA strand of hybrid RNA and DNA. In some embodiments, the modified site-directed peptide is adapted to cleave the DNA strand of hybrid RNA and DNA. In some embodiments, the composition further comprises a modified nucleic acid-targeted nucleic acid, wherein the modification of the site-directed peptide is adapted to allow the site-directed peptide to bind to the modified nucleic acid-targeted nucleic acid. In some embodiments, the modified nucleic acid-targeted nucleic acid and the modified site-directed peptide comprise a compensating mutation.

[0065] On one hand, this disclosure provides a method for enriching target nucleic acids for sequencing, comprising: contacting the target nucleic acid with a complex comprising a nucleic acid targeting a nucleic acid and a site-directed polypeptide; enriching the target nucleic acid using the complex; and determining the sequence of the target nucleic acid. In some embodiments, the method does not include an amplification step. In some embodiments, the method further includes analyzing the sequence of the target nucleic acid. In some embodiments, the method further includes cleaving the target nucleic acid before enrichment. In some embodiments, the nucleic acid targeting a nucleic acid comprises RNA. In some embodiments, the nucleic acid targeting a nucleic acid in the method comprises two RNA molecules. In some embodiments, a portion of each of the two RNA molecules hybridizes together. In some embodiments, one of the two RNA molecules in the method comprises a CRISPR repeat sequence. In some embodiments, the CRISPR repeat sequence is homologous to crRNA in 6 adjacent nucleotides. In some embodiments, the CRISPR repeat sequence contains at least 60% identity with crRNA in 6 adjacent nucleotides. In some embodiments, one of the two RNA molecules comprises a tracrRNA sequence. In some embodiments, the tracRNA sequence is homologous to tracrRNA in 6 adjacent nucleotides. In some embodiments, the tracRNA sequence contains at least 60% identity with tracrRNA in 6 adjacent nucleotides. In some embodiments, the nucleic acid-targeting nucleic acid is a dual-guided nucleic acid. In some embodiments, the nucleic acid-targeting nucleic acid comprises a continuous RNA molecule, wherein the continuous RNA molecule further comprises two domains and a linker. In some embodiments, a portion of each of the two domains of the continuous RNA molecule hybridizes together. In some embodiments, the continuous RNA molecule comprises a CRISPR repeat sequence. In some embodiments, the CRISPR repeat sequence is homologous to crRNA in 6 adjacent nucleotides. In some embodiments, the CRISPR repeat sequence contains at least 60% identity with crRNA in 6 adjacent nucleotides. In some embodiments, the continuous RNA molecule comprises a tracrRNA sequence. In some embodiments, the tracRNA sequence is homologous to tracrRNA in 6 adjacent nucleotides. In some embodiments, the tracRNA sequence contains at least 60% identity with tracrRNA in 6 adjacent nucleotides. In some embodiments, the nucleic acid targeted is a single-target nucleic acid. In some embodiments, the contact involves hybridizing a portion of the nucleic acid targeted with a portion of the target nucleic acid. In some embodiments, the nucleic acid targeted hybridizes with the target nucleic acid in a region comprising 6-20 nucleotides. In some embodiments, the site-directed polypeptide comprises Cas9.In some embodiments, the site-directed peptide contains at least 20% homology with the nuclease domain of Cas9. In some embodiments, the site-directed peptide contains at least 60% homology with Cas9. In some embodiments, the site-directed peptide contains an engineered nuclease domain, wherein the nuclease domain contains reduced nuclease activity compared to a site-directed peptide containing an unengineered nuclease domain. In some embodiments, the site-directed peptide introduces single-strand breaks in the target nucleic acid. In some embodiments, the engineered nuclease domain contains a mutation in a conserved aspartic acid. In some embodiments, the engineered nuclease domain contains a D10A mutation. In some embodiments, the engineered nuclease domain contains a mutation in a conserved histidine. In some embodiments, the engineered nuclease domain contains an H840A mutation. In some embodiments, the site-directed peptide contains an affinity label. In some embodiments, the affinity label is located at the N-terminus of the site-directed peptide, the C-terminus of the site-directed peptide, a surface-accessible region, or any combination thereof. In some embodiments, the affinity marker is selected from biotin, FLAG, His6x, His9x, and fluorescent proteins, or any combination thereof. In some embodiments, the nucleic acid targeted by the nucleic acid comprises a nucleic acid affinity marker. In some embodiments, the nucleic acid affinity marker is located at the 5' end, the 3' end, the surface accessible region, or any combination thereof of the nucleic acid targeted by the nucleic acid. In some embodiments, the nucleic acid affinity marker is selected from small molecules, fluorescent markers, radiolabels, or any combination thereof. In some embodiments, the nucleic acid affinity marker comprises a sequence constructed to bind to Csy4, Cas5, Cas6, or any combination thereof. In some embodiments, the nucleic acid affinity marker comprises 50% identity to 5'-GUUCACUGCCGUAUAGGCAGCUAAGAAA-3'. In some embodiments, the method further comprises diagnosing a disease and making patient-specific treatment decisions, or any combination thereof. In some embodiments, the determination comprises determining a genotype. In some embodiments, the method further comprises transmitting the sequence from a memory storage system to a remote computer. In some embodiments, the enrichment comprises contacting an affinity label of the complex with a capture agent. In some embodiments, the capture agent comprises an antibody. In some embodiments, the capture agent comprises a solid support. In some embodiments, the capture agent is selected from Csy4, Cas5, and Cas6. In some embodiments, the capture agent comprises reduced enzyme activity in the absence of imidazole. In some embodiments, the capture agent comprises an activatable enzyme domain, wherein the activatable enzyme domain is activated upon contact with imidazole. In some embodiments, the capture agent is a member of the Cas6 family. In some embodiments, the capture agent comprises an affinity label.In some embodiments, the capture agent comprises a conditionally non-enzymatic ribonuclease containing a mutation in its nuclease domain. In some embodiments, a conserved histidine mutation. In some embodiments, the mutation comprises an H29A mutation. In some embodiments, the target nucleic acid binds to the complex. In some embodiments, the target nucleic acid is a resected nucleic acid that does not bind to the complex. In some embodiments, multiple complexes are contacted with multiple target nucleic acids. In some embodiments, the multiple target nucleic acids differ by at least one nucleotide. In some embodiments, the multiple complexes comprise multiple nucleic acids that target nucleic acids and differ by at least one nucleotide.

[0066] On one hand, this disclosure provides a method for excising nucleic acids, comprising: contacting a target nucleic acid with two or more complexes, each complex comprising a site-directed polypeptide and a nucleic acid targeting the nucleic acid; and cleaving the target nucleic acid, wherein the cleavage produces excised target nucleic acids. In some embodiments, the cleavage is performed via a nuclease domain of the site-directed polypeptide. In some embodiments, the method does not include amplification. In some embodiments, the method further comprises enriching the excised target nucleic acid. In some embodiments, the method further comprises sequencing the excised target nucleic acid. In some embodiments, the nucleic acid targeting the nucleic acid is RNA. In some embodiments, the nucleic acid targeting the nucleic acid comprises two RNA molecules. In some embodiments, a portion of each of the two RNA molecules hybridizes together. In some embodiments, one of the two RNA molecules comprises a CRISPR repeat sequence. In some embodiments, the CRISPR repeat sequence comprises a sequence homologous to crRNA in six adjacent nucleotides. In some embodiments, the CRISPR repeat sequence comprises a sequence having at least 60% identity with crRNA in six adjacent nucleotides. In some embodiments, one of the two RNA molecules comprises a tracrRNA sequence. In some embodiments, the tracRNA sequence is homologous to crRNA in 6 adjacent nucleotides. In some embodiments, the tracRNA sequence contains at least 60% identity with crRNA in 6 adjacent nucleotides. In some embodiments, the nucleic acid-targeted nucleic acid is a dual-guided nucleic acid. In some embodiments, the nucleic acid-targeted nucleic acid comprises a continuous RNA molecule, wherein the continuous RNA molecule further comprises two domains and a linker. In some embodiments, a portion of each of the two domains of the continuous RNA molecule hybridizes together. In some embodiments, the continuous RNA molecule comprises a CRISPR repeat sequence. In some embodiments, the CRISPR repeat sequence is homologous to crRNA in 6 adjacent nucleotides. In some embodiments, the CRISPR repeat sequence contains at least 60% identity with crRNA in 6 adjacent nucleotides. In some embodiments, the continuous RNA molecule comprises a tracrRNA sequence. In some embodiments, the tracRNA sequence is homologous to crRNA in 6 adjacent nucleotides. In some embodiments, the tracRNA sequence contains at least 60% identity with crRNA in 6 adjacent nucleotides. In some embodiments, the nucleic acid targeted is a single-target nucleic acid. In some embodiments, the nucleic acid targeted hybridizes with the target nucleic acid. In some embodiments, the nucleic acid targeted hybridizes with the target nucleic acid in a region containing at least 6 nucleotides and at most 20 nucleotides.In some embodiments, the site-directed polypeptide is Cas9. In some embodiments, the site-directed polypeptide comprises a polypeptide containing at least 20% homology to the nuclease domain of Cas9. In some embodiments, the site-directed polypeptide comprises a polypeptide containing at least 60% homology to Cas9. In some embodiments, the site-directed polypeptide comprises an affinity label. In some embodiments, the affinity label is located at the N-terminus of the site-directed polypeptide, the C-terminus of the site-directed polypeptide, a surface-accessible region, or any combination thereof. In some embodiments, the affinity label is selected from: biotin, FLAG, His6x, His9x, and fluorescent proteins, or any combination thereof. In some embodiments, the nucleic acid-targeted nucleic acid comprises a nucleic acid affinity label. In some embodiments, the nucleic acid affinity label is located at the 5' end of the nucleic acid-targeted nucleic acid, the 3' end of the nucleic acid-targeted nucleic acid, a surface-accessible region, or any combination thereof. In some embodiments, the nucleic acid affinity label is selected from small molecules, fluorescent labels, radiolabels, or any combination thereof. In some embodiments, the nucleic acid affinity marker is a sequence that can bind to Csy4, Cas5, Cas6, or any combination thereof. In some embodiments, the nucleic acid affinity marker contains 50% identity with GUUCACUGCCGUAUAGGCAGCUAAGAAA. In some embodiments, the target nucleic acid is a cleaved nucleic acid that has not bound to the two or more complexes. In some embodiments, the two or more complexes are contacted with multiple target nucleic acids. In some embodiments, the multiple target nucleic acids differ by at least one nucleotide. In some embodiments, the two or more complexes comprise nucleic acids that are nucleic acid-targeted and differ by at least one nucleotide.

[0067] On one hand, this disclosure provides a method for generating a target nucleic acid library, comprising: contacting a plurality of target nucleic acids with a complex comprising a site-directed polypeptide and a nucleic acid targeting a nucleic acid; lysing the plurality of target nucleic acids; and purifying the plurality of target nucleic acids to generate a target nucleic acid library. In some embodiments, the method further comprises screening the target nucleic acid library.

[0068] On one hand, this disclosure provides a composition comprising: a first complex comprising a first site-directed polypeptide and a first nucleic acid targeting a nucleic acid; and a second complex comprising a second site-directed polypeptide and a second nucleic acid targeting a nucleic acid, wherein the first and second nucleic acids targeting nucleic acids are different. In some embodiments, the composition further comprises a target nucleic acid bound by the first or second complex. In some embodiments, the first site-directed polypeptide and the second site-directed polypeptide are the same. In some embodiments, the first site-directed polypeptide and the second site-directed polypeptide are different.

[0069] On the one hand, this disclosure provides a vector comprising a multinucleotide sequence encoding two or more nucleic acids that are separated by at least one nucleotide and targeting nucleic acids, and a site-directed polypeptide.

[0070] On the one hand, this disclosure provides a genetically modified host cell comprising: a vector containing a polynucleotide sequence encoding two or more nucleic acids that are separated by at least one nucleotide and targeting nucleic acids, and a site-directed polypeptide.

[0071] On one hand, this disclosure provides a kit comprising: a vector containing a multinucleotide sequence encoding two or more nucleic acids differing by at least one nucleotide from each other, a site-directed peptide, and a suitable buffer. In some embodiments, the kit further comprises: a capture agent, a solid support, sequencing adaptors, and a positive control, or any combination thereof. In some embodiments, the kit further comprises instructions for use.

[0072] On one hand, this disclosure provides a kit comprising: a site-targeting peptide with reduced enzyme activity compared to a wild-type site-targeting peptide, a nucleic acid targeting a nucleic acid, and a capture agent. In some embodiments, the kit further comprises: instructions for use. In some embodiments, the kit further comprises a buffer selected from wash buffer, stabilization buffer, reconstitution buffer, or dilution buffer.

[0073] On one hand, this disclosure provides a method for cleaving a target nucleic acid using two or more nickases, comprising: contacting the target nucleic acid with a first complex and a second complex, wherein the first complex comprises a first nickase and a first nucleic acid targeting a nucleic acid, and wherein the second complex comprises a second nickase and a second nucleic acid targeting a nucleic acid, wherein the target nucleic acid comprises a first pre-intercalation sequence adjacent motif on a first strand and a second pre-intercalation sequence adjacent motif on a second strand, wherein the first nucleic acid targeting a nucleic acid is adapted to hybridize with the first pre-intercalation sequence adjacent motif, and wherein the second nucleic acid targeting a nucleic acid is adapted to hybridize with the second pre-intercalation sequence adjacent motif; and cleaving the first and second strands of the target nucleic acid, wherein the cleavage generates cleaved target nucleic acid. In some embodiments, the first and second nickases are the same. In some embodiments, the first and second nickases are different. In some embodiments, the first and second nucleic acid targeting nucleic acids are different. In some embodiments, there is less than 125 nucleotides between the first pre-intercalation sequence adjacent motif and the second pre-intercalation sequence adjacent motif. In some embodiments, the first and second pre-intercalation sequence adjacent motifs comprise the sequence NGG, where N is any nucleotide. In some embodiments, the first or second nickase comprises at least one substantially inactive nuclease domain. In some embodiments, the first or second nickase comprises a mutation in a conserved aspartate amino acid. In some embodiments, the mutation is a D10A mutation. In some embodiments, the first or second nickase comprises a mutation in a conserved histidine amino acid. In some embodiments, the mutation is an H840A mutation. In some embodiments, there are fewer than 15 nucleotides between adjacent motifs of the first and second interstitial regions. In some embodiments, there are fewer than 10 nucleotides between adjacent motifs of the first and second interstitial regions. In some embodiments, there are fewer than 5 nucleotides between adjacent motifs of the first and second interstitial regions. In some embodiments, the adjacent motifs of the first and second interstitial regions are adjacent to each other. In some embodiments, the cleavage comprises the first nickase cleaving the first strand and the second nickase cleaving the second strand. In some embodiments, the cleavage produces a sticky end nick. In some embodiments, the cleavage produces a blunt end nick. In some embodiments, the method further comprises inserting a donor polynucleotide into the cleaved target nucleic acid.

[0074] On one hand, this disclosure provides a composition comprising: a plurality of nucleic acid molecules, wherein each nucleic acid molecule contains a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid targeting a nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide, and a fusion polypeptide, wherein the fusion polypeptide contains a plurality of nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites. In some embodiments, one or more of the plurality of nucleic acid-binding proteins contain a non-natural sequence. In some embodiments, the non-natural sequence is located at a position selected from: an N-terminus, a C-terminus, a surface-accessible region, or any combination thereof. In some embodiments, the non-natural sequence encodes a nuclear localization signal. In some embodiments, the plurality of nucleic acid-binding proteins are separated by a linker. In some embodiments, some of the plurality of nucleic acid-binding proteins are the same nucleic acid-binding protein. In some embodiments, all of the plurality of nucleic acid-binding proteins are the same nucleic acid-binding protein. In some embodiments, the plurality of nucleic acid-binding proteins are different nucleic acid-binding proteins. In some embodiments, the plurality of nucleic acid-binding proteins comprise RNA-binding proteins. In some embodiments, the RNA-binding protein is selected from: type I clustered regularly spaced short palindromic repeat ribonucleases, type II clustered regularly spaced short palindromic repeat ribonucleases, or type III clustered regularly spaced short palindromic repeat ribonucleases, or any combination thereof. In some embodiments, the RNA-binding protein is selected from: Cas5, Cas6, and Csy4, or any combination thereof. In some embodiments, the plurality of nucleic acid-binding proteins comprises a DNA-binding protein. In some embodiments, the nucleic acid-binding protein binding site is configured to bind nucleic acid-binding proteins selected from: type I, type II, and type III clustered regularly spaced short palindromic repeat ribonucleases, or any combination thereof. In some embodiments, the nucleic acid-binding protein binding site is configured to bind nucleic acid-binding proteins selected from: Cas6, Cas5, and Csy4, or any combination thereof. In some embodiments, some of the plurality of nucleic acid molecules contain the same nucleic acid-binding protein binding site. In some embodiments, the plurality of nucleic acid molecules contain the same nucleic acid-binding protein binding site. In some embodiments, none of the plurality of nucleic acid molecules contains the same nucleic acid-binding protein binding site. In some embodiments, the site-directed polypeptide contains at least 20% sequence identity with the nuclease domain of Cas9. In some embodiments, the site-directed polypeptide is Cas9. In some embodiments, at least one of the nucleic acid molecules encodes a cluster of regularly spaced short palindromic repeat ribonucleases.In some embodiments, the clustered, regularly spaced short palindromic repeat ribonucleases contain at least 20% sequence similarity to Csy4. In some embodiments, the clustered, regularly spaced short palindromic repeat ribonucleases contain at least 60% sequence similarity to Csy4. In some embodiments, the clustered, regularly spaced short palindromic repeat ribonucleases are Csy4. In some embodiments, the plurality of nucleic acid-binding proteins contain reduced enzymatic activity. In some embodiments, the plurality of nucleic acid-binding proteins are adapted to bind to the nucleic acid-binding protein binding site but cannot cleave the nucleic acid-binding protein binding site. In some embodiments, the nucleic acid-targeting nucleic acid comprises two RNA molecules. In some embodiments, a portion of each of the two RNA molecules hybridizes together. In some embodiments, the first molecule of the two RNA molecules contains a sequence in 8 adjacent nucleotides containing at least 60% identity with a clustered, regularly spaced short palindromic repeat RNA sequence, and the second molecule of the two RNA molecules contains a sequence in 6 adjacent nucleotides containing at least 60% identity with a trans-activation-clustered, regularly spaced short palindromic repeat RNA sequence. In some embodiments, the nucleic acid-targeted nucleic acid comprises a continuous RNA molecule, wherein the continuous RNA molecule further comprises two domains and a linker. In some embodiments, portions of the two domains of the continuous RNA molecule hybridize together. In some embodiments, the first portion of the continuous RNA molecule contains a sequence in 8 adjacent nucleotides containing at least 60% identity with a clustered, regularly spaced short palindromic repeat RNA sequence, and the second portion of the continuous RNA molecule contains a sequence in 6 adjacent nucleotides containing at least 60% identity with a trans-activation-clustered, regularly spaced short palindromic repeat RNA sequence. In some embodiments, the nucleic acid-targeted nucleic acid is adapted to hybridize with the target nucleic acid in 6-20 nucleotides. In some embodiments, the composition is constructed for delivery to cells. In some embodiments, the composition is configured to deliver equal amounts of the plurality of nucleic acid molecules to cells. In some embodiments, the composition further comprises a donor polynucleotide molecule, wherein the donor polynucleotide molecule contains a nucleic acid-binding protein binding site, wherein the binding site is bound by a nucleic acid-binding protein of the fusion polypeptide.

[0075] On one hand, this disclosure provides a method for delivering nucleic acids to a subcellular location within a cell, comprising: introducing a composition into a cell, the composition comprising: a plurality of nucleic acid molecules, wherein each nucleic acid molecule contains a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid targeting a nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide; and a fusion polypeptide, wherein the fusion polypeptide contains a plurality of nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites to stoichiometrically deliver the composition to the subcellular location, forming a unit comprising a site-directed polypeptide translated from a nucleic acid molecule encoding the site-directed polypeptide and the nucleic acid targeting a nucleic acid; and cleaving the target nucleic acid, wherein the site-directed polypeptide of the unit cleaves the target nucleic acid. In some embodiments, the plurality of nucleic acid-binding proteins bind to their associated nucleic acid-binding protein binding sites. In some embodiments, a ribonuclease cleaves one or more of the nucleic acid-binding protein binding sites. In some embodiments, a ribonuclease cleaves the nucleic acid-binding protein binding sites of the nucleic acid encoding the nucleic acid targeting a nucleic acid, thereby releasing the nucleic acid targeting a nucleic acid. In some embodiments, the subcellular location is selected from: nucleases, ER, Golgi apparatus, mitochondria, cell wall, lysosomes, and nucleus. In some embodiments, the subcellular location is the nucleus.

[0076] On one hand, this disclosure provides a vector comprising: a polynucleotide sequence encoding a composition comprising: a plurality of nucleic acid molecules, wherein each nucleic acid molecule contains a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid targeting a nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide; and a fusion polypeptide comprising a plurality of nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites to stoichiometrically deliver the composition to the subcellular location. In some embodiments, the vector further comprises a polynucleotide encoding a promoter. In some embodiments, the promoter is operatively linked to the polynucleotide. In some embodiments, the promoter is an inducible promoter.

[0077] On one hand, this disclosure provides a genetically modified organism comprising a vector comprising: a polynucleotide sequence encoding a plurality of nucleic acid molecules, wherein each nucleic acid molecule includes a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid targeting a nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide; and a fusion polypeptide, wherein the fusion polypeptide includes a plurality of nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites to stoichiometrically deliver the composition to the subcellular location.

[0078] On one hand, this disclosure provides a genetically modified organism comprising a composition comprising: a plurality of nucleic acid molecules, wherein each nucleic acid molecule contains a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid targeting a nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide; and a fusion polypeptide, wherein the fusion polypeptide contains a plurality of nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites.

[0079] On one hand, this disclosure provides a kit comprising: a composition comprising: a plurality of nucleic acid molecules, wherein each nucleic acid molecule contains a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid targeting a nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide, and a fusion polypeptide, wherein the fusion polypeptide contains a plurality of nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites, and a buffer.

[0080] On one hand, this disclosure provides a kit comprising: a vector comprising: a multinucleotide sequence encoding a plurality of nucleic acid molecules, wherein each nucleic acid molecule includes a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid targeting a nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide; and a fusion polypeptide, wherein the fusion polypeptide includes a plurality of nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites to stoichiometrically deliver the composition to the subcellular site; and a buffer. In some embodiments, the kit further includes instructions for use. In some embodiments, the buffer is selected from: dilution buffer, reconstitution buffer, and stabilization buffer, or any combination thereof.

[0081] On one hand, this disclosure provides a donor polynucleotide comprising: a genetic factor of interest and a reporter element, wherein the reporter element comprises a polynucleotide sequence encoding a site-directed polypeptide and one or more nucleic acids, wherein the one or more nucleic acids comprise a sequence containing at least 50% sequence identity with crRNA in six adjacent nucleotides and a sequence containing at least 50% sequence identity with tracrRNA in six adjacent nucleotides. In some embodiments, the genetic factor of interest comprises a gene. In some embodiments, the genetic factor of interest comprises a non-coding nucleic acid selected from: microRNA, siRNA, and long non-coding RNA or any combination thereof. In some embodiments, the genetic factor of interest comprises a non-coding gene. In some embodiments, the genetic factor of interest comprises a non-coding nucleic acid selected from: microRNA, siRNA, and long non-coding RNA or any combination thereof. In some embodiments, the reporter element comprises a gene selected from: a gene encoding a fluorescent protein, a gene encoding a chemiluminescent protein, and an antibiotic resistance gene or any combination thereof. In some embodiments, the reporter element comprises a gene encoding a fluorescent protein. In some embodiments, the fluorescent protein comprises green fluorescent protein. In some embodiments, the reporter element is operatively linked to a promoter. In some embodiments, the promoter comprises an inducible promoter. In some embodiments, the promoter comprises a tissue-specific promoter. In some embodiments, the site-directed polypeptide contains at least 15% amino acid sequence identity with the nuclease domain of Cas9. In some embodiments, the site-directed polypeptide contains at least 95% amino acid sequence identity with Cas9 within 10 amino acids. In some embodiments, the nuclease domain is selected from: HNH domain, HNH-like domain, RuvC domain, and RuvC-like domain, or any combination thereof.

[0082] On one hand, this disclosure provides an expression vector comprising a multinucleotide sequence encoding a genetic factor of interest; and a reporter element, wherein the reporter element comprises a multinucleotide sequence encoding a site-directed polypeptide and one or more nucleic acids, wherein the one or more nucleic acids comprise a sequence in six adjacent nucleotides containing at least 50% sequence identity with crRNA and a sequence in six adjacent nucleotides containing at least 50% sequence identity with tracrRNA.

[0083] On one hand, this disclosure provides a genetically modified cell comprising a donor polynucleotide, the donor polynucleotide comprising: a genetic factor of interest; and a reporter element, wherein the reporter element comprises a polynucleotide sequence encoding a site-directed polypeptide and one or more nucleic acids, wherein the one or more nucleic acids comprise a sequence in six adjacent nucleotides containing at least 50% sequence identity with crRNA and a sequence in six adjacent nucleotides containing at least 50% sequence identity with tracrRNA.

[0084] On one hand, this disclosure provides a kit comprising: a donor polynucleotide, the donor polynucleotide comprising: a genetic factor of interest; and a reporter element, wherein the reporter element comprises a polynucleotide sequence encoding a site-directed polypeptide and one or more nucleic acids, wherein the one or more nucleic acids comprise a sequence comprising at least 50% sequence identity with crRNA in six adjacent nucleotides and a sequence comprising at least 50% sequence identity with tracrRNA in six adjacent nucleotides; and a buffer. In some embodiments, the kit further comprises: a polypeptide comprising at least 10% amino acid sequence identity with Cas9; and a nucleic acid, wherein the nucleic acid binds to the polypeptide and hybridizes with a target nucleic acid. In some embodiments, the kit further comprises instructions for use. In some embodiments, the kit further comprises a polynucleotide encoding a polypeptide, wherein the polypeptide comprises at least 15% amino acid sequence identity with Cas9. In some embodiments, the kit further comprises a polynucleotide encoding a nucleic acid, wherein the nucleic acid comprises a sequence comprising at least 50% sequence identity with crRNA in six adjacent nucleotides and a sequence comprising at least 50% sequence identity with tracrRNA in six adjacent nucleotides.

[0085] On one hand, this disclosure provides a method for selecting cells using a reporter element and excising the reporter element from the cells, comprising: contacting a target nucleic acid with a complex comprising a site-directed polypeptide and a nucleic acid targeting the nucleic acid; cleaving the target nucleic acid with the site-directed polypeptide to generate cleaved target nucleic acid; inserting a donor polynucleotide comprising a genetic factor of interest; and a reporter element into the cleaved target nucleic acid, wherein the reporter element comprises a polynucleotide sequence encoding the site-directed polypeptide and one or more nucleic acids, wherein the one or more nucleic acids comprise a sequence comprising at least 50% sequence identity with crRNA in six adjacent nucleotides and a sequence comprising at least 50% sequence identity with tracrRNA in six adjacent nucleotides; and selecting cells based on the donor polynucleotide to generate selected cells. In some embodiments, selection includes selecting cells from a subject undergoing disease treatment. In some embodiments, selection includes selecting cells from a subject undergoing disease diagnosis. In some embodiments, after selection, the cells comprise the donor polynucleotide. In some embodiments, the method further comprises excising all, some, or no reporter elements, thereby generating a second selected cell. In some embodiments, excision includes contacting the 5' end of the reporter element with a complex comprising a site-directed peptide and a nucleic acid targeting a nucleic acid, wherein the complex cleaves the 5' end. In some embodiments, excision includes contacting the 3' end of the reporter element with a complex comprising a site-directed peptide and a nucleic acid targeting a nucleic acid, wherein the complex cleaves the 3' end. In some embodiments, excision includes contacting the 5' and 3' ends of the reporter element with one or more complexes comprising a site-directed peptide and a nucleic acid targeting a nucleic acid, wherein the complex cleaves the 5' and 3' ends. In some embodiments, the method further includes screening a second selected cell. In some embodiments, screening includes observing the absence of all or some reporter elements.

[0086] On one hand, this disclosure provides a composition comprising: a nucleic acid comprising: a spacer region, wherein the spacer region is 12-30 nucleotides (including 12 and 30) and wherein the spacer region is adapted to hybridize with a 5' sequence in a PAM; a first double strand, wherein the first double strand is at the 3' of the spacer region; a protrusion, wherein the protrusion comprises at least 3 unpaired nucleotides on a first strand of the first double strand and at least 1 unpaired nucleotide on a second strand of the first double strand; a linker, wherein the linker connects the first and second strands of the double strand and is at least 3 nucleotides in length; a P-domain; and a second double strand, wherein the second double strand is at the 3' of the P-domain and is adapted to bind to a site-directed polypeptide. In some embodiments, the 5' sequence in the PAM is at least 18 nucleotides long. In some embodiments, the 5' sequence in the PAM is adjacent to the PAM. In some embodiments, the PAM comprises 5'-NGG-3'. In some embodiments, the first double strand is adjacent to the spacer region. In some embodiments, the P-domain begins 1-5 nucleotides downstream of the duplex, comprises at least 4 nucleotides, and is adapted to hybridize with sequences selected from: a 5'-NGG-3' pre-interstitial motif sequence, a sequence containing at least 50% identity with amino acids 1096-1225 of Cas9 from Streptococcus pyogenes, or any combination thereof. In some embodiments, the site-directed polypeptide comprises at least 15% identity with the nuclease domain of Cas9 from Streptococcus pyogenes. In some embodiments, the nucleic acid is RNA. In some embodiments, the nucleic acid is type A RNA. In some embodiments, the first duplex is at least 6 nucleotides long. In some embodiments, the three unpaired nucleotides of the protrusion comprise 5'-AAG-3'. In some embodiments, a nucleotide forming a wobble pair with a nucleotide on the second strand of the first duplex is adjacent to the three unpaired nucleotides. In some embodiments, the polypeptide binds to a nucleic acid region selected from: the first duplex, the second duplex, and the P-domain, or any combination thereof.

[0087] On one hand, this disclosure provides a method for modifying a target nucleic acid, comprising: contacting the target nucleic acid with a composition comprising: a nucleic acid comprising: a spacer region, wherein the spacer region is 12-30 nucleotides (including 12 and 30) and wherein the spacer region is adapted to hybridize with a 5' sequence in a PAM; a first double strand, wherein the first double strand is at the 3' of the spacer region; a protrusion, wherein the protrusion comprises at least 3 unpaired nucleotides on a first strand of the first double strand and at least 1 unpaired nucleotide on a second strand of the first double strand; a linker, wherein the linker connects the first and second strands of the double strand and is at least 3 nucleotides in length; a P-domain; and a second double strand, wherein the second double strand is at the 3' of the P-domain and is adapted to bind to a site-directed polypeptide; and modifying the target nucleic acid. In some embodiments, the method further comprises contacting with a site-directed polypeptide. In some embodiments, the contact comprises contacting the spacer region with the target nucleic acid. In some embodiments, the modification comprises cleaving the target nucleic acid to produce cleaved target nucleic acid. In some embodiments, the cleavage is performed via the site-directed polypeptide. In some embodiments, the method further includes inserting a donor polynucleotide into the cleaved target nucleic acid. In some embodiments, the modification includes modifying the transcription of the target nucleic acid.

[0088] On one hand, this disclosure provides a vector comprising a polynucleotide sequence encoding a nucleic acid comprising: a spacer region of 12-30 nucleotides (including 12 and 30) and wherein the spacer region is adapted to hybridize with a sequence at the 5' of a PAM; a first duplex, wherein the first duplex is at the 3' of the spacer region; a protrusion, wherein the protrusion comprises at least three unpaired nucleotides on a first strand of the first duplex and at least one unpaired nucleotide on a second strand of the first duplex; a linker, wherein the linker connects the first and second strands of the duplex and is at least three nucleotides in length; a P-domain; and a second duplex, wherein the second duplex is at the 3' of the P-domain and is adapted to bind to a site-directed polypeptide.

[0089] On one hand, this disclosure provides a kit comprising: a composition comprising: a nucleic acid comprising: a spacer region, wherein the spacer region is 12-30 nucleotides (including 12 and 30) and wherein the spacer region is adapted to hybridize with a 5' sequence in a PAM; a first duplex, wherein the first duplex is 3' of the spacer region; a protrusion, wherein the protrusion comprises at least 3 unpaired nucleotides on a first strand of the first duplex and at least 1 unpaired nucleotide on a second strand of the first duplex; a linker, wherein the linker connects the first and second strands of the duplex and is at least 3 nucleotides in length; a P-domain; a second duplex, wherein the second duplex is 3' of the P-domain and adapted to bind to a site-directed polypeptide; and a buffer. In some embodiments, the kit further comprises a site-directed polypeptide. In some embodiments, the kit further comprises a donor polynucleotide. In some embodiments, the kit further comprises instructions for use.

[0090] On one hand, this disclosure provides a method for manufacturing synthetically designed nucleic acid targeting nucleic acids, comprising: designing a composition comprising: a nucleic acid comprising: a spacer region, wherein the spacer region is 12-30 nucleotides (including 12 and 30) and wherein the spacer region is adapted to hybridize with a sequence at the 5' of a PAM; a first duplex, wherein the first duplex is at the 3' of the spacer region; a protrusion, wherein the protrusion comprises at least 3 unpaired nucleotides on a first strand of the first duplex and at least 1 unpaired nucleotide on a second strand of the first duplex; a linker, wherein the linker connects the first and second strands of the duplex and is at least 3 nucleotides in length; a P-domain; and a second duplex, wherein the second duplex is at the 3' of the P-domain and is adapted to bind to a site-directed polypeptide.

[0091] On one hand, this disclosure provides a pharmaceutical composition comprising an engineered nucleic acid targeting a nucleic acid, said engineered nucleic acid targeting a nucleic acid selected from: engineered nucleic acid targeting a nucleic acid with a mutation in the P-domain of said nucleic acid targeting a nucleic acid; or engineered nucleic acid targeting a nucleic acid with a mutation in the protrusion region of said nucleic acid targeting a nucleic acid.

[0092] On one hand, this disclosure provides a pharmaceutical composition comprising a composition selected from the following: a composition comprising: an engineered nucleic acid targeting a nucleic acid including a 3' hybridization overhang, and a donor polynucleotide, wherein the donor polynucleotide hybridizes to the 3' hybridization overhang; a composition comprising: an effector protein and a nucleic acid, wherein the nucleic acid comprises: at least 50% sequence identity with crRNA in 6 adjacent nucleotides, at least 50% sequence identity with tracrRNA in 6 adjacent nucleotides, and a non-natural sequence, wherein the nucleic acid is adapted to bind to the effector protein; a composition comprising: a multiple genetic target, wherein... A multiple genetic targeting agent comprises one or more nucleic acid modules, wherein the nucleic acid modules comprise a non-natural sequence, and wherein the nucleic acid modules are configured to bind to a polypeptide comprising at least 10% amino acid sequence identity with a nuclease domain of Cas9, and wherein the nucleic acid modules are configured to hybridize with a target nucleic acid; a composition comprising: a modified site-directed polypeptide, wherein the polypeptide is modified to be suitable for targeting a motif adjacent to a second pre-interstitial region sequence compared to a wild-type site-directed polypeptide; a composition comprising: a modified site-directed polypeptide, wherein the polypeptide is modified to be suitable for targeting a second nucleic acid-targeted nucleic acid compared to a wild-type site-directed polypeptide; a composition comprising: [the following is incomplete and requires further context: SEQ] Compared to SEQ ID:8, a modified site-directed polypeptide containing a modification in the bridged helix; a composition comprising: a modified site-directed polypeptide containing a modification in the high-alkalinity patch compared to SEQ ID:8; a composition comprising: a modified site-directed polypeptide containing a modification in the polymerase domain compared to SEQ ID:8; a composition comprising: a modified site-directed polypeptide containing a modification in the bridged helix, the high-alkalinity patch, the nuclease domain, and the polymerase domain compared to SEQ ID:8, or any combination thereof ... compared to SEQ ID:8, a modified site-directed polypeptide containing a modification in the bridged helix, the high-alkalinity patch, the nuclease domain, and the polymerase domain; ID:8 Compared to a site-directed polypeptide containing a modified nuclease domain; a composition comprising: a first complex comprising a first site-directed polypeptide and a first nucleic acid targeting a nucleic acid, a second complex comprising a second site-directed polypeptide and a second nucleic acid targeting a nucleic acid, wherein the first and second nucleic acids targeting a nucleic acid are different; a composition comprising: a plurality of nucleic acid molecules, wherein each nucleic acid molecule comprises a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid targeting a nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide, and a fusion polypeptide, wherein the fusion polypeptide comprises a plurality of the nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites;And a composition comprising: a nucleic acid comprising: a spacer region, wherein the spacer region is 12-30 nucleotides (inclusive) and wherein the spacer region is adapted to hybridize with a 5' sequence in a PAM; a first duplex, wherein the first duplex is 3' of the spacer region; a protrusion, wherein the protrusion comprises at least 3 unpaired nucleotides on a first strand of the first duplex and at least 1 unpaired nucleotide on a second strand of the first duplex; a linker, wherein the linker connects the first and second strands of the duplex and is at least 3 nucleotides in length; a P-domain and a second duplex, wherein the second duplex is 3' of the P-domain and is adapted to bind to a site-directed polypeptide; or any combination thereof.

[0093] On one hand, this disclosure provides a pharmaceutical composition comprising a modified site-directed polypeptide comprising: a first nuclease domain, a second nuclease domain, and an inserted nuclease domain.

[0094] On one hand, this disclosure provides a pharmaceutical composition comprising a donor polynucleotide, the donor polynucleotide comprising: a genetic factor of interest and a reporter element, wherein the reporter element comprises a polynucleotide sequence encoding a site-directed polypeptide and one or more nucleic acids, wherein the one or more nucleic acids comprise a sequence in six adjacent nucleotides containing at least 50% sequence identity with crRNA and a sequence in six adjacent nucleotides containing at least 50% sequence identity with tracrRNA. On one hand, this disclosure provides a pharmaceutical composition comprising a vector selected from the following: a vector comprising a polynucleotide sequence encoding an engineered nucleic acid targeting a nucleic acid, wherein the engineered nucleic acid targeting a nucleic acid comprises a mutation in the P-domain of the nucleic acid targeting a nucleic acid; a vector comprising a polynucleotide sequence encoding an engineered nucleic acid targeting a nucleic acid, wherein the engineered nucleic acid targeting a nucleic acid comprises a mutation in a protrusion region of the nucleic acid targeting a nucleic acid; and a vector modifying the target nucleic acid; a vector comprising a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence; a vector comprising a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence; and a vector comprising a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence. The vector contains a nucleic acid targeting a nucleic acid comprising a sequence and a site-directed polypeptide constructed to bind to an effector protein; a vector comprising a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence, a site-directed polypeptide, and an effector protein; a vector comprising a polynucleotide sequence encoding a multiple genetic target, wherein the multiple genetic target comprises one or more nucleic acid modules, wherein the nucleic acid modules comprise a non-natural sequence, and wherein the nucleic acid modules are constructed to bind to a polypeptide comprising at least 10% amino acid sequence identity with the nuclease domain of Cas9, and wherein the nucleic acid modules are constructed to hybridize with the target nucleic acid; and a vector comprising a polynucleotide sequence encoding a modified site-directed polypeptide, wherein the modified site-directed polypeptide is identical to SEQ. ID:8 compared to modifications contained in bridging helices, high-basic patches, nuclease domains, and polymerase domains, or any combination thereof; a vector comprising a polynucleotide sequence encoding two or more nucleic acids targeting nucleic acids differing by at least one nucleotide and a site-directed polypeptide; a vector comprising: a polynucleotide sequence encoding a composition comprising: a plurality of nucleic acid molecules, wherein each nucleic acid molecule contains a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid-target nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide; and a fusion polypeptide comprising a plurality of nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites to stoichiometrically deliver the composition to the subcellular location;Expression vectors comprising a polynucleotide sequence encoding a genetic factor of interest and a reporter element, wherein the reporter element comprises a polynucleotide sequence encoding a site-directed polypeptide and one or more nucleic acids, wherein the one or more nucleic acids comprise a sequence containing at least 50% sequence identity with crRNA in 6 adjacent nucleotides and a sequence containing at least 50% sequence identity with tracrRNA in 6 adjacent nucleotides; and a vector comprising a polynucleotide sequence encoding a nucleic acid comprising: a spacer region, wherein the spacer region is 12-30 nucleotides (inclusive) and wherein the spacer region is adapted to hybridize with a sequence at the 5' of a PAM; a first duplex, wherein the first duplex is at the 3' of the spacer region; a protrusion, wherein the protrusion comprises at least 3 unpaired nucleotides on the first strand of the first duplex and at least 1 unpaired nucleotide on the second strand of the first duplex; a linker, wherein the linker connects the first and second strands of the duplex and is at least 3 nucleotides in length; a P-domain; and a second duplex, wherein the second duplex is at the 3' of the P-domain and adapted to bind to a site-directed polypeptide; or any combination thereof.

[0095] On one hand, this disclosure provides a method for treating a disease, comprising administering to a subject: an engineered nucleic acid targeting a nucleic acid, comprising: a mutation in the P-domain of the nucleic acid targeting a nucleic acid; an engineered nucleic acid targeting a nucleic acid, comprising: a mutation in a protrusion region of the nucleic acid targeting a nucleic acid; a composition comprising: an engineered nucleic acid targeting a nucleic acid containing a 3' hybridization overhang and a donor polynucleotide, wherein the donor polynucleotide hybridizes with the 3' hybridization overhang; a composition comprising: an effector protein and a nucleic acid, wherein the nucleic acid comprises: at least 50% sequence identity with crRNA in 6 adjacent nucleotides, at least 50% sequence identity with tracrRNA in 6 adjacent nucleotides, and a non-natural sequence, wherein the nucleic acid is adapted to bind to the effector protein; a composition A composition comprising: a multiple genetic target, wherein the multiple genetic target comprises one or more nucleic acid modules, wherein the nucleic acid modules comprise a non-natural sequence, and wherein the nucleic acid modules are configured to bind to a polypeptide comprising at least 10% amino acid sequence identity with a nuclease domain of Cas9, and wherein the nucleic acid modules are configured to hybridize with a target nucleic acid; a composition comprising: a modified site-directed polypeptide, wherein the polypeptide is modified to be adapted to target a motif adjacent to a second pre-interstitial region sequence compared to a wild-type site-directed polypeptide; a composition comprising: a modified site-directed polypeptide, wherein the polypeptide is modified to be adapted to target a second nucleic acid target compared to a wild-type site-directed polypeptide; a composition comprising: a modified site-directed polypeptide comprising a modified portion of a bridged helix compared to SEQ ID:8; a composition comprising: a modified site-directed polypeptide comprising, compared to SEQ ID:8; Compared to SEQ ID:8, a modified site-directed polypeptide containing a modified polymerase domain; a composition comprising: a modified site-directed polypeptide containing a modified polymerase domain compared to SEQ ID:8; a composition comprising: a modified site-directed polypeptide containing a modified bridge helix, a high-basic patch, a nuclease domain, and a polymerase domain compared to SEQ ID:8, or any combination thereof; a composition comprising: a modified site-directed polypeptide containing a modified nuclease domain compared to SEQ ID:8; a composition comprising: a first complex comprising a first site-directed polypeptide and a first nucleic acid targeting a nucleic acid, a second complex comprising a second site-directed polypeptide and a second nucleic acid targeting a nucleic acid, wherein the first and second nucleic acid targeting nucleic acids are different; a composition comprising: a plurality of nucleic acid molecules, wherein each nucleic acid molecule contains a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid targeting a nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide, and a fusion polypeptide, wherein the fusion polypeptide contains a plurality of the nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites;A composition comprising: a nucleic acid comprising: a spacer region, wherein the spacer region is 12-30 nucleotides (inclusive) and wherein the spacer region is adapted to hybridize with a sequence at the 5' of a PAM; a first duplex, wherein the first duplex is at the 3' of the spacer region; a protrusion, wherein the protrusion comprises at least 3 unpaired nucleotides on a first strand of the first duplex and at least 1 unpaired nucleotide on a second strand of the first duplex; a linker, wherein the linker connects the first and second strands of the duplex and is at least 3 nucleotides in length; a P-domain and a second duplex, wherein the second duplex is a... The 3' of the P-domain is adapted to bind to a site-directed polypeptide; a modified site-directed polypeptide comprising: a first nuclease domain, a second nuclease domain, and an inserted nuclease domain; a donor polynucleotide comprising: a genetic factor of interest and a reporter element, wherein the reporter element comprises a polynucleotide sequence encoding the site-directed polypeptide and one or more nucleic acids, wherein the one or more nucleic acids comprise a sequence in six adjacent nucleotides containing at least 50% sequence identity with crRNA and a sequence in six adjacent nucleotides containing at least 50% sequence identity with tracrRNA; comprising a polynucleotide sequence encoding an engineered nucleic acid targeting a nucleic acid. The vector, wherein the engineered nucleic acid-targeting nucleic acid comprises: a mutation in the P-domain of the nucleic acid-targeting nucleic acid; a vector comprising a polynucleotide sequence encoding the engineered nucleic acid-targeting nucleic acid, wherein the engineered nucleic acid-targeting nucleic acid comprises: a mutation in the protrusion region of the nucleic acid-targeting nucleic acid; and modification of the target nucleic acid; a vector comprising a polynucleotide sequence encoding the modified nucleic acid-targeting nucleic acid, wherein the modified nucleic acid-targeting nucleic acid comprises a non-natural sequence; a vector comprising a polynucleotide sequence encoding the modified nucleic acid-targeting nucleic acid, wherein the modified nucleic acid-targeting nucleic acid comprises a sequence constructed to bind to an effector protein, and site-directed... The following are not directly related to the preceding text: a vector comprising a polynucleotide sequence encoding a modified nucleic acid targeting a nucleic acid, wherein the modified nucleic acid targeting a nucleic acid comprises a non-natural sequence, a site-directed peptide, and an effector protein; a vector comprising a polynucleotide sequence encoding a multiple genetic target, wherein the multiple genetic target comprises one or more nucleic acid modules, wherein the nucleic acid modules comprise a non-natural sequence, and wherein the nucleic acid modules are configured to bind to a peptide comprising at least 10% amino acid sequence identity with a nuclease domain of Cas9, and wherein the nucleic acid modules are configured to hybridize with a target nucleic acid; a vector comprising a polynucleotide sequence encoding a modified site-directed peptide, wherein the modified site-directed peptide contains modifications in a bridging helix, a high-basic patch, a nuclease domain, and a polymerase domain, or any combination thereof, compared to SEQ ID:8; and a vector comprising a polynucleotide sequence encoding two or more nucleic acids targeting a nucleic acid differing from each other by at least one nucleotide and a site-directed peptide.A vector comprising: a polynucleotide sequence encoding a composition comprising: a plurality of nucleic acid molecules, wherein each nucleic acid molecule contains a nucleic acid-binding protein binding site, wherein at least one of the plurality of nucleic acid molecules encodes a nucleic acid targeting a nucleic acid and one of the plurality of nucleic acid molecules encodes a site-directed polypeptide; and a fusion polypeptide comprising a plurality of nucleic acid-binding proteins, wherein the plurality of nucleic acid-binding proteins are adapted to bind to their associated nucleic acid-binding protein binding sites to stoichiometrically deliver the composition to the subcellular location; and an expression vector comprising a polynucleotide sequence encoding a genetic factor of interest and a reporter element, wherein the reporter element comprises a polynucleotide sequence encoding a site-directed polypeptide and one or more nucleic acids, wherein the one or more nucleic acids contain at least 50% sequence identity with crRNA in 6 adjacent nucleotides. The sequence and a sequence containing at least 50% sequence identity with tracrRNA within 6 adjacent nucleotides; and a vector containing a polynucleotide sequence encoding a nucleic acid comprising: a spacer region, wherein the spacer region is 12-30 nucleotides (inclusive) and wherein the spacer region is adapted to hybridize with a sequence at the 5' of PAM; a first duplex, wherein the first duplex is at the 3' of the spacer region; a protrusion, wherein the protrusion comprises at least 3 unpaired nucleotides on the first strand of the first duplex and at least 1 unpaired nucleotide on the second strand of the first duplex; a linker, wherein the linker connects the first and second strands of the duplex and is at least 3 nucleotides in length; a P-domain; and a second duplex, wherein the second duplex is at the 3' of the P-domain and adapted to bind to a site-directed polypeptide; or any combination thereof. In some embodiments, the administration comprises administration via viral delivery. In some embodiments, the administration comprises administration via electroporation. In some embodiments, the administration comprises administration via nanoparticle delivery. In some embodiments, the administration comprises administration via liposome delivery. In some embodiments, the administration includes administration by a method selected from: intravenous administration, subcutaneous administration, intramuscular administration, oral administration, rectal administration, aerosol administration, parenteral administration, ocular administration, pulmonary administration, transdermal administration, vaginal administration, ocular administration, intranasal administration, and topical administration, or any combination thereof. In some embodiments, the methods of this disclosure are performed in cells selected from: plant cells, microbial cells, and fungal cells, or any combination thereof.

[0096] This is incorporated through this citation.

[0097] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference, as if each publication, patent or patent application were expressly and individually indicated to be incorporated herein by reference. Brief description of the attached diagram

[0099] The novel features of the invention are set forth in particular in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description (which illustrates exemplary embodiments utilizing the principles of the invention) and the accompanying drawings, wherein:

[0100] Figure 1A An exemplary implementation of the single-target nucleic acid targeting method disclosed herein is described.

[0101] Figure 1B An exemplary implementation of the single-target nucleic acid targeting method disclosed herein is described.

[0102] Figure 2 An exemplary implementation of the dual-targeted nucleic acid of this disclosure is described.

[0103] Figure 3 An exemplary embodiment of the sequence enrichment method of this disclosure utilizing target nucleic acid lysis is described.

[0104] Figure 4 An exemplary implementation of the sequence enrichment method of this disclosure using target nucleic acid enrichment is described.

[0105] Figure 5 An exemplary embodiment of the method of this disclosure for determining the off-target binding sites of a site-directed peptide by purifying the peptide is described.

[0106] Figure 6 An exemplary embodiment of the method of this disclosure is described, which uses the purification of nucleic acids targeting nucleic acids to determine the off-target binding sites of site-directed peptides.

[0107] Figure 7 This diagram illustrates an exemplary implementation of the array-based sequencing method for site-directed peptides disclosed herein.

[0108] Figure 8 This illustration shows an exemplary embodiment of the array-based sequencing method for site-directed peptides using the present disclosure, wherein the lysis products are sequenced.

[0109] Figure 9 This diagram illustrates an exemplary implementation of a next-generation sequencing-based method using the site-directed peptides disclosed herein.

[0110] Figure 10 An exemplary labeled single-guided nucleic acid targeting nucleic acid is described.

[0111] Figure 11 An exemplary labeled dual-guided nucleic acid targeting nucleic acid is described.

[0112] Figure 12This diagram illustrates an exemplary implementation of a method for targeting nucleic acids using labels with a separation system (e.g., a separation fluorescence system).

[0113] Figure 13 Some exemplary data depicting the effect of 5'-labeled nucleic acid-targeted nucleic acids on the cleavage of target nucleic acids.

[0114] Figure 14 The diagram illustrates an exemplary 5' labeled nucleic acid targeting a nucleic acid, with a label adapter sequence between the nucleic acid and the label.

[0115] Figure 15 An exemplary implementation of a multiple target nucleic acid lysis method is described.

[0116] Figure 16 An exemplary implementation of a chemometric delivery method for RNA nucleic acids.

[0117] Figure 17 An exemplary implementation of a chemometric delivery method for nucleic acids.

[0118] Figure 18 An exemplary embodiment is depicted that uses the site-directed peptide of this disclosure to seamlessly insert a reporter element into a target nucleic acid.

[0119] Figure 19 An exemplary implementation of removing the reporter element from the target nucleic acid is described.

[0120] Figure 20 The complementary portions of the pre-CRISPR nucleic acid sequence and the tracr nucleic acid sequence from Streptococcus pyogenes SF370 were depicted.

[0121] Figure 21 An exemplary secondary structure of a synthetic single-guided nucleic acid targeting a nucleic acid is depicted.

[0122] Figure 22A Figures B and B show exemplary single-guided nucleic acid backbone variants targeting nucleic acids. The nucleotides in the boxes correspond to nucleotides that have been altered relative to the CRISPR sequence labeled FL-tracr-crRNA.

[0123] Figure 23A -C shows exemplary data from in vitro lysis analysis. The results confirm that more than one synthesized nucleic acid-targeted nucleotide backbone sequence can be cleaved using site-directed peptides (e.g., Cas9).

[0124] Figure 24This shows an exemplary synthetic single-guided nucleic acid sequence targeting a nucleic acid and containing a variant in the complementary region / double strand. The nucleotides in the box correspond to nucleotides that have been altered relative to the CRISPR sequence labeled FL-tracr-crRNA.

[0125] Figure 25 An exemplary variant of the single-guided nucleic acid structure targeting nucleic acid is shown within the 3' region of the complementary region / double strand. The nucleotides in the box correspond to nucleotides whose pairing with the naturally occurring Streptococcus pyogenes SF370 CRISPR and tracr nucleic acid sequences has been altered.

[0126] Figure 26A -B indicates an exemplary variant of the single-guided nucleic acid structure targeting nucleic acid within the 3' region of the complementary region / double strand. The nucleotides in the box correspond to nucleotides whose pairing with the naturally occurring Streptococcus pyogenes SF370 CRISPR and tracr nucleic acid sequences has been altered.

[0127] Figure 27A -B shows an exemplary variant of a nucleic acid-targeted nucleic acid structure containing an additional hairpin structure sequence of a CRISPR repeat fragment derived from Pseudomonas aeruginosa (PA14). The sequence in the box can bind to the ribonuclease Csy4 from PA14.

[0128] Figure 28 Exemplary data from an in vitro lysis assay are shown, confirming that multiple synthetic nucleic acid-targeted nucleotide backbone sequences support Cas9 lysis. The upper and lower gel images represent two independent replicates of this assay.

[0129] Figure 29 Exemplary data from an in vitro lysis assay are shown, confirming that multiple synthetic nucleic acid-targeted nucleotide backbone sequences support Cas9 lysis. The upper and lower gel images represent two independent replicates of this assay.

[0130] Figure 30 The present disclosure describes an exemplary method for delivering donor polynucleotides to modification sites in target nucleic acids.

[0131] Figure 31 Describe a system for storing and sharing electronic information.

[0132] Figure 32 An exemplary embodiment depicting two nickases generating blunt-ended nicks in target nucleic acids is shown. No site-specific modified peptides complexing with nucleic acids targeting nucleic acids are shown.

[0133] Figure 33An exemplary embodiment is depicted using two cleavage enzymes to generate staggered cutting of the target nucleic acid with sticky ends. No site-specific modified peptides complexed with the nucleic acid targeting the nucleic acid are shown.

[0134] Figure 34 An exemplary embodiment is depicted using two cleavage enzymes to generate staggered cutting of the target nucleic acid with medium-sized sticky ends. No site-modified peptides complexed with the nucleic acid-targeting nucleic acid are shown.

[0135] Figure 35 This diagram illustrates the sequence alignment of Cas9 orthologs. Amino acids marked with an "X" below them are considered similar. Amino acids marked with a "Y" below them are considered highly conserved or identical across all sequences. Amino acid residues without an "X" or "Y" may not be conserved.

[0136] Figure 36 This demonstrates the ability of nucleic acid variants targeting nucleic acids to cleave target nucleic acids. Figure 36 The variant tested in the middle corresponds to Figure 22. Figure 24 and Figure 25 The variant depicted in the text.

[0137] Figure 37A -D indicates the in vitro lysis assay using variant nucleic acids targeting nucleic acids.

[0138] Figure 38 An exemplary amino acid sequence of Csy4 from wild-type Pseudomonas aeruginosa is depicted.

[0139] Figure 39 An exemplary amino acid sequence of a non-enzymatic ribonuclease (e.g., Csy4) is depicted.

[0140] Figure 40 An exemplary amino acid sequence of Csy4 from Pseudomonas aeruginosa is depicted.

[0141] Figure 41A -J depicts an exemplary Cas6 amino acid sequence.

[0142] Figure 42A -C depicts an exemplary Cas6 amino acid sequence. Invention Details

[0144] definition

[0145] As used in this article, "affinity marker" can refer to peptide affinity markers or nucleic acid affinity markers. An affinity marker is generally a protein or nucleic acid sequence that can bind to a molecule (e.g., via small molecules, proteins, or covalent bonds). Affinity markers can be non-natural sequences. Peptide affinity markers may contain peptides. Peptide affinity markers can be affinity markers that can be part of a separation system (e.g., two inactive peptide fragments can be trans-coupled to form an active affinity marker). Nucleic acid affinity markers may contain nucleic acids. Nucleic acid affinity markers can be sequences that selectively bind to known nucleic acid sequences (e.g., through hybridization). Nucleic acid affinity markers can be sequences that selectively bind to proteins. Affinity markers can be fused to native proteins. Affinity markers can be fused to nucleotide sequences. Sometimes, one, two, or more affinity markers can be fused to native proteins or nucleotide sequences. Affinity markers can be introduced into nucleic acids targeting nucleic acids using in vitro or in vivo transcription methods. Nucleic acid affinity tags may include, for example, chemical tags, RNA-binding protein binding sequences, DNA-binding protein binding sequences, sequences that can hybridize to affinity-tagged polynucleotides, synthetic RNA aptamers, or synthetic DNA aptamers. Examples of chemical nucleic acid affinity tags may include, but are not limited to, ribo-nucleotriphosphates containing biotin, fluorescent dyes, and digoxigenin. Examples of nucleic acid affinity tags for binding proteins may include, but are not limited to, MS2 binding sequences, U1A binding sequences, stem-loop binding protein sequences, boxB sequences, eIF4A sequences, or any sequences recognized by RNA-binding proteins. Examples of nucleic acid affinity-tagged oligonucleotides may include, but are not limited to, biotinylated oligonucleotides, 2,4-dinitrophenyl oligonucleotides, fluorescein oligonucleotides, and primary amine-conjugated oligonucleotides.

[0146] Nucleic acid affinity markers can be RNA aptamers. Aptamers may include those binding to theophylline, streptavidin, dextran B512, adenosine, guanosine, guanine / xanthine, and 7-methyl-GTP; amino acid aptamers, such as those binding to arginine, citrulline, valine, tryptophan, cyanocobalamin, N-methylporphyrin IX, flavin, and NAD; and antibiotic aptamers, such as those binding to tobramycin, neomycin, penicillin, kanamycin, streptomycin, purpuric acid, and chloramphenicol.

[0147] Nucleic acid affinity tags may comprise RNA sequences that can be bound by site-directed peptides. These site-directed peptides may be conditionally enzymatically inactive. The RNA sequence may comprise sequences that can be bound by members of type I, II, and / or III CRISPR systems. The RNA sequence may be bound by RAMP family member proteins. The RNA sequence may be bound by Cas6 family member proteins (e.g., Csy4, Cas6). The RNA sequence may be bound by Cas5 family member proteins (e.g., Cas5). For example, Csy4 can bind to a specific RNA hairpin sequence with high affinity (Kd ~ 50 pM) and cleave the RNA at the 3' site of the hairpin structure. Cas5 or Cas6 family member proteins can bind RNA sequences containing at least about or at most about 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% sequence identity and / or sequence similarity to the following nucleotide sequences:

[0148] 5'-GUUCACUGCCGUAUAGGCAGCUAAGAAA-3';

[0149]

[0150] and 5'-GUCGCGCCCCGCAUGGGGCGCGUGGAUUGAAA-3',

[0151] Nucleic acid affinity tags may contain DNA sequences that can be bound by site-directed peptides. These site-directed peptides may be conditionally enzymatically inactive. The DNA sequence may contain sequences that can be bound by members of type I, II, and / or III CRISPR systems. The DNA sequence may be bound by Argonaut proteins. The DNA sequence may be bound by proteins containing zinc finger domains, TALE domains, or any other DNA-binding domains.

[0152] Nucleic acid affinity markers may include ribozyme sequences. Suitable ribozymes may include peptidyl transferase 23S rRNA, RNase P, type I introns, type II introns, GIR1 branched ribozymes, leadzymes, hairpin ribozymes, hammerhead ribozymes, HDV ribozymes, CPEB3 ribozymes, VS ribozymes, glmS ribozymes, CoTC ribozymes, and synthetic ribozymes.

[0153] Peptide affinity labeling may include labels that can be used for tracing or purification (e.g., fluorescent protein, green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato, his label (e.g., 6XHis label), hemagglutinin (HA) label, FLAG label, Myc label, GST label, MBP label, chitin-binding protein label, calmodulin label, V5 label, streptavidin-binding label, etc.).

[0154] Nucleic acid and peptide affinity labeling can both include small molecule labels, such as biotin or digitalisin, and fluorescent labels, such as fluorescein, rhodamine, Alexa fluorite dye, anthocyanin 3 dye, and anthocyanin 5 dye.

[0155] Nucleic acid affinity tags can be located at the 5' end of a nucleic acid (e.g., a nucleic acid targeting a nucleic acid). Nucleic acid affinity tags can be located at the 3' end of a nucleic acid. Nucleic acid affinity tags can be located at both the 5' and 3' ends of a nucleic acid. Nucleic acid affinity tags can be located within a nucleic acid. Peptide affinity tags can be located at the N-terminus of a polypeptide sequence. Peptide affinity tags can be located at the C-terminus of a polypeptide sequence. Peptide affinity tags can be located at both the N-terminus and C-terminus of a polypeptide sequence. Multiple affinity tags can be fused to nucleic acid and / or polypeptide sequences.

[0156] As used herein, "capture agent" generally refers to a reagent capable of purifying peptides and / or nucleic acids. Capture agents can be bioactive molecules or materials (e.g., any biological substance found or synthesized in nature, including but not limited to cells, viruses, subcellular particles, proteins, and more specifically antibodies, immunoglobulins, antigens, lipoproteins, glycoproteins, peptides, polypeptides, protein complexes, (streptavidin-biotin complexes), ligands, receptors or small molecules, aptamers, nucleic acids, DNA, RNA, peptide nucleic acids, oligosaccharides, polysaccharides, lipopolysaccharides, cell metabolites, haptens, pharmacologically active substances, alkaloids, steroids, vitamins, amino acids, and sugures). In some embodiments, the capture agent may contain an affinity label. In some embodiments, the capture agent may preferentially bind to the target peptide or nucleic acid of interest. The capture agent may be free-floating in a mixture. The capture agent may bind to particles (e.g., beads, microbeads, nanoparticles). The capture agent may bind to a solid or semi-solid surface. In some cases, the capture agent binds irreversibly to the target. In other cases, the trapping agent can be reversibly bound to the target (e.g., if the target is elutable, or by using chemicals such as imidazole).

[0157] As used herein, "Cas5" generally refers to a polypeptide having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas5 polypeptide (e.g., Cas5 from D. vulgaris, and / or any sequence depicted in Figure 42). Cas5 generally refers to a polypeptide having at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas5 polypeptide (e.g., Cas5 from D. vulgaris). Cas5 can refer to the wild-type or modified form of the Cas5 protein, which may include amino acid variations such as deletions, insertions, substitutions, mutations, fusions, chimeras, or any combination thereof.

[0158] As used herein, "Cas6" generally refers to a polypeptide having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas6 polypeptide (e.g., Cas6 from *T. thermophilus*, and / or the sequence depicted in Figure 41). Cas6 generally refers to a polypeptide having at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas6 polypeptide (e.g., from *T. thermophilus*). Cas6 can refer to the wild-type or modified form of the Cas6 protein, which may include amino acid variations such as deletions, insertions, substitutions, mutations, fusions, chimeras, or any combination thereof.

[0159] As used herein, "Cas9" generally refers to a polypeptide having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from Streptococcus pyogenes (SEQ ID NO:8, SEQ ID NO:1-256, SEQ ID NO:795-1346)). Cas9 can refer to a polypeptide having up to about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas9 polypeptide (e.g., from Streptococcus pyogenes). Cas9 can refer to a wild-type or modified form of the Cas9 protein, which may include amino acid variations such as deletions, insertions, substitutions, mutations, fusions, chimeras, or any combination thereof.

[0160] As used herein, “cell” generally refers to a biological cell. A cell can be the basic structural, functional, and / or biological unit of a living organism. Cells can originate from any organism having one or more cells. Some non-limiting examples include: prokaryotic cells, eukaryotic cells, bacterial cells, archaea cells, cells of unicellular eukaryotes, protozoan cells, cells from plants (e.g., cells from crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, lycophytes, hornwort, liverwort, mosses), algal cells (e.g., *Botryococcus braunii*, *Chlamydomonas reinhardtii*, *Nannochloropsis gaditana*, *Chlorella pyrenoidosa*, *Sargassum patens*). Cells can be derived from various organisms, including seaweed (such as kelp), fungal cells (such as yeast cells, mushroom cells), animal cells, cells from invertebrates (such as fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (such as fish, amphibians, reptiles, birds, mammals), and cells from mammals (such as pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Cells are sometimes not derived from natural organisms (for example, cells can be synthetic, sometimes called artificial cells).

[0161] Cells can be in vitro. Cells can be in vivo. Cells can be isolated cells. Cells can be cells within an organism. Cells can be cells in a cell culture. Cells can be one of a group of cells. Cells can be prokaryotic cells or cells derived from prokaryotic cells. Cells can be bacterial cells or cells that may be derived from bacterial cells. Cells can be archaea cells or cells derived from archaea cells. Cells can be eukaryotic cells or cells derived from eukaryotic cells. Cells can be plant cells or cells derived from plant cells. Cells can be animal cells or cells derived from animal cells. Cells can be invertebrate cells or cells derived from invertebrate cells. Cells can be vertebrate cells or cells derived from vertebrate cells. Cells can be mammalian cells or cells derived from mammalian cells. Cells can be rodent cells or cells derived from rodent cells. Cells can be human cells or cells derived from human cells. Cells can be microbial cells or cells derived from microbial cells. Cells can be fungal cells or cells derived from fungal cells.

[0162] Cells can be stem cells or progenitor cells. Cells can include stem cells (e.g., adult stem cells, embryonic stem cells, iPS cells) and progenitor cells (e.g., cardiac progenitor cells, neural progenitor cells, etc.). Cells can include mammalian stem cells and progenitor cells, including rodent stem cells, rodent progenitor cells, human stem cells, human progenitor cells, etc. Cloned cells may contain cell progeny. Cells may contain target nucleic acids. Cells may be in a living organism. Cells may be genetically modified cells. Cells may be host cells.

[0163] The cell may be a totipotent stem cell, but in some embodiments of this disclosure, the term "cell" may be used but may not refer to a totipotent stem cell. The cell may be a plant cell, but in some embodiments of this disclosure, the term "cell" may be used but may not refer to a plant cell. The cell may be a pluripotent cell. For example, the cell may be a pluripotent hematopoietic cell that can differentiate into other cells in the hematopoietic cell lineage but may not differentiate into any other non-hematopoietic cells. The cell may be capable of developing into the entire organism. The cell may or may not be capable of developing into the entire organism. The cell may be the entire organism.

[0164] The cells can be primary cells. For example, a culture of primary cells can be passaged 0, 1, 2, 4, 5, 10, 15, or more times. The cells can be single-celled organisms. The cells can grow in culture.

[0165] The cells can be diseased cells. Diseased cells may have altered metabolic, gene expression, and / or morphological characteristics. Diseased cells can be cancer cells, diabetic cells, and apoptotic cells. Diseased cells can originate from a diseased object. Exemplary diseases include blood disorders, cancer, metabolic disorders, eye diseases, organ disorders, musculoskeletal disorders, heart diseases, etc.

[0166] If the cells are primary cells, they can be harvested from an individual by any method. For example, leukocytes can be harvested via plasma exchange, leukocyte removal, density gradient separation, etc. Cells can be harvested from tissues such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, stomach, etc., via biopsy. Appropriate solutions can be used to disperse or suspend the harvested cells. Such solutions are typically balanced salt solutions (e.g., physiological saline, phosphate-buffered saline (PBS), Hank's balanced salt solution, etc.), conveniently supplemented with fetal bovine serum or other naturally occurring factors, coupled with low concentrations of acceptable buffers. Buffers may include HEPES, phosphate buffer, lactate buffer, etc. Cells can be used immediately, or they can be stored (e.g., by freezing). Frozen cells can be thawed and reused. Cells can be thawed in DMSO, serum, culture medium buffers (e.g., 10% DMSO, 50% serum, 40% buffered medium), and / or some other such common solutions used for preserving cells at freezing temperatures.

[0167] As used herein, "conditionally enzymatically inactive site-directed peptides" generally refer to peptides that can bind to nucleic acid sequences in polynucleotides in a sequence-specific manner, but cleave target polynucleotides only under one or more conditions that activate the enzymatic domain. Conditionally enzymatically inactive site-directed peptides may contain enzymatically inactive domains that can be conditionally activated. Conditionally enzymatically inactive site-directed peptides can be conditionally activated in the presence of imidazole. Conditionally enzymatically inactive site-directed peptides may contain a mutant active site that cannot bind to its associated ligand to produce an enzymatically inactive site-directed peptide. This mutant active site can be programmed to bind to a ligand analog so that the ligand analog binds to the mutant active site and reactivates the site-directed peptide. For example, an ATP-binding protein may contain a mutant active site that inhibits protein activity but is programmed to specifically bind to an ATP analog. Binding to an ATP analog, rather than ATP, can reactivate the protein. Conditionally enzymatically inactive site-directed peptides may contain one or more non-natural sequences (e.g., fusion complexes, affinity markers).

[0168] As used herein, “crRNA” generally refers to a nucleic acid having at least approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary crRNA (e.g., crRNA from Streptococcus pyogenes, e.g., SEQ ID NO: 569, SEQ ID NO: 563-679)). crRNA generally refers to a nucleic acid having at most approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary crRNA (e.g., crRNA from Streptococcus pyogenes). crRNA can refer to a modified form of crRNA, which may include nucleotide changes such as deletions, insertions or substitutions, variations, mutations, or chimeras. crRNA can be a nucleic acid that has at least approximately 60% identity with a wild-type exemplary crRNA sequence (e.g., crRNA from Streptococcus pyogenes) in a sequence of at least 6 adjacent nucleotides. For example, the crRNA sequence can be at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 95%, at least approximately 98%, at least approximately 99%, or 100% identical to a wild-type exemplary crRNA sequence (e.g., crRNA from Streptococcus pyogenes) in a sequence of at least 6 adjacent nucleotides.

[0169] The “CRISPR repeat fragment” or “CRISPR repeat sequence” used in this article may refer to the smallest CRISPR repeat sequence.

[0170] As used herein, "Csy4" generally refers to a peptide containing a wild-type exemplary Csy4 polypeptide (e.g., Csy4 from Pseudomonas aeruginosa, see [link]). Figure 40 Csy4 can generally refer to a polypeptide having at least approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary Csy4 polypeptide (e.g., Csy4 from Pseudomonas aeruginosa). Csy4 can refer to a wild-type or modified form of the Csy protein, which may include amino acid variations such as deletions, insertions, substitutions, mutations, fusions, chimeras, or any combination thereof.

[0171] As used herein, "ribonuclease" generally refers to a polypeptide capable of cleaving RNA. In some embodiments, the ribonuclease may be a site-directed polypeptide. Ribonucleases may be members of the CRISPR system (e.g., type I, type II, type III). Ribonucleases may refer to proteins of the repeat-associated doubtful protein (RAMP) superfamily (e.g., Cas6, Cas5 family). Ribonucleases may also include RNase A, RNase H, RNase I, RNase III family members (e.g., Drosha, Dicer, RNase N), RNase L, RNase P, RNase PhyM, RNase T1, RNase T2, RNase U2, RNase V1, and RNase V. Ribonucleases may refer to conditionally inactive ribonucleases. Ribonucleases may refer to catalytically inactive ribonucleases.

[0172] The term "donor polynucleotide" as used in this article can refer to nucleic acids that can be integrated into a site during genome engineering or target nucleic acid engineering.

[0173] The terms "fixative" or "crosslinking agent" used in this article generally refer to reagents that can fix or crosslink cells. Fixed or crosslinked cells can stabilize the protein-nucleic acid complexes within them. Suitable fixatives and crosslinking agents may include formaldehyde, glutaraldehyde, ethanol-based fixatives, methanol-based fixatives, acetone, acetic acid, osmium tetroxide, potassium dichromate, chromic acid, potassium permanganate, mercury preparations, picrates, formalin, oligooxymethylene, and amine-reactive NHS-ester crosslinking agents such as bis[sulfosuccinimide] octanoate (BS3), 3,3′-dithiobis[sulfosuccinimide] propionate (DTSSP), and ethylene glycol bis[sulfosuccinimide] succinate (sulfon-EGS). Disuccinimidyl glutarate (DSG), dithiobis[succinimidyl propionate] (DSP), disuccinimidyl octanoate (DSS), ethylene glycol bis[succinimidyl succinate] (EGS), NHS-ester / bisacrylidine crosslinking agents, such as NHS-bisacrylidine, NHS-LC-bisacrylidine, NHS-SS-bisacrylidine, sulfon-NHS-bisacrylidine, sulfon-NHS-LC-bisacrylidine, and sulfon-NHS-SS-bisacrylidine.

[0174] As used herein, "fusion body" can refer to a protein and / or nucleic acid containing one or more non-natural sequences (e.g., portions). Fusion bodies may contain one or more identical non-natural sequences. Fusion bodies may contain one or more different non-natural sequences. Fusion bodies can be chimeras. Fusion bodies may contain nucleic acid affinity tags. Fusion bodies may contain barcodes. Fusion bodies may contain peptide affinity tags. Fusion bodies can provide subcellular localization of site-directed peptides (e.g., nuclear localization signals (NLS) for targeting the nucleus, mitochondrial localization signals for targeting mitochondria, chloroplast localization signals for targeting chloroplasts, endoplasmic reticulum (ER) retention signals, etc.). Fusion bodies can provide non-natural sequences (e.g., affinity tags) that can be used for tracking or purification. Fusion bodies can be small molecules, such as biotin or dyes, such as Alexa fluor dye, anthocyanin 3 dye, anthocyanin 5 dye. Fusion bodies can provide increased or decreased stability.

[0175] In some implementations, the fusion may include a detectable tag, including a portion that can provide a detectable signal. Suitable detectable tags and / or portions that can provide a detectable signal may include, but are not limited to, enzymes, radioisotopes, members of specific binding pairs; fluorophores; fluorescent proteins; quantum dots; etc.

[0176] The fusion may contain members of a FRET pair. Applicable FRET pairs (donor / acceptor) may include, but are not limited to, EDANS / fluorescein, IAEDANS / fluorescein, fluorescein / tetramethylrhodamine, fluorescein / Cy 5, IEDANS / DABCYL, fluorescein / QSY-7, fluorescein / LC Red 640, fluorescein / Cy 5.5, and fluorescein / LC Red 705.

[0177] Fluorophore / quantum dot donor / acceptor pairs can be used as fusion compounds. Suitable fluorophores (“fluorescent tags”) can include any molecule detectable by its inherent fluorescent properties (which may include fluorescence detectable upon excitation). Suitable fluorescent tags can include, but are not limited to, fluorescein, rhodamine, tetramethylrhodamine, eosin, phycoerythrone, coumarin, methyl-coumarin, pyrene, malachite green, stilbene, fluorescein, and cascade blue. TM Texas Red, IAEDANS, EDANS, BODIPY FL, LC Red 640, Cy 5, Cy 5.5, LC Red 705, and Oregon Green.

[0178] The fusion may contain an enzyme. Suitable enzymes may include, but are not limited to, horseradish peroxidase, luciferase, β-galactosidase, etc.

[0179] The fusion may contain a fluorescent protein. Suitable fluorescent proteins may include, but are not limited to, green fluorescent protein (GFP) (e.g., GFP from Aequoria victoria, fluorescent protein from Anguilla japonica, or mutants or derivatives thereof), red fluorescent protein, yellow fluorescent protein, and any of the various fluorescent and colored proteins.

[0180] The fusion may contain nanoparticles. Suitable nanoparticles may include fluorescent or luminescent nanoparticles and magnetic nanoparticles. Any optical or magnetic properties or characteristics of the nanoparticles can be detected.

[0181] The fusion may contain quantum dots (QDs). The QDs are made water-soluble by applying a coating comprising various different materials. For example, amphiphilic polymers can be used to solubilize the QDs. Exemplary polymers used may include octylamine-modified low molecular weight polyacrylic acid, polyethylene glycol (PEG)-derived phospholipids, polyanhydrides, block copolymers, etc. The QDs may be conjugated to the peptide via any of a number of different functional groups or linkers that can be directly or indirectly attached to the coating. QDs with a wide variety of absorption and emission spectra are available, for example, from Quantum Dot Corp. (Hayward Calif.; now owned by Invitrogen) or Evident Technologies (Troy, NY). For example, QDs with peak emission wavelengths of approximately 525, 535, 545, 565, 585, 605, 655, 705, and 800 nm are available. Therefore, QDs may have a range of different colors in the visible portion of the spectrum and, in some cases, even beyond that range.

[0182] Suitable radioactive isotopes may include, but are not limited to, those that are not limited to. 14 C 3 H, 32 P, 33 P, 35 S and 125 I.

[0183] As used herein, "genetically modified cell" generally refers to a genetically modified cell. Some non-limiting examples of gene modification may include: insertion, deletion, inversion, translocation, gene fusion, or alteration of one or more nucleotides. Genetically modified cells may contain target nucleic acids with introduced double-strand breaks (e.g., DNA breaks). Genetically modified cells may contain exogenously introduced nucleic acids (e.g., vectors). Genetically modified cells may contain exogenously introduced polypeptides and / or nucleic acids of this disclosure. Genetically modified cells may contain donor polynucleotides. Genetically modified cells may contain exogenous nucleic acids integrated into the genome of the gene-modified cell. Genetically modified cells may contain DNA deletions. Genetically modified cells may also refer to cells with modified mitochondrial or chloroplast DNA.

[0184] As used herein, "genome engineering" can refer to methods of modifying target nucleic acids. Genome engineering can refer to the integration of non-natural nucleic acids into natural nucleic acids. Genome engineering can involve site-directed peptides and nucleic acids targeting target nucleic acids without integration or deletion of the target nucleic acid. Genome engineering can refer to the cleavage and rejoining of target nucleic acids without integration of exogenous sequences or deletions in the target nucleic acid. The natural nucleic acid may contain genes. The non-natural nucleic acid may contain donor polynucleotides. In the methods disclosed herein, site-directed peptides (e.g., Cas9) can introduce double-strand breaks into nucleic acids (e.g., genomic DNA). These double-strand breaks can mimic endogenous DNA repair pathways in cells (e.g., homologous recombination (HR) and / or non-homologous end joining (NHEJ) or A-NHEJ (alternative non-homologous end joining)). Mutations, deletions, alterations, and integrations of foreign, exogenous, and / or alternative nucleic acids can be introduced into the sites of double-strand DNA breaks.

[0185] As used herein, the term "isolation" can refer to the separation by human hands from nucleic acids or peptides that exist in their natural environment and are therefore not natural products. Isolation can mean essentially pure. Isolated nucleic acids or peptides may exist in purified form and / or may exist in non-natural environments, such as transgenic cells.

[0186] As used herein, "non-natural" can refer to nucleic acid or polypeptide sequences not found in native nucleic acids or proteins. Non-natural can refer to affinity markers. Non-natural can refer to fusions. Non-natural can refer to naturally occurring nucleic acid or polypeptide sequences containing mutations, insertions, and / or deletions. Non-natural sequences may exhibit and / or encode activities that nucleic acid and / or polypeptide sequences fused to the non-natural sequence may also exhibit (e.g., enzyme activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.). Non-natural nucleic acid or polypeptide sequences can be genetically engineered to link to naturally occurring nucleic acid or polypeptide sequences (or their variants) to generate chimeric nucleic acid and / or polypeptide sequences encoding chimeric nucleic acids and / or polypeptides. Non-natural sequences can refer to 3' hybridization overhang sequences.

[0187] As used herein, “nucleic acid” generally refers to a polynucleotide sequence or a fragment thereof. Nucleic acids may contain nucleotides. Nucleic acids may be exogenous or endogenous. Nucleic acids may exist in a cell-free environment. Nucleic acids may be genes or fragments thereof. Nucleic acids may be DNA. Nucleic acids may be RNA. Nucleic acids may contain one or more analogs (e.g., modified backbone, sugar, or nucleobase). Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acids, heteronucleic acids, morpholino, locked nucleic acids, glycol nucleic acids, threonine, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugars), nucleotide-containing thiols, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, Q-nucleotides, and G-nucleotides.

[0188] As used herein, "nucleic acid sample" generally refers to a sample derived from a biological entity. A nucleic acid sample may contain nucleic acids. The nucleic acids from a nucleic acid sample may be purified and / or enriched. The nucleic acid sample may exhibit whole-body properties. Nucleic acid samples can originate from a variety of sources. Nucleic acid samples may originate from one or more individuals. One or more nucleic acid samples may originate from the same individual. A non-limiting example is if one sample is from an individual's blood and a second sample is from the individual's tumor biopsy. Examples of nucleic acid samples may include, but are not limited to, blood, serum, plasma, nasal swabs or nasopharyngeal washes, saliva, urine, gastric juice, cerebrospinal fluid, tears, feces, mucus, sweat, cerumen, oil, glandular secretions, cerebrospinal fluid, tissue, semen, vaginal secretions, interstitial fluid, including interstitial fluid derived from tumor tissue, intraocular fluid, cerebrospinal fluid, throat swabs, cheek swabs, breath, hair, nails, skin, biopsy, amniotic fluid, amniotic fluid, umbilical cord blood, emphatic fluids, cavity fluid, sputum, pus, micropiota, meconium, breast milk, oral samples, nasopharyngeal washes, other secretions, or any combination thereof. Nucleic acid samples may be derived from tissues. Examples of tissue samples may include, but are not limited to, connective tissue, muscle tissue, nerve tissue, epithelial tissue, cartilage, cancer or tumor samples, bone marrow, or bone. Nucleic acid samples may be provided by humans or animals. Nucleic acid samples can be provided by mammals, vertebrates such as rodents, apes, humans, farm animals, racing animals, or pets. Nucleic acid samples can be collected from living or dead objects. Nucleic acid samples can be collected fresh from the object or may have undergone some form of pretreatment, storage, or transportation. Nucleic acid samples may contain target nucleic acids. Nucleic acid samples may be derived from cell lysates. Cell lysates may be derived from cells.

[0189] As used herein, "nucleic acid targeting nucleic acid" can refer to a nucleic acid capable of hybridizing with another nucleic acid. A nucleic acid targeting nucleic acid can be RNA. A nucleic acid targeting nucleic acid can be DNA. A nucleic acid targeting nucleic acid can be programmed to bind to a nucleic acid sequence specifically at a site. The target nucleic acid may contain nucleotides. A portion of the target nucleic acid may be complementary to a portion of the nucleic acid targeting nucleic acid. A nucleic acid targeting nucleic acid may contain a polynucleotide chain and may be referred to as a "single-guided nucleic acid" (i.e., "nucleic acid targeting single-guided nucleic acid"). A nucleic acid targeting nucleic acid may contain two polynucleotide chains and may be referred to as a "double-guided nucleic acid" (i.e., "nucleic acid targeting double-guided nucleic acid"). Unless otherwise specified, the term "nucleic acid targeting nucleic acid" may be inclusive, referring to both single-guided and double-guided nucleic acids.

[0190] Nucleic acids that target nucleic acids can contain fragments that can be called "nucleic acid-targeted fragments" or "nucleic acid-targeted sequences." Nucleic acids that target nucleic acids can also contain fragments that can be called "protein-binding fragments" or "protein-binding sequences."

[0191] Nucleic acids targeting nucleic acids may contain one or more modifications (e.g., base modifications, backbone modifications) to provide new or enhanced characteristics (e.g., improved stability). Nucleic acids targeting nucleic acids may contain nucleic acid affinity tags. Nucleosides may be base-sugar combinations. The base moiety of a nucleoside may be a heterocyclic base. The two most common types of such heterocyclic bases are purines and pyrimidines. Nucleotides may further include a phosphate group covalently linked to the sugar moiety of the nucleoside. For those nucleosides including pentofuranosyl sugar, the phosphate group may be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming nucleic acids targeting nucleic acids, phosphate groups may covalently link adjacent nucleosides to form linear polymers. The respective ends of such linear polymers may be further linked to form cyclic compounds; however, linear compounds are generally preferred. Furthermore, linear compounds may have internal nucleotide base complementarity and thus fold in a manner that produces fully or partially double-stranded compounds. In nucleic acids targeting nucleic acids, phosphate groups are often mentioned as forming the internucleotide backbone of the nucleic acid targeting nucleic acid. The bonds or backbone of the nucleic acid targeting nucleic acid can be 3' to 5' phosphodiester bonds.

[0192] Nucleic acids targeting nucleic acids may contain modified backbones and / or modified internucleotide bonds. Modified backbones may include those with phosphorus atoms in the backbone and those without phosphorus atoms in the backbone.

[0193] Suitable nucleic acid-targeted modified nucleic acid backbones containing phosphorus atoms may include, for example, thiophosphates, chiral thiophosphates, dithiophosphates, phosphate triesters, aminoalkyl phosphate triesters, methyl and other alkylphosphonates, such as 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphonites, aminophosphates including 3'-aminoaminophosphates and aminoalkylaminophosphates, diaminophosphates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates and boranophosphates with normal 3'-5' bonds, 2'-5' linked analogs and those with antipolarity, wherein one or more nucleotide bonds are 3' to 3', 5' to 5', or 2' to 2' bonds. Suitable antipolar nucleic acids targeting nucleic acids can be contained in a single 3' to 3' bond between the 3' terminal nucleotides (i.e., a single antinucleoside residue, where a nucleobase is missing or replaced by a hydroxyl group). They can also include various salts (e.g., potassium chloride or sodium chloride), mixed salts, and free acid forms.

[0194] Nucleic acids targeting nucleic acids may contain one or more thiophosphate esters and / or heteroatom nucleoside bonds, particularly -CH2-NH-O-CH2-, -CH2-N(CH3)-O-CH2- (i.e., methylene (methylimino) or MMI backbone), -CH2-ON(CH3)-CH2-, -CH2-N(CH3)-N(CH3)-CH2- and -ON(CH3)-CH2-CH2- (where the natural phosphodiester nucleotide bond is represented as -OP(=O)(OH)-O-CH2-).

[0195] Nucleic acids targeting nucleic acids may contain a morpholine backbone structure. For example, the nucleic acid may contain a 6-membered morpholine ring replacing the ribose ring. In some of these embodiments, diaminophosphate or other non-phosphodiester nucleoside bonds may replace the phosphodiester bonds.

[0196] Nucleic acids targeting nucleic acids may comprise a polynucleotide backbone formed by short-chain alkyl or cycloalkyl nucleoside bonds, mixed heteroatom and alkyl or cycloalkyl nucleoside bonds, or one or more short-chain heteroatom or heterocyclic nucleoside bonds. These may include those with morpholino bonds (partially formed from the sugar moiety of the nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formyl and thioformyl backbones; methyleneformyl and thioformyl backbones; riboacetyl backbones; olefin-containing backbones; aminosulfonate backbones; methyleneimino and methylenehydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and other backbones having mixed N, O, S, and CH2 components.

[0197] Nucleic acids targeting nucleic acids can include nucleic acid mimetics. The term "mimetic" is intended to include polynucleotides in which only the furanose ring or the furanose ring and the intermolecular bonds are replaced by non-furanose groups; substitution of only the furanose ring can also be called a sugar surrogate. The heterocyclic base moiety can be retained or modified to hybridize with a suitable target nucleic acid. One such nucleic acid can be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of the polynucleotide can be replaced by an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotide can be retained and directly or indirectly bonded to the aza-nitrogen atom of the amide moiety of the backbone. The backbone in a PNA compound can contain two or more linked aminoethylglycine units, which gives the PNA an amide-containing backbone. The heterocyclic base moiety can be directly or indirectly bonded to the aza-nitrogen atom of the amide moiety of the backbone.

[0198] Nucleic acids targeting nucleic acids can contain morpholino units (i.e., morpholinonucleotides) linked to heterocyclic bases attached to a morpholino ring. Linkers can connect the morpholino monomer units within morpholinonucleotides. Nonionic morpholino oligomers exhibit fewer undesirable interactions with cellular proteins. Morpholino polynucleotides can be nonionic mimics of nucleic acids targeting nucleic acids. Various compounds within the morpholino group can be linked using different linkers. Another class of polynucleotide mimics can be called cyclohexenyl nucleic acids (CeNA). The furanose ring typically present in nucleic acid molecules can be replaced by a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared and used in the synthesis of oligomers using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid chains can improve the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complement exhibiting stability similar to the natural complexes. Another modification may include locked nucleic acids (LNAs), in which a 2'-hydroxyl group is attached to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxomethylene bond, thus forming a bicyclic sugar moiety. This bond can be a methylene (-CH2-) group, a group bridging the 2' oxygen atom and the 4' carbon atom, where n is 1 or 2. LNAs and LNA analogs can exhibit extremely high duplex thermal stability (Tm = +3 to +10 °C) with complementary nucleic acids, stability against 3'-exonuclease degradation, and good solubility properties.

[0199] Nucleic acids targeting nucleic acids may contain one or more substituted sugar moieties. Suitable polynucleotides may contain sugar substituents selected from the following: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-ynyl; or O-alkyl-O-alkyl, wherein the alkyl, alkenyl, and ynyl groups may be substituted or unsubstituted C1 to C2 groups. 10 Alkyl or C2 to C 10 Alkenyl and alkynyl groups. Particularly suitable are O((CH2)nO)mCH3 and O(CH2). n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON((CH2) n CH3)2, where n and m are 1 to about 10. Sugar substituents can be selected from: C1 to C2. 10Low-carbon alkyl, substituted low-carbon alkyl, alkenyl, alkynyl, aryl, aralkyl, O-aryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocyclic alkyl, heterocyclic aryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving group, reporter group, intercalating agent, group used to improve the pharmacokinetic properties of nucleic acids targeting nucleic acids or group used to improve the pharmacodynamic properties of nucleic acids targeting nucleic acids, and other substituents with similar properties. Suitable modifications may include 2'-methoxyethoxy (2'-O-CH2CH2OCH3, also known as 2'-O-(2-methoxyethyl) or 2'-MOE, i.e., alkoxyalkoxy). Another suitable modification may include 2'-dimethylaminoethoxy (i.e., the O(CH2)2ON(CH3)2 group, also known as 2'-DMAOE) and 2'-dimethylaminoethoxyethoxy (also known as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), i.e. 2'-O-CH2-O-CH2-N(CH3)2.

[0200] Other suitable sugar substituents may include methoxy (-O-CH3), aminopropoxy (-OCH2CH2CH2NH2), allyl (-CH2-CH=CH2), -O-allyl (-O--CH2—CH=CH2), and fluorine (F). The 2'-sugar substituent can be at the arabinose (top) or ribose (bottom) position. A suitable 2'-arabinose modification is 2'-F. Similar modifications can also be made at other positions on the oligomer, particularly at the 3' terminal nucleotide or the 3' position of the sugar and the 5' position of the 5' terminal nucleotide in a 2'-5' linked nucleotide. The oligomer may also have sugar mimics, such as a cyclobutyl moiety, to replace the furanopentose.

[0201] Nucleic acids that target nucleic acids may also include nucleobase modifications or substitutions (often simply referred to as "bases"). The "unmodified" or "natural" nucleobases used in this article may include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases may include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3)uracil and cytosine, and other alkynyl derivatives of pyrimidine bases, 6-azouracil. Cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halogenated, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxy and other 8-substituted adenine and guanine, 5-halogenated, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracil and cytosine, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deadenine and 7-deadenine and 3-deadenine and 3-deadenine. Modified nucleobases may include tricyclic pyrimidines, such as phenoxazincytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenthiazincytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps, such as substituted phenoxazincytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazolecytidine (2H-pyrimido(4,5-b)indol-2-one), and pyridoindolcytidine (H-pyrido(3',2':4,5)pyrrolo(2,3-d)pyrimido-2-one).

[0202] The heterocyclic base moiety may include those in which the purine or pyrimidine base is replaced by another heterocycle, such as 7-deazo-adenine, 7-deazoguanosine, 2-aminopyridine, and 2-pyridone. Nucleotides can be used to improve the binding affinity of polynucleotide compounds. These may include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-methylcytosine substitution can improve the stability of the nucleic acid duplex by 0.6–1.2 °C and can be a suitable base substitution (e.g., when combined with 2'-O-methoxyethyl sugar modification).

[0203] Modification of nucleic acids targeting nucleic acids may include chemically linking one or more moieties or conjugates that enhance the activity, cellular distribution, or cellular uptake of the target nucleic acid to the target nucleic acid. These moieties or conjugates may include conjugate groups covalently bonded to functional groups, such as primary or secondary hydroxyl groups. Conjugate groups may include, but are not limited to, intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, groups that enhance the pharmacodynamic properties of oligomers, and groups that enhance the pharmacokinetic properties of oligomers. Conjugate groups may include, but are not limited to, cholesterol, lipids, phospholipids, biotin, phenazine, folic acid, phenanthridine, anthraquinones, acridine, fluorescein, rhodamine, coumarin, and dyes. Groups that enhance pharmacodynamic properties include groups that improve uptake, enhance resistance to degradation, and / or enhance sequence-specific hybridization with the target nucleic acid. Groups that enhance pharmacokinetic properties include groups that improve the uptake, distribution, metabolism, or secretion of nucleic acids. The conjugated portion may include, but is not limited to, lipid portions such as cholesterol portions, bile acids, thioethers (e.g., hexyl-S-triphenylmethanethiol), mercaptocholesterol, aliphatic chains (e.g., dodecyl glycol or undecyl residues), phospholipids (e.g., di-hexadecyl-rac-glycerol or 1,2-di-O-hexadecyl-rac-glycerol-3-H-phosphonate triethylammonium), polyamines or polyethylene glycol chains, or adamantaneacetic acid, palmityl portions, or octadecylamine or hexylamino-carbonyloxy cholesterol portions.

[0204] One modification may include a “protein transduction domain” or PTD (i.e., cell-penetrating peptide (CPP)). A PTD can refer to a polypeptide, polynucleotide, carbohydrate, or an organic or inorganic compound that facilitates passage through a lipid bilayer, micelles, cell membrane, organelle membrane, or vesicle membrane. A PTD can be attached to another molecule (ranging from small polar molecules to large polymers and / or nanoparticles) and facilitate that molecule's passage through membranes, such as from the extracellular space to the intracellular space, or from the cytosol to the organelle. A PTD can be covalently attached to the amino terminus of a polypeptide. A PTD can be covalently attached to the carboxyl terminus of a polypeptide. A PTD can be covalently attached to a nucleic acid. Exemplary PTDs may include, but are not limited to, a minimal peptide protein transduction domain; a polyarginine sequence (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginine residues) containing a number of arginine residues sufficient for direct introduction into the cell; a VP22 domain; a Drosophila tentacles transduction domain; a truncated human calcitonin peptide; polylysine and a transport protein; and an arginine homopolymer of 3 to 50 arginine residues. The PTD may be an activatable CPP (ACPP). The ACPP may comprise a polycationic CPP (e.g., Arg9 or "R9") linked to a matched polyanion (e.g., Glu9 or "E9") via a cleavable linker, which reduces the net charge to near zero and thereby inhibits adhesion and uptake into the cell. Upon linker cleavage, the polyanion can be released, locally exposing the polyarginine and its inherent adhesiveness, thereby "activating" the ACPP to cross the membrane.

[0205] The term "nucleotide" generally refers to a base-sugar-phosphate combination. Nucleotides may include synthetic nucleotides. Nucleotides may include synthetic nucleotide analogs. Nucleotides can be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleosides triphosphate, adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleosides triphosphate, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-denitro-dGTP, and 7-denitro-dATP, and nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide may refer to dideoxynucleotide triphosphates (ddNTPs) and their derivatives. Exemplary examples of dideoxynucleotide triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides can be unlabeled or detectably labeled using known techniques. They can also be labeled with quantum dots. Detectable labels can include, for example, radioactive isotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labeling of nucleotides can include, but is not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'-dimethylaminophenylazo)benzoic acid (DABCYL), cascade blue, Oregon green, Texas red, anthocyanin, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides may include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, all available from Perkin Elmer, Foster City, Calif.; and FluoroLink deoxynucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink FluorX-dCTP, FluoroLink Cy3-dUTP, and FluoroLink, all available from Amersham, Arlington Heights, Ill. Cy5-dUTP; luciferin-15-dATP, luciferin-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, luciferin-12-ddUTP, luciferin-12-UTP, and luciferin-15-2'-dATP, available from Boehringer Mannheim, Indianapolis, Inc.; and chromosome marker nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, and BODIPY-TMR-14-, available from Molecular Probes, Eugene, Oreg. UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides can also be chemically modified or spiked. Chemically modified mononucleotides can be biotin-dNTPs. Some non-limiting examples of biotinylated dNTPs may include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0206] As used herein, "P-domain" can refer to a region within a nucleic acid targeting a nucleic acid. This P-domain can interact with pre-intercalation adjacent motifs (PAMs), site-directed peptides, and / or nucleic acids targeting nucleic acids. The P-domain can interact directly or indirectly with pre-intercalation adjacent motifs (PAMs), site-directed peptides, and / or nucleic acids targeting nucleic acids. The terms "PAM-interacting region," "anti-repetitive fragment adjacent region," and "P-domain" are used interchangeably in this document.

[0207] As used herein, “purified” can refer to at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the molecules (e.g., site-directed peptides, nucleic acids targeting nucleic acids) constituting the composition. For example, a sample containing 10% site-directed peptides but containing 60% site-directed peptides after a purification step can be described as purified. A purified sample can refer to an enriched sample or a sample that has been subjected to methods to remove particles other than the particles of interest.

[0208] The term "reactivating agent" as used in this article generally refers to any reagent that can convert a non-enzymatic peptide into an enzymatic peptide. Imidazole can be a reactivating agent. Ligand analogs can be reactivating agents.

[0209] As used herein, "recombinant" can refer to a sequence originating outside a specific host (e.g., a cell), or, if from the same source, modified from its original form. Recombinant nucleic acids in cells can include nucleic acids that are endogenous to a specific cell but have been modified, for example, using site-directed mutagenesis. The term can include non-naturally generated multiple copies of a naturally occurring DNA sequence. Therefore, the term can refer to nucleic acids originating outside the cell or from a different source than the cell, or homologous to the cell but whose intracellular location or form is not the original location or form of the nucleic acid. Similarly, when used with polypeptide or amino acid sequences, exogenous polypeptide or amino acid sequences can be from a specific extracellular source or, if from the same source, modified from their original form.

[0210] The term "site-directed polypeptide" as used in this article generally refers to nucleases, site-directed nucleases, ribonucleases, conditionally inactive ribonucleases, Argonauts, and nucleic acid-binding proteins. Site-directed polypeptides or proteins may include nucleases such as homing endonucleases, such as PI-TliII, H-DreI, I-DmoI and I-CreI, I-SceI, LAGLIDADG family nucleases, broad-spectrum nucleases, GIY-YIG family nucleases, His-Cys box family nucleases, Vsr-like nucleases, ribonucleases, ribonucleases, endonucleases, and exonucleases. Site-directed polypeptides may refer to Cas gene members of type I, II, III, and / or U-type CRISPR / Cas systems. Site-directed polypeptides may refer to members of the repeat-associated doubt protein (RAMP) superfamily (e.g., Cas5, Cas6 subfamilies). Site-directed polypeptides may refer to Argonaute proteins.

[0211] A site-directed peptide can be a type of protein. A site-directed peptide can refer to a nuclease. A site-directed peptide can refer to a ribonuclease. A site-directed peptide can be a peptide sequence or homolog of any modified (e.g., shortened, mutated, elongated) site-directed peptide. A site-directed peptide can be codon-optimized. A site-directed peptide can be a codon-optimized homolog of a site-directed peptide. A site-directed peptide can be enzymatically inactive, partially active, constitutively active, fully active, inducible, and / or more active (e.g., more active than the wild-type homolog of the protein or peptide). A site-directed peptide can be Cas9. A site-directed peptide can be Csy4. A site-directed peptide can be Cas5 or a member of the Cas5 family. A site-directed peptide can be Cas6 or a member of the Cas6 family.

[0212] In some cases, the site-directed polypeptide (e.g., variants, mutations, inactive, and / or conditionally inactive site-directed polypeptides) can be the target nucleic acid. The site-directed polypeptide (e.g., variants, mutations, inactive, and / or conditionally inactive ribonucleases) can target RNA. RNA-targeting ribonucleases may include members of other CRISPR subfamily families, such as Cas6 and Cas5.

[0213] As used herein, the term "specific" can refer to the interaction between two molecules, in which one molecule specifically binds to the second molecule by, for example, chemical or physical means. Exemplary specific binding interactions can refer to antigen-antibody binding, avidin-biotin binding, carbohydrate and lectin interactions, complementary nucleic acid sequences (e.g., hybridization), complementary peptide sequences, including those formed by recombination, effector and receptor molecules, enzyme cofactors and enzymes, enzyme inhibitors and enzymes, etc. "Non-specific" can refer to the interaction between two molecules that is not specific.

[0214] As used herein, "solid support" generally refers to any insoluble or partially soluble material. Solid supports can refer to test strips, well trays, etc. These solid supports can contain a variety of substances (e.g., glass, polystyrene, polyvinyl chloride, polypropylene, polyethylene, polycarbonate, dextran, nylon, amylose, natural and modified cellulose, polyacrylamide, agarose, and magnetite) and can be provided in various forms, including agarose beads, polystyrene beads, latex beads, magnetic beads, colloidal metal particles, glass and / or silicon wafers and surfaces, nitrocellulose strips, nylon membranes, sheets, pores in reaction trays (e.g., well plates), plastic tubes, etc. Solid supports can be solid, semi-solid, beads, or surfaces. The support can be flowable in solution or can be immobilized. Solid supports can be used to capture peptides. Solid supports may contain capturing agents.

[0215] As used herein, “target nucleic acid” generally refers to the nucleic acid used in the methods of this disclosure. Target nucleic acid can refer to a chromosomal sequence or an extrachromosomal sequence (e.g., episome sequence, microcircular sequence, mitochondrial sequence, chloroplast sequence, etc.). Target nucleic acid can be DNA. Target nucleic acid can be RNA. Target nucleic acid may be used interchangeably herein with “polynucleotide,” “nucleotide sequence,” and / or “target polynucleotide.” Target nucleic acid can be a nucleic acid sequence that is independent of any other sequence in the nucleic acid sample by substituting a single nucleotide. Target nucleic acid can be a nucleic acid sequence that is independent of any other sequence in the nucleic acid sample by substituting 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In some embodiments, this substitution cannot occur within 5, 10, 15, 20, 25, 30, or 35 nucleotides at the 5' end of the target nucleic acid. In some embodiments, this substitution cannot occur within 5, 10, 15, 20, 25, 30, or 35 nucleotides at the 3' end of the target nucleic acid.

[0216] As used herein, “tracrRNA” generally refers to a nucleic acid having at least approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from Streptococcus pyogenes (SEQ ID 433), SEQ IDs 431-562). tracrRNA may refer to a nucleic acid having up to approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from Streptococcus pyogenes). tracrRNA may refer to a modified form of tracrRNA that may contain nucleotide variations, such as deletions, insertions or substitutions, mutations, mutations, or chimeras. A tracrRNA can refer to a nucleic acid sequence that is at least approximately 60% identical to a wild-type exemplary tracrRNA (e.g., tracrRNA from Streptococcus pyogenes) sequence in at least six adjacent nucleotides. For example, a tracrRNA sequence can be at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 95%, at least approximately 98%, at least approximately 99%, or 100% identical to a wild-type exemplary tracrRNA (e.g., tracrRNA from Streptococcus pyogenes) sequence in at least six adjacent nucleotides. A tracrRNA can refer to an intermediate tracrRNA. A tracrRNA can refer to a minimal tracrRNA sequence.

[0217] CRISPR system

[0218] CRISPR (clustered, regularly spaced short palindromic repeats) are genomic loci found in the genomes of many prokaryotes, such as bacteria and archaea. CRISPR loci can provide resistance to foreign invaders (such as viruses and bacteriophages) in prokaryotes. Thus, the CRISPR system can be considered as a type of immune system to help prokaryotes defend against foreign invaders. The function of CRISPR loci involves three stages: integration of new sequences into the locus, biogenesis of CRISPR RNA (crRNA), and silencing of the foreign invader's nucleic acid. There are four types of CRISPR systems (e.g., type I, type II, type III, and type U).

[0219] CRISPR loci may comprise numerous short repetitive sequences called “repetitions.” These repetitions can form hairpin structures and / or can be unstructured single-stranded sequences. Repetitions can exist in clusters. The sequence of repetitions often varies by species. Repetitions can be regularly spaced by unique insertion sequences (called “spacer regions”) to create a repetition-spacer-repetition locus structure. Spacer regions may be identical to or highly homologous to known foreign intruder sequences. Spacer-repetition units may encode crprRNA (crRNA). crRNA can refer to the mature form of this spacer-repetition unit. crRNA may contain a “seed” sequence that can participate in targeting target nucleic acids (e.g., potentially as a surveillance mechanism against foreign nucleic acids). The seed sequence may be located at the 5’ or 3’ end of the crRNA.

[0220] CRISPR loci may contain polynucleotide sequences encoding CRISPR-associated genes (Cas). Cas genes may be involved in the biogenesis and / or interference phases of crRNA function. Cas genes may exhibit extreme sequence (e.g., primary sequence) divergence between species and homologs. For example, Cas1 homologs may contain less than 10% primary sequence identity among homologs. Some Cas genes may contain homologous secondary and / or tertiary structures. For example, despite extreme sequence divergence, many members of the Cas6 family of CRISPR proteins contain N-terminal ferroredoxin-like folds. Cas genes can be named according to their organism of origin. For example, the Cas gene in *Staphylococcus epidermidis* may be designated Csm, the Cas gene in *Streptococcus thermophilus* may be designated Csn, and the Cas gene in *Pyrococcus furiosus* may be designated Cmr.

[0221] Integration

[0222] The integration phase of a CRISPR system refers to the ability of a CRISPR locus to integrate a new spacer region into the crRNA array upon infection by a foreign intruder. Acquisition of the foreign intruder spacer region helps provide immunity against subsequent invasion by the same foreign intruder. Integration can occur at the leader of a CRISPR locus. Cas proteins (such as Cas1 and Cas2) can participate in the integration of the new spacer region sequence. Integration can occur similarly in some types of CRISPR systems (e.g., types I-III).

[0223] Biogenesis

[0224] Mature crRNA can be processed from long polycistronic CRISPR locus transcripts (i.e., pre-crRNA arrays). Pre-crRNA arrays can contain multiple crRNAs. The repetitive segments in the pre-crRNA array can be recognized by Cas genes. Cas genes can bind to and cleave the repetitive segments. This action releases multiple crRNAs. Further events can occur on the crRNA to produce the mature crRNA form, such as pruning (e.g., with the aid of exonucleases). crRNA may contain all, some, or no CRISPR repetitive sequences.

[0225] Interference can refer to the functionally responsible phase of a CRISPR system that combats infection by foreign invaders. CRISPR interference can follow a mechanism similar to RNA interference (RNAi, where a short interfering RNA (siRNA) targets (e.g., hybridizes) a target RNA), which can cause degradation and / or instability of the target RNA. The CRISPR system can interfere with target nucleic acids by coupling crRNA and the Cas gene, thereby forming a CRISPR ribonucleoprotein (crRNP). The crRNA of the crRNP can direct the crRNP toward the foreign invader nucleic acid (e.g., by recognizing the foreign invader nucleic acid via hybridization). The hybridized target foreign invader nucleic acid-crRNA unit can undergo Cas protein cleavage. Target nucleic acid interference may require a spacer adjacent motif (PAM) in the target nucleic acid.

[0226] There are four types of CRISPR systems: type I, type II, type III, and type U. More than one type of CRISPR system can be found in organisms. CRISPR systems can be complementary to each other and / or provide trans-functional units to facilitate CRISPR locus processing.

[0227] CRISPR system type I

[0228] In type I CRISPR systems, crRNA biogenesis involves cleavage by ribonucleases of repetitive segments in the pre-crRNA array, yielding multiple crRNAs. Type I crRNAs may not undergo crRNA pruning. crRNAs can be processed from the pre-crRNA array via a multi-protein cascade (derived from CRISPR-associated complexes used for antiviral defense). The cascade may contain protein subunits (e.g., CasA-CasE). Some subunits may be members of the repeat-associated doubtful protein (RAMP) superfamily (e.g., the Cas5 and Cas6 families). The cascade-crRNA complex (i.e., the interference complex) recognizes the target nucleic acid through hybridization of crRNA with the target nucleic acid. The cascade interference complex recruits Cas3 helicases / nucleases that can transact to promote the cleavage of the target nucleic acid. Cas3 nucleases can cleave the target nucleic acid (e.g., using their HD nuclease domain). Target nucleic acids in type I CRISPR systems may contain PAM. Target nucleic acids in type I CRISPR systems can be DNA.

[0229] The Type I system can be further subdivided by their source species. The Type I system may include: Type IA (Aeropyrum pernix or CASS5); Type IB (Thermotoga neapolitana-Haloarculamarismortui or CASS7); Type IC (Desulfovibrio vulgaris or CASS1); Type ID; Type IE (Escherichia coli or CASS2); and the Type IF (Yersinia pestis or CASS3) subfamily.

[0230] CRISPR system type II

[0231] In a type II CRISPR system, crRNA biogenesis may involve trans-activating CRISPR RNA (tracrRNA). tracrRNA can be modified by endogenous RNase III. The tracrRNA complex hybridizes to crRNA repeat fragments in a pre-crRNA array. Endogenous RNase III can be recruited to cleave the pre-crRNA. The cleaved crRNA undergoes exonuclease pruning to produce a mature crRNA form (e.g., 5' pruning). tracrRNA maintains hybridization with crRNA. tracrRNA and crRNA can bind to a site-directed peptide (e.g., Cas9). The crRNA in the crRNA-tracrRNA-Cas9 complex directs the complex to a target nucleic acid to which the crRNA can hybridize. Hybridization of the crRNA with the target nucleic acid activates Cas9 to cleave the target nucleic acid. The target nucleic acid in a type II CRISPR system may contain a PAM. In some embodiments, the PAM is essential for facilitating the binding of the site-directed peptide (e.g., Cas9) to the target nucleic acid. Type II systems can be further subdivided into II-A (Nmeni or CASS4) and II-B (Nmeni or CASS4).

[0232] CRISPR system type III

[0233] In type III CRISPR systems, crRNA biogenesis may involve a ribonuclease cleavage step of the repetitive fragment in the pre-crRNA array, which can produce multiple crRNAs. The repetitive fragment in a type III CRISPR system can be an unstructured single-stranded region. The repetitive fragment can be recognized and cleaved by members of the RAMP superfamily of ribonucleases (e.g., Cas6). In type III (e.g., type III-B) systems, crRNA can undergo crRNA trimming (e.g., 3' trimming). Type III systems may contain polymerase-like proteins (e.g., Cas10). Cas10 may contain domains homologous to the palm domain.

[0234] The type III system can process pre-crRNA using a complex containing multiple RAMP superfamily member proteins and one or more CRISPR polymerase-like proteins. The type III system can be divided into III-A and III-B. The interference complex (Csm complex) of the type III-A system targets plasmid nucleic acids. The HD nuclease domain of the polymerase-like protein in this complex can be used to cleave plasmid nucleic acids. The interference complex (Cmr complex) of the type III-B system targets RNA.

[0235] U-shaped CRISPR system

[0236] U-type CRISPR systems may not contain the characteristic genes of any of the type I-III CRISPR systems (e.g., Cas3, Cas9, Cas6, Cas1, Cas2). Examples of U-type CRISPR Cas genes may include, but are not limited to, Csf1, Csf2, Csf3, and Csf4. U-type Cas genes can be very different homologs of type I-III Cas genes. For example, Csf3 may be highly deviated from, but functionally similar to, members of the Cas5 family. U-type systems can interact trans-complementarily with type I-III systems. In some cases, U-type systems may be independent of the processing CRISPR array. U-type systems may represent another foreign intrusion defense system.

[0237] RAMP Superfamily

[0238] Repeat-associated suspected proteins (RAMP proteins) are characterized by protein folding with a βαββββ[beta-alpha-beta-beta-alpha-beta] motif containing a β-chain (β) and an α-helix (α). RAMP proteins may contain RNA recognition motifs (RRMs) (which may contain ferrugin or ferrugin-like folds). RAMP proteins may contain an N-terminal RRM. The C-terminal domain of RAMP proteins is variable but may also contain an RRM. RAMP family members can recognize structured and / or unstructured nucleic acids. RAMP family members can recognize single-stranded and / or double-stranded nucleic acids. RAMP proteins are involved in the biogenesis and / or interference phases of type I and type III CRISPR systems. RAMP superfamily members may include members of the Cas7, Cas6, and Cas5 families. RAMP superfamily members may be ribonucleases.

[0239] The RRM domains in the RAMP superfamily can be extremely divergent. The RRM domains may contain at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or 100% sequence or structural homology with wild-type exemplary RRM domains (e.g., RRM domains from Cas7). The RRM domain may contain up to approximately 5%, up to approximately 10%, up to approximately 15%, up to approximately 20%, up to approximately 25%, up to approximately 30%, up to approximately 35%, up to approximately 40%, up to approximately 45%, up to approximately 50%, up to approximately 55%, up to approximately 60%, up to approximately 65%, up to approximately 70%, up to approximately 75%, up to approximately 80%, up to approximately 85%, up to approximately 90%, up to approximately 95%, or 100% sequence or structural homology with wild-type exemplary RRM domains (such as RRM domains from Cas7).

[0240] Cas7 family

[0241] Cas7 family members can be a subclass of RAMP family proteins. Cas7 family proteins can be classified in the type I CRISPR system. Cas7 family members may not contain the glycine-rich ring familiar to some RAMP family members. Cas7 family members may contain an RRM domain. Cas7 family members may include, but are not limited to, Cas7 (COG1857), Cas7 (COG3649), Cas7 (CT1975), Csy3, Csm3, Cmr6, Csm5, Cmr4, Cmr1, Csf2, and Csc2.

[0242] Cas6 family

[0243] The Cas6 family can be a subfamily of RAMPs. Cas6 family members may contain two RNA recognition motif (RRM)-like domains. Cas6 family members (e.g., Cas6f) may contain an N-terminal RRM domain and a distinct C-terminal domain that exhibits weak sequence similarity or structural homology to the RRM domain. Cas6 family members may contain catalytic histidine residues involved in ribonuclease activity. Equivalent motifs can be found in the Cas5 and Cas7 RAMP families. Cas6 family members may include, but are not limited to, Cas6, Cas6e, and Cas6f (e.g., Csy4).

[0244] Cas5 family

[0245] The Cas5 family can be a subfamily of RAMP. The Cas5 family can be divided into two subgroups: one subgroup can contain two RRM domains, and the other subgroup can contain one RRM domain. The Cas5 family may include, but is not limited to, Csm4, Csx10, Cmr3, Cas5, Cas5(BH0337), Csy2, Csc1, and Csf3.

[0246] Cas gene

[0247] Exemplary CRISPR Cas genes may include Cas1, Cas2, Cas3' (Cas3-prime), Cas3" (Cas3-double-prime), Cas4, Cas5, Cas6, Cas6e (formerly known as CasE, Cse3), Cas6f (i.e., Csy4), Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Csy1, Csy2, Csy3, Cse1, and Cs. e2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4. Table 1 provides an exemplary classification of CRISPR Cas genes by CRISPR system type.

[0248] Since the discovery of Cas genes, the CRISPR-Cas gene nomenclature system has been largely rewritten. For the purposes of this application, the Cas gene names used herein are based on the nomenclature system briefly described in Makarova et al., Evolution and classification of the CRISPR-Cas systems. Nature Reviews Microbiology. June 2011; 9(6):467-477. Doi:10.1038 / nrmicro2577.

[0249] Table 1: Exemplary classification of CRISPR Cas genes by CRISPR type

[0250]

[0251] Site-directed peptides can be peptides that can bind to target nucleic acids. Site-directed peptides can be nucleases.

[0252] Site-directed peptides may contain a nucleic acid binding domain. This nucleic acid binding domain may contain a region that contacts nucleic acid. The nucleic acid binding domain may contain nucleic acid. The nucleic acid binding domain may contain protein-like material. The nucleic acid binding domain may contain both nucleic acid and protein-like material. The nucleic acid binding domain may contain RNA. A single nucleic acid binding domain may be present. Examples of nucleic acid binding domains may include, but are not limited to, helix-turn-helix domains, zinc finger domains, leucine zipper (bZIP) domains, winged helix domains, winged helix-turn-helix domains, helix-loop-helix domains, HMG-box domains, Wor3 domains, immunoglobulin domains, B3 domains, TALE domains, RNA-recognition motif domains, double-stranded RNA-binding motif domains, double-stranded nucleic acid binding domains, single-stranded nucleic acid binding domains, KH domains, PUF domains, RGG box domains, DEAD / DEAH box domains, PAZ domains, Piwi domains, and cold shock domains.

[0253] Nucleic acid binding domains can be domains of argonaute proteins. Argonaute proteins can be eukaryotic or prokaryotic. Argonaute proteins can bind RNA, DNA, or both. Argonaute proteins can cleave RNA or DNA, or both. In some cases, argonaute proteins bind to DNA and cleave the target DNA.

[0254] In some cases, two or more nucleic acid binding domains can be linked together. Linking multiple nucleic acid binding domains together can provide enhanced polynucleotide targeting specificity. Two or more nucleic acid binding domains can be linked via one or more adapters. The adapter can be a flexible adapter. The adapter may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40 or more amino acids in length. The adapter may contain at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or 100% glycine. The connector may contain up to 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% glycine. The connector may contain at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% serine. The connector may contain up to 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% serine.

[0255] Nucleic acid binding domains can bind to nucleic acid sequences. Nucleic acid binding domains can bind to nucleic acids through hybridization. Nucleic acid binding domains can be engineered (e.g., engineered to hybridize with sequences in the genome). Nucleic acid binding domains can be engineered using molecular cloning techniques (e.g., directed evolution, site-specific mutations, and rational mutagenesis).

[0256] Site-directed peptides may contain a nucleic acid cleavage domain. The nucleic acid cleavage domain can be derived from any nucleic acid cleavage protein. The nucleic acid cleavage domain may be derived from a nuclease. Suitable nucleic acid cleavage domains include those of endonucleases (e.g., AP endonuclease, RecBCD endonuclease, T7 endonuclease, T4 endonuclease IV, Bal 31 endonuclease, endonuclease I (endo I), micrococcal endonuclease, endonuclease II (endo VI, exo III)), exonucleases, restriction nucleases, ribonucleases, ribonucleases, and RNases (e.g., RNase I, II, or III). In some cases, the nucleic acid cleavage domain may be derived from FokI endonuclease. Site-directed peptides may contain multiple nucleic acid cleavage domains. Nucleic acid cleavage domains may be linked together. Two or more nucleic acid cleavage domains may be linked via a linker. In some embodiments, the linker may be a flexible linker. The linker may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40 or more amino acids in length. In some embodiments, the site-directed polypeptide may contain multiple nucleic acid cleavage domains.

[0257] Site-directed peptides (e.g., Cas9, argonaute) may contain two or more nuclease domains. Cas9 may contain an HNH or HNH-like nuclease domain and / or a RuvC or RuvC-like nuclease domain. The HNH or HNH-like domain may contain a McrA-like fold. The HNH or HNH-like domain may contain two antiparallel β-strands and an α-helix. The HNH or HNH-like domain may contain a metal-binding site (e.g., a divalent cation-binding site). The HNH or HNH-like domain may cleave one strand of the target nucleic acid (e.g., the complementary strand of the target strand of crRNA). Proteins containing HNH or HNH-like domains may include endonucleases, clicins, restriction endonucleases, transposases, and DNA packaging factors.

[0258] RuvC or RuvC-like domains may contain RNase H or RNase H-like folds. The RuvC / RNase H domain can participate in a diverse set of nucleic acid-based functions, including interactions with RNA and DNA. The RNase H domain may contain five β-strands surrounded by multiple α-helices. The RuvC / RNase H or RuvC / RNase H-like domain may contain metal-binding sites (e.g., divalent cation-binding sites). The RuvC / RNase H or RuvC / RNase H-like domain can cleave one strand of a target nucleic acid (e.g., the non-complementary strand of the target strand of crRNA). Proteins containing RuvC, RuvC-like, or RNase H-like domains may include RNase H, RuvC, DNA transposases, retroviral integrases, and Argonaut proteins.

[0259] The site-directed polypeptide can be a ribonuclease. The site-directed polypeptide can be a site-directed polypeptide without enzymatic activity. The site-directed polypeptide can be a site-directed polypeptide that is conditionally without enzymatic activity.

[0260] Site-directed peptides can introduce double-strand breaks or single-strand breaks into nucleic acids (e.g., genomic DNA). Double-strand breaks can stimulate endogenous DNA repair pathways in cells (e.g., homologous recombination and non-homologous end joining (NHEJ) or alternative non-homologous end joining (A-NHEJ)). NHEJ can repair cleaved target nucleic acids without requiring a homologous template. This can lead to the deletion of the target nucleic acid. Homologous recombination (HR) can occur using a homologous template. This homologous template can contain sequences homologous to sequences adjacent to the cleavage site of the target nucleic acid. After the target nucleic acid is cleaved by the site-directed peptide, the cleavage site can be destroyed (e.g., the site may become unusable for another round of cleavage by the original target nucleic acid and the site-directed peptide).

[0261] In some cases, homologous recombination can insert a foreign polynucleotide sequence into the cleavage site of a target nucleic acid. The foreign polynucleotide sequence can be referred to as a donor polynucleotide. In some cases of the methods disclosed herein, a donor polynucleotide, a portion of a donor polynucleotide, a copy of a donor polynucleotide, or a portion of a copy of a donor polynucleotide can be inserted into the cleavage site of the target nucleic acid. The donor polynucleotide can be a foreign polynucleotide sequence. The donor polynucleotide can be a sequence that is not naturally present at the cleavage site of the target nucleic acid. The vector may contain the donor polynucleotide. Modifications to the target DNA, attributable to NHEJ and / or HR, can result in, for example, mutations, deletions, alterations, integration, gene corrections, gene substitutions, gene markers, transgenic insertions, nucleotide deletions, gene disruptions, and / or gene mutations. Methods for integrating non-natural nucleic acids into genomic DNA can be referred to as genome engineering. In some cases, the site-directed polypeptide may comprise an amino acid sequence having amino acid sequence identity of up to 10%, up to 15%, up to 20%, up to 30%, up to 40%, up to 50%, up to 60%, up to 70%, up to 75%, up to 80%, up to 85%, up to 90%, up to 95%, up to 99%, or 100% of the wild-type exemplary site-directed polypeptide (e.g., Cas9 from Streptococcus pyogenes, SEQ ID NO: 8).

[0262] In some cases, the site-directed polypeptide may comprise an amino acid sequence having at least 10%, at least 15%, 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity with a wild-type exemplary site-directed polypeptide (e.g., Cas9 from Streptococcus pyogenes, SEQ ID NO: 8).

[0263] In some cases, the site-directed polypeptide may comprise an amino acid sequence having up to 10%, up to 15%, up to 20%, up to 30%, up to 40%, up to 50%, up to 60%, up to 70%, up to 75%, up to 80%, up to 85%, up to 90%, up to 95%, up to 99%, or 100% amino acid sequence identity with the nuclease domain of a wild-type exemplary site-directed polypeptide (e.g., Cas9 from Streptococcus pyogenes, SEQ ID NO: 8).

[0264] The site-directed polypeptide may contain at least 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity with the wild-type site-directed polypeptide (e.g., Cas9 from Streptococcus pyogenes, SEQ ID NO:8) in 10 adjacent amino acids. The site-directed polypeptide may contain at most 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity with the wild-type site-directed polypeptide (e.g., Cas9 from Streptococcus pyogenes, SEQ ID NO:8) in 10 adjacent amino acids. The site-directed polypeptide may contain at least 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity with the wild-type site-directed polypeptide (e.g., Cas9 from Streptococcus pyogenes, SEQ ID NO:8) in 10 adjacent amino acids within the HNH nuclease domain of the site-directed polypeptide. The site-directed peptide may contain at most 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity with the wild-type site-directed peptide (e.g., Cas9 from Streptococcus pyogenes, SEQ ID NO:8) in the 10 adjacent amino acids of the HNH nuclease domain of the site-directed peptide. The site-directed peptide may contain at least 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity with the wild-type site-directed peptide (e.g., Cas9 from Streptococcus pyogenes, SEQ ID NO:8) in the 10 adjacent amino acids of the RuvC nuclease domain of the site-directed peptide. The site-directed peptide may contain at most 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity with the wild-type site-directed peptide (e.g., Cas9 from Streptococcus pyogenes, SEQ ID NO:8) in the 10 adjacent amino acids of the RuvC nuclease domain of the site-directed peptide.

[0265] In some cases, the site-directed polypeptide may comprise an amino acid sequence having at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity with the nuclease domain of a wild-type exemplary site-directed polypeptide (e.g., Cas9 from Streptococcus pyogenes).

[0266] The site-directed peptide may comprise a modified form of the wild-type exemplary site-directed peptide. The modified form of the wild-type exemplary site-directed peptide may comprise amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the site-directed peptide. For example, the modified form of the wild-type exemplary site-directed peptide may have less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nucleic acid cleavage activity of the wild-type exemplary site-directed peptide (e.g., Cas9 from Streptococcus pyogenes). The modified form of the site-directed peptide may have essentially no nucleic acid cleavage activity. When the site-directed peptide is a modified form with essentially no nucleic acid cleavage activity, it may be referred to as "enzymatically inactive."

[0267] The modified form of the wild-type exemplary site-directed peptide may have nucleic acid cleavage activity greater than 90%, greater than 80%, greater than 70%, greater than 60%, greater than 50%, greater than 40%, greater than 30%, greater than 20%, greater than 10%, greater than 5%, or greater than 1% of that of the wild-type exemplary site-directed peptide (e.g., Cas9 from Streptococcus pyogenes).

[0268] The modified form of a site-directed peptide may include mutations. These mutations may enable the peptide to induce single-strand breaks (SSBs) on the target nucleic acid (e.g., by cleaving only one sugar-phosphate backbone of the target nucleic acid). This mutation may result in less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of one or more of the plurality of nucleic acid cleavage domains of the wild-type site-directed peptide (e.g., Cas9 from Streptococcus pyogenes). This mutation may cause one or more of the plurality of nucleic acid cleavage domains to retain the ability to cleave the complementary strand of the target nucleic acid but reduce its ability to cleave the non-complementary strand of the target nucleic acid. This mutation may also cause one or more of the plurality of nucleic acid cleavage domains to retain the ability to cleave the non-complementary strand of the target nucleic acid but reduce its ability to cleave the complementary strand of the target nucleic acid. For example, residues in the wild-type exemplary *Streptococcus pyogenes* Cas9 polypeptide, such as Asp10, His840, Asn854, and Asn856, can be mutated to inactivate one or more of the plurality of nucleic acid cleavage domains (e.g., nuclease domains). The residues to be mutated may correspond to Asp10, His840, Asn854, and Asn856 in the wild-type exemplary *Streptococcus pyogenes* Cas9 polypeptide (e.g., determined by sequence and / or structural alignment). Non-limiting examples of mutations may include D10A, H840A, N854A, or N856A. Those skilled in the art will recognize that mutations other than alanine substitution are suitable.

[0269] The D10A mutation can bind to one or more of the H840A, N854A, or N856A mutations to produce site-directed peptides that are essentially devoid of DNA cleavage activity. The H840A mutation can bind to one or more of the D10A, N854A, or N856A mutations to produce site-directed peptides that are essentially devoid of DNA cleavage activity. The N854A mutation can bind to one or more of the H840A, D10A, or N856A mutations to produce site-directed peptides that are essentially devoid of DNA cleavage activity. The N856A mutation can bind to one or more of the H840A, N854A, or D10A mutations to produce site-directed peptides that are essentially devoid of DNA cleavage activity. Site-directed peptides containing a essentially inactive nuclease domain are called nickases.

[0270] The mutations disclosed herein can be generated through site-directed mutagenesis. Mutations can include substitutions, additions, and deletions, or any combination thereof. In some cases, the mutation converts the mutated amino acid to alanine. In some cases, the mutation converts the mutated amino acid to another amino acid (e.g., glycine, serine, threonine, cysteine, valine, leucine, isoleucine, methionine, proline, phenylalanine, tyrosine, tryptophan, aspartic acid, glutamic acid, asparagine, glutamine, histidine, lysine, or arginine). The mutation can convert the mutated amino acid to a non-natural amino acid (e.g., selenomethionine). The mutation can convert the mutated amino acid to an amino acid mimic (e.g., phosphomimics). The mutation can be a conserved mutation. For example, the mutation can convert the mutated amino acid to an amino acid of similar size, shape, charge, polarity, conformation, and / or a rotational isomer of the mutated amino acid (e.g., cysteine / serine mutation, lysine / aspartic acid mutation, histidine / phenylalanine mutation).

[0271] In some cases, site-directed peptides (e.g., variants, mutations, inactive, and / or conditionally inactive site-directed peptides) can target nucleic acids. These site-directed peptides (e.g., variants, mutations, inactive, and / or conditionally inactive ribonucleases) can target RNA. RNA-targeting site-directed peptides may include members of other CRISPR subfamilies, such as Cas6 and Cas5.

[0272] The site-directed polypeptide may contain one or more non-natural sequences (e.g., fusion sequences).

[0273] The site-directed polypeptide may contain an amino acid sequence with at least 15% amino acid identity to Cas9 from bacteria (e.g., Streptococcus pyogenes), a nucleic acid binding domain, and two nucleic acid cleavage domains (i.e., the HNH domain and the RuvC domain).

[0274] The site-directed polypeptide may contain an amino acid sequence with at least 15% amino acid identity to Cas9 from bacteria (e.g., Streptococcus pyogenes) and two nucleic acid cleavage domains (i.e., the HNH domain and the RuvC domain).

[0275] The site-directed polypeptide may comprise an amino acid sequence containing at least 15% amino acid identity to Cas9 from bacteria (e.g., Streptococcus pyogenes) and two nucleic acid cleavage domains, wherein one or both of the nucleic acid cleavage domains contain at least 50% amino acid identity to the nuclease domain of Cas9 from bacteria (e.g., Streptococcus pyogenes).

[0276] The site-directed peptide may contain an amino acid sequence with at least 15% amino acid identity to Cas9 from bacteria (e.g., Streptococcus pyogenes), two nucleic acid cleavage domains (i.e., the HNH domain and the RuvC domain), and a linker connecting the site-directed peptide to a non-natural sequence.

[0277] The site-directed peptide may contain an amino acid sequence with at least 15% amino acid identity to Cas9 from bacteria (e.g., Streptococcus pyogenes), two nucleic acid cleavage domains (i.e., the HNH domain and the RuvC domain), wherein the site-directed peptide contains a mutation in one or both of the nucleic acid cleavage domains that reduces the cleavage activity of the nuclease domain by at least 50%.

[0278] The site-directed polypeptide may comprise an amino acid sequence containing at least 15% amino acid identity to Cas9 from bacteria (e.g., Streptococcus pyogenes) and two nucleic acid cleavage domains (i.e., the HNH domain and the RuvC domain), wherein one of the nuclease domains contains a mutation of aspartic acid 10, and / or wherein one of the nuclease domains contains a mutation of histidine 840, and wherein these mutations reduce the cleavage activity of the nuclease domain by at least 50%.

[0279] In some implementations, the site-directed polypeptide can be a ribonuclease.

[0280] In some cases, the ribonuclease may contain an amino acid sequence having up to about 20%, up to about 30%, up to about 40%, up to about 50%, up to about 60%, up to about 70%, up to about 75%, up to about 80%, up to about 85%, up to about 90%, up to about 95%, up to about 99%, or 100% amino acid sequence identity and / or homology with a wild-type reference ribonuclease (e.g., Csy4 from Pseudomonas aeruginosa). The ribonuclease may contain an amino acid sequence having up to about 20%, up to about 30%, up to about 40%, up to about 50%, up to about 60%, up to about 70%, up to about 75%, up to about 80%, up to about 85%, up to about 90%, up to about 95%, or 100% amino acid sequence identity and / or homology with a wild-type reference ribonuclease (e.g., Csy4 from Pseudomonas aeruginosa). The reference ribonuclease can be a member of the Cas6 family (e.g., Csy4, Cas6). The reference ribonuclease can be a member of the Cas5 family (e.g., Cas5 from *D. vulgaris*). The reference ribonuclease can be a member of the type I CRISPR family (e.g., Cas3). The reference ribonuclease can be a member of the type II family. The reference ribonuclease can be a member of the type III family (e.g., Cas6). The reference ribonuclease can be a member of the repeat-associated-suspect (RAMP) superfamily (e.g., Cas7).

[0281] The ribonuclease may contain amino acid modifications (e.g., substitutions, deletions, additions, etc.). The ribonuclease may contain one or more non-natural sequences (e.g., fusions, affinity tags). The amino acid modifications may not substantially alter the activity of the ribonuclease. Ribonucleases containing amino acid modifications and / or fusions may retain at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 95%, at least approximately 97%, or 100% of the activity of wild-type ribonucleases.

[0282] This modification can alter the enzymatic activity of ribonucleases. The modification can result in less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the activity of the ribonuclease. In some cases, this modification occurs in the nuclease domain of the ribonuclease. Such modification can result in less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nucleic acid cleavage activity of one or more of the plurality of nucleic acid cleavage domains of the wild-type ribonuclease.

[0283] Conditionally inactive ribonucleases

[0284] In some implementations, the ribonuclease may be conditionally inactive. Conditionally inactive ribonucleases can bind to polynucleotides in a sequence-specific manner. Conditionally inactive ribonucleases can bind to polynucleotides in a sequence-specific manner but cannot cleave the target polynucleotide.

[0285] In some cases, the conditionally non-enzymatic ribonuclease may contain up to about 20%, up to about 30%, up to about 40%, up to about 50%, up to about 60%, up to about 70%, up to about 75%, up to about 80%, up to about 85%, up to about 90%, up to about 95%, up to about 99%, or 100% amino acid sequence identity and / or homology with a reference conditionally non-enzymatic ribonuclease (e.g., Csy4 from Pseudomonas aeruginosa). In some cases, the conditionally non-enzymatic ribonuclease may comprise at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% amino acid sequence identity and / or homology with a reference conditionally non-enzymatic ribonuclease (e.g., Csy4 from Pseudomonas aeruginosa).

[0286] The conditionally enzyme-free ribonuclease may contain modified forms of the ribonuclease. These modifications may include amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the cleavage activity of the ribonuclease. For example, the modified form of the conditionally enzyme-free ribonuclease may have less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the cleavage activity of a conditionally enzyme-free ribonuclease (e.g., Csy4 from Pseudomonas aeruginosa), as referenced for example, a wild-type conditionally enzyme-free ribonuclease. The modified form of the conditionally enzyme-free ribonuclease may also have essentially no cleavage activity. When a conditionally enzyme-free ribonuclease is a modified form with essentially no cleavage activity, it may be referred to as "enzyme-free."

[0287] The modified form of the conditionally inactive ribonuclease may include a mutation that results in reduced nucleic acid cleavage activity (i.e., rendering the conditionally inactive ribonuclease inactive in one or more of its nucleic acid cleavage domains). This mutation may result in less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nucleic acid cleavage activity in one or more of the plurality of nucleic acid cleavage domains of a wild-type ribonuclease (e.g., Csy4 from Pseudomonas aeruginosa). This mutation may occur in the nuclease domain of the ribonuclease. This mutation may occur in ferrugin-like folds. This mutation may include mutations in conserved aromatic amino acids. This mutation may include mutations in catalytic amino acids. This mutation may include mutations in histidine. For example, the mutation may include an H29A mutation in Csy4 (e.g., Csy4 from Pseudomonas aeruginosa), or any residue corresponding to H29A as determined by sequence and / or structural alignment. Other residues may be mutated to achieve the same effect (i.e., inactivation of one or more of the plurality of nuclease domains).

[0288] The mutations of this invention can be generated through site-directed mutagenesis. Mutations can include substitutions, additions, and deletions, or any combination thereof. In some cases, the mutation converts the mutated amino acid to alanine. In some cases, the mutation converts the mutated amino acid to another amino acid (e.g., glycine, serine, threonine, cysteine, valine, leucine, isoleucine, methionine, proline, phenylalanine, tyrosine, tryptophan, aspartic acid, glutamic acid, asparagine, glutamine, histidine, lysine, or arginine). The mutation can convert the mutated amino acid to a non-natural amino acid (e.g., selenomethionine). The mutation can convert the mutated amino acid to an amino acid mimic (e.g., a phosphate mimic). The mutation can be a conserved mutation. For example, the mutation can convert the mutated amino acid to an amino acid of similar size, shape, charge, polarity, conformation, and / or a rotatable isomer of the mutated amino acid (e.g., cysteine / serine mutation, lysine / aspartic acid mutation, histidine / phenylalanine mutation).

[0289] Conditionally inactive ribonucleases can be inactive in the absence of a reactivating agent (e.g., imidazole). The reactivating agent can be a reagent mimicking histidine residues (e.g., having an imidazole ring). Conditionally inactive ribonucleases can be activated by contact with a reactivating agent. This reactivating agent may contain imidazole. For example, the conditionally inactive ribonuclease can be activated by contacting it with imidazole at concentrations from approximately 100 mM to approximately 500 mM. Imidazole can be present at concentrations of approximately 100 mM, approximately 150 mM, approximately 200 mM, approximately 250 mM, approximately 300 mM, approximately 350 mM, approximately 400 mM, approximately 450 mM, approximately 500 mM, approximately 550 mM, or approximately 600 mM. The presence of imidazole (e.g., in the concentration range of about 100 mM to about 500 mM) can reactivate the conditionally inactive ribonuclease to make it enzymatically active, for example, the conditionally inactive ribonuclease exhibiting at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or greater than 95% of a reference conditionally inactive ribonuclease (e.g., Csy4 from Pseudomonas aeruginosa containing the H29A mutation).

[0290] Conditionally inactive ribonucleases may contain at least 20% amino acid identity of Csy4 from Pseudomonas aeruginosa and a mutation at histidine 29, wherein the mutation results in a reduction of at least 50% of the nuclease activity of the ribonuclease, and wherein at least 50% of the lost nuclease activity can be restored by culturing the ribonuclease with at least 100 mM imidazole.

[0291] Codon optimization of polynucleotides encoding site-directed peptides and / or ribonucleases can be performed. This type of optimization may require mutations in exogenous (e.g., recombinant) DNA to mimic the codon preferences of the intended host organism or cell while encoding the same protein. Thus, the codon can be changed, but the encoded protein remains unchanged. For example, if the intended target cell is a human cell, the human codon-optimized polynucleotide Cas9 can be used to generate a suitable site-directed peptide. As another non-limiting example, if the intended host cell is a mouse cell, a mouse codon-optimized polynucleotide encoding Cas9 can be a suitable site-directed peptide. Polynucleotides encoding site-directed peptides can be codon-optimized for many host cells of interest. The host cell can be a cell from any organism (e.g., bacterial cells, archaea cells, cells of single-celled eukaryotes, plant cells, algal cells such as *Botryococcus braunii*, *Chlamydomonas reinhardtii*, *Nannochloropsis gaditana*, *Chlorella pyrenoidosa*, *Sargassum patens C. Agardh*, etc.), fungal cells (e.g., yeast cells), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cattle, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Codon optimization may not be necessary. In some cases, codon optimization is preferred.

[0292] Nucleic acid targeting nucleic acid

[0293] This disclosure provides nucleic acids that target nucleic acids, which can direct the activity of a related polypeptide (e.g., a site-directed polypeptide) to a specific target sequence within the target nucleic acid. The nucleic acid targeting nucleic acid may comprise nucleotides. The nucleic acid targeting nucleic acid may be RNA. The nucleic acid targeting nucleic acid may comprise a single-target nucleic acid. Figure 1AAn exemplary single-target nucleic acid is depicted. Spacer region overhang 105 and tracrRNA overhang 135 may contain elements that contribute to additional functionality (e.g., stability) of the nucleic acid targeting the nucleic acid. In some embodiments, spacer region overhang 105 and tracrRNA overhang 135 are optional. Spacer region sequence 110 may contain a sequence capable of hybridizing with the target nucleic acid sequence. Spacer region sequence 110 may be a variable portion of the nucleic acid targeting the nucleic acid. The sequence of spacer region sequence 110 may be sequence-engineered to hybridize with the target nucleic acid sequence. CRISPR repeat fragment 115 (i.e., referred to as the minimal CRISPR repeat fragment in this exemplary embodiment) may contain nucleotides capable of hybridizing with tracrRNA sequence 125 (i.e., referred to as the minimal tracrRNA sequence in this exemplary embodiment). Minimal CRISPR repeat fragment 115 and minimal tracrRNA sequence 125 may interact, the interacting molecule comprising a base-paired double-stranded structure. Minimal CRISPR repeat fragment 115 and minimal tracrRNA sequence 125 may together promote binding to a site-directed peptide. The minimum CRISPR repeat fragment 115 and the minimum tracrRNA sequence 125 can be joined together via a single guide adapter 120 to form a hairpin structure. The 3' tracrRNA sequence 130 may contain a motif recognition sequence adjacent to the pre-interstitial region sequence. The 3' tracrRNA sequence 130 may be identical or similar to a portion of the tracrRNA sequence. In some embodiments, the 3' tracrRNA sequence 130 may contain one or more hairpin structures.

[0294] In some implementation schemes, nucleic acids targeting nucleic acids may include, for example: Figure 1BThe nucleic acid targeted by the target nucleic acid is depicted. The nucleic acid targeted by the target nucleic acid may include a spacer sequence 140. The spacer sequence 140 may include a sequence capable of hybridizing with the target nucleic acid sequence. The spacer sequence 140 may be a variable portion of the nucleic acid targeted by the target nucleic acid. The spacer sequence 140 may be the 5' of a first double strand 145. The first double strand 145 includes a hybridization region between a minimal CRISPR repeat fragment 146 and a minimal tracrRNA sequence 147. The first double strand 145 may be interrupted by a protrusion 150. The protrusion 150 may contain unpaired nucleotides. The protrusion 150 may facilitate the recruitment of a site-directed polypeptide to the nucleic acid targeted by the target nucleic acid. The protrusion 150 is followed by a first stem 155. The first stem 155 includes a linker sequence connecting the minimal CRISPR repeat fragment 146 and the minimal tracrRNA sequence 147. The last pair of nucleotides at the 3' end of the first double strand 145 may be linked to a second linker sequence 160. The second linker 160 may include a P-domain. The second connector 160 can link the first double strand 145 to the intermediate-tracrRNA 165. In some embodiments, the intermediate-tracrRNA 165 may include one or more hairpin regions. For example, the intermediate-tracrRNA 165 may include a second stem 170 and a third stem 175.

[0295] In some implementations, the nucleic acid targeting a nucleic acid may include a dual-guided nucleic acid structure. Figure 2 An exemplary dual-guide RNA structure targeting nucleic acids is depicted. Similar to the single-guide RNA structure of Figure 1, this dual-guide RNA structure may include a spacer overhang 205, a spacer region 210, a minimal CRISPR repeat 215, a minimal tracrRNA sequence 230, a 3' tracrRNA sequence 235, and a tracrRNA overhang 240. However, the nucleic acid-targeting dual-guide RNA may not contain a single-guide adapter 120. Instead, the minimal CRISPR repeat sequence 215 may include a 3' CRISPR repeat sequence 220, which may be similar to or identical to a portion of the CRISPR repeat. Similarly, the minimal tracrRNA sequence 230 may include a 5' tracrRNA sequence 225, which may be similar to or identical to a portion of the tracrRNA. The dual-guide RNA may hybridize together via the minimal CRISPR repeat 215 and the minimal tracrRNA sequence 230.

[0296] In some embodiments, the first fragment (i.e., the nucleic acid-targeted fragment) may include a spacer overhang (e.g., 105 / 205) and a spacer region (e.g., 110 / 210). The nucleic acid-targeted nucleic acid can guide the bound polypeptide to a specific nucleotide sequence within the target nucleic acid via the aforementioned nucleic acid-targeted fragment.

[0297] In some embodiments, the second fragment (i.e., the protein-binding fragment) may comprise a minimal CRISPR repeat (e.g., 115 / 215), a minimal tracrRNA sequence (e.g., 125 / 230), a 3' tracrRNA sequence (e.g., 130 / 235), and / or a tracrRNA overhang sequence (e.g., 135 / 240). The protein-binding fragment of a nucleic acid targeting a nucleic acid can interact with a site-directed polypeptide. The protein-binding fragment of a nucleic acid targeting a nucleic acid may comprise two nucleotides that can hybridize. The nucleotides of the protein-binding fragment can hybridize to form a double-stranded nucleic acid duplex. This double-stranded nucleic acid duplex can be RNA. This double-stranded nucleic acid duplex can be DNA.

[0298] In some cases, nucleic acids targeting nucleic acids can contain, in a 5' to 3' order, a spacer overhang, a spacer region, a minimal CRISPR repeat, a single guide adapter, a minimal tracrRNA, a 3' tracrRNA sequence, and a tracrRNA overhang. In other cases, nucleic acids targeting nucleic acids can contain, in any order, a tracrRNA overhang, a 3' tracrRNA sequence, a minimal tracrRNA, a single guide adapter, a minimal CRISPR repeat, a spacer region, and a spacer overhang.

[0299] Nucleic acids targeting nucleic acids and site-directed peptides can form complexes. The nucleic acid targeting the nucleic acid can provide target specificity to the complex by including a nucleotide sequence that hybridizes to the sequence of the target nucleic acid. In other words, the site-directed peptide can be directed to the nucleic acid sequence by binding to at least a protein-binding fragment of the nucleic acid targeting the nucleic acid. The nucleic acid targeting the nucleic acid can direct the activity of the Cas9 protein. The nucleic acid targeting the nucleic acid can also direct the activity of the non-enzymatically active Cas9 protein.

[0300] The methods disclosed herein can provide genetically modified cells. Genetically modified cells may contain exogenous nucleic acids targeting nucleic acids and / or exogenous nucleic acids, said exogenous nucleic acids comprising nucleotide sequences encoding nucleic acids targeting nucleic acids.

[0301] Spacer region overhang sequences can provide stability and / or modification sites for nucleic acids targeting nucleic acids. Spacer region overhang sequences can be approximately 1 nucleotide to approximately 400 nucleotides in length. Spacer region overhang sequences can also be greater than 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 1000, 2000, 3000, 4000, 5000, 6000, or 7000 or more nucleotides in length. The spacer region overhang sequence can be less than 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 1000, 2000, 3000, 4000, 5000, 6000, 7000 or more nucleotides in length. The spacer region overhang sequence can be less than 10 nucleotides in length. The spacer region overhang sequence can be 10 to 30 nucleotides in length. The spacer region overhang sequence can be 30 to 70 nucleotides in length.

[0302] The overhanging sequence of this spacer region may contain a portion (e.g., a stability control sequence, a ribonuclease-binding sequence, or a ribozyme). This portion can affect the stability of the nucleic acid targeting the RNA. The portion may be a transcription terminator fragment (i.e., a transcription termination sequence). The portion of the nucleic acid targeting the RNA may have a total length of approximately 10 nucleotides to approximately 100 nucleotides, approximately 10 nucleotides (nt) to approximately 20 nucleotides, approximately 20 nucleotides (nt) to approximately 30 nucleotides, approximately 30 nucleotides (nt) to approximately 40 nucleotides, approximately 40 nucleotides (nt) to approximately 50 nucleotides, approximately 50 nucleotides (nt) to approximately 60 nucleotides, approximately 60 nucleotides (nt) to approximately 70 nucleotides, approximately 70 nucleotides (nt) to approximately 80 nucleotides, approximately 80 nucleotides (nt) to approximately 90 nucleotides, or approximately 90 nucleotides (nt) to approximately 100 nucleotides, approximately 15 nucleotides (nt) to approximately 80 nucleotides, approximately 15 nucleotides (nt) to approximately 50 nucleotides, approximately 15 nucleotides (nt) to approximately 40 nucleotides, approximately 15 nucleotides (nt) to approximately 30 nucleotides, or approximately 15 nucleotides (nt) to approximately 25 nucleotides. This portion may be a functional portion in eukaryotic cells. In some cases, this part can be the part that functions in prokaryotic cells. This part can function in both eukaryotic and prokaryotic cells.

[0303] Non-limiting examples of suitable portions may include: a 5' cap (e.g., a 7-methylguanylate cap (m7G)), a riboswitch sequence (e.g., to allow regulated stability and / or accessibility regulated by proteins and protein complexes), a sequence forming a dsRNA double helix (i.e., a hairpin structure), a sequence targeting RNA to subcellular locations (e.g., the nucleus, mitochondria, chloroplasts, etc.), modifications or sequences providing tracking (e.g., direct conjugation to fluorescent molecules, conjugation to a portion that promotes fluorescence detection, sequences that allow fluorescence detection, etc.), modifications or sequences providing binding sites for proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.), modifications or sequences providing enhanced, degraded, and / or controllable stability, or any combination thereof. The spacer region overhang sequence may contain primer binding sites, molecular indicators (e.g., barcode sequences). The spacer region overhang sequence may contain nucleic acid affinity markers.

[0304] Spacer regions, which are segments of nucleic acids that target nucleic acids, may contain nucleotide sequences (e.g., spacer regions) that can hybridize with sequences in the target nucleic acid. The spacer regions of nucleic acids that target nucleic acids can interact with the target nucleic acid in a sequence-specific manner through hybridization (i.e., base pairing). Therefore, the nucleotide sequence of the spacer region is variable and can determine the position within the target nucleic acid where the nucleic acid targets the nucleic acid and the target nucleic acid can interact.

[0305] This spacer region sequence can hybridize with the target nucleic acid located at the 5' of the adjacent motif (PAM). Different organisms may contain different PAM sequences. For example, in Streptococcus pyogenes, the PAM can be a sequence in the target nucleic acid containing the sequence 5'-XRR-3', where R can be A or G, and X is any nucleotide and X is immediately adjacent to the 3' of the target nucleic acid sequence targeted by this spacer region sequence.

[0306] The target nucleic acid sequence can be 20 nucleotides. The target nucleic acid can be less than 20 nucleotides. The target nucleic acid can be at least 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The target nucleic acid can be up to 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The target nucleic acid sequence can be the 20 bases of the 5' of the first nucleotide immediately adjacent to the PAM. For example, in a sequence containing 5'-NNNNNNNNNNNNNNNNNNNNXRR-3', the target nucleic acid can be a sequence equivalent to N's, where N is any nucleotide.

[0307] The spacer region sequence targeting the target nucleic acid can have a length of at least approximately 6 nt. For example, the spacer region sequence capable of hybridizing with the target nucleic acid can have lengths of at least approximately 6 nt, at least approximately 10 nt, at least approximately 15 nt, at least approximately 18 nt, at least approximately 19 nt, at least approximately 20 nt, at least approximately 25 nt, at least approximately 30 nt, at least approximately 35 nt, or at least approximately 40 nt, approximately 6 nt to approximately 80 nt, approximately 6 nt to approximately 50 nt, approximately 6 nt to approximately 45 nt, approximately 6 nt to approximately 40 nt, approximately 6 nt to approximately 35 nt, approximately 6 nt to approximately 30 nt, approximately 6 nt to approximately 25 nt, approximately 6 nt to approximately 20 nt, approximately 6 nt to approximately 19 nt, approximately 10 nt to approximately 50 nt, approximately 10 nt to approximately 45 nt, approximately 10 nt to approximately 40 nt, approximately... The length can range from 10 nt to approximately 35 nt, approximately 10 nt to approximately 30 nt, approximately 10 nt to approximately 25 nt, approximately 10 nt to approximately 20 nt, approximately 10 nt to approximately 19 nt, approximately 19 nt to approximately 25 nt, approximately 19 nt to approximately 30 nt, approximately 19 nt to approximately 35 nt, approximately 19 nt to approximately 40 nt, approximately 19 nt to approximately 45 nt, approximately 19 nt to approximately 50 nt, approximately 19 nt to approximately 60 nt, approximately 20 nt to approximately 25 nt, approximately 20 nt to approximately 30 nt, approximately 20 nt to approximately 35 nt, approximately 20 nt to approximately 40 nt, approximately 20 nt to approximately 45 nt, approximately 20 nt to approximately 50 nt, or approximately 20 nt to approximately 60 nt. In some cases, the spacer region sequence of the hybridizable target nucleic acid can be 20 nucleotides in length. The spacer region of the hybridizable target nucleic acid can be 19 nucleotides in length.

[0308] The complementarity percentage between the spacer region sequence and the target nucleic acid can be at least approximately 30%, at least approximately 40%, at least approximately 50%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 95%, at least approximately 97%, at least approximately 98%, at least approximately 99%, or 100%. The complementarity percentage between the spacer region sequence and the target nucleic acid can be at most approximately 30%, at most approximately 40%, at most approximately 50%, at most approximately 60%, at most approximately 65%, at most approximately 70%, at most approximately 75%, at most approximately 80%, at most approximately 85%, at most approximately 90%, at most approximately 95%, at most approximately 97%, at most approximately 98%, at most approximately 99%, or 100%. In some cases, the complementarity percentage between the spacer region sequence and the target nucleic acid can be 100% of the 6 adjacent 5' terminal nucleotides of the target sequence on the complementary strand of the target nucleic acid. In some cases, the complementarity percentage between the spacer region sequence and the target nucleic acid can be at least 60% in approximately 20 adjacent nucleotides. In some cases, the complementarity percentage between the spacer region sequence and the target nucleic acid can be 100% in the 14 adjacent 5' terminal nucleotides of the target sequence on the complementary strand of the target nucleic acid and as low as 0% in the remainder. In this case, the spacer region sequence can be considered to be 14 nucleotides long. In some cases, the complementarity percentage between the spacer region sequence and the target nucleic acid can be 100% in the 6 adjacent 5' terminal nucleotides of the target sequence on the complementary strand of the target nucleic acid and as low as 0% in the remainder. In this case, the spacer region sequence can be considered to be 6 nucleotides long. The target nucleic acid can be more than approximately 50%, 60%, 70%, 80%, 90%, or 100% complementary to the seed region of the crRNA. The target nucleic acid can be less than approximately 50%, 60%, 70%, 80%, 90%, or 100% complementary to the seed region of the crRNA.

[0309] Spacer regions of nucleic acids targeted by a specific nucleic acid can be modified (e.g., through genetic engineering) to hybridize with any desired sequence within the target nucleic acid. For example, spacer regions can be engineered (e.g., designed, programmed) to hybridize with sequences in target nucleic acids relating to cancer, cell growth, DNA replication, DNA repair, HLA genes, cell surface proteins, T-cell receptors, immunoglobulin superfamily genes, tumor suppressor genes, microRNA genes, long non-coding RNA genes, transcription factors, globins, viral proteins, mitochondrial genes, etc.

[0310] Spacer sequences can be identified using computer programs (e.g., machine-readable codes). These programs can utilize variables such as predicted melting temperature, secondary structure formation and predicted annealing temperature, sequence identity, genomic background, chromatin accessibility, %GC, genomic frequency, methylation status, and the presence of SNPs.

[0311] Minimal CRISPR repeat sequence

[0312] The minimal CRISPR repeat sequence can be a sequence having at least approximately 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% sequence identity and / or sequence homology with a reference CRISPR repeat sequence (e.g., crRNA from Streptococcus pyogenes). The minimal CRISPR repeat sequence can also be a sequence having at most approximately 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% sequence identity and / or sequence homology with a reference CRISPR repeat sequence (e.g., crRNA from Streptococcus pyogenes). The minimal CRISPR repeat fragment may contain nucleotides capable of hybridizing with the minimal tracrRNA sequence. The minimal CRISPR repeat fragment and the minimal tracrRNA sequence may form a base-paired double-stranded structure. Minimal CRISPR repeats and minimal tracrRNA sequences can work together to promote site-directed peptide binding. A portion of the minimal CRISPR repeat sequence can hybridize with the minimal tracrRNA sequence. A portion of the minimal CRISPR repeat sequence can be at least approximately 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to the minimal tracrRNA sequence. A portion of the minimal CRISPR repeat sequence can be up to approximately 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to the minimal tracrRNA sequence.

[0313] The minimum CRISPR repeat sequence can be from approximately 6 nucleotides to approximately 100 nucleotides in length. For example, the minimum CRISPR repeat sequence can be from approximately 6 nucleotides (nt) to approximately 50 nt, from approximately 6 nt to approximately 40 nt, from approximately 6 nt to approximately 30 nt, from approximately 6 nt to approximately 25 nt, from approximately 6 nt to approximately 20 nt, from approximately 6 nt to approximately 15 nt, from approximately 8 nt to approximately 40 nt, from approximately 8 nt to approximately 30 nt, from approximately 8 nt to approximately 25 nt, from approximately 8 nt to approximately 20 nt, or from approximately 8 nt to approximately 15 nt, from approximately 15 nt to approximately 100 nt, from approximately 15 nt to approximately 80 nt, from approximately 15 nt to approximately 50 nt, from approximately 15 nt to approximately 40 nt, from approximately 15 nt to approximately 30 nt, or from approximately 15 nt to approximately 25 nt. In some embodiments, the minimum CRISPR repeat sequence is approximately 12 nucleotides in length.

[0314] The minimum CRISPR repeat sequence may be at least approximately 60% identical to a reference minimum CRISPR repeat sequence (e.g., wild-type crRNA from Streptococcus pyogenes) in a sequence of at least 6, 7, or 8 adjacent nucleotides. For example, the minimum CRISPR repeat sequence may be at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 95%, at least approximately 98%, at least approximately 99%, or 100% identical to the reference minimum CRISPR repeat sequence in a sequence of at least 6, 7, or 8 adjacent nucleotides.

[0315] Minimal tracrRNA sequence

[0316] The minimal tracrRNA sequence can be a sequence having at least approximately 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% sequence identity and / or sequence homology with a reference tracrRNA sequence (e.g., wild-type tracrRNA from Streptococcus pyogenes). The minimal tracrRNA sequence can also be a sequence having at most approximately 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% sequence identity and / or sequence homology with a reference tracrRNA sequence (e.g., wild-type tracrRNA from Streptococcus pyogenes). The minimal tracrRNA sequence may contain nucleotides capable of hybridizing with the minimal CRISPR repeat sequence. The minimal tracrRNA sequence and the minimal CRISPR repeat sequence may form a base-paired double-stranded structure. Minimal tracrRNA sequences and minimal CRISPR repeats can work together to promote site-directed peptide binding. A portion of the minimal tracrRNA sequence can hybridize with the minimal CRISPR repeat sequence. A portion of the minimal tracrRNA sequence can be 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to the minimal CRISPR repeat sequence.

[0317] The minimum tracrRNA sequence can be approximately 6 nucleotides to approximately 100 nucleotides in length. For example, the minimum tracrRNA sequence can be approximately 6 nucleotides (nt) to approximately 50 nt, approximately 6 nt to approximately 40 nt, approximately 6 nt to approximately 30 nt, approximately 6 nt to approximately 25 nt, approximately 6 nt to approximately 20 nt, approximately 6 nt to approximately 15 nt, approximately 8 nt to approximately 40 nt, approximately 8 nt to approximately 30 nt, approximately 8 nt to approximately 25 nt, approximately 8 nt to approximately 20 nt, or approximately 8 nt to approximately 15 nt, approximately 15 nt to approximately 100 nt, approximately 15 nt to approximately 80 nt, approximately 15 nt to approximately 50 nt, approximately 15 nt to approximately 40 nt, approximately 15 nt to approximately 30 nt, or approximately 15 nt to approximately 25 nt. In some embodiments, the minimum tracrRNA sequence is approximately 14 nucleotides in length.

[0318] The minimum tracrRNA sequence may be at least approximately 60% identical to the reference minimum tracrRNA sequence (e.g., wild-type tracrRNA from Streptococcus pyogenes) in a sequence of at least 6, 7, or 8 adjacent nucleotides. For example, the minimum tracrRNA sequence may be at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 95%, at least approximately 98%, at least approximately 99%, or 100% identical to the reference minimum tracrRNA sequence in a sequence of at least 6, 7, or 8 adjacent nucleotides.

[0319] The double strand between the smallest CRISPR RNA and the smallest tracrRNA (i.e. Figure 1B The first double strand of the double strand may contain a double helix. The first base of the first strand of the double strand (e.g., ...) Figure 1B The smallest CRISPR repeat fragment in the bilayer (e.g., guanine) can be guanine. The first base of the first strand of the bilayer (e.g., Figure 1B The smallest CRISPR repeat fragment in the RNA can be adenine. The double strand between the smallest CRISPR RNA and the smallest tracrRNA (i.e., Figure 1B The first double strand in the CRISPR RNA (i.e., the smallest CRISPR RNA) may contain at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotides. The double strand between the smallest CRISPR RNA and the smallest tracrRNA (i.e., the smallest tracrRNA) Figure 1B The first double strand in a nucleotide may contain up to approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotides.

[0320] The duplex may contain mismatches. The duplex may contain at least approximately 1, 2, 3, 4, or 5 or more mismatches. The duplex may contain at most approximately 1, 2, 3, 4, or 5 or more mismatches. In some cases, the duplex may contain no more than 2 mismatches.

[0321] protrusion

[0322] A protrusion can refer to an unpaired region of nucleotides within a duplex consisting of a minimal CRISPR repeat and a minimal tracrRNA sequence. This protrusion pair is important for site-directed polypeptide binding. A protrusion may contain an unpaired 5'-XXXY-3' region on one side of the duplex, where X is any purine and Y can be a nucleotide capable of forming a wobble pair with a nucleotide on the opposite strand, and an unpaired nucleotide region on the other side of the duplex.

[0323] For example, the protrusion may contain an unpaired purine (e.g., adenine) on the minimum CRISPR repeat strand of the protrusion. In some embodiments, the protrusion may contain an unpaired 5'-AAGY-3' of the minimum tracrRNA sequence strand of the protrusion, where Y may be a nucleotide capable of forming a wobble pair with a nucleotide on the minimum CRISPR repeat strand.

[0324] The protrusions on the first side of the double strand (e.g., the minimal CRISPR repeat side) may contain at least 1, 2, 3, 4, or 5 or more unpaired nucleotides. The protrusions on the first side of the double strand (e.g., the minimal CRISPR repeat side) may contain up to 1, 2, 3, 4, or 5 or more unpaired nucleotides. The protrusions on the first side of the double strand (e.g., the minimal CRISPR repeat side) may contain 1 unpaired nucleotide.

[0325] A protrusion on the second side of the duplex (e.g., the side with the smallest tracrRNA sequence of the duplex) may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more unpaired nucleotides. A protrusion on the second side of the duplex (e.g., the side with the smallest tracrRNA sequence of the duplex) may contain up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more unpaired nucleotides. A protrusion on the second side of the duplex (e.g., the side with the smallest tracrRNA sequence of the duplex) may contain 4 unpaired nucleotides. Regions on each strand of the duplex with different numbers of unpaired nucleotides may pair together. For example, a protrusion may contain 5 unpaired nucleotides from the first strand and 1 unpaired nucleotide from the second strand. A protrusion may contain 4 unpaired nucleotides from the first strand and 1 unpaired nucleotide from the second strand. A protrusion may contain 3 unpaired nucleotides from the first strand and 1 unpaired nucleotide from the second strand. A protrusion may contain two unpaired nucleotides from the first strand and one unpaired nucleotide from the second strand. A protrusion may contain one unpaired nucleotide from the first strand and one unpaired nucleotide from the second strand. A protrusion may contain one unpaired nucleotide from the first strand and two unpaired nucleotides from the second strand. A protrusion may contain one unpaired nucleotide from the first strand and three unpaired nucleotides from the second strand. A protrusion may contain one unpaired nucleotide from the first strand and four unpaired nucleotides from the second strand. A protrusion may contain one unpaired nucleotide from the first strand and five unpaired nucleotides from the second strand.

[0326] In some cases, the protrusion may contain at least one wobble pair. In some cases, the protrusion may contain at most one wobble pair. The protrusion sequence may contain at least one purine nucleotide. The protrusion sequence may contain at least three purine nucleotides. The protrusion sequence may contain at least five purine nucleotides. The protrusion sequence may contain at least one guanine nucleotide. The protrusion sequence may contain at least one adenine nucleotide.

[0327] P-domain

[0328] The P-domain can be a region of a target nucleic acid that recognizes a pre-interstitial adjacent motif (PAM) in the target nucleic acid. The P-domain can hybridize with the PAM in the target nucleic acid. Therefore, the P-domain can contain a sequence complementary to the PAM. The P-domain can be located at the 3' of the minimal tracrRNA sequence. The AP-domain can be located within the 3' tracrRNA sequence (i.e., the intermediate tracrRNA sequence).

[0329] The p-domain begins at at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 or more nucleotides at the 3' end of the last nucleotide pair in both the minimum CRISPR repeat and the minimum tracrRNA sequence duplex. The p-domain may begin at most approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotides at the 3' end of the last nucleotide pair in both the minimum CRISPR repeat and the minimum tracrRNA sequence duplex.

[0330] The P-domain may contain at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 or more consecutive nucleotides.

[0331] In some cases, the P-domain may contain a CC dinucleotide (i.e., two consecutive cytosine nucleotides). This CC dinucleotide can interact with the GG dinucleotide of the PAM, where the PAM contains a 5'-XGG-3' sequence.

[0332] The P-domain can be a nucleotide sequence located within a 3' tracrRNA sequence (i.e., the intermediate tracrRNA sequence). The P-domain can contain double-stranded nucleotides (e.g., nucleotides in a hairpin structure that hybridize together). For example, the P-domain can contain a CC dinucleotide that hybridizes with a GG dinucleotide in a hairpin double strand of the 3' tracrRNA sequence (i.e., the intermediate tracrRNA sequence). The activity of the P-domain (e.g., the ability of a nucleic acid targeting a target nucleic acid) can be modulated by the hybridization state of the P-domain. For example, if the P-domain hybridizes, the nucleic acid targeting the target nucleic acid may not recognize its target. If the P-domain does not hybridize, the nucleic acid targeting the target nucleic acid can recognize its target.

[0333] The P-domain can interact with P-domain interaction regions within a site-directed polypeptide. The P-domain can interact with arginine-rich basic patches within the site-directed polypeptide. The P-domain interaction region can interact with PAM sequences. The P-domain may contain stem-loops. The P-domain may contain protrusions.

[0334] 3' tracrRNA sequence

[0335] The 3' tracr RNA sequence may be a sequence having at least approximately 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% sequence identity and / or sequence homology with a reference tracr RNA sequence (e.g., tracr RNA from Streptococcus pyogenes). The 3' tracr RNA sequence may also be a sequence having at most approximately 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% sequence identity and / or sequence homology with a reference tracr RNA sequence (e.g., tracr RNA from Streptococcus pyogenes).

[0336] The 3' tracrRNA sequence can be from approximately 6 nucleotides to approximately 100 nucleotides in length. For example, the 3' tracrRNA sequence can be from approximately 6 nucleotides (nt) to approximately 50 nt, from approximately 6 nt to approximately 40 nt, from approximately 6 nt to approximately 30 nt, from approximately 6 nt to approximately 25 nt, from approximately 6 nt to approximately 20 nt, from approximately 6 nt to approximately 15 nt, from approximately 8 nt to approximately 40 nt, from approximately 8 nt to approximately 30 nt, from approximately 8 nt to approximately 25 nt, from approximately 8 nt to approximately 20 nt, or from approximately 8 nt to approximately 15 nt, from approximately 15 nt to approximately 100 nt, from approximately 15 nt to approximately 80 nt, from approximately 15 nt to approximately 50 nt, from approximately 15 nt to approximately 40 nt, from approximately 15 nt to approximately 30 nt, or from approximately 15 nt to approximately 25 nt. In some embodiments, the 3' tracrRNA sequence is approximately 14 nucleotides in length.

[0337] The 3' tracrRNA sequence may be at least approximately 60% identical to a reference 3' tracrRNA sequence (e.g., a wild-type 3' tracrRNA sequence from Streptococcus pyogenes) in a sequence of at least 6, 7, or 8 adjacent nucleotides. For example, the 3' tracrRNA sequence may be at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 95%, at least approximately 98%, at least approximately 99%, or 100% identical to the reference 3' tracrRNA sequence (e.g., a wild-type 3' tracrRNA sequence from Streptococcus pyogenes) in a sequence of at least 6, 7, or 8 adjacent nucleotides.

[0338] A 3' tracrRNA sequence may contain more than one double-stranded region (e.g., hairpin structure, hybridization region). A 3' tracrRNA sequence may contain two double-stranded regions.

[0339] 3' tracrRNA sequences can also be called intermediate tracrRNA (see [link]). Figure 1B The intermediate tracrRNA sequence may contain stem-loop structures. In other words, the intermediate tracrRNA sequence may contain, for example, stem-loop structures. Figure 1BThe hairpin structure depicted is different from the second or third stem. The stem-loop structure in this intermediate tracrRNA (i.e., 3' tracrRNA) may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 or more nucleotides. The stem-loop structure in this intermediate tracrRNA (i.e., 3' tracrRNA) may contain up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotides. The stem-loop structure may contain functional parts. For example, the stem-loop structure may contain aptamers, ribozymes, hairpin structures for protein-protein interactions, CRISPR arrays, introns, and exons. The stem-loop structure may contain at least about 1, 2, 3, 4, or 5 or more functional parts. The stem-loop structure may contain up to about 1, 2, 3, 4, or 5 or more functional parts.

[0340] The hairpin structure in the intermediate tracrRNA sequence may contain a P-domain. This P-domain may be included within the double-stranded region of the hairpin structure.

[0341] tracrRNA overhang sequence

[0342] The overhang sequence of tracrRNA can provide stability to nucleic acids targeting nucleic acids and / or provide modification sites. The overhang sequence of tracrRNA can be approximately 1 nucleotide to approximately 400 nucleotides in length. The overhang sequence of tracrRNA can be longer than 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400 or more nucleotides. The overhang sequence of tracrRNA can be approximately 20 to approximately 5000 or more nucleotides in length. The overhang sequence of tracrRNA can be longer than 1000 nucleotides. The overhanging sequence of tracrRNA can be less than 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, or 400 nucleotides in length. The overhanging sequence of tracrRNA can be less than 1000 nucleotides in length. The overhanging sequence of tracrRNA can be less than 10 nucleotides in length. The overhanging sequence of tracrRNA can be 10 to 30 nucleotides in length. The overhanging sequence of tracrRNA can be 30 to 70 nucleotides in length.

[0343] The overhanging sequence of the tracrRNA may contain a portion (e.g., a stability control sequence, a ribozyme, or a ribonuclease-binding sequence). This portion can affect the stability of the nucleic acid targeting the RNA. The portion may be a transcription terminator fragment (i.e., a transcription termination sequence). The portion of the nucleic acid targeting the RNA may have a total length of approximately 10 nucleotides to approximately 100 nucleotides, approximately 10 nucleotides (nt) to approximately 20 nucleotides, approximately 20 nucleotides to approximately 30 nucleotides, approximately 30 nucleotides to approximately 40 nucleotides, approximately 40 nucleotides to approximately 50 nucleotides, approximately 50 nucleotides to approximately 60 nucleotides, approximately 60 nucleotides to approximately 70 nucleotides, approximately 70 nucleotides to approximately 80 nucleotides, approximately 80 nucleotides to approximately 90 nucleotides, or approximately 90 nucleotides to approximately 100 nucleotides, approximately 15 nucleotides (nt) to approximately 80 nucleotides, approximately 15 nucleotides to approximately 50 nucleotides, approximately 15 nucleotides to approximately 40 nucleotides, approximately 15 nucleotides to approximately 30 nucleotides, or approximately 15 nucleotides to approximately 25 nucleotides. This portion may be a functional part in eukaryotic cells. In some cases, this part can be the part that functions in prokaryotic cells. This part can function in both eukaryotic and prokaryotic cells.

[0344] Non-limiting examples of suitable tracrRNA overhangs include: 3' poly-adenosyl tails, riboswitch sequences (e.g., to allow regulated stability and / or accessibility regulated by proteins and protein complexes), sequences forming dsRNA duplexes (i.e., hairpin structures), sequences targeting RNA to subcellular locations (e.g., nuclei, mitochondria, chloroplasts, etc.), modifications or sequences providing tracking (e.g., direct conjugation to fluorescent molecules, conjugation to portions promoting fluorescence detection, sequences allowing fluorescence detection, etc.), modifications or sequences providing binding sites for proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.), modifications or sequences providing enhanced, degraded, and / or controllable stability, or any combination thereof. The tracrRNA overhang sequence may contain primer binding sites or molecular indicators (e.g., barcode sequences). In some embodiments of this disclosure, the tracrRNA overhang sequence may contain one or more affinity markers.

[0345] Single-target nucleic acid

[0346] The nucleic acid targeting a nucleic acid can be a single-target nucleic acid. This single-target nucleic acid can be RNA. The single-target nucleic acid can contain a linker (i.e., derived from...) between the minimal CRISPR repeat sequence and the minimal TRACR RNA sequence. Figure 1A Entry 120), which can be referred to as a single-guided connector sequence.

[0347] The single-director linker of a single-director nucleic acid can be from about 3 nucleotides to about 100 nucleotides in length. For example, the linker can be from about 3 nucleotides (nt) to about 90 nt, from about 3 nt to about 80 nt, from about 3 nt to about 70 nt, from about 3 nt to about 60 nt, from about 3 nt to about 50 nt, from about 3 nt to about 40 nt, from about 3 nt to about 30 nt, from about 3 nt to about 20 nt, or from about 3 nt to about 10 nt in length. For example, the adapter may have a length of approximately 3 nt to approximately 5 nt, approximately 5 nt to approximately 10 nt, approximately 10 nt to approximately 15 nt, approximately 15 nt to approximately 20 nt, approximately 20 nt to approximately 25 nt, approximately 25 nt to approximately 30 nt, approximately 30 nt to approximately 35 nt, approximately 35 nt to approximately 40 nt, approximately 40 nt to approximately 50 nt, approximately 50 nt to approximately 60 nt, approximately 60 nt to approximately 70 nt, approximately 70 nt to approximately 80 nt, approximately 80 nt to approximately 90 nt, or approximately 90 nt to approximately 100 nt. In some embodiments, the adapter of the single-guided nucleic acid may be 4 to 40 nucleotides. The adapter may have a length of at least approximately 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, or 7000 or more nucleotides. The adapter can have a length of up to approximately 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, or 7000 or more nucleotides.

[0348] The adapter sequence may contain functional parts. For example, the adapter sequence may contain aptamers, ribozymes, hairpin structures for protein-protein interactions, CRISPR arrays, introns, and exons. The adapter sequence may contain at least about 1, 2, 3, 4, or 5 or more functional parts.

[0349] In some implementations, the single-guide connector can link the 3' end of the smallest CRISPR repeat fragment to the 5' end of the smallest tracrRNA sequence. Alternatively, the single-guide connector can link the 3' end of the tracrRNA sequence to the 5' end of the smallest CRISPR repeat fragment. That is, the single-guide nucleic acid may contain a 5' DNA-binding fragment linked to a 3' protein-binding fragment.

[0350] The nucleic acid targeting the target nucleic acid may include a spacer region overhang sequence of 10-5000 nucleotides in length; a spacer region sequence of 12-30 nucleotides in length, wherein the spacer region is at least 50% complementary to the target nucleic acid; a minimal CRISPR repeat fragment containing at least 60% identity with crRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages in 6, 7, or 8 adjacent nucleotides, wherein the minimal CRISPR repeat fragment is 5-30 nucleotides in length; and a minimal tracrRNA containing at least 60% identity with tracrRNA from bacteria (e.g., Streptococcus pyogenes) in 6, 7, or 8 adjacent nucleotides. The sequence comprises: a minimal tracrRNA sequence having a length of 5-30 nucleotides; a linker sequence connecting the minimal CRISPR repeat and the minimal tracrRNA and containing a length of 3-5000 nucleotides; a 3' tracrRNA containing at least 60% identity in 6, 7, or 8 adjacent nucleotides to a tracrRNA from a prokaryote (e.g., Streptococcus pyogenes) or bacteriophage, wherein the 3' tracrRNA contains a length of 10-20 nucleotides and includes a double-stranded region; and / or a tracrRNA overhang containing a length of 10-5000 nucleotides, or any combination thereof. Such a nucleic acid targeting a nucleic acid may be referred to as a nucleic acid-targeted single-directed nucleic acid.

[0351] Nucleic acid targeting a nucleic acid may contain a spacer region overhang sequence of 10-5000 nucleotides in length; a spacer region sequence of 12-30 nucleotides in length, wherein the spacer region is at least 50% complementary to the target nucleic acid; and a duplex comprising: 1) a minimal CRISPR repeat fragment containing at least 60% identity in 6 adjacent nucleotides to a crRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages, wherein the minimal CRISPR repeat fragment is 5-30 nucleotides in length; 2) a minimal tracrRNA sequence containing at least 60% identity in 6 adjacent nucleotides to a tracrRNA from bacteria (e.g., Streptococcus pyogenes), wherein the minimal tracrRNA sequence is 5-30 nucleotides in length; and 3) a protrusion, wherein the protrusion contains at least 3 unpaired nucleotides on the minimal CRISPR repeat strand of the duplex and on the minimal tracrRNA of the duplex. At least one unpaired nucleotide on the crRNA sequence strand; a linker sequence connecting the minimum CRISPR repeat and the minimum tracrRNA and containing a length of 3-5000 nucleotides; a 3' tracrRNA containing at least 60% identity with tracrRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages in 6 adjacent nucleotides, wherein the 3' tracrRNA contains a length of 10-20 nucleotides and contains a double-stranded region; a P-domain starting 1-5 nucleotides downstream of the double strand containing the minimum CRISPR repeat and the minimum tracrRNA, containing 1-10 nucleotides, containing a sequence that can hybridize with a motif adjacent to the pre-interstitial region sequence in the target nucleic acid, forming a hairpin structure and located in the 3' tracrRNA region; and / or a tracrRNA overhang containing a length of 10-5000 nucleotides, or any combination thereof.

[0352] A dual-target nucleic acid can be a nucleic acid that targets a specific nucleic acid. This dual-target nucleic acid can be RNA. It can contain two separate nucleic acid molecules (i.e., polynucleotides). Each of the two nucleic acid molecules in a dual-target nucleic acid can contain a nucleotide segment that can hybridize with each other, allowing the complementary nucleotides of the two nucleic acid molecules to hybridize and form a double-stranded protein-binding fragment. Unless otherwise specified, the term "nucleic acid targeting a specific nucleic acid" can be inclusive, referring to both single-molecule and bimolecular nucleic acids targeting a specific nucleic acid.

[0353] A dual-target nucleic acid may comprise: 1) a first nucleic acid molecule containing a spacer region overhang sequence of 10-5000 nucleotides in length; a spacer region sequence of 12-30 nucleotides in length, wherein the spacer region is at least 50% complementary to the target nucleic acid; and a minimum CRISPR repeat fragment containing at least 60% identity with crRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages in 6 adjacent nucleotides, wherein the minimum CRISPR repeat fragment has a length of 5-30 nucleotides; and 2) a second nucleic acid molecule of the dual-target nucleic acid may comprise a spacer region sequence of 10-5000 nucleotides in length, wherein the spacer region is at least 50% complementary to the target nucleic acid; and a minimum CRISPR repeat fragment containing at least 60% identity with crRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages in 6 adjacent nucleotides. The minimal tracrRNA sequence having at least 60% identity with tracrRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages and wherein the minimal tracrRNA sequence has a length of 5-30 nucleotides; a 3' tracrRNA containing at least 60% identity with tracrRNA from bacteria (e.g., Streptococcus pyogenes) in 6 adjacent nucleotides and wherein the 3' tracrRNA contains a length of 10-20 nucleotides and contains a double-stranded region; and / or a tracrRNA overhang containing a length of 10-5000 nucleotides, or any combination thereof.

[0354] In some cases, a dual-target nucleic acid targeting nucleic acid may comprise: 1) a first nucleic acid molecule containing a spacer region protruding sequence of 10-5000 nucleotides in length; a spacer region sequence of 12-30 nucleotides in length, wherein the spacer region is at least 50% complementary to the target nucleic acid; a minimum CRISPR repeat fragment containing at least 60% identity in 6 adjacent nucleotides to a crRNA from a prokaryote (e.g., Streptococcus pyogenes) or bacteriophage, wherein the minimum CRISPR repeat fragment is 5-30 nucleotides in length, and at least 3 protruding unpaired nucleotides; and 2) a second nucleic acid molecule of the dual-target nucleic acid targeting nucleic acid may comprise a minimum tracrRNA sequence having at least 60% identity in 6 adjacent nucleotides to a tracrRNA from a prokaryote (e.g., Streptococcus pyogenes) or bacteriophage, wherein the minimum tracrRNA sequence is 5-30 nucleotides in length. The following are considered as separate components: a length and a protrusion of at least one unpaired nucleotide, wherein the unpaired nucleotide of the protrusion is located in a protrusion identical to the three unpaired nucleotides of the minimum CRISPR repeat; a 3' tracrRNA containing at least 60% identity in six adjacent nucleotides to a tracrRNA from a prokaryote (e.g., Streptococcus pyogenes) or bacteriophage, wherein the 3' tracrRNA contains a length of 10-20 nucleotides and includes a double-stranded region; a region containing 1-5 nucleotides downstream of the double-stranded region containing the minimum CRISPR repeat and the minimum tracrRNA, containing 1-10 nucleotides, including a sequence that can hybridize with a motif adjacent to the pre-interstitial region sequence in the target nucleic acid, a P-domain that can form a hairpin structure and is located in the 3' tracrRNA region; and / or a tracrRNA protrusion of 10-5000 nucleotides in length, or any combination thereof.

[0355] Nucleic acid and site-directed peptide complexes targeting nucleic acids

[0356] Nucleic acids targeting nucleic acids can interact with site-directed peptides (e.g., nucleic acid-guided nucleases, Cas9) to form complexes. These nucleic acids can then direct the site-directed peptides to the target nucleic acid.

[0357] In some implementations, the nucleic acid targeting the nucleic acid can be engineered so that the complex (e.g., comprising a site-directed peptide and the nucleic acid targeting the nucleic acid) can bind outside the cleavage site of the site-directed peptide. In this case, the target nucleic acid may not interact with the complex, and the target nucleic acid may be cleavable (e.g., detached from the complex).

[0358] In some implementations, the target nucleic acid can be engineered so that the complex can bind within the cleavage site of the site-directed peptide. In this case, the target nucleic acid can interact with the complex, and the target nucleic acid can be bound (e.g., bound to the complex).

[0359] In some cases, the complex may contain a site-directed polypeptide, wherein the site-directed polypeptide may contain an amino acid sequence containing at least 15% amino acid identity with Cas9 from Streptococcus pyogenes, and two nucleic acid cleavage domains (i.e., an HNH domain and a RuvC domain); and a dual-target nucleic acid comprising: 1) a first nucleic acid molecule containing a spacer region overhang sequence of 10-5000 nucleotides in length; a spacer region sequence of 12-30 nucleotides in length, wherein the spacer region is at least 50% complementary to the target nucleic acid; and a minimal CRISPR repeat fragment containing at least 60% identity in 6 adjacent nucleotides with crRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages, wherein the minimal CRISPR repeat fragment has 5-30 nucleotides in length. The length of the second nucleic acid molecule of the nucleic acid-targeted dual-guided nucleic acid may include: a minimum tracrRNA sequence containing at least 60% identity in 6 adjacent nucleotides with a tracrRNA from a prokaryote (e.g., Streptococcus pyogenes) or bacteriophage, wherein the minimum tracrRNA sequence has a length of 5-30 nucleotides; a 3' tracrRNA containing at least 60% identity in 6 adjacent nucleotides with a tracrRNA from a bacterium (e.g., Streptococcus pyogenes), wherein the 3' tracrRNA contains a length of 10-20 nucleotides and includes a double-stranded region; and / or a tracrRNA overhang of 10-5000 nucleotides, or any combination thereof.

[0360] In some cases, the complex may comprise a site-directed polypeptide containing an amino acid sequence with at least 15% amino acid identity to Cas9 from Streptococcus pyogenes and two cleavage domains (i.e., an HNH domain and a RuvC domain); and a dual-target nucleic acid comprising: 1) a first nucleic acid molecule containing a spacer region protruding at least 10-5000 nucleotides in length; a spacer region sequence of 12-30 nucleotides in length, wherein the spacer region is at least 50% complementary to the target nucleic acid; a minimum CRISPR repeat fragment containing at least 60% identity to crRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages in 6 adjacent nucleotides, wherein the minimum CRISPR repeat fragment is 5-30 nucleotides in length, and at least 3 protruding unpaired nucleotides; and 2) a second nucleic acid molecule of the dual-target nucleic acid comprising at least 60% identity to tracrRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages in 6 adjacent nucleotides. A single minimal tracrRNA sequence having a length of 5-30 nucleotides and at least one unpaired nucleotide of a protrusion, wherein the unpaired nucleotide of the protrusion is located in a protrusion identical to the three unpaired nucleotides of the minimal CRISPR repeat; a 3' tracrRNA containing at least 60% identity with a tracrRNA from a prokaryote (e.g., Streptococcus pyogenes) or bacteriophage within 6 adjacent nucleotides, wherein the 3' tracrRNA contains a length of 10-20 nucleotides and contains a double-stranded region; a 1-5 nucleotide downstream of the double strand containing the minimal CRISPR repeat and the minimal tracrRNA, containing 1-10 nucleotides, containing a sequence that can hybridize with a motif adjacent to the pre-interstitial region sequence in the target nucleic acid, a P-domain that can form a hairpin structure and is located in the 3' tracrRNA region; and / or a tracrRNA protrusion of 10-5000 nucleotides, or any combination thereof.

[0361] In some cases, the complex may contain a site-directed polypeptide, wherein the site-directed polypeptide may contain an amino acid sequence containing at least 15% amino acid identity with Cas9 from Streptococcus pyogenes and two nucleic acid cleavage domains (i.e., the HNH domain and the RuvC domain); and a nucleic acid targeting the nucleic acid, comprising a spacer region overhang sequence of 10-5000 nucleotides in length; a spacer region sequence of 12-30 nucleotides in length, wherein the spacer region is at least 50% complementary to the target nucleic acid; a minimum CRISPR repeat fragment containing at least 60% identity with crRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages in 6, 7, or 8 adjacent nucleotides, wherein the minimum CRISPR repeat fragment is 5-30 nucleotides in length; and a minimum CRISPR repeat fragment containing 6, 7, or 8 adjacent nucleotides. The following are nucleic acids: a minimal tracrRNA sequence that is at least 60% identical to a tracrRNA from bacteria (e.g., Streptococcus pyogenes) and has a length of 5-30 nucleotides; a linker sequence connecting the minimal CRISPR repeat and the minimal tracrRNA and containing a length of 3-5000 nucleotides; a 3' tracrRNA containing at least 60% identity to a tracrRNA from prokaryotes (e.g., Streptococcus pyogenes) or bacteriophages in 6, 7, or 8 a...

Claims

1. An engineered single-target nucleic acid that targets nucleic acids and forms a complex with a type II CRISPR Cas9 protein, said engineered single-target nucleic acid comprising a spacer region sequence capable of hybridizing with a target sequence, and an sg backbone encoded by any one of the following sequences (i) to (iv): (i)GTTTTAGAGGAAACAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCT; (ii)GTTTTAGAGACAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCT; (iii)GTTTTAGGAGAAACTTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCT; (iv)GTTTTAGATACTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGT GGCACCGAGTCGGTGCT, The spacer sequence is connected to the 5' end of the sg skeleton sequence.

2. The engineered nucleic acid-targeted single-direction nucleic acid as described in claim 1, wherein the sg backbone sequence further comprises modifications selected from base modifications, backbone modifications, and modified nucleoside inter-bonds.

3. The engineered single-target nucleic acid polynucleotide of claim 1 or 2.

4. The polynucleotide of claim 3 further comprises a promoter operatively linked to the polynucleotide.

5. A composition comprising: The engineered, nucleic acid-targeting, single-guided nucleic acid as described in claim 1 or 2; and Cas9 protein, The engineered, nucleic acid-targeting single-target nucleic acid and the Cas9 protein form a complex.

6. The composition of claim 5, wherein the Cas9 protein is the Cas9 protein of Streptococcus pyogenes.

7. An in vitro method for cleaving target nucleic acids, the method comprising: Contacting the nucleic acid containing the target nucleic acid with the composition of claim 5 promotes the binding of the complex to the target nucleic acid, resulting in the cleavage of the target nucleic acid.

8. An in vitro method for binding target nucleic acids, the method comprising: Contacting the nucleic acid containing the target nucleic acid with the composition of claim 5 promotes the binding of the complex to the target nucleic acid.

9. A kit comprising: The engineered nucleic acid-targeting single-target nucleic acid as described in claim 1 or 2, or the polynucleotide encoding the engineered nucleic acid-targeting single-target nucleic acid; and Buffer solution.

10. The kit of claim 9, further comprising: Cas9 protein or polynucleotides encoding Cas9 protein.