Engineered crisper-cas9 nucleases with altered pam specificity

CN107532161BActive Publication Date: 2026-09-29THE GENERAL HOSPITAL CORP
View PDF 58 Cites 0 Cited by

Patent Information

Application Number
CN201680024041.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-11-20
Filing Date
2016-03-03
Publication Date
2026-09-29
Estimated Expiration
2036-03-03

AI Technical Summary

Technical Problem

然而,Cas9还必须识别位于与sgRNA碱基配对的DNA的近端的特异性原型间隔子邻近基序(PAM)(Mojica等人,Microbiology[微生物学]155,733-740(2009);Shah等人,RNA Biol[RNA生物学]10,891-899(2013);Jinek等人,Science[科学]337,816-821(2012);Sapranauskas等人,Nucleic Acids Res[核酸研究]39,9275-9282(2011);Horvath等人,J Bacteriol[细菌学杂志]190,1401-1412(2008)),这是启动序列特异性识别所需的要求(Sternberg等人,Nature[自然]507,62-67(2014)),但也可以限制这些核酸酶用于基因组编辑的靶向范围

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN107532161B_ABST
    Figure CN107532161B_ABST
Patent Text Reader

Abstract

Engineered crispr-cas9 nucleases with altered and improved pam specificity and uses thereof in genome engineering, epigenetic engineering, and genome targeting.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Statement

[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 61 / 127,634, filed March 3, 2015; 62 / 165,517, filed May 22, 2015; 62 / 239,737, filed October 9, 2015; and 62 / 258,402, filed November 20, 2015. The entire contents of these documents are incorporated herein by reference.

[0003] Federally funded research or development

[0004] This invention was completed with government support, based on authorization numbers DP1 GM105378, NIH R01 GM107427, and R01 GM088040 granted by the National Institutes of Health. The government enjoys certain rights in this invention. Technical Field

[0005] This invention relates at least in part to engineered regular spacer clustered short palindromic repeat (CRISPR) / CRISPR-associated protein 9 (Cas9) nucleases with modified and improved prototype spacer neighbor motif (PAM) specificity and their use in genome engineering, epigenome engineering and genome targeting. Background Technology

[0006] CRISPR-Cas9 nucleases enable efficient and customizable genome editing in a wide variety of organisms and cell types (Sander & Joung, Nat Biotechnol 32, 347-355 (2014); Hsu et al., Cell 157, 1262-1278 (2014); Doudna & Charpentier, Science 346, 1258096 (2014); Barrangou & May, Expert Opinion on Biotherapy 15, 311-314 (2015)). Cas9 target site recognition is guided by two short RNAs called crRNA and tracrRNA (Deltcheva et al., Nature 471, 602-607 (2011); Jinek et al., Science 337, 816-821 (2012)), which can fuse into a chimeric single guide RNA (sgRNA) (Jinek et al., Science 337, 816-821 (2012); Jinek et al., Elife 2, e00471 (2013); Mali et al., Science 339, 823-826 (2013); Cong et al., Science 339, 819-823 (2013)). The 5' end of sgRNA (derived from crRNA) can pair with the bases of the target DNA site, thus allowing direct reprogramming of site-specific cleavage via the Cas9 / sgRNA complex (Jinek et al., Science 337, 816-821 (2012)). However, Cas9 must also recognize specific prototypical spacer neighbor motifs (PAMs) located near the proximal end of the DNA that pairs with the sgRNA bases (Mojica et al., Microbiology 155, 733-740 (2009); Shah et al., RNA Biol 10, 891-899 (2013); Jinek et al., Science 337, 816-821 (2012); Sapranauskas et al., Nucleic Acids Research 39, 9275-9282 (2011); Horvath et al., J Bacteriol 190, 1401-1412 (2008)), which is a requirement for initiating sequence-specific recognition (Sternberg et al., Nature 507, 62-67 (2014)), but may also limit the target range of these nucleases for genome editing.The widely used Streptococcus pyogenes Cas9 (SpCas9) recognizes short NGG PAM (Jinek et al., Science 337, 816-821 (2012); Jiang et al., Nat Biotechnol 31, 233-239 (2013)), occurring once every 8 bp of random DNA sequence. In contrast, other Cas9 orthologs characterized to date can recognize longer PAMs (Horvath et al., J Bacteriol 190, 1401-1412 (2008); Fonfara et al., Nucleic Acids Res 42, 2577-2590 (2014); Esvelt et al., Nat Methods 10, 1116-1121 (2013); Ran et al., Nature 520, 186-191 (2015); Zhang et al., Mol Cell 50, 488-503 (2013)). For example, Staphylococcus aureus Cas9 (SaCas9), one of several smaller Cas9 orthologs more suited to viral delivery (Horvath et al., J Bacteriol 190, 1401-1412 (2008); Ran et al., Nature 520, 186-191 (2015); Zhang et al., Mol Cell 50, 488-503 (2013)), recognizes the longer NNGRRT (SEQ ID NO:46) PAM, which is expected to occur once every 32 bp of random DNA. Expanding the targeting range of Cas9 orthologs is important for a variety of applications, including modifying small genetic elements (e.g., transcription factor binding sites (Canver et al., Nature; 527(7577):192-7(2015); Vierstra et al., Nat Methods 12(10):927-30(2015)) or performing allele-specific alterations by localizing sequence differences within PAMs (Courtney, DG et al., Gene Ther. 23(1):108-12(2015)). Summary of the Invention

[0007] As described herein, the commonly used *Streptococcus pyogenes* Cas9 (SpCas9) and *Staphylococcus aureus* Cas9 (SaCas9) were engineered to recognize novel PAM sequences using structural information, directed evolution based on bacterial selection, and combinatorial design. These altered PAM-specific variants enabled robust editing of endogenous gene sites in zebrafish and human cells that were not effectively targeted by wild-type SpCas9 or SaCas9. Furthermore, we identified and characterized another SpCas9 variant that exhibited improved PAM specificity in human cells with reduced activity at sites with atypical NAG and NGA PAMs. Additionally, we found that two smaller Cas9 orthologs with completely different PAM specificities, *Streptococcus thermophilus* Cas9 (St1Cas9) and *Staphylococcus aureus* Cas9 (SaCas9), functioned effectively in our bacterial selection system and human cells, suggesting that our engineering strategy can be extended to Cas9 from other species. Our results provide a wide range of useful SpCas9 and SaCas9 variants, collectively referred to herein as “variants” or “these variants”.

[0008] In a first aspect, the present invention provides isolated Streptococcus pyogenes Cas9 (SpCas9) protein having mutations at one or more of the following locations: G1104, S1109, L1111, D1135, S1136, G1218, N1317, R1335, T1337, for example, comprising a sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO:1. In some embodiments, the variant SpCas9 comprises one or more of the following mutations: G1104K; S1109T; L1111H; D1135V; D1135E; D1135N; D1135Y; S1136N; G1218R; N1317K; R1335E; R1335Q; and T1337R. In some embodiments, the SpCas9 variant contains the following mutations: D1135; D1135V / R1335Q / T1337R (VQR variant); D1135E / R1335Q / T1337R (EQR variant); D1135V / G1218 / R1335Q / T1337R (VRQR variant); D1135N / G1218R / R1335Q / T1337R (NRQR variant); D1135Y / G1218R / R1335Q / T1337R (YRQR variant); G1104K / D1135V / G1218R / R1335Q / T133 7R (KVRQR variant); S1109T / D1135V / G1218R / R1335Q / T1337R (TVRQR variant); L1111H / D1135V / G1218R / R1335Q / T1337R (HVRQR variant); D1135V / S1136N / G1218R / R1335Q / T1337R (VNRQR variant); D1135V / G1218R / N1317K / R1335Q / T1337R (VRKQR variant); or D1135V / G1218R / R1335E / T1337R (VRER variant).

[0009] In some embodiments, the variant SpCas9 protein contains one or more mutations selected from the group consisting of the following: mutations at D10, E762, D839, H983, or D986; and mutations at H840 or N863 that reduce nuclease activity.

[0010] In some embodiments, the mutations are: (i) D10A or D10N and (ii) H840A, H840N, or H840Y.

[0011] Isolated Staphylococcus aureus Cas9 (SaCas9) proteins with mutations at one or more of the following locations are also provided herein: E782, N968, and / or R1015, for example, comprising a sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO:2. Isolated Staphylococcus aureus Cas9 (SaCas9) proteins with mutations at two or more of the following locations are also provided herein: E735, E782, K929, N968, A1021, K1044, and / or R1015. In some embodiments, the variant SaCas9 protein comprises one or more of the following mutations: R1015Q, R1015H, E782K, N968K, E735K, K929R, A1021T, K1044N. In some embodiments, the variant SaCas9 protein contains one or more mutations selected from the group consisting of mutations at D10, D556, H557 and / or N580 that reduce nuclease activity.

[0012] In some embodiments, the variant SaCas9 protein contains mutations at D10A, D556A, H557A, and N580A, for example, D10A / H557A and / or D10A / D556A / H557A / N580A.

[0013] The SpCas9 variants described herein may include the amino acid sequence of SEQ ID NO:1 with mutations at one or more of the following positions: D1135, G1218, R1335, T1337. In some embodiments, the SpCas9 variants may include one or more of the following mutations: D1135V; D1135E; G1218R; R1335E; R1335Q; and T1337R. In some embodiments, the SpCas9 variants may include one of the following group of mutations: D1135V / R1335Q / T1337R (VQR variant); D1135V / G1218R / R1335Q / T1337R (VRQR variant); D1135E / R1335Q / T1337R (EQR variant); or D1135V / G1218R / R1335E / T1337R (VRER variant).

[0014] The SaCas9 variants described herein may include the amino acid sequence of SEQ ID NO:2 with mutations at one or more of the following positions: E735, E782, K929, N968, R1015, A1021, and / or K1044. In some embodiments, the SaCas9 variants may include one or more of the following mutations: R1015Q, R1015H, E782K, N968K, E735K, K929R, A1021T, K1044N. In some embodiments, the SaCas9 variants may include a mutation from the following group: E782K / N968K / R1015H (KKH variant); E782K / K929R / R1015H (KRH variant); or E782K / K929R / N968K / R1015H (KRKH variant).

[0015] This document also provides fusion proteins comprising isolated variants of SaCas9 or SpCas9 protein fused with an optional intervening linker to a heterologous functional domain, wherein the linker does not interfere with the activity of the fusion protein. In some embodiments, the heterologous functional domain is a transcriptional activation domain. In some embodiments, the transcriptional activation domain is derived from VP64 or NF-κBp65. In some embodiments, the heterologous functional domain is a transcriptional silencer or a transcriptional repression domain. In some embodiments, the transcriptional repression domain is a Krueppel-associated box (KRAB) domain, an ERF repression factor domain (ERD), or an mSin3A interacting domain (SID). In some embodiments, the transcriptional silencer is heterochromatin protein 1 (HP1), such as HP1α or HP1β. In some embodiments, the heterologous functional domain is an enzyme that modifies the methylation state of DNA. In some embodiments, the enzyme that modifies the methylation state of DNA is a DNA methyltransferase (DNMT) or a TET protein. In some embodiments, the TET protein is TET1. In some embodiments, the heterologous functional domain is an enzyme that modifies histone subunits. In some embodiments, the enzyme modifying the histone subunit is a histone acetyltransferase (HAT), histone deacetylase (HDAC), histone methyltransferase (HMT), or histone demethylase. In some embodiments, the heterologous functional domain is a biological chain. In some embodiments, the biological chain is MS2, Csy4, or λN protein. In some embodiments, the heterologous functional domain is FokI.

[0016] Also provided herein are isolated nucleic acids encoding the variant SaCas9 or SpCas9 protein described herein, and vectors comprising the isolated nucleic acid optionally operably linked to one or more regulatory domains for expression of the variant SaCas9 or SpCas9 protein described herein. Host cells, such as mammalian host cells, containing the nucleic acid described herein and optionally expressing the variant SaCas9 or SpCas9 protein described herein are also provided herein.

[0017] Methods for altering the cellular genome are also provided herein by expressing in the cell the isolated variant SaCas9 or SpCas9 protein described herein, along with a guide RNA having a region complementary to a selected portion of the cellular genome.

[0018] Methods for altering (e.g., selectively altering) the genome of a cell by expressing variant proteins in the cell and by having guide RNAs having regions complementary to selected portions of the cell’s genome.

[0019] Methods are also provided for altering (e.g., selectively altering) the cell genome by contacting the cell with the protein variants described herein and with a guide RNA having a region complementary to a selected portion of the cell's genome.

[0020] In some embodiments, the isolated protein or fusion protein comprises one or more of a nuclear localization sequence, a cell-penetrating peptide sequence, and / or an affinity tag.

[0021] In some embodiments of the methods described herein, the cells are stem cells, such as embryonic stem cells, mesenchymal stem cells, or induced pluripotent stem cells; in a living animal; or in an embryo, such as a mammalian, insect, or fish (e.g., zebrafish) embryo or embryonic cell.

[0022] Furthermore, methods for altering double-stranded DNA (dsDNA) molecules, such as in vitro methods, are provided herein. These methods involve contacting the dsDNA molecule with one or more variant proteins described herein and a guide RNA having a region complementary to a selected portion of the dsDNA molecule.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Methods and materials used in this invention are described herein; other suitable methods and materials known in the art may also be used. These materials, methods, and examples are illustrative only and are not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated herein by reference in their entirety. In case of conflict, this specification (including definitions) shall prevail.

[0024] Other features and advantages of the invention will be apparent from the following detailed description and drawings, as well as from the claims. Attached Figure Description

[0025] This patent or application document contains at least one color drawing. Upon request and payment of the necessary fees, the official authority will provide a copy of this patent or patent application publication with one or more color drawings.

[0026] Figure 1A -J| Evolution and characterization of SpCas9 variants with altered PAM specificity. A, Reasonable mutations in SpCas9 residues that form base-specific contacts with PAM bases are insufficient to alter PAM specificity in an enhanced green fluorescent protein (EGFP) disruption assay based on U2OS human cells. Disruption frequency was quantified by flow cytometry; for this and subsequent plots (C, G, H, and J), the mean degree of disruption observed with background controls is indicated by a red dashed line; error bars indicate sem, n=3. B, Schematic diagram of a dual-plasmid positive selection assay used to alter PAM specificity of SpCas9. When bacteria are plated on selective media, cleavage of the target site within the positive selection plasmid by a functional Cas9 / sgRNA complex is essential for survival (see also...). Figure 12A -B). C, Assembly and testing of mutant combinations obtained from positive selection of SpCas9 variants that can cleave target sites containing NGA PAM. SpCas9 variants were paired with sgRNAs targeting sites containing NGG or NGA PAM, and activity was assessed using an EGFP disruption assay. Error bars represent sem, n=3. D, Schematic diagram of a negative selection assay where cleavage of the selection plasmid resulted in cell death when bacteria were plated on selective media. This system was adapted to analyze the PAM specificity of Cas9 by generating plasmid libraries containing randomized sequences adjacent to the 3' end of the prototype spacer (see also...). Figure 13B E, a scatter plot of post-selection PAM depletion values ​​(PPDV) for wild-type SpCas9 with two randomized PAM libraries (each with a different prototype spacer). PAMs are grouped and plotted according to their 2nd / 3rd / 4th positions. The red dashed line indicates the statistically significant depletion cutoff value (obtained from the dCas9 control experiment, see [link]). Figure 13C), and the gray dashed line represents five-fold depletion (PPDV = 0.2). F, PPDV scatter plot for VQR and EQR SpCas9 variants used to identify PAMs different from those recognized by wild-type SpCas9. G, EGFP destruction frequency of wild-type, VQR, and EQR SpCas9 at sites with NGAN and NGNG PAMs. Error bars represent sem, n = 3. H, Combinatorial assembly and testing of mutations obtained from positive selection of SpCas9 variants that can cleave target sites containing NGCG PAM. EGFP destruction assays were used to evaluate sgRNAs targeting sites containing NGGG or NGCG PAMs for Cas9 targeting. Error bars represent sem, n = 3. I, PPDV scatter plot for VRER variants. J, EGFP destruction frequency of wild-type and VRER SpCas9 at sites with NGCN and NGNG PAMs. Error bars represent sem, n = 3.

[0027] Figure 2 | SpCas9 variants with evolved PAM specificity robustly modify endogenous sites in zebrafish embryos and human cells. A, Quantification of wild-type or VQR SpCas9-induced mutagenesis frequencies at endogenous gene sites carrying NGAG PAM in zebrafish embryos. Mutation frequencies were determined using T7E1 assay; error bars represent sem, n = 5 to 9 individual embryos. B, Mutation frequencies of VQR variants quantified by T7E1 at 16 target sites in four endogenous human genes with sgRNAs targeting sites containing NGAG, NGAT, and NGAA PAMs. Error bars represent sem, n = 3. C, Mutation frequencies of wild-type SpCas9 at endogenous human gene target sites with NGA PAM. For ease of comparison, mutation frequencies of VQR variants using the same sgRNA are presented again here (same data shown in Figure B). Error bars represent sem, n = 3; nd, not detectable by T7E1. D. Mutation frequencies of wild-type, VRER, and VQR SpCas9 at nine target sites containing NGCGPAM in three endogenous human genes, quantified by T7E1 assay. The complementary lengths of the sgRNAs used were 19 and 20 nt; error bars represent sem, n=3. E. Representation of the number of sites in the human genome with 20 nt spacers that can be targeted by wild-type, VQR, and VRER SpCas9. F. Number of off-target cleavage sites identified by GUIDE-seq for VQR and VRER SpCas9 variants using sgRNAs from Figures B and D.

[0028] Figure 3 | A. The D1135E mutation improves PAM recognition and spacer specificity of SpCas9. A, PPDV scatter plots for two randomized PAM libraries, wild-type and D1135E SpCas9 (left and right plots, respectively). PAMs are grouped and plotted according to their 2nd / 3rd / 4th positions. Data for wild-type SpCas9 are shown compared with those from... Figure 1D The same graph is presented here again for ease of comparison. The red dashed line indicates statistically significant depletion of PAM (see Figure 1). Figure 13C (A) and the gray dashed line indicates a five-fold depletion cutoff value (PPDV = 0.2). B, EGFP disruption activity of wild-type and D1135E SpCas9 at sites containing NGG, NAG, and NGA PAM in human cells. Disruption frequency was quantified by flow cytometry; for this figure and (D), the mean degree of disruption observed with background controls is indicated by a red dashed line; error bars represent sem, n = 3; showing the mean fold change in activity. C, Mutagenicity frequency of wild-type and D1135E SpCas9 detected by T7E1 at six endogenous sites in human cells. Error bars represent sem, n = 3; showing the mean fold change in activity. D, Titration of the amount of wild-type or D1135E SpCas9 encoding plasmid transfected for EGFP disruption experiments in human cells. The amount of sgRNA plasmid used for all these experiments was fixed at 250 ng. Two sgRNAs targeting different EGFP sites were used; error bars represent sem, n = 3. E. Target depth sequencing of the mid-target and off-target sites of the three sgRNAs using wild-type and D1135E SpCas9. Mid-target sites are shown at the top, and off-target sites listed below are highlighted to show mismatches with mid-target sites. The fold reduction in D1135E activity at off-target sites relative to wild-type SpCas9 at the fold is greater than the change in activity at the mid-target site, highlighted in green; control indel levels for each amplicon are reported. F. Summary of target depth sequencing data, plotted as the fold reduction in activity at mid-target and off-target sites using D1135E relative to the indel frequency observed with wild-type SpCas9. G. Summary of specificity changes between wild-type and D1135E at off-target sites detected by GUIDE-seq, plotted as the normalized fold change in specificity using D1135E versus the read counts at that off-target site using wild-type SpCas9 (see also...). Figure 18C For D1135E sites where no readings were plotted, the specificity estimate increases by fold (see [reference]). Figure 18C ).

[0029] Figure 4 | Characterization of St1Cas9 and SaCas9 orthogonal homologs in bacterial and human cells. A, Scatter plot of PPDV for St1Cas9 using two randomized PAM libraries. PAMs were grouped and plotted according to their 3rd / 4th / 5th / 6th positions. St1Cas9 was programmed for the two libraries (left and right panels, respectively) using sgRNA complement lengths of 20 and 21 nucleotides. Red dashed lines indicate statistically significantly depleted PAMs (see Figure 4). Figure 13C ), and the gray dashed line indicates five-fold depletion (PPDV of 0.2); α, PAM previously predicted by bioinformatics methods 27 ;β, PAM previously identified under rigorous experimental conditions 20 *, A novel PAM discovered in this study; γ, A PAM previously identified under moderate experimental conditions. 20 B. Scatter plot of PPDV for SaCas9 using two randomized PAM libraries. PAMs were grouped and plotted according to their 3rd / 4th / 5th / 6th positions. SaCas9 was programmed for the two libraries (left and right plots, respectively) using sgRNA complement lengths of 21 and 23 nucleotides. PAMs identified against SaCas9 are shown, across all combinations of spacers and spacer lengths used in these experiments, with PAMs 1–3 consistently depleted. C. Percentage of St1Cas9 and SaCas9 survival in bacterial positive selection when excited with selection plasmids indicating different target sites and PAMs on the x-axis. Highly depleted PAMs against St1Cas9 and SaCas9 used for target sites in positive selection plasmids from Figures (A) and (B). D and E show the EGFP disrupting activity of St1Cas9 (Fig. D) or SaCas9 (Fig. E) at sites in EGFP containing NNAGAA (SEQ ID NO:3) or NNGGGT (SEQ ID NO:4) / NNGAGT (SEQ ID NO:5) PAM, respectively. Matching sgRNAs of different lengths at the same site are indicated; disruption frequency was quantified by flow cytometry; the mean frequency of EGFP disruption obtained with negative controls is represented by a red dashed line; error bars represent sem, n=3. F and G show the mutation frequencies of St1Cas9 (Fig. F) and SaCas9 (Fig. G) at four endogenous human gene sites containing NNAGAA (SEQ ID NO:3) or NNGGGT (SEQ ID NO:4) / NNGAGT (SEQ ID NO:5) / NNGAAT (SEQ ID NO:6) PAM, respectively, quantified by T7E1 assay. Error bars represent sem, n=3; nd indicates undetectable by T7E1.

[0030] Figure 5A -J. Sequences and Maps- Plasmids used in this study

[0031]

[0032]

[0033]

[0034] Figure 6 | Comparison of Cas9 orthologs to predict PAM interaction residues in SaCas9. PAM interaction domains of SpCas9, SaCas9, and 11 other Cas9 orthologs were compared to identify PAM contact residues in SaCas9 based on those known for SpCas9. The amino acids are 1229-1368 of SEQ ID NO:1, SEQ ID NO:29-40, respectively.

[0035] Figure 7 The activity of substitutions in SaCas9 against different PAMs was evaluated in bacterial screening. Based on Figure 6 In comparison, single amino acid substitutions were tested in a bacterial positive selection to screen for their effects on the activity of typical NNGAGT (SEQ ID NO:5) as well as atypical NNAAGT (SEQ ID NO:41) and NNAGGT (SEQ ID NO:42) PAMs. Bacterial colonies on selective media indicated that the SaCas9 variant possessed activity against sites containing the indicative PAM.

[0036] Figure 8A -B| Summary of amino acid substitutions that enable SaCas9 variants to target NNARRT (SEQ ID NO:43) PAM. Amino acid sequences of the PAM interaction domain of 52 selected mutant SaCas9 clones that enable survival in bacteria against sites containing NNARRT (SEQ ID NO:43) PAM; the sequences presented are partial sequences of SEQ ID NO:53-104 shown in Table 6.

[0037] Figure 9 Human cell activity of wild-type and engineered SaCas9 variants. The activity of wild-type, KKQ, and KKH SaCas9 against sites containing NNRRRT (SEQ ID NO:45) PAM was evaluated in human cell EGFP reporter assays.

[0038] Figure 10 SaCas9 activity against atypical PAMs in bacteria, and how directional mutations at R1015 affect activity against the same atypical PAMs.

[0039] Figure 11 Engineered variants can recognize PAM in the NNNRRT form.

[0040] Figure 12A -B| Bacterial-based positive selection for engineered SpCas9 modified PAM-specific variants. A, from Figure 1B A schematic diagram of positive selection (left) and verification of the expected behavior of SpCas9 in positive selection (right). Spacer 1, SEQ ID NO:105; Spacer 2, SEQ ID NO:106. B, Schematic diagram of how positive selection is tuned to select SpCas9 variants with altered PAM recognition specificity. SpCas9 clone libraries with randomized PAM interaction (PI) domains (residues 1097-1368) are activated by a selection plasmid containing altered PAM. SpCas9 variants that survive selection by cleaving the positive selection plasmid are sequenced to identify mutations that alter PAM specificity.

[0041] Figure 13A -D| A site depletion assay based on bacterial cells for analyzing global PAM specificity of Cas9 nucleases. A, showing the results from... Figure 1D A schematic diagram of negative selection (left) and validation of expected wild-type SpCas9 behavior in screening sites with functional (NGG) and non-functional (NGA) PAMs (right). B, Schematic diagram of how negative selection is used to screen sites for functional PAMs by constructing a negative selection plasmid library containing 6 randomized base pairs (instead of PAMs). The selection plasmid containing PAMs cleaved by the Cas9 / sgRNA of interest is depleted, while uncleaved (or poorly cleaved) PAMs are retained. The frequency of the selected PAMs is compared with the preselected frequency in the starting library to calculate the PAM depletion value (PPDV). Spacer 1, SEQ ID NO:105; Spacer 2, SEQ ID NO:106. C, D, Determining statistically significant PPDV cutoff values ​​by plotting the PPDVs of PAMs (grouped and plotted according to their 2nd / 3rd / 4th positions) of SpCas9 (dCas9) that are catalytically inactive to two randomized PAM libraries (C). The threshold of 3.36 standard deviations of the mean PPDV for both libraries was calculated (red line in (D)). Any PPDV deviation below 0.85 was determined to be statistically significant compared to the dCas9 treatment (red dashed line in (C)). The gray dashed line in (D) indicates five-fold depletion in the assay (PPDV of 0.2).

[0042] Figure 14| The concordance between site depletion assays and EGFP destruction activity. Data points represent the average EGFP destruction at two NGAN and NGNG PAM sites in the VQR and EQR SpCas9 variants. Figure 1G ), for libraries 1 and 2, the mean PPDV of the corresponding PAM ( Figure 1F The red dashed line indicates statistically significant depletion of PAM (PPDV of 0.85, see [reference]). Figure 13C The gray dashed line indicates five times exhaustion (PPDV = 0.2). Means are plotted with 95% confidence intervals.

[0043] Figure 15 | Insertion or deletion mutations induced by the VQR SpCas9 variant at endogenous zebrafish loci containing NGAG PAM. For each target locus, the wild-type sequence is shown at the top, the prototype spacer is highlighted in yellow (or green if present on the complementary strand), and PAM is marked as text with a red underline. Deletions are shown as red dashed lines highlighted in gray, and insertions are shown as lowercase letters highlighted in blue. The net change in length caused by each indel mutation is shown on the right (+, insertion; –, deletion). Note that some changes have both insertions and deletions of the sequence, and in these cases, these changes will be enumerated in parentheses. The number of times each mutant allele was restored (if more than once) is shown in parentheses.

[0044] Figure 16A -B|Endogenous genes targeted by wild-type and evolved variants of SpCas9. A, Sequences targeted by wild-type, VQR, and VRERS spCas9 are shown in blue, red, and green, respectively. The sequences of sgRNA and primers used to amplify these loci targeting T7E1 are provided in Tables 1 and 2 below. B, Mean mutagenicity of wild-type SpCas9 by T7E1 at eight NGG PAM-carrying target sites in four different endogenous human genes (corresponding to the notes in the top figure). Error bars represent sem, n=3.

[0045] Figure 17A-B| GUIDE-seq was used to determine the specificity characteristics of VQR and VRER SpCas9 variants. Expected target sites are marked with black squares, and mismatch sites within off-target sites are highlighted. A, The specificity of VQR variants in human cells was assessed by targeting endogenous sites containing NGA PAM: EMX1 site 4 (SEQ ID NO:142), FANCF site 1 (SEQ ID NO:143), FANCF site 3 (SEQ ID NO:144), FANCF site 4 (SEQ ID NO:145), RUNX1 site 1 (SEQ ID NO:146), RUNX1 site 3 (SEQ ID NO:147), VEGFA site 1 (SEQ ID NO:148), and ZSCAN2 (SEQ ID NO:149). B. Specificity of VRER variants in human cells was evaluated by targeting endogenous sites containing NGCG PAM: FANCF site 3 (SEQ ID NO:150), FANCF site 4 (SEQ ID NO:151), RUNX1 site 1 (SEQ ID NO:152), VEGFA site 1 (SEQ ID NO:153), and VEGFA site 2 (SEQ ID NO:154).

[0046] Figure 18A -C| Activity differences between D1135E and wild-type SpCas9 at off-target sites detected by GUIDE-seq. A, Mean frequency of oligonucleotide tag integration at intermediate target sites estimated by restriction fragment length polymorphism analysis. Error bars represent sem, n=4. B, Mean mutagenic frequency at intermediate target sites detected by T7E1. Error bars represent sem, n=4. C, Difference in GUIDE-seq read counts between wild-type SpCas9 and D1135E at three endogenous human cell sites (EMX1 site 3 (SEQ ID NO:155), ZNF629 site (SEQ ID NO:156), VEGFA site 3 (SEQ ID NO:157)). Intermediate target sites are shown at the top, and off-target sites are listed below, highlighting mismatches. In the table, the ratio of off-target activity to intermediate target activity was compared between wild-type and D1135E to calculate the normalized fold change in specificity (increased specificity is highlighted in green). For sites without detectable GUIDE-seq reads, a value of 1 has been assigned to calculate the estimated specific variation (indicated in orange). Figure 3E The off-target sites in the deep sequencing analysis were numbered to the left of the off-target sites at EMX1 site 3 and VEGFA site 3.

[0047] Figure 19A-F| Additional PAMs for St1Cas9 and SaCas9 and their activity in human cells based on spacer length. A, PPDV scatter plot for St1Cas9, comparing the 20 and 21 nucleotide sgRNA complementary sequence lengths obtained with randomized PAM libraries using spacer 1 (top plot) or spacer 2 (bottom plot). PAMs are grouped and plotted according to their 3rd / 4th / 5th / 6th positions. Red dashed lines indicate statistically significant depletion of PAMs (see [link]). Figure 13C ), and the gray dashed line indicates five times depletion (PPDV of 0.2). B, under each of the four test conditions, the PPDV of St1Cas9 is less than 0.2 for PAMs in the table. The PAM numbering shown in the left figure is related to Figure 4A Same as in (A). C, Scatter plot of PPDV for SaCas9, comparing the lengths of 21 and 23 nucleotide sgRNA complementary sequences obtained with randomized PAM libraries using spacer 1 (top plot) or spacer 2 (bottom plot). PAMs are grouped and plotted according to their 3rd / 4th / 5th / 6th positions. Red and gray dashed lines are the same as in (A). D, Table of PAMs with PPDV less than 0.2 for SaCas9 under each of the four tested conditions. PAM numbers are the same as in (A). Figure 4B The same applies in Figures E and F. Human cell activity of St1Cas9 and SaCas9 across different spacer lengths, disrupted by EGFP (Figure E, from...). Figure 4D , 4E Data) and endogenous gene mutagenesis detected by T7E1 (Figure F, from Figure 4F , 4G The data is shown. It shows the activity of all replicates (n=3 or 4); the bar chart shows the mean and 95% confidence intervals; and indicates the number of sites per spacer.

[0048] Figure 20A-B|Structural and functional roles of D1135, G1218, and T1337 in PAM recognition in SpCas9. A, Structural representation of the six residues involved in PAM recognition. The left panel shows the proximity of D1135 to S1136, a residue that forms a water-mediated minor groove contact with the 3rd base position of PAM15. The right panel shows the proximity of G1218, E1219, and T1337 to R1335, a residue that forms a direct base-specific major groove contact with the 3rd base position of PAM15. Angstrom distances are indicated by yellow dashed lines; the non-target strand guanine bases dG2 and dG3 of PAM are shown in blue; other DNA bases are shown in orange; water molecules are shown in red; images were generated from PDB:4UN3 using PyMOL. B, Mutational analysis of the six residues involved in PAM recognition in SpCas9. Using two sgRNAs targeting sites containing NGG PAM, clones containing one of three types of mutations at each site were tested for EGFP disruption. For each site, we established alanine substitutions and two non-conserved mutations. Previously reported, S1136 and R1335 mediate contact with the third guanine of PAM15, while this study reports D1135, G1218, E1219, and T1337. EGFP disruption activity was quantified by flow cytometry; background controls are indicated by red dashed lines; error bars represent sem, n=3.

[0049] Figure 21A -F Selection and assembly of SaCas9 variants with altered PAM specificity (A) Phylogenetic tree of Cas9 orthologs, highlighting SpCas9 and SaCas9. (B) Evaluation of the activity of SaCas9 variants with single amino acid substitutions in bacterial positive selection assays (see also...) Figure 31B Error bars represent sem, n=3; NS = no survival. (C) Human cell activity of wild-type and R1015H SaCas9. EGFP destruction activity was quantified by flow cytometry; error bars represent sem, n=3, with the mean level of background EGFP loss indicated by the red dashed line (for this figure and Figure E). (D) Total number of substitutions observed at each amino acid position when selected for SaCas9 variants with altered PAM specificity. The initiation mutation at R1015 is not included. (E) Human cell EGFP destruction activity of variants containing mutations observed when selected for altered PAM specificity. (F) Scatter plot of mean PAM depletion value (PPDV) after selection of wild-type SaCas9 and KKH variants (n=2, see also) Figure 34CTwo libraries with different prototype spacers and 8 randomized base pairs (instead of PAMs) were used to determine which PAMs each Cas9 could target. (See red dashed line for the dCas9 control). Figure 34A and 34B The gray dashed line indicates statistically significant depletion, and the gray dashed line indicates 5-fold depletion.

[0050] Figure 22A -F. Activity of SaCas9 KKH variants targeting endogenous sites in human cells (A) Mutagenic frequencies induced by KKH SaCas9 across 55 distinct sites carrying NNNRRT PAM were determined by T7E1 assay. Error bars represent sem, n=3, ND, undetectable by T7E1 assay. (B) KKH variant preference for the third position of PAM. The mean activity from data in Figure A is shown for this figure, as well as Figures B and C. (C) KKH variant preference for the fourth and fifth positions of PAM. (D) Spacer length preference of KKH SaCas9 variants. (E) Comparison of EGFP destruction activity in wild-type and KKH SaCas9-targeted sites containing NNNRRT PAM in human cells. EGFP destruction was quantified by flow cytometry; error bars represent sem, n=3, with the mean level of background EGFP loss indicated by the red dashed line. (F) For each of the 16 possible NNNRRT sites from Figure A, the mutagenic frequency of wild-type SaCas9 at one site (selected for the site with the highest KKH activity). Error bars represent sem, n=3, ND, which cannot be detected by T7E1 assay.

[0051] Figure 23AGenome-specific features of wild-type and KKH SaCas9 (A) and (B) are directly compared to those targeting sites containing NNGRRT (SEQ ID NO:46) PAM, represented by the total number of off-target events at EMX site 1 (SEQ ID NO:158) and VEGF site 8 (SEQ ID NO:159) (Fig. A) and the number of mismatches observed at each off-target site (Fig. B). For Figs. B and E, GUIDE-seq read counts at each site are indicated; the target sequence is marked with a black box; mismatches within off-target sites are highlighted; sequences corrected for cell type-specific SNPs are shown; sites with potential sgRNA or DNA protrusion nucleotides are indicated by small red borders or dashed lines, respectively. (C) The Venn diagram highlights the differences in off-target site cleavage between wild-type and KKH SaCas9 at VEGFA site 8. (D) and (E) Specific characteristics of KKH variants targeting sites containing NNHRRT (SEQ ID NO:44) PAM: EMX site 1 (SEQ ID NO:160), EMX site 4 (SEQ ID NO:161), EMX site 10 (SEQ ID NO:162), FANCF site 9 (SEQ ID NO:163), and FANCF site 16 (SEQ ID NO:164), represented by the total number of off-targets (Fig. D) and the mismatches observed at each off-target site (Fig. E).

[0052] Figure 24 : Activity of VQR derivative clones in bacterial 2-plasmid screening. Twenty-four different VQR derivative variants were tested against sites in bacteria containing NGAN PAM. Survival indicators on selective plates, relative to non-selective plates, indicate PAM activity.

[0053] Figure 25 EGFP-disrupting activity of SpCas9-VQR derivatives in human cells. The EGFP-disrupting activity of SpCas9 variants is a measure of activity at sites containing indicators of PAM.

[0054] Figure 26 EGFP disruption activity of SpCas9-VQR and -VRQR variants in human cells. The EGFP disruption activity of SpCas9 variants is a measure of activity containing sites indicating PAM.

[0055] Figure 27 Activity of SpCas9-VRQR derivative variants in bacterial plasmid screening. Twelve different VQR derivative variants were tested against sites in bacteria containing NGAN PAM, compared to VQR and VRQR variants. Survival indicators on selective plates, relative to non-selective plates, indicated PAM activity.

[0056] Figure 28 EGFP-disrupting activity of the SpCas9-VRQR variant in human cells. The EGFP-disrupting activity of the SpCas9 variant is a measure of the activity of sites containing indicators of PAM.

[0057] Figure 29 Cas9 orthologs (from) Figure 21A The protein domain alignments of SpCas9 are shown. The domain structure of SpCas9 is shown at the top (based on PDB:4UN3; Anders et al., 2014); PAM contact residues of SpCas9 are highlighted; and the SaCas9 region selected for mutagenesis to target altered PAM-specific variants is shown.

[0058] Figure 30 shows the primary sequence alignment of Cas9 orthogonal homologs used to identify PAM-interacting residues; SEQ ID NO: 165-176, respectively. Previously identified SpCas9 residues important for PAM contact (Anders et al., 2014; Examples 1-2) are highlighted in blue, residues regulating SaCas9 PAM specificity (identified in this study) are highlighted in orange, and positively charged residues adjacent to R1015 are highlighted in yellow. The structurally predicted PAM-interacting domains of SpCas9 are highlighted with blue dashed lines (based on PDB: 4UN3; Anders et al., 2014), and conservative estimates of the SaCas9 PAM-interacting domains used as PCR mutagenesis boundaries are indicated by orange dashed lines.

[0059] Figure 31A -B Schematic diagram of bacterial positive selection assay (A) The selection plasmid can be modified to screen for Cas9 variants that can recognize alternative PAM sequences. (B) Schematic diagram of positive selection plasmid (left) and expected results (right) when screening for functional or non-functional Cas9 / sgRNA pairs in positive selection.

[0060] Figure 32 The K929R mutation was added to the KNH and KKH variants. EGFP destruction activity was quantified by flow cytometry; error bars represent sem, n=3, and the mean value of background EGFP loss is represented by the red dashed line.

[0061] Figure 33Schematic diagram of bacterial site depletion assay. The plasmid with 8 randomized nucleotides (replacing PAMs that are difficult to cleave by wild-type or KKH SaCas9) was sequenced. Spacer sequence of library 1, SEQ ID NO:105; spacer sequence of library 2, SEQ ID NO:106. Targetable PAMs were inferred from their depletion relative to the input library, calculated as the selected PAM depletion value (PPDV).

[0062] Figure 34A (A) PPDV values ​​for the dCas9 control experiment of the two libraries. Red dashed lines indicate statistical significance (PPDV = 0.794, see Figure B); gray dashed lines indicate 5-fold depletion; PPDVs for windows containing the 3rd / 4th / 5th / 6th positions of PAM are plotted (for this figure and Figure C). (B) Statistically significant post-selection PAM depletion values ​​(PPDV) determined from the dCas9 control experiment in Figure a. Statistical significance was determined by setting the threshold to 3.36 times the standard deviation. (C) Comparison of PPDVs for wild-type and KKH SaCas9 for each of the two libraries containing 8 randomized nucleotides (in place of PAM). (D) and (E) are the corresponding PPDV values ​​for PAM and all depleted PAMs of wild-type and KKH SaCas9, respectively, greater than 5-fold. Sequence motifs for PAM are shown in two categories: 1) greater than 10-fold or 2) 5 to 10-fold depletion.

[0063] Figure 35A Additional features of the endogenous sites targeted by -D KKH SaCas9 (A) are based on the binaries of 16 possible NRR motifs of NNNRRT PAM, and the activity of each of the 55 endogenous sgRNA sites. For this figure, as well as Figures B and C, the activity of the sgRNAs from the NNNRRT PAM is shown. Figure 2A Mean activity. (B) and (C) Relationship between endogenous gene disruption activity and GC content of spacers and PAMs, respectively. (D) Sequence identifiers of spacers and PAMs at target sites based on activity compartments. Based on mean mutation frequency (from...) Figure 2A The sites were categorized into low activity (0-10%, 17 sites), moderate activity (10%-30%, 17 sites), or high activity (>30%, 21 sites).

[0064] Figure 36A -B Target tag integration and mutagenesis frequencies in GUIDE-seq experiments (A) Restricted fragment length polymorphism (RFLP) analysis to determine the mean GUIDE-seq tag integration frequency. Error bars represent sem, n=3 (for this figure and Figure B). (B) Mean mutagenesis detected by T7E1 determination.

[0065] Figure 37A -B Truncated repeats: Anti-repeated sgRNAs outperformed full-length sgRNAs, similar to previous results (Ran et al., 2015). (A) EGFP-disrupting activity of wild-type SaCas9 in human cells at four sites containing NNGRRT (SEQ ID NO:46) PAM. EGFP-disrupting activity was quantified by flow cytometry; error bars represent sem, n=3, and the mean level of background EGFP loss is indicated by the red dashed line (for this figure and Figure B). (B) EGFP-disrupting activity of KKH SaCas9 in human cells at eight sites containing NNNRRT PAM. Detailed Implementation

[0066] Although CRISPR-Cas9 nuclease is widely used in genome editing 1-4 However, the range of sequences that Cas9 can cut is limited by the requirement for a specific prototype spacer neighbor motif (PAM) at the target site. 5,6 For example, SpCas9, the most robust and widely used Cas9 to date, primarily recognizes NGG PAMs. Therefore, it is often difficult to target double-strand breaks (DSBs) with the precision required for various genome editing applications. Furthermore, imperfect PAM recognition by Cas9 can lead to unwanted off-target mutations. 7,8 The ability to evolve Cas9 derivatives with purposefully altered or improved PAM specificity would overcome these limitations, but to the inventors' knowledge, no such Cas9 variant has been described.

[0067] A potential strategy to improve the targeting range of Cas9 orthologs that recognize extended PAM sequences is to alter their PAM recognition specificity. As described herein, the PAM recognition specificity of SpCas9 can be altered using a combination of structure-guided design and directed evolution using a bacterial cell-based selection system; see Examples 1 and 2. Variants that have evolved to have relaxed or partially relaxed specificity for certain locations within the PAM are also described herein; see Example 3. These variants extend the utility of Cas9 orthologs that specify longer PAM sequences.

[0068] Engineered Cas9 variants with altered PAM specificity

[0069] The engineered SpCas9 variants in this study significantly increased the number of sites available for wild-type SpCas9, further enhancing the opportunity to implement efficient HDR using the CRISPR-Cas9 platform to target small genetic elements with NHEJ-mediated indels and leverage the requirement for PAM to distinguish two distinct alleles in the same cell. The modified PAM-specific SpCas9 variants effectively disrupted endogenous gene sites that SpCas9 currently cannot target in zebrafish embryos and human cells, suggesting they can function in a wide variety of cell types and organisms. Importantly, GUIDE-seq experiments showed that the global profiles of the VQR and VRERSpCas9 variants were similar to or better than those observed with wild-type SpCas9. Furthermore, the improved specific D1135E variant that we identified and characterized provides an excellent alternative to the widely used wild-type SpCas9. D1135E exhibits similar activity to wild-type SpCas9 at sites with typical NGG PAM, but reduces genome-wide cleavage at off-target sites carrying mismatched spacer sequences and typical or atypical PAMs.

[0070] All the SpCas9 and SaCas9 variants described herein can be rapidly incorporated into existing and widely used vectors, for example, through simple site-directed mutagenesis. And because they require only a small number of mutations within the PAM interacting domain, these variants should also function in conjunction with other previously described improvements to the SpCas9 platform (e.g., truncated sgRNAs (Tsai et al., Nat Biotechnol 33, 187-197 (2015); Fu et al., Nat Biotechnol 32, 279-284 (2014)), nickase mutations (Mali et al., Nat Biotechnol 31, 833-838 (2013); Ran et al., Cell 154, 1380-1389 (2013)), and dimeric FokI-dCas9 fusions (Guilinger et al., Nat Biotechnol 31, 833-838 (2013)). Biotechnol [Nature Biotechnology] 32, 577-582 (2014); Tsai et al., Nat Biotechnol [Nature Biotechnology] 32, 569-576 (2014)).

[0071] Besides the mutation at R1335 that may contact the third PAM base position, the SpCas9 variants evolved in this study carry amino acid substitutions at D1135, G1218, and T1337. All of these positions are close to or adjacent to residues that directly or indirectly contact the third PAM position in the SpCas9-PAM structure, but they themselves do not mediate contact with the PAM base (Anders et al., Nature 513, 569-573 (2014)). Figure 20A Consistent with this, we found that various mutations at these locations did not affect SpCas9-mediated cleavage of sites carrying NGG PAM. Figure 20B These results, along with the nature of the amino acid substitutions at G1218 and T1337 in the VQR and VRER SpCas9 variants, suggest that changes at these two positions can be gain-of-function mutations. For example, the T1337R mutation may form a backbone or base-specific contact near or toward position 4 of the PAM, particularly in the case of the VRER variant. The mechanical effects of mutations at D1135 remain less clear, but they may affect the activity of adjacent S1136 residues, which has been involved in water-mediated contacts with guanine at position 3 of the PAM via minor grooves (Anders et al., Nature 513, 569-573 (2014)). The D1135E mutation may improve specificity by disrupting this network, possibly by reducing the overall interaction energy of the SpCas9 / gRNA complex with the target site; a mechanism we previously proposed could reduce off-target effects by making the cleavage of these unwanted sequences more energy-unfavorable (Fu et al., Nat Biotechnol 32, 279-284 (2014)).

[0072] The current findings clearly establish the feasibility of engineered Cas9 nucleases with altered PAM specificity. Characterization of other Cas9 orthologs, as previously described (Esvelt et al., Nat Methods 10, 1116-1121 (2013); Fonfara et al., Nucleic Acids Res 42, 2577-2590 (2014)) or the generation of Cas9 chimeras with domain exchanges (Nishimasu et al., Cell 156(5):935-49 (2014)) also provide potential pathways for targeting different PAMs. The engineering strategies described here can also be carried out with these orthologs or synthetic hybrid Cas9s to further diversify the range of targetable PAMs. St1Cas9 and SaCas9, due to their smaller size relative to SpCas9, and the robust genome editing activity we have demonstrated in our bacterial selection system and human cells, provide a particularly attractive framework for future engineering efforts.

[0073] Our results strongly suggest that R1015 in wild-type SaCas9 contacts G at the third PAM position. Without being bound by theory, R1015H substitution could eliminate this contact and relax the specificity at the third position; however, it is conceivable that the loss of R1015 contact with G could also reduce the energy associated with target site binding, which could explain why the R1015H mutation alone is insufficient for robust activity at the NNNRRT site in human cells. Because both E782K and N968K substitutions add a positive charge, they may interact nonspecifically with the DNA phosphate backbone to compensate for the loss of R1015 contact with guanine.

[0074] The genetic approach described here requires no structural information and should therefore be applicable to many other Cas9 orthologs. The only requirement for evolving Cas9 nucleases with extended PAM specificity is that they function in bacterial-based selection. While previous studies have demonstrated that PAM recognition can be altered by exchanging PAM interaction domains of highly correlated Cas9 orthologs (Nishimasu et al., Cell (2014)), it remains to be determined whether this strategy is generalizable or effective when using more divergent orthologs. In contrast, the evolutionary strategy we describe here can be used to engineer PAM recognition specificity beyond those encoded within naturally occurring Cas9 orthologs. This overall strategy can be used to broaden the targeting range and expand the utility of the numerous Cas9 orthologs present in nature.

[0075] SpCas9 variants with altered specificity

[0076] Therefore, the spCas9 variant is provided here. The wild-type spCas9 sequence is as follows:

[0077]

[0078]

[0079]

[0080] The SpCas9 variants described herein may include mutations at one or more of the following positions: D1135, G1218, R1335, T1337 (or similar positions). In some embodiments, the SpCas9 variant includes one or more of the following mutations: D1135V; D1135E; G1218R; R1335E; R1335Q; and T1337R. In some embodiments, the SpCas9 variant is at least 80%, for example, at least 85%, 90%, or 95% identical to the amino acid sequence of SEQ ID NO:1, for example, differing at up to 5%, 10%, 15%, or 20% of the substituted residues of SEQ ID NO:1, for example, having conserved mutations. In a preferred embodiment, the variant retains the desired activity of the parent, such as nuclease activity (except when the parent is a nicking enzyme or a dead Cas9), and / or the ability to interact with guide RNA and target DNA.

[0081] To determine the percentage similarity between two nucleic acid sequences, the sequences are aligned for optimal alignment purposes (e.g., vacancies are introduced in one or both of the first and second amino acid sequences or nucleic acid sequences for optimal alignment, and non-homologous sequences can be ignored for comparison purposes). For comparison purposes, the length of the reference sequence aligned for comparison purposes is at least 80% of the length of the reference sequence, and in some embodiments at least 90% or 100%. The nucleotides at the corresponding amino acid positions or nucleotide positions are then compared. When a position in the first sequence is occupied by the same nucleotide as the corresponding position in the second sequence, then the molecule is consistent at that position (as used herein, nucleic acid “consistency” is equivalent to nucleic acid “homology”). The percentage similarity between two sequences is a function of the number of consistent positions shared by these sequences, taking into account the number of vacancies and the length of each vacancy, which needs to be introduced for optimal alignment of the two sequences. The percentage similarity between two polypeptide or nucleic acid sequences is determined in various ways within the scope of the art, such as using publicly available computer software, such as Smith Waterman Alignment (Smith, TF and MS Waterman (1981) J Mol Biol [Journal of Molecular Biology] 147: 195-7); “Best Fit” (Smith and Waterman, Advances in Applied Mathematics [Advances in Applied Mathematics], 482-489 (1981)), such as in combination with GeneMatcher Plus. TMSchwarz and Dayhof (1979) Atlas of Protein Sequence and Structure, Dayhof, MO (ed.), pp. 353-358; BLAST program (basic local alignment search tool); (Altschul, SF, W. Gish, et al. (1990) J Mol Biol [Journal of Molecular Biology] 215:403-10), BLAST-2, BLAST-P, BLAST-N, BLAST-X, WU-BLAST-2, ALIGN, ALIGN-2, CLUSTAL, or Megalign (DNASTAR) software. Additionally, those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithm required to achieve maximum alignment across the length of the sequences being compared. Typically, for proteins or nucleic acids, the comparison length can be any length, up to and including the full length (e.g., 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100%). For the purposes of the compositions and methods of the present invention, at least 80% of the full length of the sequence is aligned using the BLAST algorithm and default parameters.

[0082] For the purposes of this invention, the comparison of sequences and the determination of percentage consistency between two sequences can be accomplished using a Blossum 62 scoring matrix with a 12-point penalty for vacancy, a 4-point penalty for vacancy extension, and a 5-point penalty for shifted vacancy.

[0083] Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine.

[0084] In some embodiments, the SpCas9 variant includes a mutation from one of the following groups: D1135V / R1335Q / T1337R (VQR variant); D1135V / G1218R / R1335Q / T1337R (VRQR variant); D1135E / R1335Q / T1337R (EQR variant); or D1135V / G1218R / R1335E / T1337R (VRER variant).

[0085] In some embodiments, SpCas9 variants also include one of the following mutations that reduce or disrupt the nuclease activity of Cas9: D10, E762, D839, H983 or D986 and H840 or N863, such as D10A / D10N and H840A / H840N / H840Y, to render the nuclease portion of the protein non-catalytically active; the substitutions at these positions may be alanine (as in Nishimasu et al., Cell 156, 935–949 (2014)) or other residues, such as glutamine, asparagine, tyrosine, serine or aspartic acid, such as E762Q, H983N, H983Y, D986N, N863D, N863S or N863H (see WO 2014 / 152432). In some embodiments, the variant includes a mutation at D10A or H840A (which produces a single-stranded nickase) or a mutation at both D10A and H840A (which eliminates nuclease activity; such a mutant is called dead Cas9 or dCas9).

[0086] The SaCas9 variant is also provided here. The wild-type SaCas9 sequence is as follows:

[0087]

[0088]

[0089] The SaCas9 variants described herein include mutations at one or more of the following positions: E782, N968, and / or R1015 (or at similar positions). In some embodiments, these variants include one or more of the following mutations: R1015Q, R1015H, E782K, N968K, E735K, K929R, A1021T, K1044N. In some embodiments, the SaCas9 variants include mutations E782K, K929R, N968K, and R1015X, where X is any amino acid other than R. In some embodiments, the SaCas9 variants are at least 80%, such as at least 85%, 90%, or 95%, identical to the amino acid sequence of SEQ ID NO:2, for example, by substitutions of residues in SEQ ID NO:2 (e.g., conserved mutations) with a difference of up to 5%, 10%, 15%, or 20%. In a preferred embodiment, the variant retains the desired activities of the parent, such as nuclease activity (except when the parent is a nicking enzyme or dead Cas9), and / or the ability to interact with guide RNA and target DNA.

[0090] In some embodiments, SaCas9 variants also include one of the following mutations that can reduce or disrupt the nuclease activity of SaCas9: D10A, D556A, H557A, N580A, such as D10A / H557A and / or D10A / D556A / H557A / N580A, to render the nuclease portion of the protein non-catalytically active; the substitutions at these positions may be alanine (as in Nishimasu et al., Cell 156, 935–949 (2014)) or other residues, such as glutamine, asparagine, tyrosine, serine, or aspartic acid. In some embodiments, the variant includes mutations at D10A, D556A, H557A, or N580A (which can produce a single-stranded nickase) or mutations at D10A / H557A and / or D10A / D556A / H557A / N580A (which can eliminate nuclease activity, similar to SpCas9; these are called dead Cas9 or dCas9).

[0091] Also provided herein are isolated nucleic acids encoding SpCas9 and / or SaCas9 variants, vectors for expressing these variant proteins comprising the isolated nucleic acids optionally operably linked to one or more regulatory domains, and host cells (e.g., mammalian host cells) comprising the nucleic acids and optionally expressing these variant proteins.

[0092] The variants described herein can be used to alter the genome of a cell; methods typically involve expressing these variant proteins in the cell, along with a guide RNA having a region complementary to a selected portion of the cell's genome. Methods for selectively altering the cellular genome are known in the art, see, for example, US 8,697,359; US 2010 / 0076057; US 2011 / 0189776; US 2011 / 0223638; US 2013 / 0130248; WO / 2008 / 108989; WO / 2010 / 054108; WO / 2012 / 164565; WO / 2013 / 098244; WO / 2013 / 176772;; US 20150050699; US 20150045546; US20150031134; US 20150024500; US 20140377868; US 20140357530; US 20140349400; US20140335620; US 20140335063; US 20140315985; US 20140310830; US 20140310828; US20140309487; US 20140304853; US 20140298547; US 20140295556; US 20140294773; US20140287938; US 20140273234; US 20140273232; US 20140273231; US 20140273230; US20140271987; US 20140256046; US 20140248702; US 20140242702; US 20140242700; US20140242699; US 20140242664; US 20140234972; US 20140227787; US 20140212869; US20140201857; US 20140199767; US 20140189896; US 20140186958; US 20140186919; US20140186843; US 20140179770; US 20140179006;US 20140170753;Makarova et al., “Evolution and classification of the CRISPR-Cas system” 9(6) Nature Reviews Microbiology 467-477(1-23) (June 2011); Wiedenheft et al., “RNA-guided genetic silencing systems in bacteria and archaea” 482 Nature 331-338 (February 16, 2012); Gasinas et al., “Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria” 109(39) Proceedings of the National Academy of Sciences USA [Proceedings of the National Academy of Sciences] E2579-E2586 (September 4, 2012); Jinek et al., “A Programmable Dual-RNA-Guided DNA Endonuclease in Adaptive Bacterial Immunity” 337 Science 816-821 (August 17, 2012); Carroll, “A CRISPR Approach to Gene Targeting” 20(9) Molecular Therapy 1658-1660 (September 2012); U.S. Application No. 61 / 652,086, filed May 25, 2012;Al-Attar et al., Clustered Regularly Interspaced Short Palindromic Repeats (CRISPRs): The Hallmark of an Ingenious Antiviral Defense Mechanism in Prokaryotes, Biol Chem. (2011), Vol. 392, No. 4, pp. 277-289; Hale et al., Essential Features and Rational Design of CRISPR RNAs That Function With the Cas RAMP Module Complex to Cleave RNAs, Molecular Cell, (2012), Vol. 45, No. 3, pp. 292-302.

[0093] The variant protein described herein can be used in place of the SpCas9 protein described in the foregoing references, which has a guide RNA that targets a sequence having a PAM sequence according to Table 4 below.

[0094]

[0095] Furthermore, the variants described herein can be used for fusion proteins, replacing wild-type Cas9 or other Cas9 mutations known in the art (e.g., dCas9 or Cas9 nickases described above), such as fusion proteins with heterologous functional domains, as described in WO2014 / 124284. For example, variants preferably containing one or more nuclease-reducing or kill mutations can be fused at the N or C terminus of Cas9 to a transcriptional activation domain or other heterologous functional domain (e.g., transcriptional repressors such as KRAB, ERD, SID, etc., e.g., amino acids 473–530 of the repressor domain (ERD) of the ets2 repressor factor (ERF), amino acids 1–97 of the KRAB domain of KOX1, or amino acids 1–36 of the Mad mSIN3 interaction domain (SID); see Beerli et al., PNAS). USA [Proceedings of the National Academy of Sciences of the United States of America] 95:14628-14633 (1998)) or silencers such as heterochromatin protein 1 (HP1, also known as swi6), for example, HP1α or HP1β; or proteins or peptides known in the art that can recruit long non-coding RNAs (lncRNAs) fused to a fixed RNA-binding sequence (e.g., those bound by MS2 capsid protein, endonuclease Csy4, or λN protein); enzymes that modify the methylation state of DNA (e.g., DNA methyltransferase (DNMT) or TET protein); or enzymes that modify histone subunits (e.g., histone acetyltransferase (HAT), histone deacetylase (HDAC), histone methyltransferase (e.g., for methylation of lysine or arginine residues) or histone demethylase (e.g., for demethylation of lysine or arginine residues)). Multiple sequences of such domains are known in the art, for example, domains that catalyze the hydroxylation of methylated cytosine in DNA. Exemplary proteins include the deca-eleven translocation (TET) 1-3 family of enzymes that convert 5-methylcytosine (5-mC) to 5-hydroxymethylcytosine (5-hmC) in DNA.

[0096] The sequences of human TET1-3 are known in the art and are shown in the table below:

[0097]

[0098] *Variant (1) represents a longer transcript and encodes a longer isotype (a). Variant (2) differs from variant 1 in the 5' UTR and the 3' UTR, as well as in the coding sequence. The resulting isotype (b) is shorter and has a different C-terminus compared to isotype a.

[0099] In some embodiments, the catalytic domain may include all or part of its full-length sequence, such as a catalytic module comprising a cysteine-rich extension and a 2OGFeDO domain encoded by seven highly conserved exons, for example, the Tet1 catalytic domain comprising amino acids 1580-2052, the Tet2 domain comprising amino acids 1290-1905, and the Tet3 domain comprising amino acids 966-1678. See, for example, Figure 1, Iyer et al., Cell Cycle. June 1, 2009; 8(11):1698-710. The electronic publication of June 27, 2009 illustrates the alignment of key catalytic residues in all three Tet proteins, and supplemental material for their full-length sequences (see, for example, seq 2c); in some embodiments, the sequence comprises amino acids 1418-2136 of Tet1 or the corresponding region in Tet2 / 3.

[0100] Other catalytic modules may be derived from proteins identified by Iyer et al., 2009.

[0101] In some embodiments, the heterologous functional domain is a biological chain and comprises all or part of the MS2 coat protein, the endonuclease Csy4, or the λN protein (e.g., a DNA-binding domain). These proteins can be used to recruit RNA molecules containing a specific stem-loop structure to a location specified by a dCas9 gRNA targeting sequence. For example, dCas9 variants fused with the MS2 coat protein, the endonuclease Csy4, or λN can be used to recruit long non-coding RNAs (lncRNAs) such as XIST or HOTAIR; see, for example, Keryer-Bibens et al., Biol. Cell 100:125–138 (2008), which are linked to the Csy4, MS2, or λN binding sequence. Alternatively, the Csy4, MS2, or λN protein binding sequence can be linked to another protein, such as Keryer-Bibens et al., as described above, and the protein can be targeted to the dCas9 variant binding site using the methods and compositions described herein. In some embodiments, Csy4 is non-catalytically active. In some embodiments, the Cas9 variant (preferably the dCas9 variant) is fused with FokI, as described in WO 2014 / 204578.

[0102] In some embodiments, the fusion protein includes a linker between a dCas9 variant and a heterologous functional domain. The linker that can be used for these fusion proteins (or between fusion proteins in a linked structure) can include any sequence that does not interfere with the function of the fusion protein. In preferred embodiments, the linker is short, for example, 2-20 amino acids, and is typically flexible (i.e., includes amino acids with high degrees of freedom, such as glycine, alanine, and serine). In some embodiments, the linker includes one or more units consisting of GGGS (SEQ ID NO:188) or GGGGS (SEQ ID NO:189), such as two, three, four, or more repetitions of GGGS (SEQ ID NO:188) or GGGGS (SEQ ID NO:189) units. Other linker sequences may also be used.

[0103] Expression System

[0104] To use the Cas9 variants described herein, it may be desirable to express them from the nucleic acids encoding them. This can be done in several ways. For example, the nucleic acid encoding the Cas9 variant can be cloned into an intermediate vector for transformation into prokaryotic or eukaryotic cells for replication and / or expression. Intermediate vectors are typically prokaryotic vectors, such as plasmids, shuttle vectors, or insect vectors, used for the storage or manipulation of the nucleic acid encoding the Cas9 variant or for the production of the Cas9 variant. Alternatively, the nucleic acid encoding the Cas9 variant can be cloned into an expression vector for administration to plant cells, animal cells, preferably mammalian or human cells, fungal cells, bacterial cells, or protozoan cells.

[0105] To achieve expression, the sequence encoding the Cas9 variant is typically subcloned into an expression vector containing a promoter to direct transcription. Suitable bacterial and eukaryotic promoters are well known in the art and are described, for example, in Sambrook et al., *Molecular Cloning: A Laboratory Manual* (3rd edition, 2001); Kriegler, *Gene Transfer and Expression: A Laboratory Manual* (1990); and *Current Protocols in Molecular Biology* (edited by Ausubel et al., 2010). Bacterial expression systems for expressing engineered proteins are available, for example, in *Escherichia coli*, *Bacillus* species, and *Salmonella* (Palva et al., 1983, *Gene* 22:229-235). Kits for such expression systems are commercially available. Eukaryotic expression systems for mammalian cells, yeast, and insect cells are well known in the art and are also commercially available.

[0106] The promoter used to guide nucleic acid expression depends on the specific application. For example, strong constitutive promoters are typically used for the expression and purification of fusion proteins. In contrast, when administering Cas9 variants in vivo for gene regulation, either constitutive or inducible promoters can be used, depending on the specific purpose of the Cas9 variant. Furthermore, preferred promoters for administering Cas9 variants can be weak promoters, such as HSV-TK, or promoters with similar activity. The promoter may also include elements that respond to transcriptional activation, such as hypoxia-responsive elements, Gal4-responsive elements, lac repressor-responsive elements, and small molecule control systems, such as tetracycline regulatory systems and RU-486 systems (see, for example, Gossen & Bujard, 1992, Proc. Natl. Acad. Sci. USA, 89:5547; Oligino et al., 1998, Gene Ther., 5:491-496; Wang et al., 1997, Gene Ther., 4:432-441; Neering et al., 1996, Blood, 88:1147-55; and Rendahl et al., 1998, Nat. Biotechnol., 16:757-761).

[0107] In addition to the promoter, expression vectors typically contain a transcription unit or expression cassette containing all the additional elements required for nucleic acid expression in a host cell (prokaryotic or eukaryotic). Thus, a typical expression cassette contains a promoter operatively linked to, for example, a nucleic acid sequence encoding a Cas9 variant, and any signals required for effective polyadenylation of the transcript, transcription termination, ribosome binding site, or translation termination. Additional elements of the cassette may include, for example, enhancers and heterosplicing intrinsic signals.

[0108] The intended use of the Cas9 variant is determined by selecting a specific expression vector for transporting genetic information into cells, such as for expression in plants, animals, bacteria, fungi, protozoa, etc. Standard bacterial expression vectors include plasmids, such as pBR322-based plasmids, pSKF, pET23D, and commercially available tag fusion expression systems, such as GST and LacZ.

[0109] Expression vectors containing regulatory elements derived from eukaryotic viruses are commonly used for eukaryotic expression vectors, such as SV40 vectors, papillomavirus vectors, and vectors derived from E. coli and cyclophosphamide viruses. Other exemplary eukaryotic vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo-5, baculovirus pDSVE, and any other vectors that allow protein expression under the guidance of promoters such as: SV40 early promoter, SV40 late promoter, metallothionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrosis protein promoter, or other promoters shown to be effective for expression in eukaryotic cells.

[0110] Vectors used to express Cas9 variants may include RNA Pol III promoters to drive the expression of guide RNA, such as H1, U6, or 7SK promoters. These human promoters allow the expression of Cas9 variants in mammalian cells after plasmid transfection.

[0111] Some expression systems have markers for selecting stable transfected cell lines, such as thymidine kinase, hygromycin B phosphotransferase, and dihydrofolate reductase. High-yield expression systems are also suitable, such as those using baculovirus vectors in insect cells, with gRNA coding sequences directed by polyhedrosis protein promoters or other strong baculovirus promoters.

[0112] Elements typically included in expression vectors also include replicons that function in E. coli, genes encoding antibiotic resistance that allow selection of bacteria to accommodate recombinant plasmids, and unique restriction enzyme sites in non-essential regions of the plasmid that allow insertion of recombinant sequences.

[0113] Standard transfection methods were used to generate bacterial, mammalian, yeast, or insect cell lines expressing high levels of proteins, which were then purified using standard techniques (see, for example, Colley et al., 1989, J. Biol. Chem., 264:17619-22; Guide to Protein Purification, Methods in Enzymology, Vol. 182 (edited by Deutscher, 1990)). Transformation into eukaryotic and prokaryotic cells was performed according to standard techniques (see, for example, Morrison, 1977, J. Bacteriol. 132:349-351; Clark-Curtiss & Curtiss, Methods in Enzymology, 101:347-362 (edited by Wu et al., 1983).

[0114] Any known procedure for introducing exogenous nucleotide sequences into host cells may be used. These include transfection with calcium phosphate, polybrene, protoplast fusion, electroporation, nuclear transfection, liposomes, microinjection, naked DNA, plasmid vectors, viral vectors, both appendage- and integrative vectors, and any other well-known method for introducing cloned genomic DNA, cDNA, synthetic DNA, or other exogenous genetic material into host cells (see, for example, Sambrook et al., ibid.). Only a specific genetic engineering procedure capable of successfully introducing at least one gene into host cells expressing a Cas9 variant needs to be used.

[0115] The present invention includes a carrier and cells containing the carrier.

[0116] Example

[0117] The invention is further illustrated in the following examples, which are not intended to limit the scope of the invention as described in the claims.

[0118] method

[0119] Use the following materials and methods in Examples 1 and 2.

[0120] plasmids and oligonucleotides

[0121] Schematic diagrams and DNA sequences of the parental constructs used in this study can be found in [the database]. Figure 5A-J and SEQ ID NO:7-20 were found. The sequences of the oligonucleotides used to generate the positive select plasmid, negative select plasmid, and site-depleted library are available in Table 1. The sequences of all gRNA targets in this study are available in Table 2. Point mutations in Cas9 were generated by PCR.

[0122]

[0123]

[0124] Table 2

[0125]

[0126] Table 2

[0127]

[0128] Table 2

[0129]

[0130] Table 2

[0131]

[0132] Table 2

[0133]

[0134] Table 2

[0135]

[0136] Bacterial Cas9 / sgRNA expression plasmids were constructed using two T7 promoters to express Cas9 and sgRNA, respectively. These plasmids encode the SpCas9 sequence targeting Streptococcus pyogenes (BPK764, subcloned from JDS246). 17 ), from CRISPR locus 1, thermophilic streptococcus Cas9 (MSP1673, modified from the previously disclosed St1Cas9 sequence). 20 The human codon-optimized form of Cas9 from *Staphylococcus aureus* (BPK2101, SaCas9 sequence codon optimized from Uniprot J7RUA5) and *SpCas9*. Previously described sgRNA sequences are used for SpCas9. 34,35 and St1Cas9 20 The sequence HE980450 was searched for in the European nucleotide archive using CRISPRfinder targeting crRNA repeat sequences and bioinformatics methods similar to those previously described. 36The tracrRNA was identified to determine the SaCas9 sgRNA sequence. Annealed oligonucleotides used to complete the spacer complementary region of the sgRNA were ligated to BsaI-cleaved BPK764 and BPK2101 or BspMI-cleaved MSP1673 (5'-ATAG was attached to the spacer to produce the top oligonucleotide and 5'-AAAC was attached to the inverse complement of the spacer sequence to produce the bottom oligonucleotide).

[0137] Residues 1097-1368 of SpCas9 were randomly mutagenized at a rate of approximately 5.2 substitutions / kilobase using Mutazyme II (Agilent Technologies) to generate mutagenized PAM interaction (PI) domain libraries. Based on the number of transformed organisms obtained, the theoretical complexity of each PI domain library is estimated to be greater than 10. 7 One clone. Annealed target site oligonucleotides were ligated into p11-lacY-wtx1 cleaved by XbaI / SphI or EcoRI / SphI, respectively. 17 Positive and negative selection plasmids were generated. Two randomized PAM libraries (each with a different prototype spacer sequence) were constructed using Klenow(-exo) to fill the bottom strand of an oligonucleotide containing six randomized nucleotides directly adjacent to the 3' end of the prototype spacer (see Table 1). The double-stranded product was cleaved with EcoRI to ligate the EcoRI / SphI end to the cleaved p11-lacY-wtx1. Based on the number of transformants obtained, the theoretical complexity of each randomized PAM library was estimated to be greater than 10. 6 .

[0138] SpCas9 and SpCas9 variants are expressed in human cells from a vector derived from JDS246. 16 For St1Cas9 and SaCas9, Cas9 ORFs from MSP1673 and BPK2101 were subcloned into CAG promoter vectors to generate MSP1594 and BPK2139, respectively. Using the sgRNA sequences for SpCas9 sgRNA (BPK1520), St1Cas9 sgRNA (BPK2301), and SaCas9 gRNA (VVT1) described above, plasmids for U6 expression of the sgRNAs were generated (in which desired spacer oligonucleotides could be cloned). Annealed oligonucleotides for completing the spacer complementary regions of the sgRNAs were ligated to the BsmBI overhangs of these vectors (5'-CACC was appended to the spacer to generate the top oligonucleotide and 5'-AAAC was appended to the inverse complement of the spacer sequence to generate the bottom oligonucleotide).

[0139] Bacterial-based positive selection assay for evolving SpCas9 variants

[0140] Competent *E. coli* BW25141(λDE3) were transformed with a Cas9 / sgRNA-encoded plasmid containing a positive selection plasmid (with an insertion target site). 23 After recovering in SOB medium for 60 minutes, the transformants were plated on LB plates containing chloramphenicol (non-selective) or chloramphenicol + 10 mM arabinose (selective). The cleavage of the positive selection plasmid was estimated by calculating the survival frequency: colonies on the selective plate / colonies on the non-selective plate (see also Figure 12).

[0141] To select SpCas9 variants capable of cleaving novel PAM, a library of PI-domain-mutated Cas9 / sgRNA plasmids was electroporated into *E. coli* BW25141(λDE3) cells containing a positive-selection plasmid encoding the target site of interest + PAM. Typically, approximately 50,000 clones were screened to obtain between 50 and 100 survivors. The PI domain of the surviving clones was subcloned into a fresh backbone plasmid and retested in the positive selection. Clones with a survival rate greater than 10% in this secondary activity screening were sequenced. Sequencing was performed based on their frequency in the surviving clones, substitution type, and PAM bases in the SpCas9 / sgRNA crystal structure (PDB:4UN3). 14 The proximity of the mutations and (in some cases) their activity in human cell-based EGFP disruption assays were selected for further evaluation by selecting mutations observed in the sequencing clones.

[0142] Bacterial site depletion assay for analyzing Cas9 PAM specificity

[0143] Competent *E. coli* BW25141(λDE3) containing the Cas9 / sgRNA expression plasmid were transformed with a negative selection plasmid containing either a cleavable or non-cleavable target site. After recovery in SOB medium for 60 minutes, the transformants were plated on LB plates containing chloramphenicol and carbenicillin. The cleavage of the negative selection plasmid was estimated by calculating colony-forming units / μg of transformed DNA (see also Figure 13).

[0144] Negative selection was adjusted to determine the PAM-specific characteristics of the Cas9 nuclease by electroporating each randomized PAM library into *E. coli* BW25141(λDE3) cells already containing a suitable Cas9 / sgRNA plasmid. Colonies of 80,000–100,000 were spread at a low density on LB+chloramphenicol+carbenicillin plates. Viable colonies containing negative selection plasmids that were difficult to cleave by Cas9 were harvested, and plasmid DNA was isolated using maxi-prep (Qiagen). The resulting plasmid libraries were amplified by PCR using Phusion Hot-Start Flex DNA polymerase (New England BioLabs), followed by an Agencourt Ampure XP cleanup step (Beckman Coulter Genomics). Double-exponential Tru-Seq Illumina deep sequencing libraries were prepared using the KAPA HTP Library Preparation Kit (KAPA BioSystems) from approximately 500 ng of clean PCR product from each site depletion experiment. Paired-end sequencing of 150 bp was performed on an Illumina MiSeq sequencer at the Dana-Farber Cancer Institute Molecular Biology Core.

[0145] The raw FASTQ files from each MiSeq run output were analyzed using a Python program to determine relative PAM depletion. The program (see Methods) operates as follows: First, a file dialog box is presented to the user, from which all FASTQ read files for a given experiment can be selected. For these files, each FASTQ entry is scanned against a fixed spacer subregion on both strands. If a spacer subregion is found, six variable nucleotides flanking the spacer are captured and added to a counter. From this set of detected variable regions, the counts and frequencies for each window of length 2–6 nt at each possible location are listed. Site depletion data for the two randomized PAM libraries are analyzed by calculating the post-selection PAM depletion value (PPDV): the post-selection frequency of the PAM in the selected population divided by the pre-selected library frequency of that PAM. PPDV analysis is performed for all possible 2–6 length windows across a 6 bp randomized region for each experiment. The windows we used to visualize PAM preferences were: a 3nt window representing the 2nd, 3rd, and 4th PAM positions in wild-type and variant SpCas9 experiments; and a 4nt window representing the 3rd, 4th, 5th, and 6th PAM positions in St1Cas9 and SaCas9.

[0146] Two significance thresholds for PPDV were determined based on the following: 1) a statistical significance threshold based on the distribution of dCas9 and the ratio of log read counts in the preselected library (see [reference]). Figure 13C and 13D ), and 2) a bioactivity threshold based on the empirical correlation between depletion value and activity in human cells. The statistical threshold was set at 3.36 standard deviations of the mean PPDV of dCas9 (equivalent to a relative PPDV of 0.85), corresponding to a two-sided p-value of 0.05 after multiple comparisons adjustment for a normal distribution (i.e., p = 0.05 / 64). The bioactivity threshold was set at 5-fold depletion (equivalent to a PPDV of 0.2) because this level of depletion serves as a reasonable predictor of activity in human cells (see also...). Figure 14 ). Figure 14 The 95% confidence interval is calculated by dividing the standard deviation of the mean by the product of the square root of the sample size and 1.96.

[0147] Human cell culture and transfection

[0148] At 37°C and 5% CO2, the constitutively expressed EGFP-PEST reporter gene was... 15 Single-integrated copy U2OS.EGFP cells were cultured in advanced DMEM medium (Life Technologies) supplemented with 10% FBS, 2 mM GlutaMax (Life Technologies), penicillin / streptomycin, and 400 μg / ml G418. Cells were co-transfected with 750 ng of Cas9 plasmid and 250 ng of sgRNA plasmid (unless otherwise noted) using the DN-100 program on a Lonza 4D nuclear transfection instrument, according to the manufacturer's protocol. The Cas9 plasmid transfected with an empty U6 promoter plasmid served as a negative control for all human cell experiments. Target sites for endogenous gene experiments were selected within the 200 bp NGG site cleavable by wild-type SpCas9 (see [link to documentation]). Figure 16A (and Table 2).

[0149] Zebrafish care and injection

[0150] Zebrafish care and use were approved by the Massachusetts General Hospital Subcommittee on Research Animal Care. Cas9 mRNA was transcribed from PmeI-digested JDS246 (wild-type SpCas9) or MSP469 (VQR variant) using the mMESSAGE mMACHINE T7 ULTRA kit (Lifetech Corporation), as previously described. 21All sgRNAs in this study were generated using a cloning-independent sgRNA generation method. 24 Prepared. sgRNA was transcribed using the MEGAscript SP6 transcription kit (Lifetech Corporation), purified using RNA Clean & Concentrator-5 (Zymo Research), and eluted with RNase-free water.

[0151] The mRNAs encoding sgRNA and Cas9 were co-injected into single-cell stage zebrafish embryos. Each embryo was injected with approximately 2–4.5 nL of a solution containing 30 ng / μL gRNA and 300 ng / μL Cas9 mRNA. The next day, the normal morphological development of the injected embryos was examined under a stereomicroscope, and genomic DNA was extracted from 5 to 9 embryos.

[0152] Human cell EGFP destruction assay

[0153] As previously stated 16 EGFP disruption assays were performed. EGFP expression in transfected cells was analyzed approximately 52 hours post-transfection using a Fortessa flow cytometer (BD Biosciences). The background EGFP loss gating for all experiments was approximately 2.5% (indicated by the red dashed line in the image).

[0154] T7E1 assay, targeted deep sequencing, and GUIDE-seq were used to quantify the nuclease-induced mutation rate.

[0155] As previously targeted at human cells 15 and zebrafish 21 The T7E1 assay was performed. For U2OS.EGFP human cells, genomic DNA was extracted from transfected cells approximately 72 hours post-transfection using the Agencourt DNAdvance Genomic DNA Isolation Kit (Beckman Coulter Genomics). Target loci from zebrafish or human cell genomic DNA were amplified using primers listed in Table 1. Approximately 200 ng of purified PCR product was denatured, annealed, and digested with T7E1 (New England Biolabs). Mutagenesis frequencies were quantified using a Qiaxcel capillary electrophoresis system (Qiagen), as previously performed on human cells. 15 and zebrafish 21 As stated above.

[0156] For targeted deep sequencing, previously characterized mid- and off-target sites were amplified using Phusion Hot-start Flex with primers listed in Table 1 (Tsai et al., Nat Biotechnol [Nature Biotechnology] 33, 187-197 (2015); Fu et al., Nat Biotechnol [Nature Biotechnology] 31, 822-826 (2013); Fu et al., Nat Biotechnol [Nature Biotechnology] 32, 279-284 (2014)). Genomic loci were amplified for control conditions (empty sgRNA), wild-type, and D1135E SpCas9. An Agencourt Ampure XP cleanup step (Beckman Coulter Genomics) was performed before pooling approximately 500 ng of DNA from each condition in library preparation. Double-index Tru-Seq Illumina deep sequencing libraries were generated using the KAPA HTP Library Preparation Kit (KAPA Biosystems). Dana-Farber Cancer Institute Molecular Biology Core Cancer Institute Core) 150 bp paired-end sequencing was performed on an Illumina MiSeq sequencer. Mutation analysis of the targeted deep sequencing data was performed as previously described (Tsai et al., Nat Biotechnol 32, 569-576 (2014)). Briefly, the Illumina MiSeq paired-end reads were mapped to the human genome reference GRChr37 using bwa (Li et al., Bioinformatics 25, 1754-1760 (2009)). High-quality reads (quality score >= 30) of indel mutations overlapping with target or off-target sites were evaluated. 1-bp indel mutations were excluded from the analysis unless they occurred within 1-bp of the predicted breakpoint. The activity changes of D1135E and wild-type SpCas9 at target and off-target sites were calculated and compared by comparing the indel frequencies from both conditions (the ratio of indel levels exceeding those of the background control amplicon).

[0157] GUIDE-seq experiments were performed as previously described (Tsai et al., Nat Biotechnol, 33, 187-197 (2015)). In short, as described above, phosphorylated phosphate-thioester-modified double-stranded oligodeoxynucleotides (dsODNs) were transfected into U2OS cells along with the Cas9 nuclease and Cas9 and sgRNA expression plasmids. dsODN-specific amplification, high-throughput sequencing, and mapping were performed to identify genomic regions containing DSB activity. For wild-type and D1135E experiments, off-target read counts were normalized to target read counts to correct for sequencing depth differences between samples. The normalized ratios of wild-type and D1135E SpCas9 were then compared to calculate the fold change in activity at off-target sites. To determine whether the wild-type and D1135E samples from GUIDE-seq had similar oligonucleotide tag integration rates at the expected target sites, restriction fragment length polymorphism (RFLP) assays were performed using primers listed in Table 1 to amplify the expected target loci from 100 ng of genomic DNA (isolated as described above) using a Phusion Hot-Start Flex. Approximately 150 ng of PCR product was digested with 20 U NdeI (New England Biolabs) at 37°C for 3 hours prior to cleansing using the Agencourt Ampure XP kit. RFLP results were quantified using a Qiaxcel capillary electrophoresis system (Kiagen) to roughly estimate the oligonucleotide tag integration rate. T7E1 assays were performed for similar purposes as described above.

[0158] Software - for analyzing PAM-depleted MiSeq data

[0159] Run the command "python PAM_depletion.py" in the command prompt (the directory containing the file).

[0160]

[0161]

[0162]

[0163]

[0164] Example 1

[0165] One potential solution to address targeting limitations is to engineer Cas9 variants with novel PAM specificity. Previous attempts to alter PAM specificity have utilized structural information about base-specific SpCas9-PAM interactions to mutate arginine residues (R1333 and R1335) that contact guanine nucleotides at the second and third PAM positions, respectively (Anders et al., Nature 513, 569-573 (2014)). Replacing the two arginine residues with glutamine (whose side chain is expected to interact with adenine) does not produce SpCas9 variants capable of cleaving targets containing the intended NAA PAM in vitro (Anders et al., Nature 513, 569-573 (2014)). Using a human cell-based U2OS EGFP reporter gene disruption assay, in which nuclease-induced indels lead to fluorescence loss (Reyon et al., Nat Biotechnol [Nature Biotechnology] 30, 460-465 (2012); Fu et al., Nat Biotechnol [Nature Biotechnology] 31, 822-826 (2013)), we confirmed that the R1333Q / R1335Q SpCas9 variant cannot effectively cleave target sites with NAA PAM. Figure 1A Additionally, we found that individual R1333Q and R1335Q SpCas9 variants could not effectively cleave target sites with their intended NAG and NGA sites, respectively. Figure 1A Therefore, we infer that reengineering PAM specificity may require additional mutations at locations other than R1333 and R1335. For example, available structural information shows that K1107 and S1136 form direct and indirect minor groove contacts with the second and third bases in PAM, respectively (Anders et al., Nature 513, 569-573 (2014)). Therefore, further alterations at or near these locations may be necessary to change PAM specificity.

[0166] To identify additional sites that may be crucial for the specificity of modified PAM, we modified a bacterial selection system previously used to study the properties of homing endonucleases (hereinafter referred to as positive selection) (Chen & Zhao, Nucleic Acids Res [Nucleic Acid Research] 33, e154 (2005); Doyon et al., J Am Chem Soc [Journal of the American Chemical Society] 128, 2477-2484 (2006)). In our modification of the system, Cas9-mediated cleavage of the positive selection plasmid encoding an inducible virulence gene enabled cell survival, attributed to the subsequent degradation and loss of the linearized plasmid. Figure 1B and Figure 12AAfter determining that SpCas9 could function in the positive selection system, we tested the ability of wild-type and R1335Q variants to cleave selection plasmids containing target sites with NGAPAM, and as expected, no survival was observed. Figure 12A To screen for gain-of-function mutations, we generated wild-type and R1335Q SpCas9 libraries carrying randomly induced PAM interaction domains (amino acid positions 1097-1368), with a mean ratio of 5.2 mutations / kilobase. Figure 12B (Methods). These libraries were introduced into bacteria using positive selection plasmids containing target sites with NGA PAM and plated on selective media. Sequences from surviving clones of the R1335Q-based library revealed that, in addition to the previously present R1335Q mutation, the most common substitutions were D1135V / Y / N / E and T1337R (Table 3). We obtained fewer survivors by selecting with wild-type SpCas9-based libraries, but these clones also included D1135V / Y / N and R1335Q mutations. We then assembled and tested all possible single, dual, and triple combinations of the D1135V / Y / N / E, R1335Q, and T1337R mutations using a human cell-based EGFP destruction assay. This analysis showed that SpCas9 variants with substitutions at all three positions exhibited the highest activity on NGA PAM but the lowest activity on NGG PAM. Figure 1C We selected two SpCas9 variants, D1135V / R1335Q / T1337R and D1135E / R1335Q / T1337R (hereinafter referred to as VQR and EQR SpCas9 variants, respectively), because they have the greatest difference between NGA and NGG PAM. Figure 1C ), for further characterization.

[0167] To evaluate the global PAM-specific characteristics of our novel SpCas9 variant, we used a bacteria-based negative selection system ( Figure 1D and Figure 13A Previous studies have used similar selection systems to identify cleavage site preferences of Cas9 nucleases (Jiang et al., Nat Biotechnol 31, 233-239 (2013); Esvelt et al., Nat Methods 10, 1116-1121 (2013)). In this form of assay (which we call the site depletion assay), plasmid libraries carrying randomized 6 bp sequences of prototypically adjacent spacers were tested for cleavage by the Cas9 / sgRNA complex in E. coli. Figure 13BPlasmids carrying cleavable sequences adjacent to the prototype spacer containing Cas9 / sgRNA complexes are able to survive in cells due to the presence of antibiotic resistance genes, while plasmids carrying cleavable sequences are degraded and thus depleted from the library. Figure 13B High-throughput sequencing of approximately 100,000 untargetable sequences enabled us to calculate the post-selection PAM depletion value (PPDV) for any given PAM. The PPDV of a PAM (or PAM group) is defined as the frequency of the PAM in the post-selection population divided by its frequency in the pre-selected library. This quantitative value provides an estimate of the Cas9 activity of that PAM. Characterization of catalytically inactive Cas9 (dCas9) obtained on two randomized PAM libraries (each with different prototype spacers) allowed us to define what constitutes a statistically significant change in the PPDV of any given PAM or PAM group. Figure 13C Then, we validated our site depletion assay by demonstrating that the PPDV of wild-type SpCas9 obtained with two randomized PAM libraries recapitulated the previously described characteristics of targetable PAMs (Jiang et al., Nat Biotechnol, 31, 233-239 (2013)). Figure 1E ).

[0168] Using this site depletion assay, we obtained PAM-specific signatures for the VQR and EQR SpCas9 variants using two randomized PAM libraries. The VQR variant strongly depleted sites containing NGAN and NGCG PAMs, and weakly depleted sites containing NGGG, NGTG, and NAAG PAMs. Figure 1F In contrast, the EQR variant strongly depletes NGAG PAM and weakly depletes NGAT, NGAA, and NGCG PAM. Figure 1F This demonstrates a potentially more limited targeting range compared to the VQR variant. To test whether PAMs identified by the site depletion assay could also be recognized in human cells, we evaluated cleavage at target sites by the VQR and EQR SpCas9 variants using the EGFP destruction assay. The VQR variant robustly cleaved sites carrying NGANPAM (relative efficiency NGAG>NGAT=NGAA>NGAC) and sites carrying NGCG, NGCG, and NGTG PAMs (generally less efficient) in EGFP. Figure 1G The EQR variant also reiterated its preference for NGAG and NGNG PAMs over other NGAN PAMs in human cells, while exhibiting lower activity than the VQR variant. Figure 1G In summary, these results in human cells strongly reflect those observed in bacterial site depletion assays. Figure 14), and it indicates that 0.2 PPDV (representing five-fold depletion) in bacterial assays provides a reasonable predictive threshold for activity in human cells. Figure 14 ).

[0169] We then attempted to extend the generalizability of our engineering strategy by trying to identify SpCas9 variants capable of recognizing NGC PAM. We first engineered Cas9 mutants carrying R1335 amino acid substitutions (D, E, S, or T) that were expected to interact with cytosine, and found no activity at the NGC PAM site using a positive selection system. We then randomly mutated the PAM-interacting domain of each of these single-substituted SpCas9 variants, but still did not obtain viable colonies in positive selection. This was because the T1337R mutation increased the activity of our VQR and EQR SpCas9 variants. Figure 1C Therefore, we combined this mutation with R1335 substitutions of A, D, E, S, T, or V, and then randomly mutagenized its PAM interaction domain. Selection using two of these six mutagenic libraries (carrying pre-existing R1335E / T1337R and R1335T / T1337R substitutions) yielded surviving colonies with a variety of additional mutations (Table 3). Characterization of different selected clones using both bacterial and human cell-based assays showed that the manifestation of substitutions at four positions (D1135V, G1218R, R1335E, and T1337R) was important for cleavage of NGC PAM. Using an EGFP destruction assay, we assembled and tested all potential single, double, triple, and quadruple combinations of these mutations, identifying the quadruple VRER variant as exhibiting the highest activity on NGCG PAM and the lowest activity on NGGG PAM. Figure 1H Site depletion assays revealed that the VRER variant exhibits high specificity for NGCG PAM. Figure 1I Consistent with these results, EGFP disruption assays in human cells using VRER variants revealed efficient cleavage at sites with NGCG PAM, significantly reduced and inconsistent cleavage at sites with NGCA, NGCC, and NGCT PAM, and essentially no activity at sites with NGAG, NGTG, and NGCG PAM. Figure 1J ).

[0170] To directly demonstrate that our VQR and VRER SpCas9 variants can target sites that are currently unmodifiable by wild-type SpCas9, we tested their activity against endogenous genes in zebrafish embryos and human cells. In single-cell zebrafish embryos, we found that the VQR variants effectively modified endogenous gene sites carrying NGAG PAM, with a mean mutagenicity rate of 20% to 43%. Figure 2A ), and the indel begins at the predicted cleavage site ( Figure 15 In human cells, we found that the VQR variant robustly modified 16 sites across four different endogenous genes containing NGAG, NGAT, and NGAA PAM (ranging from 6% to 53%, with a mean of 33%). Figure 2B and Figure 16A Importantly, we demonstrated that wild-type SpCas9 cannot effectively alter most of the same sites with NGAG and NGAT PAM in zebrafish and human cells. Figure 2A and 2C However, it can effectively modify nearby sites carrying NGG PAM. Figure 16B Similarly, when examining VRER variant activity across three endogenous human genes at nine sites with NGCG PAM, we also observed robust mean disruption frequencies (ranging from 5% to 36%, with a mean of 21%). Figure 2D Consistent with our site depletion data. Figure 1E and 1F The VQR variant altered the NGCG PAM site efficiency observed in the VRER variant, which was not seen in wild-type SpCas9. Figure 2D Computational analysis referencing human genome sequences showed that adding our VQR and VRER SpCas9 variants doubled the range of potential target sites compared to previous possibilities with only wild-type SpCas9. Figure 2E In summary, these results demonstrate that our engineered SpCas9 variant extends the targeting range of SpCas9 by enabling modification of previously inaccessible endogenous sites in zebrafish embryos and human cells.

[0171] To determine the genome-wide specificity of our VQR and VRER SpCas9 nucleases, we used the recently described GUIDE-seq (genome-wide unbiased identification of double-strand breaks via sequencing) method. 10 To analyze off-target cleavage events of these SpCas9 variants in human cells, we used a total of 13 different sgRNAs (8 VQRs and 5 VRERs, derived from...). Figure 2B and 2DWe analyzed the genome-wide activity of VQR and VRER SpCas9 variants and showed that they can induce highly efficient modifications at the expected target sites. These GUIDE-seq experiments yielded several important observations: the number of off-target DSBs induced by our SpCas9 variant in human cells was comparable to (or perhaps even better than) those previously observed with wild-type SpCas9. Figure 2F We note that the high genome-wide specificity observed with VRER may be due to its restriction specificity to NGCG PAM, and possibly due to the relative depletion of sites with NGCG PAM in the human genome. Figure 2E ) 21 Furthermore, the observed off-target sites typically exhibit the expected PAM sequence predicted by our site depletion experiments, including some tolerance to a 1-base "shift" in the PAM 3' region (which will...). Figure 1F and 1I (The PAMs are compared with those in Figure 17). Finally, the locations and numbers of mismatches found in the off-target sites of our VQR and VRER SpCas9 variants (Figure 17) are similar in distribution to those we previously observed with wild-type SpCas9 for sgRNAs targeting non-repetitive sequences.

[0172] Previous studies have shown that imperfect PAM recognition by SpCas9 can lead to the recognition of unwanted sites containing atypical NAG, NGA, and other PAMs in human cells (Hsu et al., Nat Biotechnol 31, 827-832 (2013); Tsai et al., Nat Biotechnol 33, 187-197 (2015); Jiang et al., Nat Biotechnol 31, 233-239 (2013); Mali et al., Nat Biotechnol 31, 833-838 (2013); Zhang et al., Sci Rep 4, 5405 (2014)). Therefore, we are interested in exploring whether mutations at or near residues mediating PAM interactions might improve SpCas9 PAM specificity. When engineering the VQR variant, we noted that the D1135E SpCas9 mutant showed better differentiation between typical NGG PAM and atypical NGA PAM compared to wild-type SpCas9. Figure 1CBased on this observation, we used our site depletion assay to comprehensively evaluate the PAM recognition characteristics of this D1135E variant. This experiment revealed reduced depletion of atypical NAG, NGA, and NNGG PAMs in the D1135E SpCas9 case compared to wild-type SpCas9. Figure 3A Interestingly, this effect is more pronounced for one of the two prototype spacers we used, suggesting that the impact of D1135E substitution on atypical PAM recognition can vary to some extent in a prototype spacer-dependent manner. Importantly, we did not observe the emergence of any new atypical PAM specificities.

[0173] We then tested whether the enhanced PAM specificity of D1135E SpCas9 could also be observed in human cells. In a direct comparison of wild-type and D1135E SpCas9 at eight target sites with atypical NAG or NGA PAM, we observed that cleavage of these sites by D1135E was consistently less efficient than cleavage by wild-type SpCas9 in EGFP disruption assays. Figure 3B The average fold decrease in activity was 1.94%. Importantly, wild-type and D1135E SpCas9 showed comparable activity at four EGFP reporter loci and six endogenous human loci with typical NGG PAM (respectively). Figure 3B and 3C This demonstrates that the D1135E variant does not significantly affect cleavage at the intermediate target site with NGG PAM (mean fold reduction in activity across all ten sites was 1.04). Our titration experiments, which reduced the concentration of Cas9-encoding plasmids transfected into human cells, revealed no significant difference in activity when wild-type and D1135E SpCas9 targeted the same site. Figure 3D This means that the increased specificity observed with the D1135E variant is not solely a result of protein instability.

[0174] To more directly assess whether the introduction of D1135E could reduce the off-target cleavage effect of SpCas9, we used deep sequencing to compare the wild-type and D1135E SpCas9-induced mutation rates at 25 previously known off-target sites on three different sgRNAs (Hsu et al., Nat Biotechnol 31, 827-832 (2013); Tsai et al., Nat Biotechnol 33, 187-197 (2015); Fu et al., Nat Biotechnol 31, 822-826 (2013)). These 25 sites included off-target sites with different mismatches in both spacer sequences and typical NGG and atypical PAM. Figure 3E The results of these deep sequencing experiments revealed that, relative to the mutation frequencies observed at the three target sites, the D1135E variant exhibited reduced mutation frequencies at 19 of the 22 off-target sites, with activity exceeding the background indel ratio. Figure 3E and 3F Interestingly, these reduced off-target mutation frequencies were observed at many sites with typical PAMs, suggesting that the increased specificity in the D1135E case is not limited to sites with atypical PAMs. To assess the genome-wide specificity improvements associated with D1135E, we performed GUIDE-seq experiments using wild-type and D1135E SpCas9 with three different sgRNAs (two of which were previously known to have off-target sites with both typical and atypical PAMs) (Hsu et al., Nat Biotechnol 31, 827-832 (2013); Tsai et al., Nat Biotechnol 33, 187-197 (2015); Fu et al., Nat Biotechnol 31, 822-826 (2013)). When using the D1135E SpCas9 variant compared to wild-type SpCas9, we observed a broad improvement in genome-wide specificity. Figure 3G For all three sgRNAs we tested, these improvements in specificity were observed at off-target sites containing mismatched spacers with either typical or atypical PAMs (Fig. 18). Importantly, these GUIDE-seq experiments demonstrated that the introduction of the D1135E mutation did not increase the number of SpCas9-induced off-target effects. In summary, these results show that D1135E substitution can increase the global specificity of SpCas9.

[0175] Although all the above experiments were performed using SpCas9, there are many Cas9 orthologs from other bacteria that could be attractive candidates for characterizing and engineering Cas9 with novel PAM specificity (Fonfara et al., Nucleic Acids Res. 42, 2577-2590 (2014); Ran et al., Nature. 520, 186-191 (2015)). To explore the feasibility of doing so, we determined whether two smaller orthologous species, namely *Streptococcus thermophilus* Cas9 (St1Cas9) from the CRISPR1 locus (Deveau et al., J Bacteriol 190, 1390-1400 (2008); Horvath et al., J Bacteriol 190, 1401-1412 (2008)) and *Staphylococcus aureus* (SaCas9) (Hsu et al., Cell 157, 1262-1278 (2014); Ran et al., Nature 520, 186–191 (2015)), could also play a role in our bacterial selection assay. Although the PAM of St1Cas9 was previously characterized as NNAGAA (SEQ ID NO:3) (Esvelt et al., Nat Methods 10, 1116-1121 (2013); Fonfara et al., Nucleic Acids Res 42, 2577-2590 (2014); Deveau et al., J Bacteriol 190, 1390-1400 (2008); Horvath et al., J Bacteriol 190, 1401-1412 (2008)), our attempt to derive the SaCas9 PAM from bioinformatics using the previously described method (Fonfara et al., Nucleic Acids Res 42, 2577-2590 (2014)) failed to produce a shared sequence (data not shown). Therefore, we used our site depletion assay to determine the PAM of SaCas9, as well as the PAM of St1Cas9 as a positive control. These experiments were performed using two different prototype spacers and sgRNAs, each prototype spacer having two different complementary lengths, resulting in four choices for each Cas9.For St1Cas9, in addition to the six PAMs previously described (Esvelt et al., Nat Methods 10, 1116-1121 (2013); Fonfara et al., Nucleic AcidsRes 42, 2577-2590 (2014); Horvath et al., J Bacteriol 190, 1401-1412 (2008)), we also identified two novel PAMs. Figure 4A as well as Figure 19C and 19D This is consistent with the recent definition of SaCas9 PAM specificity (Ran et al., Nature 520, 186–191 (2015)). For SaCas9, there is PPDV variability among the four choices, mainly due to the restricted PAM preference observed with a prototype spacer. As a result, only three PAMs were more than 5-fold depleted in all four experiments: NNGGGT (SEQ ID NO:4), NNGAAT (SEQ ID NO:6), and NNGAGT (SEQ ID NO:5). Figure 4B However, we did use a second prototype spacer sublibrary to identify more targetable PAMs, meaning SaCas9 could potentially identify many additional PAMs. Figure 18C Using the PAMs identified in our site depletion experiments (NNAGAA (SEQ ID NO:3) for SaCas9 and NGAGT (SEQ ID NO:5) for St1Cas9), we found that both St1Cas9 and SaCas9 can function effectively in bacterial positive selection systems. Figure 4C This indicates that their PAM specificity can be modified through mutagenesis and selection.

[0176] Because not all Cas9 orthologs function effectively outside their original background (Esvelt et al., Nat Methods 10, 1116-1121 (2013)), we tested whether St1Cas9 and SaCas9 could robustly cleave target sites in human cells. St1Cas9 has previously been shown to act as a nuclease in human cells, but only at a few sites (Esvelt et al., 2013; Cong et al., Science 339, 819-823 (2013)). We assessed St1Cas9 activity at sites containing NNAGAA (SEQ ID NO:3) PAM using sgRNA with variable-length complementary regions, and found high activity at three of the five target sites. Figure 4DFor SaCas9, we observed potent activity at 8 sites containing NNGGGT (SEQ ID NO:4) or NNGGAGT (SEQ ID NO:5) PAM. Figure 4E For St1Cas9 and SaCas9, no significant correlation was observed between the activity and length of spacer complementarity. Figure 19E We then determined whether St1Cas9 and SaCas9 could effectively modify endogenous loci in human cells. For St1Cas9, as determined by T7E1 assay, 7 out of 11 loci across 4 genes were effectively disrupted (1% to 25%, mean 13%). Figure 4F SaCas9 showed somewhat more robust activity at 16 sites across 4 genes (1% to 37%, mean 19%). Figure 4G Furthermore, no obvious trend was observed when considering the sgRNA spacer lengths of St1Cas9 and SaCas9. Figure 19F In summary, our results demonstrate that both St1Cas9 and SaCas9 function robustly in both our bacterial-based selection and human cell-based studies, making them attractive candidates for engineering additional SpCas9 variants with novel PAM specificity.

[0177]

[0178]

[0179]

[0180]

[0181]

[0182]

[0183]

[0184]

[0185]

[0186]

[0187]

[0188]

[0189]

[0190] Example 2. PAM specificity of engineered Staphylococcus aureus Cas9

[0191] Because we know that residues in *Streptococcus pyogenes* Cas9 (SpCas9) are important for PAM recognition (R1333 and R1335), we performed a Cas9 orthogonal homolog comparison to search for homologous residues in the PAM interaction domain (PI domain) of *Staphylococcus aureus* Cas9 (SaCas9) (see [link to documentation]). Figure 6 We and others have previously shown that the PAM of SaCas9 is NNGRRT (SEQ ID NO:46) (where N is any nucleotide and R is A or G). Based on our data, a preference for G at position 3 of the PAM appears to be the most stringent requirement, so we hypothesize that positively charged residues such as lysine (K) or arginine (R) may mediate this interaction. Figure 6 As shown, there are several candidate residues in the homologous regions of SaCas9 and SpCas9, namely R1333 and R1335, including K1101, R1012, R1015, K1018 and K1023.

[0192] We generated alanine (A) and glutamine (Q) substitutions at these five positions to determine whether the mutant clone could still cleave sites containing typical NNGRRT PAM (SEQ ID NO:46) or potentially cleave previously untargetable NNARRT (SEQ ID NO:43) PAM. Figure 7 We utilized our bacterial assay (described in a previous patent application) in which Cas9 activity can be visualized by the survival of bacterial colonies when plated under selective conditions. The relative activity of Cas9 can be quantified by calculating the ratio of bacterial colonies grown on selective to non-selective media. Figure 7 In our study, we showed that only R1015A and R1015Q mutations affect SaCas9's ability to recognize typical NGAGT (SEQ ID NO:5) PAMs, while no mutations target NNARRT (SEQ ID NO:43) PAMs (NNAAGT (SEQ ID NO:41) or NNAGGT (SEQ ID NO:42)). These results indicate that R1015 plays an important role in SaCas9's PAM recognition.

[0193] We then selected wild-type SaCas9 or R1015Q variants through random mutagenesis and chose PAM-specific clones with altered sites containing NNAAGT (SEQ ID NO:41) or NNAGGT (SEQ ID NO:42) PAMs (as previously described for SpCas9). We identified, re-screened, and sequenced several mutant clones that could target these PAMs, and their amino acid sequences are shown in Figure 8 (and Table 6). In summarizing these sequences, several changes appear to be significant for altering SaCas9 specificity (R1015Q, R1015H, E782K), while many other mutations may also contribute (N968K, E735K, K929R, A1021T, K1044N).

[0194] After identifying the positions and mutations necessary to alter the PAM specificity of SaCas9 for NNARRT (SEQ ID NO:43), we evaluated the contribution of the most abundant mutations to the specificity changes by preparing combinations of single, dual, and triple mutants (Table 5). When these mutations were tested against different PAMs in our positive selections (as previously described), we observed that multiple mutations allowed activity on typical NGAGT (SEQ ID NO:5) and atypical NNAAGT (SEQ ID NO:41) or NNAGGT (SEQ ID NO:42) PAMs, while the wild-type SaCas9 enzyme exhibited very low activity on atypical PAMs. Specifically, the triple mutation appears to enable relaxed specificity at the third position of the PAM (KKQ, KKH, GKQ, GKH – mutation nomenclature based on position E782 / N968 / R1015), resulting in a shared PAM motif for NNRRRT (SEQ ID NO:45) and typical NNGRRT (SEQ ID NO:46). This relaxation required by the PAM theoretically doubles the targeting range of SpCas9. Subsequently, the variants will be named based on their identifiers at positions 782, 968, and 1015. For example, E782K / N968K / R1015H will be named the SaCas9 KKH variant.

[0195] Table 5. SaCas9 mutant activity in bacterial screening

[0196]

[0197] We then evaluated two of the triple mutants in a human cell EGFP disruption assay (as previously described) to determine whether the engineered variants could target atypical PAM in a human cell setting. Figure 9Variants that target sites within the EGFP gene containing atypical PAMs disrupt the EGFP coding framework, leading to signal loss. The results revealed that the KKQ and KKH mutants maintained similar activity to wild-type SaCas9 on the typical NNGRRT (SEQ ID NO:46) PAM, but exhibited significantly higher activity on the NNARRT (SEQ ID NO:43) PAM.

[0198] Overall, we have identified mutations in SaCas9 (KKQ or KKH variants) that relax the wild-type enzyme's preference for the third position of PAM from G to R (A or G). This effectively relaxes SaCas9's targeting from NNGRRT (SEQ ID NO:46) PAM to NNRRRT (SEQ ID NO:45) PAM.

[0199] Having successfully derived variants that target the NNARRT (SEQ ID NO:43) PAM in human cells, we then proposed whether it is possible to engineer variants specific to NNCRRT (SEQ ID NO:47) or NNTRT (SEQ ID NO:48). To this end, we first mutated R1015 to E (in the case of C specified at position 3 of the PAM) and L or M (in the case of T specified at position 3 of the PAM), and tested these variants against their intended PAM in our bacterial positive selection assay (previously described). Figure 10 We observed that wild-type SaCas9 may ineffectively cleave sites containing NNCAGT (SEQ ID NO: 511) PAM, the R1015E variant has slightly better activity at the same site, and wild-type or any other directed mutation does not convey activity against other PAMs. Figure 10 This suggests that, as we saw in the case of R1015Q, additional mutations will be needed to engineer SaCas9 variants that can target NNCRRT (SEQ ID NO:47) and NNTRRT (SEQ ID NO:48) PAM.

[0200] For the evolutionary variants of SaCas9 targeting NNARRT (SEQ ID NO:43) PAM, the E782K and N968K mutations, along with R1015 (H / Q), are necessary and essential. To test whether these mutations would increase the activity of the R1015 (E / L / M) variant against its intended PAM, we generated KKE, KKL, and KKM variants. Figure 11 As shown, KKE, KKL, and KKM all exhibit robust activity toward their expected PAM.

[0201] We were also curious whether the KKQ, KKH, KKE, KKL, or KKM variants had relaxed specificity for any nucleotide at the 3rd position of PAM, so in our bacterial positive selection assay we queried multiple sites containing NNNRRT PAM. For example... Figure 11 As shown, with a few exceptions, almost all of these variants can cleave all test sites containing NNNRRT PAM. This indicates that they have relaxed specificity at position 3 of PAM, as they can effectively target NNNRRT sites. This contrasts with the wild-type protein (ENR), which can only effectively target the NGAGT (SEQ ID NO:5) site and has very low activity at several NNNRRT sites. In summary, KKH (and Figure 11 Other similar derivative variants shown can target sites in bacteria containing NNNRRT PAM, effectively quadrupling the targeting range of SaCas9.

[0202] Therefore, the KKH variant (and Figure 6 Some other variants of SaCas9 can target NNNRRT PAM in bacteria, effectively quadrupling the targeting range of SaCas9.

[0203] Table 6

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215]

[0216] Method for Example 3

[0217] Use the following materials and methods in Example 3.

[0218] plasmids and oligonucleotides

[0219] Oligonucleotides are listed in Table 11, sgRNA target sites are listed in Table 12, and plasmids used in this study are listed in Table 10.

[0220] Bacterial Cas9 / sgRNA expression plasmids were used to express human codon-optimized forms of SaCas9 and sgRNA, each expressed under a separate T7 promoter. The bacterial expression plasmids used were derived from BPK2101 (see Examples 1-2), while those used in the site depletion assay were modified to express sgRNAs with shortened repeat:anti-repeat sequences (see below). All sgRNAs in these bacterial expression plasmids included two guanines at the 5' end of the spacer sequence for proper expression from the T7 promoter. To generate SaCas9 variant libraries, the amino acid M657-G1053 of SaCas9 was randomly mutagenized using Mutazyme II (Agilent Technologies) at a frequency of approximately 5.5 mutations / kilobase. Both wild-type and R1015Q SaCas9 were used as initiation templates for mutagenesis, yielding two variants with mutations greater than 6 x 10^6. 6 A library with estimated complexity for clones.

[0221] Positive selection plasmids were assembled by ligating the oligonucleotide duplex encoding the target site into p11-lacY-wtx1 digested with XbaI / SphI (Chen, Z. & Zhao, HA highly sensitive selection method for directed evolution of homing endonucleases. Nucleic Acids Res. 33, e154 (2005)). For site depletion experiments, two separate libraries containing different spacer sequences were generated. For each library, an oligonucleotide containing 8 randomized nucleotides adjacent to the spacer sequence (in place of PAM) was compounded with the bottom-chain primer and filled with Klenow(-exo) (see Table 11). The resulting products were digested with EcoRI and ligated into p11-lacY-wtx1 digested with EcoRI / SphI. The estimated complexity of the two site-depleted libraries was greater than 4 x 10^6. 6 One clone.

[0222] For human cell experiments, human codon-optimized wild-type and variant SaCas9 were expressed from plasmids containing the CAG promoter (Table 12). sgRNA expression plasmids (containing the U6 promoter) were generated by ligating an oligonucleotide duplex encoding the spacer sequence into BsmBI-digested VVT1 (see Examples 1-2) or BPK2660 (containing a full-length 120nt crRNA:tracrRNA sgRNA or an 84nt shortened repeat:anti-repeat form, respectively). All sgRNAs used for human expression in this study included a guanine at the 5' end of the spacer to ensure proper expression from the U6 promoter, and shortened sgRNAs similar to those previously described were also used. Figure 37A -B)(Ran, FA et al. In vivo genome editing using Staphylococcus aureus Cas9. Nature 520, 186-191 (2015)).

[0223] Bioinformatics analysis of Cas9 orthologous sequences

[0224] Similar to comparisons made in previous studies (Fonfara, I. et al. Phylogeny of Cas9 determines functional exchangeability of dual-RNA and Cas9 among orthologous type II CRISPR-Cas systems. Nucleic Acids Research 42, 2577-2590 (2014); Ran, FA et al. In vivo genome editing using Staphylococcus aureus Cas9. Nature 520, 186-191 (2015); Anders, C., Niewoehner, O., Duerst, A. & Jinek, M. Structural basis of PAM-dependent target DNA recognition by the Cas9 [Structural basis of PAM-dependent target DNA recognition of Cas9 endonucleases]. Nature [Nature] 513, 569-573 (2014) , using ClustalW2 (ebi.ac.uk / Tools / msa / clustalw2 / ) to Cas9 orthologs similar to SpCas9 and SaCas9. The resulting phylogenetic tree and protein alignments were visualized using Geneious version 8.1.6 and ESPript (espript.ibcp.fr / ESPript / ESPript / ).

[0225] Based on bacterial positive selection assay

[0226] Bacterial positive selection assays were performed as previously described (see Examples 1-2). In short, the Cas9 / sgRNA plasmid was transformed into *E. coli* BW25141(λDE3) containing the positive selection plasmid (Kleinstiver et al., *Nucleic Acids Res*, 38, 2411-2427 (2010)). The transformations were plated on non-selective (chloramphenicol) and selective (chloramphenicol + 10 mM arabinose) plates. Cas9 cleavage of the selection plasmid was estimated by calculating the survival percentage: (number of colonies on the selective plate / number of colonies on the non-selective plate) x 100. To select SaCas9 variants capable of recognizing alternative PAMs, wild-type and R1015Q libraries with mutagenic PI domains were transformed into competent *E. coli* BW25141(λDE3) containing positive selection plasmids with PAMs of NNAAGT (SEQ ID NO:41), NNAGGT (SEQ ID NO:42), NNCAGT (SEQ ID NO:511), NNCGGT (SEQ ID NO:512), NNTAGT (SEQ ID NO:513), or NNTGGT (SEQ ID NO:514). Screening was performed by plate-coating with approximately 1 x 10⁻⁶ cells under selective conditions. 5 Each clone was used to prepare mini-colonies containing the presumed cleavage selection plasmid and a SaCas9 variant (MGH DNA core). All variants were re-screened individually in a positive selection assay, and those with a survival rate greater than approximately 20% were sequenced to identify the mutations required to replace PAM.

[0227] Bacterial site depletion assay

[0228] Site depletion experiments were performed as previously described (see Examples 1-2). In short, randomized PAM libraries were electroporated into competent *E. coli* BW25141(λDE3) cells containing either the catalytically inactive wild-type (D10A / H557A) or KKH variant SaCas9 / sgRNA plasmid. A sample size greater than 1 x 10⁻⁶ was used. 5Colonies were plated on chloramphenicol / carbenicillin plates, and the selection plasmids containing Cas9-targeting resistance PAMs within the surviving colonies were isolated using maxiprep (Kaiger). PCR amplification was performed on regions of the plasmids containing spacer sequences and PAMs using primers listed in Table 11. Double-exponential Tru-seq Illumina sequencing libraries were generated using the KAPA HTP Library Preparation Kit (KAPA Biosystems) prior to Illumina MiSeq high-throughput sequencing run at the Dana-Farber Core Cancer Institute for Molecular Biology, using approximately 500 ng of purified PCR product from each site depletion condition. Data from the site depletion experiments were analyzed as previously described (see Examples 1-2), except for script modifications to analyze 8 randomized nucleotides. The ability of Cas9 to recognize PAMs was determined by calculating the post-selection PAM depletion value (PPDV) for any given PAM: the ratio of the post-selection frequency of that PAM to the pre-selected library frequency. A control experiment using non-catalytically active SaCas9 was used to determine that a PPDV of 0.794 indicates statistically significant depletion relative to the input library.

[0229] Human cell culture and transfection

[0230] U2OS cells obtained from our collaborators T. Cathomen (Freiburg) and U2OS.EGFP cells containing a single integrated copy of the EGFP-PEST reporter gene (Reyon, D. et al. FLASH assembly of TALENs for high-throughput genome editing. Nat Biotechnol 30, 460-465 (2012)) were cultured at 37°C and 5% CO2 in advanced DMEM medium (Lifetechnologies) containing 10% FBS, penicillin / streptomycin, and 2 mM GlutaMAX (Lifetechnologies). Cell line integrity was verified by STR analysis (ATCC) and deep sequencing, and mycoplasma contamination of cells was tested every two weeks. The U2OS.EGFP medium was supplemented with 400 μg / mL G418. According to the manufacturer's instructions, cells were co-transfected with 750 ng Cas9 plasmid and 250 ng sgRNA plasmid using the DN-100 program on the Lonza 4D nuclear transfection instrument.

[0231] Human cell EGFP destruction assay

[0232] EGFP disruption experiments were performed as previously described (Fu, Y. et al. High-frequency off-target mutagenesis induced by CRISPR-Cas nucleases in human cells. Nat Biotechnol 31, 822-826 (2013); Reyon, D. et al. FLASH assembly of TALENs for high-throughput genome editing. Nat Biotechnol 30, 460-465 (2012)). Approximately 52 hours post-transfection, EGFP fluorescence in transfected U2OS.EGFP cells was measured using a Fortessa flow cytometer (BD Biosciences). Negative control transfections with Cas9 and empty U6 promoter plasmids were used to determine approximately 2.5% background EGFP loss for all experiments (indicated by the red dashed line in the figure).

[0233] T7E1 determination

[0234] As previously described, T7E1 assays (Reyon, D. et al. FLASH assembly of TALENs for high-throughput genome editing. Nat Biotechnol 30, 460-465 (2012)) were performed to quantify Cas9-induced mutagenesis at endogenous loci in human cells. Approximately 72 hours post-transfection, genomic DNA was isolated using the Agencourt DNAdvance Genomic DNA Isolation Kit (Beckman Coulter Genomics). PCR amplification of the target loci was performed from approximately 100 ng of genomic DNA using primers listed in Table 11. Following an Agencourt Ampure XP cleanup step (Beckman Coulter Genomics), approximately 200 ng of purified PCR product was denatured and heterozygous before digestion with T7E1 (New England Biolabs). Following a second cleanup step, mutagenesis frequencies were quantified using a Qiaxcel capillary electrophoresis instrument (Kiagen).

[0235] GUIDE-seq experiment

[0236] As previously described, GUIDE-seq experiments were performed and analyzed (Tsai, SQ et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat Biotechnol 33, 187-197 (2015)). In short, U2OS cells were transfected with Cas9 and sgRNA plasmids and 100 pmol of phosphorylated phosphate-thioester modified double-stranded oligodeoxynucleotides (dsODN) with an NdeI site, as described above. Restriction fragment length polymorphism (RFLP) analysis was performed to determine the frequency of dsODN-tag integration (see Examples 1-2; Tsai, SQ et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases [GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases]. Nat Biotechnol [Nature Biotechnology] 33, 187-197 (2015)), and T7E1 assays were performed to quantify the frequency of targeted Cas9 mutagenesis. Prior to high-throughput sequencing using the Illumina MiSeq sequencer, dsODN tag-specific amplification and library preparation were performed (Tsai, SQ et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat Biotechnol 33, 187-197 (2015)). When mapping potential off-target sites, the cutoff values ​​for alignment with the mid-target spacer sequence were set at 8 mismatches for 21-nucleotide spacers, 9 mismatches for 22-nucleotide spacers, and 10 mismatches for 23-nucleotide spacers.Identification of off-target sites with potential DNA or RNA protrusions by manual comparison (Lin, Y. et al. CRISPR / Cas9 systems have off-target activity with insertions or deletions between target DNA and guide RNA sequences. Nucleic Acids Res. 42, 7473-7485 (2014)).

[0237] Table 10 – Plasmids used in Example 3

[0238]

[0239]

[0240]

[0241] Table 11 – Oligonucleotides used in Example 3

[0242]

[0243]

[0244] Table 12 – sgRNA target sites for Example 3

[0245]

[0246]

[0247]

[0248] *exist Figure 1C and 1E , Figure 32 Use

[0249] **For use in Figure 3, Figure 36A GUIDE-seq experiments in -B

[0250] Example 3. PAM specificity of engineered Staphylococcus aureus Cas9

[0251] Site-specific DNA cleavage by the CRISPR-Cas9 nuclease is primarily guided by RNA-DNA interactions, but also requires Cas9-mediated recognition of the prototypical spacer adjacent motif (PAM). While the commonly used *Streptococcus pyogenes* Cas9 specifies only two nucleotides within its NGG PAM, other Cas9 orthologs with the desired properties recognize longer PAMs. While potentially advantageous from a specificity standpoint, extending the PAM sequence can limit the targeting range of Cas9 orthologs for genome editing applications. One possible strategy to broaden the range of sequences that these Cas9 orthologs can target might be to evolve variants with relaxed specificity at certain positions within the PAM. Here we use molecular evolution to modify the NNGRRT (SEQ ID NO: 46) PAM specificity of *Staphylococcus aureus* Cas9 (SaCas9) for applications requiring viral delivery. One variant we identified, termed KKH SaCas9, exhibits robust genome editing activity with an endogenous human target at the NNNRRT PAM. Importantly, using the GUIDE-seq method, we show that both wild-type and KKH SaCas9 induce a significant number of off-target effects in human cells. KKH SaCas9 increases the targeting range of SaCas9 by nearly two to four times, enabling the targeting of sequences that cannot be altered by wild-type nucleases. More generally, these results demonstrate the feasibility of relaxing PAM specificity to broaden the targeting range of Cas9 orthologs. Our molecular evolutionary strategy does not require access to structural information or prior knowledge of specific residues in PAM and should therefore be applicable to a wide range of Cas9 orthologs.

[0252] result

[0253] We envision an unbiased genetic approach for engineering Cas9 variants with relaxed PAM recognition specificity that does not require structural information. We tested this strategy using SaCas9 because no structural data were available when we initiated these studies. In the initial steps, we attempted to conservatively estimate the PAM interaction domains of SaCas9 by comparing the sequence with that of well-characterized SpCas9 (Jiang et al., Science 348, 1477-1481 (2015); Anders et al., Nature 513, 569-573 (2014); Jinek et al., Science (2014); Nishimasu et al., Cell (2014)). Although SpCas9 and SaCas9 are significantly different at the major sequence level ( Figure 21A , Figure 29However, comparisons of both with 10 other orthogonal homologs allow us to conservatively define the predicted PAM interaction domains for SaCas9 (see the method for Example 3). Figure 29 and 30).

[0254] Because the guanine at the third position in SaCas9 PAM is the most strictly designated base (Ran et al., Nature 520, 186-191 (2015)), we randomly mutagenesis predicted the PI domain and used our previously described bacterial cell-based approach (see Examples 1-2) to attempt to select the nucleotides that could cleave three other possible nucleotides with the third PAM position (i.e., NN[A / C / T]RRT PAM (NNHRRT (SEQ ID NO:44)); Figure 31A Mutants at each of the sites in the SpCas9 PAM were obtained. All except one of the surviving variants from selection for sites containing NNARRT (SEQ ID NO:43) and NNCRRT (SEQ ID NO:47) PAM contained the R1015H mutation; however, no variants were obtained from selection for NNTRRT (SEQ ID NO:48) PAM. These results strongly suggest that R1015 may be involved in the recognition of guanine at the third position in the SpCas9 PAM. In fact, in our comparison, we found that R1015 of SaCas9 is near R1335 of SpCas9 (Fig. 30), which is the residue previously involved in the third base position of PAM recognition (see Examples 1-2; Anders, C., Niewoehner, O., Duerst, A. & Jinek, M. Structural basis of PAM-dependent target DNA recognition by the Cas9 endonuclease. Nature 513, 569-573 (2014)). Consistent with this, we found that when tested in our bacterial selection system ( Figure 31B When R1015 is mutated to alanine or glutamine, the activity of SaCas9 at the target site containing NNGRRT (SEQ ID NO:46) PAM is significantly reduced. Figure 21B Substitution of other positively charged residues, such as alanine or glutamine, near R1015 has no strong effect on SaCas9 activity. Figure 21B (Figure 30).

[0255] Our bacterial selection results also suggest that the R1015H mutation may at least partially relax the specificity of SaCas9 at the third position in the PAM. However, when tested against sites with any nucleotides at the third position in the NNNRRT PAM ( Figure 21C We found that the R1015H single mutant exhibited suboptimal activity in our previously described human cell-based EGFP disruption assay (Fu et al., Nat Biotechnol [Nature Biotechnology] 31, 822-826 (2013); Reyon et al., Nat Biotechnol [Nature Biotechnology] 30, 460-465 (2012)). Because this suggests that additional mutations may be needed to increase or optimize the activity of the R1015H mutant in human cells, we randomly mutagenized a region of SaCas9 containing the predicted PI domain, which also contained the R1015Q mutation. Then, using our bacterial selection system, we selected variants from this library that could cleave target sites with each of the three different NNHRRT (SEQ ID NO:44) PAMs. We used R1015Q because, unlike R1015H, this mutant did not show activity in bacteria (…). Figure 21B Although no surviving clones were again observed when selection was performed against NNTRT (SEQ ID NO:48) PAM, selection against the R1015Q variant against NNARRT (SEQ ID NO:43) or NNCRRT (SEQ ID NO:47) produced mutations at E782, K929, and N968, and unexpectedly, the Q at 1015 was mutated to H.

[0256] Combined with the selection results of wild-type SaCas9, the most common missense mutations identified across all selections were E782K, K929R, N968K, and R1015H. Figure 21D This suggests that combinations of these mutations may allow for efficient cleavage at the third position of the SaCas9 PAM containing either A or C. Therefore, we used a human cell-based EGFP disruption assay to test SaCas9 variants containing different combinations of these mutations, where the sgRNA targets a site containing each of the four bases at the third position of the PAM (i.e., on the NNNRRT PAM). Figure 21E , Figure 32 We found that variants with the triple mutant combinations E782K / N968K / R1015H and E782K / K929R / R1015H exhibited high activity at sites containing NNNRRT PAM. Figure 21E , Figure 32), while quadruple mutant variants containing all four mutations (E782K / K929R / N968K / R1015H) generally have lower activity at these sites. Figure 32 We selected E782K / N968K / R1015H (hereinafter referred to as the KKH variant) for further characterization and verified using our human cell-based EGFP disruption assay that all three substitutions of this KKH variant are required for activity. Figure 21E ).

[0257] To more comprehensively define the PAM specificity of KKH and wild-type SaCas9, we used our previously described site depletion assay based on bacterial cells (see Examples 1-2). Figure 33 This method generates a Cas9 PAM-specific signature by identifying the relative cleavage (and thus depletion) of DNA plasmids carrying randomized PAM sequences, quantified as the post-selection PAM depletion value (PPDV). We performed site depletion experiments on wild-type and KKH SaCas9 using libraries with two different spacer sequences, each containing 8 randomized bases (in place of PAM). Figure 33 Control experiments using inactive SaCas9 showed minimal depletion of any PAM sequence. Figure 34A This allows us to determine the threshold for statistically significant exhaustion, with a PPDV of 0.794. Figure 34B Previous experiments have shown that PAMs with PPDV < 0.2 can be efficiently cleaved in our cell site depletion assays (see Examples 1-2). In the case of wild-type SaCas9, as expected, the most depleted PAMs (based on the mean PPDV obtained from both libraries) were four NNGRRT (SEQ ID NO: 46) PAMs. Figure 21F and Figure 34C Interestingly, other PAMs with a mean PPDV < 0.1 include those with the form NNGRRN (SEQ ID NO: 49). Figure 34DThis indicates that in our bacterial-based assay, the final position of PAM may not always be designated as T for some spacer subsequences (although previous reports have demonstrated that thymine at position 6 of PAM is highly preferred via in vitro PAM depletion assays, ChIP-seq, and targeting of endogenous human sites (Ran, FA et al. In vivo genome editing using Staphylococcus aureus Cas9. Nature 520, 186-191 (2015))). In contrast, in the case of the KKH variant, PAMs with a mean PPDV <0.2 include not only the NNGRRT (SEQ ID NO:46) PAM, but also three of the four NNARRT (SEQ ID NO:43), all four NNCRRT (SEQ ID NO:47), and four NNTRRT (SEQ ID NO:48) PAMs. Figure 21F , Figure 34C and 34E These results indicate that KKH SaCas9 exhibits an expanded PAM targeting range relative to its wild-type counterpart.

[0258] To assess the robustness of the KKH SaCas9 variant in human cells, we tested its activity at 55 different endogenous gene target sites containing multiple NNNNRRTPAMs. Figure 22A The KKH variant showed potent activity, with a mean mutagenic frequency of 24.7% across all sites, of which 80% of sites (44 out of 55 sites) showed greater than 5% disruption. Analysis of KKH SaCas9 activity across all 55 sites revealed PAM(NN[G>A=C>T]RRT; Figure 22B The 3rd position and PAM(NNN[AG>GG>GA>AA]T; Figure 22C The 4th / 5th position ranking preference. Consistent with this, we observed differences among the 16 possible combinations of the 3rd / 4th / 5th positions of NNNRRT PAM ( Figure 35A KKH SaCas9 supports spacer lengths ranging from 21 to 23 nucleotides. Figure 22D ), spacer subsequences with variable GC content ( Figure 35B ) and PAM with variable GC content ( Figure 35CIt works effectively. Sequence identifiers derived from sites cleaved at low, medium, and high efficiencies (mean mutagenic frequencies of 0-10%, 10%-30%, and >30%, respectively) reveal a small sequence preference across the entire target site except for positions 4 and 5 of the NNNRRT PAM, and may have a slight preference for guanine at the second PAM position on sites cleaved at high efficiency. Figure 35D ).

[0259] To demonstrate that the KKH variant can modify PAMs that cannot be targeted by wild-type SaCas9, we performed direct comparisons of these nucleases at sites carrying different NNNRRT PAMs in human cells. Sixteen sites were evaluated using our EGFP disruption assay and 16 endogenous human gene targets (respectively...). Figure 22E and 22F The results showed that KKH SaCas9 robustly modified target sites carrying NNNRRT PAM, while wild-type SaCas9 effectively targeted only sites with NNGRRT (SEQ ID NO:46) PAM. For all 24 sites with NNHRRT (SEQ ID NO:44) PAM, the KKH variant induced significantly higher mutagenesis rates than wild-type SaCas9; at the 8 sites with NNGRRT (SEQ ID NO:46) PAM, KKH SaCas9 induced comparable or slightly lower levels of mutagenesis compared to wild-type. Figure 22E and 22F These results collectively demonstrate that the KKH variant can cleave sites with NNHRRT PAM, thereby enabling targeting of sites with NNHRRT (SEQ ID NO:44) PAM that are currently not effectively altered in wild-type SaCas9 in human cells.

[0260] To assess the impact of the KKH mutation on the genome-wide specificity of SaCas9, we used the GUIDE-seq (Genome-wide unbiased identification of DSBs via sequencing) method (Tsai, SQ et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases [GUIDE-seq enables genome-wide analysis of off-target cleavage by CRISPR-Cas nucleases]. Nat Biotechnol [Nature Biotechnology] 33, 187-197 (2015)) to directly compare the off-target characteristics of wild-type and KKH SaCas9 with the same sgRNA. When testing sgRNAs targeting six endogenous human gene loci containing NNGRRT (SEQ ID NO:46) PAM, we observed that wild-type and KKH SaCas9 induced nearly identical GUIDE-seq tag integration rates and on-target cleavage frequencies at all six loci (respectively...). Figure 36A and 36B Furthermore, in each of the six sgRNAs, wild-type and KKHSaCas9 induced mutations at a similar number of off-target sites. Figure 23A and 23B The off-target sites of the KKH variant are typically attached to the NNNRRT PAM motif, while the off-target sites of wild-type SaCas9 are attached to the NNGRR[T>G] motif. Figure 22B In the case of the sgRNA that induced the highest number of off-target sites among the six sgRNAs tested, we observed a similar number of off-target sites in wild-type and KKH SaCas9. However, these off-target sites only partially overlapped between wild-type and KKH SaCas9, as could be expected given their different PAM specificities. Figure 23B and 23C While we would not advocate using the KKH variant to target sites with NNGRRT (SEQ ID NO:46) PAM (because wild-type SaCas9 can exhibit higher on-target activity than KKH for these sites), these results suggest that KKH SaCas9 cleaves only off-target sites with the expected PAM and typically induces a comparable number of off-target sites as observed with wild-type SaCas9.

[0261] To further examine the genome-wide specificity of KKH SaCas9, we tested five additional sgRNAs targeting sites containing NNHRRT (SEQ ID NO:44) PAM. Figure 23D and 23EThe number of off-target sites detected by GUIDE-seq is typically low (comparable to the numbers observed in previously published experiments on wild-type SpCas9 and SpCas9 variants) (see Examples 1-2 (Tsai, SQ et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat Biotechnol 33, 187-197 (2015)), showing potential for prominent DNA and RNA off-target effects (Lin, Y. et al. CRISPR / Cas9 systems have off-target activity with insertions or deletions between target DNA and guide RNA sequences). Nucleic Acids Res [Nucleic Acid Research] 42, 7473-7485 (2014)), and contains the expected PAM sequence. In summary, our experiments demonstrate that wild-type and KKH SaCas9 are similar in genome-wide specificity and, as determined by GUIDE-seq, typically show a low number of off-target mutations in human cells.

[0262] While wild-type SaCas9 remains the best choice for targeting the NNGRRT (SEQ ID NO:46) PAM, the KKH SaCas9 variant we describe here robustly targets sites with NNARRT (SEQ ID NO:43) and NNCRRT (SEQ ID NO:47) PAMs, and has a reasonable success rate for sites with NNTRRT (SEQ ID NO:48) PAM. Therefore, we conservatively estimate that the KKH variant increases the targeting range of SaCas9 by nearly two to four times in random DNA sequences, thereby improving the prospect of wider utilization of SaCas9 in a variety of different applications requiring high-precision targeting. Using GUIDE-seq, we demonstrate that when targeting the same site containing the NNGRRT (SEQ ID NO:46) PAM, KKH SaCas9 induces a similar number of off-target mutations as wild-type SaCas9. Furthermore, KKH SaCas9 induces only a small number of off-target mutations when targeting sites carrying the NNHRRT (SEQ ID NO:44) PAM. Although KKH SaCas9 recognizes the modified PAM sequence relative to wild-type SaCas9, our findings are not entirely unexpected, given that the combined length of the prototypical spacer and PAM is still sufficiently long, and the KKH variant (24 to 26 bp) should theoretically be orthologous to the human genome. Furthermore, it is possible that modified PAM recognition could be improved by altering the energy of the interaction between Cas9 / sgRNA and its target site (similar to previously proposed mechanisms for improving the specificity of truncated sgRNAs (Fu, Y., Sander, JD, Reyon, D., Cassio, VM & Joung, JK Improving CRISPR-Cas nuclease specificity using truncated guideRNAs). Nat Biotechnol 32, 279-284 (2014)) or the specificity of the D1135E SpCas9 mutant (see Examples 1-2).

[0263] Example 4. Improved activity of the SpCas9-VQR variant

[0264] Because the SpCas9-VQR variants exhibit an NGAN PAM preference of NGAG > NGAA = NGAT > NGAC, we sought to select derivative variants with improved activity against NGAH PAM (where H = A, C, or T). Selection of R1335Q libraries (PI domain random mutagenesis) against cells containing target sites for NGAA, NGAT, or NGAC PAMs allowed us to sequence additional clones containing mutations that convey altered PAM specificity. The sequences of these clones revealed additional mutations that may be important for altering PAM specificity against NGAA, NGAT, or NGAC PAMs.

[0265] Based on the results of these selections, VQR variants and 24 other derivative variants were tested against NGAG, NGAA, NGAT, and NGAC PAM sites in bacteria. Many of these derivative variants survived better than the VQR variants at the NGAH PAM site, and most of them contained the G1218R mutation (Table 7 and...). Figure 24 ).

[0266] Table 7 lists the variants and their corresponding amino acid changes.

[0267]

[0268] Given that results from bacterial screening demonstrated that some of these additional mutations improved activity at the NGAH PAM site, we tested some of the best candidates in human cells in an EGFP disruption assay. We observed that several of these variants (including VRQR, NRQR, and YRQR variants) were superior to the VQR variant in targeting the NGAH site (Table 8 and...). Figure 25 The main difference between these clones and VQR variants is that they include the G1218R mutation.

[0269] Table 8. SpCas9-VQR derivatives and their corresponding amino acid changes.

[0270]

[0271] Since the VRQR variant appears to be the most robust among those tested, we compared its activity to that of VQR against nine different endogenous sites in human cells (two sites each for NGAA, NGAC, NGAT, and NGAG PAM, and one site for NGCGPAM). This data reveals that the VRQR variant outperformed the VQR variant at all sites tested in human cells. Figure 26 ).

[0272] Having demonstrated the improved activity of the VRQR variant compared to the VQR variant, we sought to determine whether adding additional substitutions could further improve activity. Because we observed additional mutations in the selection very close to the PAM interaction pocket of SpCas9, a subgroup of these mutations was added to both the VQR and VRQR variants, and bacteria were screened for sites containing NGAG, NGAA, NGAT, and NGAC PAM (Table 9 and...). Figure 27 Several derivative variants showed high activity against the NGAT and NGAC PAM sites, so we continued to test these variants in human cells. We tested additional variants containing mutations added to the VQR or VRQR background in human cell EGFP disruption assays. These experiments again revealed that VRQR variants exhibited more robust activity against NGAH PAM than VQR variants, and that additional mutations in the VRQR backbone were beneficial.

[0273] Table 9. Variants and their corresponding amino acid changes

[0274]

[0275] In summary, these results indicate that including additional mutations in the SpCas9-VQR variant can improve activity against sites containing NGAN PAM, particularly those containing NGAH PAM.

[0276] References

[0277] 1. Sander, JD & Joung, JK. CRISPR-Cas systems for editing, regulating and targeting genomes. Nat Biotechnol 32, 347-355 (2014).

[0278] 2. Hsu, PD, Lander, ES & Zhang, F. Development and applications of CRISPR-Cas9 for genome engineering. Cell 157, 1262-1278 (2014).

[0279] 3. Doudna, J.A.A. & Charpentier, E. Genome editing. The new frontier of genome engineering with CRISPR-Cas9. Science 346, 1258096 (2014).

[0280] 4. Barrangou, R. & May, A.P. Unraveling the potential of CRISPR-Cas9 for gene therapy. Expert Opinion on Biotherapy 15, 311-314 (2015).

[0281] 5. Jinek, M. et al. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science 337, 816-821 (2012).

[0282] 6. Sternberg, SH, Redding, S., Jinek, M., Greene, EC & Doudna, JA. DNA interrogation by the CRISPR RNA-guided endonuclease Cas9. Nature 507, 62-67 (2014).

[0283] 7. Hsu, PD et al. DNA targeting specificity of RNA-guided Cas9 nucleases. Nat Biotechnol 31, 827-832 (2013).

[0284] 8. Tsai, SQ et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat Biotechnol 33, 187-197 (2015).

[0285] 9. Hou, Z. et al. Efficient genome engineering in human pluripotent stem cells using Cas9 from Neisseria meningitidis. Proc Natl Acad Sci USA (2013).

[0286] 10. Fonfara, I. et al. Phylogeny of Cas9 determines functional exchangeability of dual-RNA and Cas9 among orthologous type II CRISPR-Cas systems. Nucleic Acids Res. 42, 2577-2590 (2014).

[0287] 11. Esvelt, KM et al. Orthogonal Cas9 proteins for RNA-guided gene regulation and editing. NatMethods 10, 1116-1121 (2013).

[0288] 12. Cong, L. et al. Multiple genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013).

[0289] 13. Horvath, P. et al. Diversity, activity, and evolution of CRISPR loci in Streptococcus thermophilus. J Bacteriol 190, 1401-1412 (2008).

[0290] 14. Anders, C., Niewoehner, O., Duerst, A. & Jinek, M. Structural basis of PAM-dependent target DNA recognition by the Cas9 endonuclease. Nature, 513, 569-573 (2014).

[0291] 15. Reyon, D. et al. FLASH assembly of TALENs for high-throughput genome editing. Nat Biotechnol 30, 460-465 (2012).

[0292] 16. Fu, Y. et al. High-frequency off-target mutagenesis induced by CRISPR-Cas nucleases in human cells. Nat Biotechnol 31, 822-826 (2013).

[0293] 17. Chen, Z. & Zhao, HA. Highly sensitive selection method for directed evolution of homing endonucleases. Nucleic Acids Res. 33, e154 (2005).

[0294] 18. Doyon, JB, Pattanayak, V., Meyer, CB & Liu, DR. Directed evolution and substrate specificity profile of homing endonuclease I-SceI. J Am Chem Soc. 128, 2477-2484 (2006).

[0295] 19. Jiang, W., Bikard, D., Cox, D., Zhang, F. & Marraffini, LA RNA-guidedediting of bacterial genomes using CRISPR-Cas systems. Nat Biotechnol 31, 233-239 (2013).

[0296] 20. Mali, P. et al. RNA-guided humangenome engineering via Cas9, Science 339, 823-826 (2013).

[0297] 21. Hwang, WY et al. Efficient genome editing in zebrafish using a CRISPR-Cas system. Nat Biotechnol 31, 227-229 (2013).

[0298] 22. Chylinski, K., Le Rhun, A. & Charpentier, E. The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems. RNA Biol 10, 726-737 (2013).

[0299] 23. Kleinstiver, BP, Fernandes, AD, Gloor, GB & Edgell, DRA. A unified genetic, computational, and experimental framework identifies functionally relevant residues of the homing endonuclease I-BmoI. Nucleic Acids Res. 38, 2411-2427 (2010).

[0300] 24. Gagnon, JA et al. Efficient mutagenesis by Cas9 protein-mediated oligonucleotide insertion and large-scale assessment of single-guide RNAs. PLoS One 9, e98186 (2014).

[0301] Other embodiments

[0302] It should be understood that although the invention has been described in conjunction with its detailed description, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the appended claims. Other aspects, advantages, and modifications are all within the scope of the following claims. sequence list <110> GE Healthcare <120> Engineered CRISPR-Cas9 nuclease with modified PAM specificity <130> 29539-0163WO1 <150> US 61 / 127,634 <151> 2015-03-03 <150> US 62 / 165,517 <151> 2015-05-22 <150> US 62 / 239,737 <151> 2015-10-09 <150> US 62 / 258,402 <151> 2015-11-20 <160> 28 <170> SIPOSequenceListing 1.0 <210> 1 <211> 1368 <212> PRT <213> Streptococcus pyogenes (PRT) <220> <221> PEPTIDE <223> PRT <400> 1 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asp With Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu With Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915,920,925 Lys Tyr Asp Ser Arg Met Asn Thr Lys Tyr Asp 930,935,940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965,970,975 Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val 980,985,990 Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys 1010 1015 1020 Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser 1025 1030 1035 1040 Asn With Asn Phe Phe Lys Thr Glu With Thr Leu Ala Asn Glu 1045 1050 1055 Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile 1060 1065 1070 Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser 1075 1080 1085 Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1090 1095 1100 Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu Ile 1105 1110 1115 1120 Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe Asp Ser 1125 1130 1135 Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val Glu Lys Gly 1140 1145 1150 Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu Gly Ile Thr Ile 1155 1160 1165 Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala 1170 1175 1180 Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys 1185 1190 1195 1200 Tyr Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser 1205 1210 1215 Ala Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr 1220 1225 1230 Val Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His 1250 1255 1260 Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val 1265 1270 1275 1280 Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys 1285 1290 1295 His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu 1300 1305 1310 Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp 1315 1320 1325 Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp 1330 1335 1340 Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg Ile 1345 1350 1355 1360 Asp Leu Ser Gln Leu Gly Gly Asp 1365 <210> 2 <211> 1053 <212> PRT <213> Staphylococcus aureus (PRT) <220> <221> PEPTIDE <223> PRT <400> 2 Met Lys Arg Asn Tyr Ile Leu Gly Leu Asp Ile Gly Ile Thr Ser Val 1 5 10 15 Gly Tyr Gly Ile Ile Asp Tyr Glu Thr Arg Asp Val Ile Asp Ala Gly 20 25 30 Val Arg Leu Phe Lys Glu Ala Asn Val Glu Asn Asn Glu Gly Arg Arg 35 40 45 Ser Lys Arg Gly Ala Arg Arg Leu Lys Arg Arg Arg Arg His Arg Ile 50 55 60 Gln Arg Val Lys Lys Leu Leu Phe Asp Tyr Asn Leu Leu Thr Asp His 65 70 75 80 Ser Glu Leu Ser Gly Ile Asn Pro Tyr Glu Ala Arg Val Lys Gly Leu 85 90 95 Ser Gln Lys Leu Ser Glu Glu Glu Phe Ser Ala Ala Leu Leu His Leu 100 105 110 Ala Lys Arg Arg Gly Val His Asn Val Asn Glu Val Glu Glu Asp Thr 115 120 125 Gly Asn Glu Leu Ser Thr Lys Glu Gln Ile Ser Arg Asn Ser Lys Ala 130 135 140 Leu Glu Glu Lys Tyr Val Ala Glu Leu Gln Leu Glu Arg Leu Lys Lys 145 150 155 160 Asp Gly Glu Val Arg Gly Ser Ile Asn Arg Phe Lys Thr Ser Asp Tyr 165 170 175 Val Lys Glu Ala Lys Gln Leu Leu Lys Val Gln Lys Ala Tyr His Gln 180 185 190 Leu Asp Gln Ser Phe Ile Asp Thr Tyr Ile Asp Leu Leu Glu Thr Arg 195 200 205 Arg Thr Tyr Tyr Glu Gly Pro Gly Glu Gly Ser Pro Phe Gly Trp Lys 210 215 220 Asp Ile Lys Glu Trp Tyr Glu Met Leu Met Gly His Cys Thr Tyr Phe 225 230 235 240 Pro Glu Glu Leu Arg Ser Val Lys Tyr Ala Tyr Asn Ala Asp Leu Tyr 245 250 255 Asn Ala Leu Asn Asp Leu Asn Asn Leu Val Ile Thr Arg Asp Glu Asn 260 265 270 Glu Lys Leu Glu Tyr Tyr Glu Lys Phe Gln Ile Ile Glu Asn Val Phe 275 280 285 Lys Gln Lys Lys Lys Pro Thr Leu Lys Gln Ile Ala Lys Glu Ile Leu 290 295 300 Val Asn Glu Glu Asp Ile Lys Gly Tyr Arg Val Thr Ser Thr Gly Lys 305 310 315 320 Pro Glu Phe Thr Asn Leu Lys Val Tyr His Asp Ile Lys Asp Ile Thr 325 330 335 Ala Arg Lys Glu Ile Ile Glu Asn Ala Glu Leu Leu Asp Gln Ile Ala 340 345 350 Lys Ile Leu Thr Ile Tyr Gln Ser Ser Glu Asp Ile Gln Glu Glu Leu 355 360 365 Thr Asn Leu Asn Ser Glu Leu Thr Gln Glu Glu Ile Glu Gln Ile Ser 370 375 380 Asn Leu Lys Gly Tyr Thr Gly Thr His Asn Leu Ser Leu Lys Ala Ile 385 390 395 400 Asn Leu Ile Leu Asp Glu Leu Trp His Thr Asn Asp Asn Gln Ile Ala 405 410 415 Ile Phe Asn Arg Leu Lys Leu Val Pro Lys Lys Val Asp Leu Ser Gln 420 425 430 Gln Lys Glu Ile Pro Thr Thr Leu Val Asp Asp Phe Ile Leu Ser Pro 435 440 445 Val Val Lys Arg Ser Phe Ile Gln Ser Ile Lys Val Ile Asn Ala Ile 450 455 460 Ile Lys Lys Tyr Gly Leu Pro Asn Asp Ile Ile Ile Glu Leu Ala Arg 465 470 475 480 Glu Lys Asn Ser Lys Asp Ala Gln Lys Met Ile Asn Glu Met Gln Lys 485 490 495 Arg Asn Arg Gln Thr Asn Glu Arg Ile Glu Glu Ile Ile Arg Thr Thr 500 505 510 Gly Lys Glu Asn Ala Lys Tyr Leu Ile Glu Lys Ile Lys Leu His Asp 515 520 525 Met Gln Glu Gly Lys Cys Leu Tyr Ser Leu Glu Ala Ile Pro Leu Glu 530 535 540 Asp Leu Leu Asn Asn Pro Phe Asn Tyr Glu Val Asp His Ile Ile Pro 545 550 555 560 Arg Ser Val Ser Phe Asp Asn Ser Phe Asn Asn Lys Val Leu Val Lys 565 570 575 Gln Glu Glu Asn Ser Lys Lys Gly Asn Arg Thr Pro Phe Gln Tyr Leu 580,585,590 Ser Ser Ser Asp Ser Lys Ile Ser Tyr Glu Thr Phe Lys Lys His Ile 595,600,605 Leu Asn Leu Ala Lys Gly Lys Gly Arg Is Thr Lys Glu 610 615 620 Tyr Leu Leu Glu Glu Arg Asp With Asn Arg Phe Serving Val Gln Lys Asp 625 630 635 640 Phe Ile Asn Arg Asn Leu Val Asp Thr Arg Tyr Ala Thr Arg Gly Leu 645,650,655 Met Asn Leu Leu Arg Ser Tyr Phe Arg Val Asn Asn Leu Asp Val Lys 660,665,670 Val Lys Ser Ile Asn Gly Gly Phe Thr Ser Phe Leu Arg Arg Lys Trp 675,680,685 Lys Phe Lys Lys Glu Arg Asn Lys Gly Tyr Lys His Ala Glu Asp 690,695,700 Only Leu Ile On Asn Only Asp Phe Ile Phe Lys Glu Trp Lys Lys 705 710 715 720 Leu Asp Lys Ala Lys Lys Val Met Glu Asn Gln Met Phe Glu Glu Lys 725 730 735 Gln Ala Glu Ser Met Pro Glu Ile Glu Thr Glu Gln Glu Tyr Lys Glu 740 745 750 Ile Phe Ile Thr Pro His Gln Ile Lys His Ile Lys Asp Phe Lys Asp 755 760 765 Tyr Lys Tyr Ser His Arg Val Asp Lys Lys Pro Asn Arg Glu Leu Ile 770 775 780 Asn Asp Thr Leu Tyr Ser Thr Arg Lys Asp Asp Lys Gly Asn Thr Leu 785 790 795 800 Ile Val Asn Asn Leu Asn Gly Leu Tyr Asp Lys Asp Asn Asp Lys Leu 805 810 815 Lys Lys Leu Ile Asn Lys Ser Pro Glu Lys Leu Leu Met Tyr His His 820 825 830 Asp Pro Gln Thr Tyr Gln Lys Leu Lys Leu Ile Met Glu Gln Tyr Gly 835 840 845 Asp Glu Lys Asn Pro Leu Tyr Lys Tyr Tyr Glu Glu Thr Gly Asn Tyr 850 855 860 Leu Thr Lys Tyr Ser Lys Lys Asp Asn Gly Pro Val Ile Lys Lys Ile 865 870 875 880 Lys Tyr Tyr Gly Asn Lys Leu Asn Ala His Leu Asp Ile Thr Asp Asp 885 890 895 Tyr Pro Asn Ser Arg Asn Lys Val Val Lys Leu Ser Leu Lys Pro Tyr 900 905 910 Arg Phe Asp Val Tyr Leu Asp Asn Gly Val Tyr Lys Phe Val Thr Val 915 920 925 Lys Asn Leu Asp Val Ile Lys Lys Glu Asn Tyr Tyr Glu Val Asn Ser 930 935 940 Lys Cys Tyr Glu Glu Ala Lys Lys Leu Lys Lys Ile Ser Asn Gln Ala 945 950 955 960 Glu Phe Ile Ala Ser Phe Tyr Asn Asn Asp Leu Ile Lys Ile Asn Gly 965 970 975 Glu Leu Tyr Arg Val Ile Gly Val Asn Asn Asp Leu Leu Asn Arg Ile 980 985 990 Glu Val Asn Met Ile Asp Ile Thr Tyr Arg Glu Tyr Leu Glu Asn Met 995 1000 1005 Asn Asp Lys Arg Pro Pro Arg Ile Ile Lys Thr Ile Ala Ser Lys Thr 1010 1015 1020 Gln Ser Ile Lys Lys Tyr Ser Thr Asp Ile Leu Gly Asn Leu Tyr Glu 1025 1030 1035 1040 Val Lys Ser Lys Lys His Pro Gln Ile Ile Lys Lys Gly 1045 1050 <210> 3 <211> 6 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Cas9 PAM recognition <220> <221> misc_feature <222> (1)..(2) <223> n = a, t, c, or g <400> 3 nnagaa 6 <210> 4 <211> 6 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Cas9 PAM recognition <220> <221> misc_feature <222> (1)..(2) <223> n = a, t, c, or g <400> 4 nngggt 6 <210> 5 <211> 6 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Cas9 PAM recognition <220> <221> misc_feature <222> (1)..(2) <223> n = a, t, c, or g <400> 5 nngagt 6 <210> 6 <211> 6 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Cas9 PAM recognition <220> <221> misc_feature <222> (1)..(2) <223> n is a, t, c, or g <400> 6 nngaat 6 <210> 7 <211> 4572 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> BPK764 plasmid <400> 7 taatacgact cactataggg gaattgtgag cggataacaa ttcccctgta gaaataattt 60 tgtttaactt taataaggag atataccatg gataaaaagt attctattgg tttagacatc 120 ggcactaatt ccgttggatg ggctgtcata accgatgaat acaaagtacc ttcaaagaaa 180 tttaaggtgt tggggaacac agaccgtcat tcgattaaaa agaatcttat cggtgccctc 240 ctattcgata gtggcgaaac ggcagaggcg actcgcctga aacgaaccgc tcggagaagg 300 tatacacgtc gcaagaaccg aatatgttac ttacaagaaa tttttagcaa tgagatggcc 360 aaagttgacg attctttctt tcaccgtttg gaagagtcct tccttgtcga agaggacaag 420 aaacatgaac ggcaccccat ctttggaaac atagtagatg aggtggcata tcatgaaaag 480 tacccaacga tttatcacct cagaaaaaag ctagttgact caactgataa agcggacctg 540 aggttaatct acttggctct tgcccatatg ataaagttcc gtgggcactt tctcattgag 600 ggtgatctaa atccggacaa ctcggatgtc gacaaactgt tcatccagtt agtacaaacc 660 tataatcagt tgtttgaaga gaaccctata aatgcaagtg gcgtggatgc gaaggctatt 720 cttagcgccc gcctctctaa atcccgacgg ctagaaaacc tgatcgcaca attacccgga 780 gagaagaaaa atgggttgtt cggtaacctt atagcgctct cactaggcct gacaccaaat 840 tttaagtcga acttcgactt agctgaagat gccaaattgc agcttagtaa ggacacgtac 900 gatgacgatc tcgacaatct actggcacaa attggagatc agtatgcgga cttatttttg 960 gctgccaaaa accttagcga tgcaatcctc ctatctgaca tactgagagt tatactgag 1020 attaccaagg cgccgttatc cgcttcaatg atcaaaaggt acgatgaaca tcaccaagac 1080 ttgacacttc tcaggccct agtccgtcag caacgctg agaaatataa ggaaatattc 1140 tttgatcagt cgaaaaacgg gtacgcaggt tatattgacg gcggagcg tcagaggag 1200 ttctacaagt ttcaacc catattagag aagatgatg ggacggaag gttgcttgta 1260 aaactcaatc gcgaagatct actgcgaaag cagcggactt tcgacaacgg tagcattcca 1320 catcaaatcc acttaggcga attgcatgct atacttagaa ggcaggagga ttttccg 1380 ttcctcaaag acatcgtga aaagattgag aaatcctaa ccttcgcat accttactat 1440 gtgggacccc tggcccgagg gaactcgg tcgcatgga tgacagaaa gtccgagaaa 1500 acgattactc cctggaattt tgaggaagtt gtcgaataag gtgcgtcagc tcaatcgttc 1560 atcgagagga tgaccaactt tgacaagaat ttaccgaacg aaaaagtatt gcctaagcac 1620 agtttacttt acgagtattt cacagtgtac atgaactca cgaagttaa gtatgtcact 1680 gagggcatgc gtaaacccgc ctttctaagc ggagaacaga agaagcaat agtagatctg 1740 ttattcaga ccaccgcaa agtgacagtt aagcattga agaggacta ctttagaaa 1800 attgaatgct tcgattctgt cgagatctcc ggggtagaag atcgatttaa tgcgtcactt 1860 ggtacgtatc atgacctcct aaagataatt aaagataagg acttcctgga taacgaag 1920 1980 attgaggaaa gactaaaaac atacgctcac ctgttcgacg ataaggttat gaaacagtta 2040 aagaggcgtc gctatacggg ctggggacga ttgtcgcgga aacttatcaa cgggataaga 2100 gacaagcaaa gtggtaaaac tattctcgat tttctaaaga gcgacggctt cgccaatagg 2160 aactttatgc agctgatcca tgatgactct ttaaccttca aagaggattat acaaaaggca 2220 caggtttccg gaaagggga ctcattgcac gaacatattg cgaatcttgc tggttcgcca 2280 gccatcaaaa agggcatact ccagacagtc aaagtagtgg atgagctagt taaggtcatg 2340 ggacgtcaca aaccggaaaa cattgtaatc gagatggcac gcgaaaatca aacgactcag 2400 aaggggcaaa aaaacagtcg agagcggatg aagagaatag aagagggtat taaagaactg 2460 ggcagccaga tcttaaagga gcatcctgtg gaaaataccc aattgcagaa cgagaaactt 2520 tacctctatt acctacaaaa tggaagggac atgtatgttg atcaggaact ggacataaac 2580 cgtttatctg attacgacgt cgatcacatt gtaccccaat cctttttgaa ggacgattca 2640 atcgacaata aagtgcttac acgctcggat aagaaccgag ggaaaagtga caatgttcca 2700 agcgaggaag tcgtaaagaa aatgaagaac tattggcggc agctcctaaa tgcgaaactg 2760 ataacgcaaa gaaagttcga taacttaact aaagctgaga ggggtggctt gtctgaactt 2820 gacaaggccg gattttata acgtcagctc gtggaaaccc gccaaatcac aaagcatgtt 2880 gcacagatac tagattcccg aatgaatacg aaatacgacg agaacgataa gctgattcgg 2940 gaagtcaaag taatcacttt aaagtcaaaa ttggtgtcgg acttcagaaa ggattttcaa 3000 ttctataaag ttagggagat aaataactac caccatgcgc acgacgctta tcttaatgcc 3060 gtcgtaggga ccgcactcat taagaaatac ccgaagctag aaagtgagtt tgtgtatggt 3120 gattacaaag tttatgacgt ccgtaagatg atcgcgaaaa gcgaacagga gataggcaag 3180 gctacagcca aatacttctt ttatctaac attatgaatt tctttaagac ggaaatcact 3240 ctggcaaacg gagagatacg caaacgacct ttaattgaaa ccaatgggga gacaggtgaa 3300 atcgtatggg ataagggccg ggacttcgcg acggtgagaa aagttttgtc catgccccaa 3360 gtcaacatag taaagaaaac tgaggtgcag accggagggt tttcaaagga atcgattctt 3420 ccaaaaagga atagtgataa gctcatcgct cgtaaaaagg actgggaccc gaaaaagtac 3480 ggtggcttcg atagccctac agttgcctat tctgtcctag tagtggcaaa agttgagaag 3540 ggaaaatcca agaaactgaa gtcagtcaaa gaattattgg ggataacgat tatggagcgc 3600 tcgtcttttg aaaagaaccc catcgacttc cttgaggcga aaggttacaa ggaagtaaaa 3660 aaggatctca taattaaact accaaagtat agtctgtttg agttagaaaa tggccgaaaa 3720 cggatgttgg ctagcgccgg agagcttcaa aaggggaacg aactcgcact accgtctaaa 3780 tacgtgaatt tcctgtattt agcgtcccat tacgagaagt tgaaaggttc acctgaagat 3840 aacgaacaga agcaactttt tgttgagcag cacaaacatt atctcgacga aatcatagag 3900 caaattcgg aattcagtaa gagagtcatc ctagctgatg ccaatctgga caaagtatta 3960 agcgcataca acaagcacag ggataaaccc atacgtgagc aggcggaaaa tattatccat 4020 ttgtttactc ttaccaacct cggcgctcca gccgcattca agtattttga cacaacgata 4080 gatcgcaaac gatacacttc taccaaggag gtgctagacg cgacactgat tcaccaatcc 4140 atcacgggat tatatgaaac tcggatagat ttgtcacagc ttgggggtga cggatccccc 4200 aagaagaaga ggaaagtctc gagcgactac aaagaccatg acggtgatta taaagatcat 4260 gacatcgatt acaaggatga cgatgacaag tgaagcggcc gcataatgct taagtcgaac 4320 agaaagtaat cgtattgtac acggccgcat aatcgaaatt aatacgactc actataggga 4380 gacccatgcc atagcgttgt tcggaacaga ttcaccaaca cctagtggtc tccgttttag 4440 agctagaaat agcaagttaa aataaggcta gtccgttatc aacttgaaaa agtggcaccg 4500 agtcggtgct ccgctgagca ataactagca taaccccttg gggcctctaa acgggtcttg 4560 aggggttttt tg 4572 <210> 8 <211> 4572 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> MSP712Flow <400> 8 taatacgact cactataggg gattgtgag cggataacaa ttcccctgta gaataattt 60 tgtttaactt taataggag atataccatg gataaaaagt attctattgg tttagccatc 120 ggcactaatt ccgttggatg ggctgtcata accgatgaat aaagtacc ttcaaagaaa 180 tttaagtgt tgggacac agaccgtcat tcgattaaaa agaatcttat cggtgccctc 240 ctattcgata gtggcgaaac ggcagaggcg actcgcctga aacgaaccgc tcggagaagg 300 tatacacgtc gcaagaaccg atatgttac ttacaagaaa tttttagcaa tgagatggcc 360 aaagttgacg attctctt tcaccgttg gaagagtcct tccttgtcga agaggacaag 420 aaacatgaac ggcacccat ctttggaac atagtagatg agtggcata tcatgaaaag 480 tacccaacga tttacacct cagaaaaag ctagttgact caacgataa agcggacctg 540 aggttaatct acttggctct tgcccatatg ataaagttcc gtgggcactt tctcattgag 600 ggtgatctaa atccggacaa ctcggatgtc ɣaaactgt tcatccagtt agtacaacc 660 tataatcagt tgtttgaaga gaaccctata aatgcaagtg gcgtggatgc gaaggctatt 720 cttagcgccc gcctctctaa atcccgacgg ctagaaaacc tgatcgcaca attacccgga 780 gagaagaaaa atgggttgtt cggtaacctt atagcgctct cactaggcct gacaccaaat 840 tttaagtcga acttcgactt agctgaagat gccaaattgc agcttagtaa ggacacgtac 900 gatgacgatc tcgacaatct actggcacaa attggagatc agtatgcgga cttatttttg 960 gctgccaaaa accttagcga tgcaatcctc ctatctgaca tactgagagt tatactgag 1020 attaccaagg cgccgttatc cgcttcaatg atcaaaaggt acgatgaaca tcaccaagac 1080 ttgacacttc tcaaggccct agtccgtcag caactgcctg agaaatataa ggaaatattc 1140 tttgatcagt cgaaaaacgg gtacgcaggt tatattgacg gcggagcgag tcaagaggaa 1200 ttctacaagt tttcaaacc catattagag aagatggatg ggacggaaga gttgcttgta 1260 aaactcaatc gcgaagatct actgcgaaag cagcggactt tcgacaacgg tagcattcca 1320 catcaaatcc acttaggcga attgcatgct atacttagaa ggcaggagga tttttatccg 1380 ttcctcaaag aaatcgtga aaagattgag aaaatcctaa cctttcgcat accttactat 1440 gtgggacccc tggcccgagg gaactctcgg ttcgcatgga tgacaagaaa gtccgaagaa 1500 acgattactc cctggaattt tgaggaagtt gtcgataaag gtgcgtcagc tcaatcgttc 1560 atcgagaga tgaccaactt tgacaagaat ttaccgaacg aaaaagtatt gcctaagcac 1620 agtttacttt acgagtattt caaggtgtac aatgaactca cgaaagttaa gtatgtcact 1680 gagggcatgc gtaaacccgc ctttctaagc ggaagaaga agaaagcaat agtagatctg 1740 ttattcaaga ccaaccgcaa agtgacagtt aagcaattga aagaggacta ctttaagaaa 1800 attgaatgct tcgattctgt cgagatctcc ggggtagaag atcgatttaa tgcgtcactt 1860 ggtacgtatc atgacctcct aaagataatt aaagataagg acttcctgga taacgaag 1920 1980 attgaggaaa gactaaaaac atacgctcac ctgttcgacg ataaggttat gaaacagtta 2040 aagaggcgtc gctatacggg ctggggacga ttgtcgcgga aacttatcaa cgggataaga 2100 gacaagcaaa gtggtaaaac tattctcgat tttctaaaga gcgacggctt cgccaatagg 2160 aactttatgc agctgatcca tgatgactct ttaaccttca aagaggattat acaaaaggca 2220 caggtttccg gaaagggga ctcattgcac gaacatattg cgaatcttgc tggttcgcca 2280 gccatcaaaa agggcatact ccagacagtc aaagtagtgg atgagctagt taaggtcatg 2340 ggacgtcaca aaccggaaaa cattgtaatc gagatggcac gcgaaaatca aacgactcag 2400 aaggggcaaa aaaacagtcg agagcggatg aagagaatag aagagggtat taaagaactg 2460 ggcagccaga tcttaaagga gcatcctgtg gaaaataccc aattgcagaa cgagaaactt 2520 tacctctatt acctacaaaa tggaagggac atgtatgttg atcaggaact ggacataaac 2580 cgtttatctg attacgacgt cgatgccatt gtaccccaat cctttttgaa ggacgattca 2640 atcgacaata aagtgcttac acgctcggat aagaaccgag ggaaaagtga caatgttcca 2700 2760 ataacgcaaa gaaagtttcga taacttaact aaagctgaga ggggtggctt gtctgaactt 2820 gawaggccg gatttatta acgtcagctc gtggaaaccc gccaaatcac aaagcatgtt 2880 gcacagatac tagattcccg aatgaatacg aaatacgacg agaacgataa gctgattcgg 2940 gaagtcaag taatcacttt aaagtcaaa ttggtgtcgg acttcagaa ggattttcaa 3000 ttcttaaag ttaggagat aaataactac caccatgcgc acgacgctta tcttaatgcc 3060 gtcgtaggga ccgcactcat taagaatac ccgaagctag aaagtgagtt tgtgtatggt 3120 gattacaag tttagacgt ccgtagatg atcgcgaaaa gcgaacagga gatggcaag 3180 gctacagcca atacttctt ttattctaac attgaatt tctttaagac ggaatcact 3240 ctggcaaacg gagagatacg haaacgacct ttaattgaa ccaatgggga gataggtgaa 3300 atcgtatggg ataagggccg ggacttcgcg acggtgagaa aagttttgtc catgccccaa 3360 gtcaacatag taaagaaac tgaggtgcag accggaggtt ttcaagga atcgattctt 3420 ccaaaaagga atagtgataa gctcatcgct cgtaaaaagg actgggaccc gaaaaagtac 3480 ggtggcttcg atagccctac agttgcctat tctgtcctag tagtggcaa agttgagaag 3540 ggaaaatcca agaaactgaa gtcagtcaaa gaattattgg ggataacgat tatggagcgc 3600 tcgtctttg aaaagaaccc catcgacttc cttgaggcga aaggttacaa ggaagtaaaa 3660 aaggatctca taattaaact accaaagtat agtctgtttg agttagaaaa tggccgaaaa 3720 cggatgttgg ctagcgccgg agagcttcaa aaggggaacg aactcgcact accgtctaaa 3780 tacgtgaatt tcctgtattt agcgtcccat tacgagaagt tgaaaggttc acctgaagat 3840 aacgaacaga agcaactttt tgttgagcag cacaaacatt atctcgacga aatcatagag 3900 caaattcgg aattcagtaa gagagtcatc ctagctgatg ccaatctgga caaagtatta 3960 agcgcataca acaagcacag ggataaaccc atacgtgagc aggcggaaaa tattatccat 4020 ttgtttactc ttaccaacct cggcgctcca gccgcattca agtattttga cacaacgata 4080 gatcgcaaac gatacacttc taccaaggag gtgctagacg cgacactgat tcaccaatcc 4140 atcacgggat tatatgaaac tcggatagat ttgtcacagc ttgggggtga cggatccccc 4200 aagaagaaga ggaaagtctc gagcgactac aaagaccatg acggtgatta taaagatcat 4260 gacatcgatt acaaggatga cgatgacaag tgaagcggcc gcataatgct taagtcgaac 4320 agaaagtaat cgtattgtac acggccgcat aatcgaaatt aatacgactc actataggga 4380 gacccatgcc atagcgttgt tcggaacaga ttcaccaaca cctagtggtc tccgttttag 4440 agctagaaat agcaagttaa aataaggcta gtccgttatc aacttgaaaa agtggcaccg 4500 agtcggtgct ccgctgagca ataactagca taaccccttg gggcctctaa acgggtcttg 4560 aggggttttt tg 4572 <210> 9 <211> 3825 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Plasmid BPK2169 <400> 9 taatacgact cactataggg gaattgtgag cggataacaa ttcccctgta gaaataattt 60 tgtttaactt taataaggag atataccatg ggcagcgacc tggtgctggg cctggacatc 120 ggcatcggca gcgtgggcgt gggcatcctg aacaaggtga ccggcgagat catccacaag 180 aacagtcgca tcttccctgc tgctcaggct gagaacaacc tggtgcgccg caccaaccgc 240 cagggtcgcc ggcttgctcg ccgcaagaag caccggcgcg tgcgcctgaa ccgcctgttc 300 gaggagagcg gcctgatcac cgacttcacc aagatcagca tcaacctgaa cccctaccag 360 ctgcgcgtga agggcctgac cgacgagctg agcaacgagg agctgttcat cgccctgaag 420 aacatggtga agcaccgcgg catcagctac ctggacgacg ccagcgacga cggcaacagc 480 agcgtgggcg actacgccca gatcgtgaag gagaacagca agcagctgga gaccaagacc 540 cccggccaga tccagctgga gcgctaccag acctacggcc agctgcgcgg cgacttcacc 600 gtggagaagg acggcaagaa gcaccgcctg atcaacgtgt tccccaccag cgcctaccgc 660 agcgaggccc tgcgcatcct gcagacccag caggagttca acccccagat caccgacgag 720 ttcatcaacc gctacctgga gatcctgacc ggcaagcgca agtactacca cggccccggc 780 aacgagaaga gccgcaccga ctacggccgc taccgcacca gcggcgagac cctggacaac 840 atcttcggca tcctgatcgg caagtgcacc ttctaccccg acgagttccg cgccgccaag 900 gccagctaca ccgcccagga gttcaatctg ctgaacgacc tgaacaacct gaccgtgccc 960 accgagacca agaagctgag caaggagcag aagaaccaga tcatcaacta cgtgaagaac 1020 gagaaggcta tgggccccgc caagctgttc aagtacatcg ccaagctgct gagctgcgac 1080 gtggccgaca tcaagggcta ccgcatcgac aagagcggca aggccgagat ccacaccttc 1140 gaggcctacc gcaagatgaa gaccctggag accctggaca tcgagcagat ggaccgagag 1200 accctggaca agctggccta cgtgctgacc ctgaacaccg agcgcgagg catccaggag 1260 gccctggagc acgagttcgc cgacggcagc ttcagccaga aacaggtgga cgagctggtg 1320 cagttccgca aggccaacag cagcatcttc ggcaagggct ggcacaactt cagcgtgaag 1380 ctgatgatgg agctgatccc cgagctgtac gagaccagcg aggagcagat gaccatcctg 1440 acccgcctgg ccaagcagaa gaccaccagc agcagcaaca agaccaagta catcgacgag 1500 aagctgctga ccgaggagat ctacaacccc gtggtggcca agagcgtgcg ccaggccatc 1560 aagatcgtga acgccgccat caaggagtac ggcgacttcg acaacatcgt gatcgagatg 1620 gcccgcgaga ccaacgag cgacgagaag aaggccatcc agaagatcca gaaggccaac 1680 1740 ctgccccaca gcgtgttcca cggccacaag cagctggcca ccaagatccg cctgtggcac 1800 cagcagggcg agcgctgcct gtacaccggc aagaccatca gcatccacga cctgatcaac 1860 aacagcaacc agttcgaggt ggaccacatc ctgcccctga gcatcacctt cgacgacagc 1920 ctggccaaca aggtgctggt gtacgccacc gccaaccagg agaagggcca gcgcacccc 1980 taccaggccc tggacagcat ggacgacgcc tggagcttcc gcgagctgaa ggccttcgtg 2040 cgcgagagca agaccctgag caacaagaag aaggatatc tgctgaccga ggaggacatc 2100 agcaagttcg acgtgcgcaa gaagttcatc gagcgcaacc tggtggacac ccgctacgcc 2160 agccgcgtgg tgctgaacgc cctgcaggag cacttccgcg cccacaagat cgacaccaag 2220 gtgagcgtgg tgcgcggcca gttcaccagc cagctgcgcc gccactgggg catcgagaag 2280 acccgcgaca cctaccacca ccacgccgtg gacgccctga tcattgcggc ttctagccag 2340 ctgaacctgt ggaaagca gaaacacc ctggtgagct acagcgagga ccagctgctg 2400 gacatcgaga ccggcgagct gacgagcgac gacgagtaca aggagagcgt gttcaaggcc 2460. ccctaccagc acttcgtgga caccctgaag agcaaggagt tcgaggacag catcctgttc agctaccagg tggacagcaa gttcaaccgc aagatcagcg acgccaccat ctacgccacc cgccaggcca aggtgggcaa ggacaaggcc gacgagacct acgtgctggg caagatcaag gacatctaca cccaggacgg ctacgacgcc ttcatgaaga tctacaaga ggacaagagc aagttcctga tgtaccgcca cgacccccag accttcgaga aggtgatcga gcccatcctg 2760 2820. gagaactacc ccaacaagca gatcaacgat aaaggcaagg aggtgccctg caaccccttc ctgaagtaca aggaggagca cggctacatc cgcaagtaca gcaagaaggg caacggcccc gagatcaaga gcctgaagta ctacgacagc aagctgggca accacatcga catcaccccc aaggacagca acaacaaggt ggtgctgcag agcgccc cctggcgcgc cgacgtgtac ttcaacaaga ccaccggcaa gtacgagatc ctggggctga agtacgccga tctgcagttt gataaaggca caggcaccta caagatcagc caggagaagt acaacgacat caagaagaag gagggcgtgg acagcgacag cgagttcaag ttcaccctgt aagaacga ccttctgctg 3180 gtgaagcaca ccgagaccaa ggagcaacag ctgttccgct tcctgagccg caccatgccc 3240 3300 gccctgatca aggtgctggg caacgtggcc aacagcggcc agtgcaagaa gggcctgggc 3360 aagagcaaca tcagcatcta caaggtgcgc accgacgtgc tgggcaacca gcacatcatc 3420 aagaacgagg gcgacaagcc caagttggac ttcagcaggg ctgaccccaa gaagaagagg 3480 aaggtgtgag cggccgcata atgcttaagt cgaacagaaa gtaatcgtat tgtacaccgg 3540 cgcataatcg aaattaatac gactcactat aggactgcag gtcatgccat agcgttgttc 3600 ggaacagatt caccaacacc tagtacctgc actcgtttt gtactctcaa gatttaagta 3660 3720 catgccgaaa tcaacaccct gtcattttat ggcagggtgt tttccgctga gcaataacta 3780 gcataacccc ttggggcctc taaacgggtc ttgaggggtt ttttg 3825 <210> 10 <211> 3674 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> BPK2101 plasmid <400> 10 taatacgact cactataggg gaattgtgag cggataacaa ttcccctgta gaaataattt 60 tgtttaactt taataaggag atataccatg ggcaaacgga actacatcct ggggcttgac 120 attgggataa ccagcgttgg ctacggaatt attgattatg agacacgcga tgtgattgac 180 gccggggtta ggctgttcaa agaggccaac gttgaaaaca acgagggaag acggagtaag 240 cgcggagcaa gaagactcaa gcgcagacgg agacatcgga ttcagagggt gaaaaagctg 300 ctcttcgatt acaatctcct gaccgatcat agtgagctga gcggaatcaa cccctacgag 360 gcgcgagtga aagggctttc ccagaagctg tccgaagagg agttctccgc cgcgttgctg 420 cacctggcca aacggagggg gttcacaatg taaacgaag tggaggagga cacgggcaat 480 gaacttagta cgaaagaaca gatcagtagg aactctaagg ctctcgaaga gaaatacgtc 540 gctgagttgc agcttgagag actgaaaaaa gacggcgaag tacgcggatc tattaatagg 600 ttcaagactt cagattacgt aaggaagcc aagcagctcc tgaagtaca gaaagcgtac 660 catcagctcg atcagagctt catcgatacc tacatagatt tgctggagac acggaggaca 720 tactacgagg gcccagggga aggatctcct tttggtgga aggacatcaa ggaatggtac 780 gagatgctta tgggacattg tacatatttt ccggaggagc tcaggagcgt caagtacgcc 840 tacaatgccg acctgtacaa tgccctcaat gacctcaata acctcgtgat taccagggac 900 gagaacgaga agctggagta ctatgaaag ttccagatta tcgagaatgt gtttaagcag 960 aagaagaagc cgacacttaa gcagattgca aaggaatcc tcgtgaatga ggaagatatc 1020 aagggataca gagtgacaag tacaggcaag cccgagttca caatctgaa ggtgtaccac 1080 gatattagg acataccgc acgaaggg ataatcgaaa acgctgagct cctcgatcag 1140 atcgcaaaaa 1200 ctgaacagtg agctcaccca agaggaatt gagcagattt siaacctgaa agctcacc 1260 gggacgcaca atctgagcct caagcaatc aacctcattc tggatgact ttggcacaca 1320 aatgacaacc aaattgccat attcaaccc ctgaactgg tgccaaaaaa agtggatctg 1380 tcacagcaaa aggaatccc tacaacttg gttgacgatt ttattctc cccgttgtc 1440 1500 1500 ccaacgata taattatcga gcttgcacga gagagaact ccaaggacgc ccagagatg 1560 attack attacgaaa tgcagaagcg aaccgccag aacgaac gcatagagg attattaag 1620 acaaccggca agagaatgc caagtatctg atcgagaaaa tcaagctgca cgacatgca 1680 gaaggcaagt gcctgtactc tctggaagct atcccactcg aagatctgct gatataatcca 1740 ttcattacg aggtggacca catcatccct agatccgtaa gcttgaca ttccttcaat 1800 aacaagttc tggttaaaca ggaggaaat tctaaaaag ggaccggac cccgttccag 1860 tacctgagct ccagtgacag tacctgactc tacgagactt tgaaca tattctgaat 1920s ctggccaaag gcaaggcag gatcagcaag accagaagg agtaccctcct cgaagaacggc 1980 gabattaaca gatttagtgt gcagaagat ttcatcacc gaaaccttgt cgatactcgg 2040 tacgccacga gaggcctgat gatctccctc aggagctact tccgcgtca taatctggac 2100 gttaaagtca agagcataaa tggggttc accagctttc tgaggagaaa gtggaagtttt 2160 aagaaggaac gaaaaagg atacaagcac catgctgagg atgctttgat catcgctaac 2220 gcggacttta tctttaagga atggaaaag ctggaagg CAagaaagt gatgaaaac 2280 cagatgttcg aggaagca ggcagagtca atgcctgaga tcgagacaga gcaggaatac 2340 aaggaaattt tcatcacccc tcatcagatt aaacacata aggactcaa agacttaaa 2400 tactctcata gggtggacaa aaaacccaat cgcgagctca ttaatgacac cctgtactca 2460 acacggaagg atgataagg taataccttg attgtgaata atcttaatgg attgtatgac 2520 aaagataacg acaagctca gaagctgatc aacaagtctc cagagaagct ccttatgtat 2580 caccacgacc cacagactta tcagaaattg aaactgatca tggagcaata cggggatgag 2640 aagaacccac tctacaaata ttatgaggaa acaggtatt acctgaccaa gtactccaag 2700 aagtaacg gaccagtgat ciaagata aagtacttg gcaaac taatgcgcat 2760 ttggacataa ctgacgatta ccccaattct cgaaacaagg ttgtgaagct ctccctgaag 2820 ccttatagat ttgacgtgta cctggatat ggggttatata aattcgtcac cgtgaaaaat 2880 ctggacgtga tcaaaagga gaactatt gaagtaact caagtgcta tgaggaggcg 2940 agaagctga agagatctc caatcaggcc gagttcatcg cttccttcta taataacgat 3000 ctcatcaaga tcaatggaga gctttatcgc gtcattggtg tgacaatga ctgctgaac 3060 aggatcgaag tcaatgat agacattacc taccggt atctcgaaaa catgaatgat 3120 aaacggccgc ctcgcatcat caacaatc gcatctaaaa ctcagtcaat aaaaagtac 3180 tctaccgata tcctggggaa tcttatgaa gtgaagtca agaagcaccc acaaatcatt 3240 aaaaaaggtg gatccccca gagagagg aaagtctcga gcgactaca agaccacatgac 3300 ggtgattata aagatcatga catcgattac aaggatgacg atgacaagta aagcggccgc 3360 ataatgctta agtcgaacag aaagtaatcg tattgtacac ggccgcataa tcgaaattaa 3420 tacgactcac tataggga cccatgccat agcgttgttc ggaacagatt caccacacc 3480 tagtggtctc cgttttagta ctctgtaatt ttaggtatga ggtagacgaa aattgtactt 3540 atacctaaaa ttacagaatc tactaaaaca aggcaaaatg ccgtgtttat ctcgtcaact 3600 tgttggcgag atccgctgag caataactag cataacccct tggggcctct aaacgggtct 3660 tgaggggttt tttg 3674 <210> 11 <211> 4206 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> JDS246 plasmid <400> 11 atggataaaa agtattctat tggtttagac atcggcacta attccgttgg atgggctgtc 60 ataaccgatg aatacaaagt accttcaaag aaatttaagg tgttggggaa cacagaccgt 120 cattcgatta aaaagaatct tatcggtgcc ctcctattcg atagtggcga aacggcagag 180 gcgactcgcc tgaaacgaac cgctcggaga aggtatacac gtcgcaagaa ccgaatatgt 240 tacttacaag aaatttttag caatgagatg gccaaagttg acgattcttt ctttcaccgt 300 ttggaagagt ccttccttgt cgaagaggac aagaaacatg aacggcaccc catctttgga 360 aacatagtag atgaggtggc atatcatgaa aagtacccaa cgatttatca cctcagaaaa 420 aagctagttg actcaactga taagcggac ctgaggttaa tctacttggc tcttgcccat 480 atgataaagt tccgtgggca ctttctcatt gagggtgatc taaatccgga caactcggat 540 gtcgacaaac tgttcatcca gttagtacaa acctataatc agttgtttga agagaaccct 600 ataaatgcaa gtggcgtgga tgcgaaggct attcttagcg cccgcctctc taaatcccga 660 cggctgaaa acctgatcgc acaattaccc ggagagaaga aaaatgggtt gttcggtaac 720 cttatagcgc tctcactagg cctgacacca aattttaagt cgaacttcga cttagctgaa 780 gatgccaaat tgcagcttag taggcacag tacgatgacg atctcgacaa tctactggca 840 caaattggag atcagtatgc ggacttattt ttggctgcca aaaaccttag cgatgcaatc 900 ctcctatctg acatactgag agttaatact gagattacca aggcgccgtt atccgcttca 960 atgatcaaaa ggtacgatga acatcaccaa gacttgacac ttctcaaggc cctagtccgt 1020 cagcaactgc ctgagaaata taggaaata ttctttgatc agtcgaaaaa cgggtacgca 1080 ggttatattg acggcggagc gagtcaagag gaattctaca agtttatcaa acccatatta 1140 1200 aagcagcgga ctttcgacaa cggtagcatt ccacatcaaa tccacttagg cgaattgcat 1260 1320 gagaaaatcc taacctttcg cataccttac tatgtgggac ccctggcccg agggaactct 1380 cggttcgcat ggatgacaag aaagtccgaa gaacgatta ctccctggaa ttttgagaa 1440 gttgtcgata aaggtgcgtc agctcaatcg ttcatcgaga ggatgaccaa ctttgacaag 1500 1560 tacaatgaac tcacgaaagt tagtatgtc actgagggca tgcgtaaacc cgcctttcta 1620 1680 gttaagcaat tgaaagga ctactttaag aaaattgaat gcttcgattc tgtcgagatc 1740 tccggggtag aagatcgatt taatgcgtca cttggtacgt atcatgacct cctaaagata 1800 Attaagata Aggactccct ggataacgaa Gagaatgaag Atacttaga Agatatagtg 1860 ttgactctta cccctttga agatcggga atgattgagg aaagactaaa aacatacgct 1920 cacctgttcg acgataggt tatgaacag ttaaagaggc gtcgctatac gggctgggga 1980 cgattgtcgc ggaacttat caacgggata agagacaagc aaagtggtaa aactattctc 2040 gattttctaa agagcgacgg cttcgccaat aggacttta tgcagctgat ccatgatgac 2100 tctttaacct tcaagagga tatacaaag gcacaggtttt ccggacagg ggactcattg 2160 cacgaacata ttgcgaatct tgctggttcg ccagccatca aaaagggcat actccagaca 2220 gtcaaagtag tggatgagct agttaaggtc atgggacgtc aaaccgga aaacattgta 2280 atcgagatgg cacgcgaaaa tcaacgact cagaggggc aaaaaaacag tcgagagcgg 2340 atgagagaa tagagagggg tattaagaa ctgggcagcc agatcttaa ggagcatccct 2400 gtggaaaata cccaattgca gaacgagaaa ctttacctct attackaca aaatggaagg 2460 gatagtatg ttgatcagga actggacata aaccgtttat ctgattacga cgtcgatcac 2520 attgtacccc aatccttttt gaaggacgat tcaatcgaca ataaagtgct tacacgctcg 2580 gataagaacc gagggaaaag tgacaatgtt ccaagcgagg aagtcgtaaa gaaaatgaag 2640 aactattggc ggcagctcct aaatgcgaaa ctgataacgc aaagaaagtt cgataactta 2700 actaaagctg agaggggtgg cttgtctgaa cttgacaagg ccggatttat taaacgtcag 2760 ctkgtggaaa cccgccaaat cacaaagcat gttgcacaga tactagattc ccgaatgaat 2820 acgaaatacg acgagaacga taagctgatt cgggaagtca aagtaatcac tttaaaagtca 2880 aaattggtgt cggacttcag aaaggatttt caattctata aagttaggga gataaataac 2940 taccaccatg cgcacgacgc ttatcttaat gccgtcgtag ggaccgcact cattaagaaa 3000 tacccgaagc tagaaagtga gtttgtgtat ggtgattaca aagtttatga cgtccgtaag 3060 atgatcgcga aaagcgaaca ggagataggc aaggctacag ccaaatactt cttttattct 3120 aacattatga atttctttaa gacggaaatc actctggcaa acggagagat acgcaaacga 3180 cctttaattg aaaccaatgg ggagacaggt gaaatcgtat gggataaggg ccgggacttc 3240 gcgacggtga gaaaagtttt gtccatgccc caagtcaaca tagtaagaa aactgaggtg 3300 cagaccggag ggttttcaaa ggaatcgatt cttccaaaaa ggaatagtga taagctcatc 3360 gctcgtaaa aggactggga cccgaaaaag tacggtggct tcgatagccc tacagttgcc 3420 tattctgtcc tagtagtggc aaaagttgag aagggaaaat ccaagaaact gaagtcagtc 3480 aaagaattat tggggataac gattatggag cgctcgtctt ttgaaaagaa ccccatcgac 3540 ttccttgagg cgaaaggtta caaggaagta aaaaaggatc tcataattaa actaccaaag 3600 tatagtctgt ttgagttaga aaatggccga aaacggatgt tggctagcgc cggagagctt 3660 caaaagggga acgaactcgc actaccgtct aaatacgtga atttcctgta tttagcgtcc 3720 cattacgaga agttgaaagg ttcacctgaa gataacgaac agaagcaact ttttgttgag 3780 cagcacaaac attatctcga cgaaatcata gagcaaattt cggaattcag taagagagtc 3840 atcctagctg atgccaatct ggacaaagta ttaagcgcat acaacaagca cagggataaa 3900 cccatacgtg agcaggcgga aaatattatc catttgttta ctcttaccaa cctcggcgct 3960 ccagccgcat tcaagtattt tgacacaacg atagatcgca aacgatacac ttctaccaag 4020 gaggtgctag acgcgacact gattcaccaa tccatcacgg gattatatga aactcggata 4080 gatttgtcac agcttggggg tgacggatcc cccaagaaga agaggaaagt ctcgagcgac 4140 tacaaagacc atgacggtga ttataaagat catgacatcg attacaagga tgacgatgac 4200 aagtga 4206 <210> 12 <211> 4206 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Plasmid MSP469 <400> 12 atggataaaa agtattctat tggtttagac atcggcacta attccgttgg atgggctgtc 60 ataaccgatg aatacaaagt accttcaaag aaatttaagg tgttggggaa cacagaccgt 120 cattcgatta aaaagaatct tatcggtgcc ctcctattcg atagtggcga aacggcagag 180 gcgactcgcc tgaaacgaac cgctcggaga aggtatacac gtcgcaagaa ccgaatatgt 240 tacttacaag aaatttttag caatgagatg gccaaagttg acgattcttt ctttcaccgt 300 360 aacatagtag atgaggtggc atatcatgaa aagtacccaa cgatttatca cctcagaaaa 420 aagctagttg actcaactga taagcggac ctgaggttaa tctacttggc tcttgcccat 480 atgataaagt tccgtgggca ctttctcatt gagggtgatc taaatccgga caactcggat 540 gtcgacaaac tgttcatcca gttagtacaa acctataatc agttgtttga agagaaccct 600 ataaatgcaa gtggcgtgga tgcgaaggct attcttagcg cccgcctctc taaatcccga 660 cggctgaaa acctgatcgc acaattaccc ggagagaaga aaaatgggtt gttcggtaac 720 cttatagcgc tctcactagg cctgacacca aattttaagt cgaacttcga cttagctgaa 780 gatgccaaat tgcagcttag taggcacag tacgatgacg atctcgacaa tctactggca 840 caaattggag atcagtatgc ggacttattt ttggctgcca aaaaccttag cgatgcaatc 900 ctcctatctg acatactgag agttaatact gagattacca aggcgccgtt atccgcttca 960 atgatcaaaa ggtacgatga acatcaccaa gacttgacac ttctcaaggc cctagtccgt 1020 cagcaactgc ctgagaaata taggaaata ttctttgatc agtcgaaaaa cgggtacgca 1080 ggttatattg acggcggagc gagtcaagag gaattctaca agtttatcaa acccatatta 1140 1200 aagcagcgga ctttcgacaa cggtagcatt ccacatcaaa tccacttagg cgaattgcat 1260 1320 gagaaaatcc taacctttcg cataccttac tatgtgggac ccctggcccg agggaactct 1380 cggttcgcat ggatgacaag aaagtccgaa gaacgatta ctccctggaa ttttgagaa 1440 gttgtcgata aaggtgcgtc agctcaatcg ttcatcgaga ggatgaccaa ctttgacaag 1500 1560 tacaatgaac tcacgaaagt tagtatgtc actgagggca tgcgtaaacc cgcctttcta 1620 1680 gttaagcaat tgaaagga ctactttaag aaaattgaat gcttcgattc tgtcgagatc 1740 tccggggtag aagatcgatt taatgcgtca cttggtacgt atcatgacct cctaagata 1800 Attaagata Aggactccct ggataacgaa Gagaatgaag Atacttaga Agatatagtg 1860 ttgactctta cccctttga agatcggga atgattgagg aaagactaaa aacatacgct 1920 cacctgttcg acgataggt tatgaacag ttaaagaggc gtcgctatac gggctgggga 1980 cgattgtcgc ggaacttat caacgggata agagacaagc aaagtggtaa aactattctc 2040 gattttctaa agagcgacgg cttcgccaat aggacttta tgcagctgat ccatgatgac 2100 tctttaacct tcaagagga tatacaaag gcacaggtttt ccggacagg ggactcattg 2160 cacgaacata ttgcgaatct tgctggttcg ccagccatca aaaagggcat actccagaca 2220 gtcaaagtag tggatgagct agttaaggtc atgggacgtc aaaccgga aaacattgta 2280 atcgagatgg cacgcgaaaa tcaacgact cagaggggc aaaaaaacag tcgagagcgg 2340 atgagagaa tagagagggg tattaagaa ctgggcagcc agatcttaa ggagcatccct 2400 gtggaaaata cccaattgca gaacgagaaa ctttacctct attackaca aaatggaagg 2460 gacatgtatg ttgatcagga actggacata aaccgtttat ctgattacga cgtcgatcac 2520 attgtacccc aatccttttt gaaggacgat tcaatcgaca ataaagtgct tacacgctcg 2580 gataagaacc gagggaaaag tgacaatgtt ccaagcgagg aagtcgtaaa gaaaatgaag 2640 aactattggc ggcagctcct aaatgcgaaa ctgataacgc aaagaaagtt cgataactta 2700 actaaagctg agaggggtgg cttgtctgaa cttgacaagg ccggatttat taaacgtcag 2760 ctkgtggaaa cccgccaaat cacaaagcat gttgcacaga tactagattc ccgaatgaat 2820 acgaaatacg acgagaacga taagctgatt cgggaagtca aagtaatcac tttaaaagtca 2880 aaattggtgt cggacttcag aaaggatttt caattctata aagttaggga gataaataac 2940 taccaccatg cgcacgacgc ttatcttaat gccgtcgtag ggaccgcact cattaagaaa 3000 tacccgaagc tagaaagtga gtttgtgtat ggtgattaca aagtttatga cgtccgtaag 3060 atgatcgcga aaagcgaaca ggagataggc aaggctacag ccaaatactt cttttattct 3120 aacattatga atttctttaa gacggaaatc actctggcaa acggagagat acgcaaacga 3180 cctttaattg aaaccaatgg ggagacaggt gaaatcgtat gggataaggg ccgggacttc 3240 gcgacggtga gaaaagtttt gtccatgccc caagtcaaca tagtaagaa aactgaggtg 3300 cagaccggag ggttttcaaa ggaatcgatt cttccaaaaa ggaatagtga taagctcatc 3360 gctcgtaaa aggactggga cccgaaaaag tacggtggct tcgtgagccc tacagttgcc 3420 tattctgtcc tagtagtggc aaaagttgag aagggaaaat ccaagaaact gaagtcagtc 3480 aaagaattat tggggataac gattatggag cgctcgtctt ttgaaaagaa ccccatcgac 3540 ttccttgagg cgaaaggtta caaggaagta aaaaaggatc tcataattaa actaccaaag 3600 tatagtctgt ttgagttaga aaatggccga aaacggatgt tggctagcgc cggagagctt 3660 caaaagggga acgaactcgc actaccgtct aaatacgtga atttcctgta tttagcgtcc 3720 cattacgaga agttgaaagg ttcacctgaa gataacgaac agaagcaact ttttgttgag 3780 cagcacaaac attatctcga cgaaatcata gagcaaattt cggaattcag taagagagtc 3840 atcctagctg atgccaatct ggacaaagta ttaagcgcat acaacaagca cagggataaa 3900 cccatacgtg agcaggcgga aaatattatc catttgttta ctcttaccaa cctcggcgct 3960 ccagccgcat tcaagtattt tgacacaacg atagatcgca aacagtacag atctaccaag 4020 gaggtgctag acgcgacact gattcaccaa tccatcacgg gattatatga aactcggata 4080 gatttgtcac agcttggggg tgacggatcc cccaagaaga agaggaaagt ctcgagcgac 4140 tacaaagacc atgacggtga ttataaagat catgacatcg attacaagga tgacgatgac 4200 aagtga 4206 <210> 13 <211> 4206 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> MSP680 plasmid <400> 13 atggataaaa agtattctat tggtttagac atcggcacta attccgttgg atgggctgtc 60 ataaccgatg aatacaaagt accttcaaag aaatttaagg tgttggggaa cacagaccgt 120 cattcgatta aaaagaatct tatcggtgcc ctcctattcg atagtggcga aacggcagag 180 gcgactcgcc tgaaacgaac cgctcggaga aggtatacac gtcgcaagaa ccgaatatgt 240 tacttacaag aaatttttag caatgagatg gccaaagttg acgattcttt ctttcaccgt 300 360 aacatagtag atgaggtggc atatcatgaa aagtacccaa cgatttatca cctcagaaaa 420 aagctagttg actcaactga taagcggac ctgaggttaa tctacttggc tcttgcccat 480 atgataaagt tccgtgggca ctttctcatt gagggtgatc taaatccgga caactcggat 540 gtcgacaaac tgttcatcca gttagtacaa acctataatc agttgtttga agagaaccct 600 ataaatgcaa gtggcgtgga tgcgaaggct attcttagcg cccgcctctc taaatcccga 660 cggctgaaa acctgatcgc acaattaccc ggagagaaga aaaatgggtt gttcggtaac 720 cttatagcgc tctcactagg cctgacacca aattttaagt cgaacttcga cttagctgaa 780 gatgccaaat tgcagcttag taggcacag tacgatgacg atctcgacaa tctactggca 840 caaattggag atcagtatgc ggacttattt ttggctgcca aaaaccttag cgatgcaatc 900 ctcctatctg acatactgag agttaatact gagattacca aggcgccgtt atccgcttca 960 atgatcaaaa ggtacgatga acatcaccaa gacttgacac ttctcaaggc cctagtccgt 1020 cagcaactgc ctgagaaata taggaaata ttctttgatc agtcgaaaaa cgggtacgca 1080 ggttatattg acggcggagc gagtcaagag gaattctaca agtttatcaa acccatatta 1140 1200 aagcagcgga ctttcgacaa cggtagcatt ccacatcaaa tccacttagg cgaattgcat 1260 1320 gagaaaatcc taacctttcg cataccttac tatgtgggac ccctggcccg agggaactct 1380 cggttcgcat ggatgacaag aaagtccgaa gaacgatta ctccctggaa ttttgagaa 1440 gttgtcgata aaggtgcgtc agctcaatcg ttcatcgaga ggatgaccaa ctttgacaag 1500 1560 tacaatgaac tcacgaaagt tagtatgtc actgagggca tgcgtaaacc cgcctttcta 1620 1680 gttaagcaat tgaaagagga ctactttaag aaaattgaat gcttcgattc tgtcgagatc 1740 tccggggtag aagatcgatt taatgcgtca cttggtacgt atcatgacct cctaagata 1800 Attaagata Aggactccct ggataacgaa Gagaatgaag Atacttaga Agatatagtg 1860 ttgactctta cccctttga agatcggga atgattgagg aaagactaaa aacatacgct 1920 cacctgttcg acgataggt tatgaacag ttaaagaggc gtcgctatac gggctgggga 1980 cgattgtcgc ggaacttat caacgggata agagacaagc aaagtggtaa aactattctc 2040 gattttctaa agagcgacgg cttcgccaat aggacttta tgcagctgat ccatgatgac 2100 tctttaacct tcaagagga tatacaaag gcacaggtttt ccggacagg ggactcattg 2160 cacgaacata ttgcgaatct tgctggttcg ccagccatca aaaagggcat actccagaca 2220 gtcaaagtag tggatgagct agttaaggtc atgggacgtc aaaccgga aaacattgta 2280 atcgagatgg cacgcgaaaa tcaacgact cagaggggc aaaaaaacag tcgagagcgg 2340 atgagagaa tagagagggg tattaagaa ctgggcagcc agatcttaa ggagcatccct 2400 gtggaaaata cccaattgca gaacgagaaa ctttacctct attacctaca aaatggaagg 2460 gacatgtatg ttgatcagga actggacata aaccgtttat ctgattacga cgtcgatcac 2520 attgtacccc aatccttttt gaaggacgat tcaatcgaca ataaagtgct tacacgctcg 2580 gataagaacc gagggaaaag tgacaatgtt ccaagcgagg aagtcgtaaa gaaaatgaag 2640 aactattggc ggcagctcct aaatgcgaaa ctgataacgc aaagaaagtt cgataactta 2700 actaaagctg agaggggtgg cttgtctgaa cttgacaagg ccggatttat taaacgtcag 2760 ctkgtggaaa cccgccaaat cacaaagcat gttgcacaga tactagattc ccgaatgaat 2820 acgaaatacg acgagaacga taagctgatt cgggaagtca aagtaatcac tttaaaagtca 2880 aaattggtgt cggacttcag aaaggatttt caattctata aagttaggga gataaataac 2940 taccaccatg cgcacgacgc ttatcttaat gccgtcgtag ggaccgcact cattaagaaa 3000 tacccgaagc tagaaagtga gtttgtgtat ggtgattaca aagtttatga cgtccgtaag 3060 atgatcgcga aaagcgaaca ggagataggc aaggctacag ccaaatactt cttttattct 3120 3180 cctttaattg aaaccaatgg ggagacaggt gaaatcgtat gggataaggg ccgggacttc 3240 gcgacggtga gaaaagtttt gtccatgccc caagtcaaca tagtaagaa aactgaggtg 3300 cagaccggag ggttttcaaa ggaatcgatt cttccaaaaa ggaatagtga taagctcatc 3360 gctcgtaaa aggactggga cccgaaaaag tacggtggct tcgagagccc tacagttgcc 3420 tattctgtcc tagtagtggc aaaagttgag aagggaaaat ccaagaaact gaagtcagtc 3480 aaagaattat tggggataac gattatggag cgctcgtctt ttgaaaagaa ccccatcgac 3540 ttccttgagg cgaaaggtta caaggaagta aaaaaggatc tcataattaa actaccaaag 3600 tatagtctgt ttgagttaga aaatggccga aaacggatgt tggctagcgc cggagagctt 3660 caaaagggga acgaactcgc actaccgtct aaatacgtga atttcctgta tttagcgtcc 3720 cattacgaga agttgaaagg ttcacctgaa gataacgaac agaagcaact ttttgttgag 3780 cagcacaaac attatctcga cgaaatcata gagcaaattt cggaattcag taagagagtc 3840 atcctagctg atgccaatct ggacaaagta ttaagcgcat acaacaagca cagggataaa 3900 cccatacgtg agcaggcgga aaatattatc catttgttta ctcttaccaa cctcggcgct 3960 ccagccgcat tcaagtattt tgacacaacg atagatcgca aacagtacag atctaccaag 4020 gaggtgctag acgcgacact gattcaccaa tccatcacgg gattatatga aactcggata 4080 gatttgtcac agcttggggg tgacggatcc cccaagaaga agaggaaagt ctcgagcgac 4140 tacaaagacc atgacggtga ttataaagat catgacatcg attacaagga tgacgatgac 4200 aagtga 4206 <210> 14 <211> 4206 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> MSP1101 plasmid <400> 14 atggataaaa agtattctat tggtttagac atcggcacta attccgttgg atgggctgtc 60 ataaccgatg aatacaaagt accttcaaag aaatttaagg tgttggggaa cacagaccgt 120 cattcgatta aaaagaatct tatcggtgcc ctcctattcg atagtggcga aacggcagag 180 gcgactcgcc tgaacgaac cgctcggaga aggtatacac gtcgcaagaa ccgaatatgt 240 tacttacaag aaatttttag caatgagatg gccaaagttg acgattcttt ctttcaccgt 300 360 aacatagtag atgaggtggc atatcatgaa aagtacccaa cgatttatca cctcagaaaa 420 aagctagttg actcaactga taagcggac ctgaggttaa tctacttggc tcttgcccat 480 atgataaagt tccgtgggca ctttctcatt gagggtgatc taaatccgga caactcggat 540 gtcgacaaac tgttcatcca gttagtacaa acctataatc agttgtttga agagaaccct 600 ataaatgcaa gtggcgtgga tgcgaaggct attcttagcg cccgcctctc taaatcccga 660 cggctgaaa acctgatcgc acaattaccc ggagagaaga aaaatgggtt gttcggtaac 720 cttatagcgc tctcactagg cctgacacca aattttaagt cgaacttcga cttagctgaa 780 gatgccaaat tgcagcttag taggcacag tacgatgacg atctcgacaa tctactggca 840 caaattggag atcagtatgc ggacttattt ttggctgcca aaaaccttag cgatgcaatc 900 ctcctatctg acatactgag agttaatact gagattacca aggcgccgtt atccgcttca 960 atgatcaaaa ggtacgatga acatcaccaa gacttgacac ttctcaaggc cctagtccgt 1020 cagcaactgc ctgagaaata taggaaata ttctttgatc agtcgaaaaa cgggtacgca 1080 ggttatattg acggcggagc gagtcaagag gaattctaca agtttatcaa acccatatta 1140 1200 aagcagcgga ctttcgacaa cggtagcatt ccacatcaaa tccacttagg cgaattgcat 1260 1320 gagaaaatcc taacctttcg cataccttac tatgtgggac ccctggcccg agggaactct 1380 cggttcgcat ggatgacaag aaagtccgaa gaacgatta ctccctggaa ttttgagaa 1440 gttgtcgata aaggtgcgtc agctcaatcg ttcatcgaga ggatgaccaa ctttgacaag 1500 1560 tacaatgaac tcacgaaagt tagtatgtc actgagggca tgcgtaaacc cgcctttcta 1620 agcggagaac agaagaaagc atatagtagat ctgttattca agaccaaccg caaagtgaca 1680 gttaagcaat tgaaagagga ctactttaag aaaattgaat gcttcgattc tgtcgagatc 1740 tccggggtag aagatcgatt taatgcgtca cttggtacgt atcatgacct cctaaagata 1800 attaaagata aggacttcct ggataacgaa gagaatgaag atatcttaga agatatagtg 1860 ttgactctta ccctctttga agatcgggaa atgattgagg aaagactaaa aacatacgct 1920 cacctgtcg acgataaggt tatgaaacag ttaaagaggc gtcgctatac gggctgggga 1980 cgattgtcgc ggaaacttat caacgggata agagacaagc aaagtggtaa aactattctc 2040 gattttctaa agagcgacgg cttcgccaat aggaacttta tgcagctgat ccatgatgac 2100 tctttaacct tcaaagagga tatacaaaag gcacaggtttt ccggacaagg ggactcattg 2160 cacgaacata ttgcgaatct tgctggttcg ccagccatca aaaagggcat actccagaca 2220 gtcaaagtag tggatgagct agttaaggtc atgggacgtc acaaaccgga aaacattgta 2280 atcgagatgg cacgcgaaaa tcaaacgact cagaaggggc aaaaaacag tcgagagcgg 2340 atgaagagaa tagaagagg tattaaagaa ctgggcagcc agatcttaaa ggagcatcct 2400 gtggaaaata cccaattgca gaacgagaaa ctttacctct attacctaca aaatggaagg 2460 gacatgtatg ttgatcagga actggacata aaccgtttat ctgattacga cgtcgatcac 2520 attgtacccc aatccttttt gaaggacgat tcaatcgaca ataaagtgct tacacgctcg 2580 gataagaacc gagggaaaag tgacaatgtt ccaagcgagg aagtcgtaaa gaaaatgaag 2640 aactattggc ggcagctcct aaatgcgaaa ctgataacgc aaagaaagtt cgataactta 2700 actaaagctg agaggggtgg cttgtctgaa cttgacaagg ccggatttat taaacgtcag 2760 ctkgtggaaa cccgccaaat cacaaagcat gttgcacaga tactagattc ccgaatgaat 2820 acgaaatacg acgagaacga taagctgatt cgggaagtca aagtaatcac tttaaaagtca 2880 aaattggtgt cggacttcag aaaggatttt caattctata aagttaggga gataaataac 2940 taccaccatg cgcacgacgc ttatcttaat gccgtcgtag ggaccgcact cattaagaaa 3000 tacccgaagc tagaaagtga gtttgtgtat ggtgattaca aagtttatga cgtccgtaag 3060 3120 aacattatga atttctttaa gacggaaatc actctggcaa acggagagat acggaaacga 3180 cctttaattg aaaccaatgg ggagacaggt gaaatcgtat gggataaggg ccgggacttc 3240 gcgacggtga gaaaagtttt gtccatgccc caagtcaaca tagtaagaa aactgaggtg 3300 cagaccggag ggttttcaaa ggaatcgatt cttccaaaaa ggaatagtga taagctcatc 3360 gctcgtaaa aggactggga cccgaaaaag tacggtggct tcgtgagccc tacagttgcc 3420 tattctgtcc tagtagtggc aaaagttgag aagggaaaat ccaagaaact gaagtcagtc 3480 aaagaattat tggggataac gattatggag cgctcgtctt ttgaaaagaa ccccatcgac 3540 ttccttgagg cgaaaggtta caaggaagta aaaaaggatc tcataattaa actaccaaag 3600 tatagtctgt ttgagttaga aaatggccga aaacggatgt tggctagcgc cagagagctt 3660 caaaagggga acgaactcgc actaccgtct aaatacgtga atttcctgta tttagcgtcc 3720 cattacgaga agttgaaagg ttcacctgaa gataacgaac agaagcaact ttttgttgag 3780 cagcacaaac attatctcga cgaaatcata gagcaaattt cggaattcag taagagagtc 3840 atcctagctg atgccaatct ggacaaagta ttaagcgcat acaacaagca cagggataaa 3900 cccatacgtg agcaggcgga aaatattatc catttgttta ctcttaccaa cctcggcgct 3960 ccagccgcat tcaagtattt tgacacaacg atagatcgca aagagtacag atctaccaag 4020 gaggtgctag acgcgacact gattcaccaa tccatcacgg gattatatga aactcggata 4080 gatttgtcac agcttggggg tgacggatcc cccaagaaga agaggaaagt ctcgagcgac 4140 tacaaagacc atgacggtga ttataaagat catgacatcg attacaagga tgacgatgac 4200 aagtga 4206 <210> 15 <211> 4206 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> MSP977 plasmid <400> 15 atggataaaa agtattctat tggtttagac atcggcacta attccgttgg atgggctgtc 60 ataaccgatg aatacaaagt accttcaaag aaatttaagg tgttggggaa cacagaccgt 120 cattcgatta aaaagaatct tatcggtgcc ctcctattcg atagtggcga aacggcagag 180 gcgactcgcc tgaacgaac cgctcggaga aggtatacac gtcgcaagaa ccgaatatgt 240 tacttacaag aaatttttag caatgagatg gccaaagttg acgattcttt ctttcaccgt 300 360 aacatagtag atgaggtggc atatcatgaa aagtacccaa cgatttatca cctcagaaaa 420 aagctagttg actcaactga taagcggac ctgaggttaa tctacttggc tcttgcccat 480 atgataaagt tccgtgggca ctttctcatt gagggtgatc taaatccgga caactcggat 540 gtcgacaaac tgttcatcca gttagtacaa acctataatc agttgtttga agagaaccct 600 ataaatgcaa gtggcgtgga tgcgaaggct attcttagcg cccgcctctc taaatcccga 660 cggctgaaa acctgatcgc acaattaccc ggagagaaga aaaatgggtt gttcggtaac 720 cttatagcgc tctcactagg cctgacacca aattttaagt cgaacttcga cttagctgaa 780 gatgccaaat tgcagcttag taggcacag tacgatgacg atctcgacaa tctactggca 840 caaattggag atcagtatgc ggacttattt ttggctgcca aaaaccttag cgatgcaatc 900 ctcctatctg acatactgag agttaatact gagattacca aggcgccgtt atccgcttca 960 atgatcaaaa ggtacgatga acatcaccaa gacttgacac ttctcaaggc cctagtccgt 1020 cagcaactgc ctgagaaata taggaaata ttctttgatc agtcgaaaaa cgggtacgca 1080 ggttatattg acggcggagc gagtcaagag gaattctaca agtttatcaa acccatatta 1140 1200 aagcagcgga ctttcgacaa cggtagcatt ccacatcaaa tccacttagg cgaattgcat 1260 1320 gagaaaatcc taacctttcg cataccttac tatgtgggac ccctggcccg agggaactct 1380 cggttcgcat ggatgacaag aaagtccgaa gaacgatta ctccctggaa ttttgagaa 1440 gttgtcgata aaggtgcgtc agctcaatcg ttcatcgaga ggatgaccaa ctttgacaag 1500 1560 tacaatgaac tcacgaaagt taagtatgtc actgagggca tgcgtaaacc cgcctttcta 1620 agcggagaac agaagaaagc atatagtagat ctgttattca agaccaaccg caaagtgaca 1680 gttaagcaat tgaaagagga ctactttaag aaaattgaat gcttcgattc tgtcgagatc 1740 tccggggtag aagatcgatt taatgcgtca cttggtacgt atcatgacct cctaaagata 1800 attaaagata aggacttcct ggataacgaa gagaatgaag atatcttaga agatatagtg 1860 ttgactctta ccctctttga agatcgggaa atgattgagg aaagactaaa aacatacgct 1920 cacctgtcg acgataaggt tatgaaacag ttaaagaggc gtcgctatac gggctgggga 1980 cgattgtcgc ggaaacttat caacgggata agagacaagc aaagtggtaa aactattctc 2040 gattttctaa agagcgacgg cttcgccaat aggaacttta tgcagctgat ccatgatgac 2100 tctttaacct tcaaagagga tatacaaaag gcacaggtttt ccggacaagg ggactcattg 2160 cacgaacata ttgcgaatct tgctggttcg ccagccatca aaaagggcat actccagaca 2220 gtcaaagtag tggatgagct agttaaggtc atgggacgtc acaaaccgga aaacattgta 2280 atcgagatgg cacgcgaaaa tcaacgact cagaggggc aaaaaaacag tcgagagcgg 2340 atgagagaa tagagagggg tattaagaa ctgggcagcc agatcttaa ggagcatccct 2400 gtggaaaata cccaattgca gaacgagaaa ctttacctct attackaca aaatggaagg 2460 gatagtatg ttgatcagga actggacata aaccgtttat ctgattacga cgtcgatcac 2520 attgtacccc aatccttttt gaaggacgat tcaatcgaca ataagtgct tacacgctcg 2580 gatagaacc gagggaaag tgacaatgtt ccaagcgagg aagtcgtaaa gaaatgaag 2640 aactattggc ggcagctcct aaatgcgaaa ctgataacgc aaagaaagtt cgataactta 2700 actaaagctg agaggggtgg cttgtctgaa cttgacaagg ccggattt taaacgtcag 2760 ctcgtggaaa cccgccaaat cacaagcat gttgcacaga tactagattc ccgaatgaat 2820 acgaaatacg acgagaacga taagctgat cgggaagtca aagtaatcac ttagtca 2880 aaatttggtgt cggactcag aaaggatttt caatctata aagttagg gataaataac 2940 taccaccatg cgcacgacgc tattacttaat gccgtcgtag taccaccgcact cattagaaaa 3000 3060 3120 aacattatga atttctttaa gacggaaatc actctggcaa acggagagat acggaaacga 3180 cctttaattg aaaccaatgg ggagacaggt gaaatcgtat gggataaggg ccgggacttc 3240 gcgacggtga gaaaagtttt gtccatgccc caagtcaaca tagtaagaa aactgaggtg 3300 cagaccggag ggttttcaaa ggaatcgatt cttccaaaaa ggaatagtga taagctcatc 3360 gctcgtaaa aggactggga cccgaaaaag tacggtggct tcgagagccc tacagttgcc 3420 tattctgtcc tagtagtggc aaaagttgag aagggaaaat ccaagaaact gaagtcagtc 3480 aaagaattat tggggataac gattatggag cgctcgtctt ttgaaaagaa ccccatcgac 3540 ttccttgagg cgaaaggtta caaggaagta aaaaaggatc tcataattaa actaccaaag 3600 tatagtctgt ttgagttaga aaatggccga aaacggatgt tggctagcgc cggagagctt 3660 caaaagggga acgaactcgc actaccgtct aaatacgtga atttcctgta tttagcgtcc 3720 cattacgaga agttgaaagg ttcacctgaa gataacgaac agaagcaact ttttgttgag 3780 cagcacaaac attatctcga cgaaatcata gagcaaattt cggaattcag taagagagtc 3840 atcctagctg atgccaatct ggacaaagta ttaagcgcat acaacaagca cagggataaa 3900 cccatacgtg agcaggcgga aaatattatc catttgttta ctcttaccaa cctcggcgct 3960 ccagccgcat tcaagtattt tgacacaacg atagatcgca aacgatacac ttctaccaag 4020 gaggtgctag acgcgacact gattcaccaa tccatcacgg gattatatga aactcggata 4080 gatttgtcac agcttggggg tgacggatcc cccaagaaga agaggaaagt ctcgagcgac 4140 tacaaagacc atgacggtga ttataaagat catgacatcg attacaagga tgacgatgac 4200 aagtga 4206 <210> 16 <211> 3402 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> plasmid MSP1393 <400> 16 atgggcagcg acctggtgct gggcctggac atcggcatcg gcagcgtggg cgtgggcatc 60 ctgaacaagg tgaccggcga gatcatccac aagaacagtc gcatcttccc tgctgctcag 120 gctgagaaca acctggtgcg ccgcaccaac cgccagggtc gccggcttgc tcgccgcaag 180 aagcaccggc gcgtgcgcct gaaccgcctg ttcgaggaga gcggcctgat caccgacttc 240 accaagatca gcatcaacct gaacccctac cagctgcgcg tgaagggcct gaccgacgag 300 ctgagcaacg aggagctgtt catcgccctg aagaacatgg tgaagcaccg cggcatcagc 360 tacctggacg acgccagcga cgacggcaac agcagcgtgg gcgactacgc ccagatcgtg 420 aaggaaca ccaagcagct ggagaccaag accccccggcc agatccagct ggagcgctac 480 cagacctacg gccagctgcg cggcgacttc accgtggaga aggacggcaa gaagcaccgc 540 ctgatcaacg tgttccccac cagcgcctac cgcagcgagg ccctgcgcat cctgcagacc 600 cagcaggagt tcaaccccca gatcaccgac gagttcatca accgctacct ggagatcctg 660 accggcaagc gcaagtacta ccacggcccc ggcaacgaga agagccgcac cgactacggc 720 cgctacgca ccagcggcga gaccctggac aacatcttcg gcatcctgat cggcaagtgc 780 accttctacc ccgacgagtt ccgcgccgcc aaggccagct acaccgccca ggagttcaat 840 ctgctgaacg acctgaacaa cctgaccgtg cccaccgaga ccaagaagct gagcaaggag 900 cagaagaacc agatcatcaa ctacgtgaag aacgagaagg ctatgggccc cgccaagctg 960 ttcaagtaca tcgccaagct gctgagctgc gacgtggccg acatcaaggg ctaccgcatc 1020 gacaagagcg gcaaggccga gatccacacc ttcgaggcct accgcaagat gaagaccctg 1080 gagaccctgg acatcgagca gatggaccga gagaccctgg acaagctggc ctacgtgctg 1140 accctgaaca ccgagcgcga gggcatccag gaggccctgg agcacgagtt cgccgacggc 1200 agcttcagcc agaaacaggt ggacgagctg gtgcagttcc gcaaggccaa cagcagcatc 1260 ttcggcaagg gctggcacaa cttcagcgtg aagctgatga tggagctgat ccccgagctg 1320 tacgagacca gcgaggagca gatgaccatc ctgacccgcc tgggcaagca gaagaccacc 1380 agcagcagca acaagaccaa gtacatcgac gagaagctgc tgaccgagga gatctacaac 1440 cccgtggtgg ccaagagcgt gcgccaggcc atcaagatcg tgaacgccgc catcaaggag 1500 tacggcgact tcgacacat cgtgatcgag atggcccgcg agaccaacga ggacgacgag 1560 aagaaggcca tccagaagat ccagaaggcc aaaaggacg agaaggacgc cgccatgctg 1620 aaggccgcca accagtaca cggcaggcc gagctgcccc acagcgtgtt ccacggccac 1680 aagcagctgg cccaccagat ccgcctgtgg caccagcagg gcgagcgctg cctgtacacc 1740 ggcaagacca tcagcatcca cgacctgatc aacaacagca accagttcga ggtggaccac 1800 atcctgcccc tgagcatcac cttcgacgac agcctggcca acaggtgct gtgtacgcc 1860 accgccaacc aggagaaggg cccgaccacc ccctaccagg cccgacag catggacgac 1920 gcctggagct tccgcgagct gaagccttc gtgcgcgaga gcagaccct gagcacaag 1980 aagaaggagt atctgctgac cgaggaggac atcagcaagt tcgacgtgcg cagaagttc 2040 atcgagcgca acctggtgga cacccgctac gccagccgcg tggtgctgaa cgccctgcag 2100 gagcacttcc gcgcccacaa gatcgacacc aaggtgagcg tggtgcgcgg ccagttcacc 2160 agccagctgc gccgccactg gggcatcgag agacccgcg acacctacca ccaccacgcc 2220 gtggacgccc tgatcattgc ggcttctagc cagctgaacc tgtggaagaa gcagaagaac 2280 accctggtga gctacagcga ggaccagctg ctggacatcg agaccggcga gctgatcagc 2340 gacgacgagt acaaggagag cgtgttcaag gccccctacc agcacttcgt ggacaccctg 2400 aagagcaagg agttcgagga cagcatcctg ttcagctacc aggtggacag caagttcaac 2460 cgcaagatca gcgacgccac catctacgcc acccgccagg ccaaggtggg caaggacaag 2520 gccgacgaga cctacgtgct gggcaagatc aaggacatct acacccagga cggctacgac 2580 gccttcatga agatctacaa gaaggacaag agcaagttcc tgatgtaccg ccacgacccc 2640 cagaccttcg agaaggtgat cgagcccatc ctggagaact accccaacaa gcagatcaac 2700 gataaaggca aggaggtgcc ctgcaacccc ttcctgaagt acaaggagga gcacggctac 2760 atccgcaagt acagcaagaa gggcaacggc cccgagatca agagcctgaa gtactacgac 2820 agcaagctgg gcaaccacat cgacatcacc cccaaggaca gcaacaacaa ggtggtgctg 2880 cagagcgtga gcccctggcg cgccgacgtg tacttcaaca agaccaccgg caagtacgag 2940 atcctggggc tgaagtacgc cgatctgcag tttgataaag gcacaggcac ctacaagatc 3000 agccaggaga agtacaacga catcaagaag aaggagggcg tggacagcga cagcgagttc 3060 aagttcaccc tgtacaagaa cgaccttctg ctggtgaagg acaccgagac caaggagcaa 3120 cagctgttcc gcttcctgag ccgcaccatg cccaagcaga agcactacgt ggagctgaag 3180 ccctacgaca agcagaagtt cgagggcggc gaggccctga tcaaggtgct gggcaacgtg 3240 gccaacagcg gccagtgcaa gaagggcctg ggcaagagca acatcagcat ctacaaggtg 3300 cgcaccgacg tgctgggcaa ccagcacatc atcaagaacg agggcgacaa gcccaagttg 3360 gacttcagca gggctgaccc caagaagaag aggaaggtgt ga 3402 <210> 17 <211> 3264 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> BPK2139 plasmid <400> 17 atggggaaac ggaactacat cctggggctt gacattggga taaccagcgt tggctacgga 60 attattgatt atgagacacg cgatgtgatt gacgccgggg ttaggctgtt caaagaggcc 120 aacgttgaaa acaacgaggg aagacggagt aagcgcggag caagaagact caagcgaga 180 cggagacatc ggattcagag ggtgaaaaag ctgctcttcg attacaatct cctgaccgat 240 catagtgagc tgagcggaat caacccctac gaggcgcgag tgaaagggct ttcccagaag 300 ctgtccgaag aggagttctc cgccgcgttg ctgcacctgg ccaaacggag gggggttcac 360 aatgtaaacg aagtggaga ggacacggc aatgaactta gtacgaaaga acagatcagt 420 aggaactcta aggctctcga agaaatac gtcgctgagt tgcagcttga gagactgaaa 480 aaagacggcg aagtacgcgg atctattaat aggttcaaga cttcagatta cgtaaaggaa 540 gccaagcagc tcctgaaagt agaaagcg taccatcagc tcgatcagag cttcatcgat 600 acctacatag atttgctgga gacacggagg acatactacg agggcccagg ggaaggatct 660 ccttttgggt ggaaggacat caaggaatgg tacgagatgc ttatgggaca ttgtacatat 720 tttccggagg agctcaggag cgtcaagtac gcctacaatg ccgacctgta caatgccctc 780 aatgacctca ataacctcgt gattaccagg gacgagaacg agaagctgga gtactatgaa 840 aagttccaga ttatcgagaa tgtgtttaag cagaagaaga agccgacact taagcagatt 900 gcaaaggaaa tcctcgtgaa tgaggaagat atcaagggat acagagtgac aagtacaggc 960 aagcccgagt tcacaaatct gaaggtgtac cacgatatta aggacataac cgcacgaaag 1020 gagataatcg aaaacgctga gctcctcgat cagatcgcaa aaattcttac catctaccag 1080 tctagtgagg acattcagga ggaactgact aatctgaaca gtgagctcac ccaagaggaa 1140 attgagcaga tttcaaacct gaaaggctac accgggacgc acaatctgag cctcaaagca 1200 atcaacctca ttctggatga actttggcac acaaatgaca accaaattgc catattcaac 1260 cgcctgaaac tggtgccaaa aaaagtggat ctgtcacagc aaaaggaaat ccctacaacc 1320 ttggttgacg attttattct gtcccccgtt gtcaagcgga gcttcatcca gtcaatcaag 1380 gtgatcaatg ccatcattaa aaaatacgga ttgccaaacg atataattat cgagcttgca 1440 cgagagaaga actcaaagga cgcccagaag atgattaacg aaatgcagaa gcgcaaccgc 1500 cagacaaacg aacgcataga ggaaattata agaacaaccg gcaaagagaa tgccaagtat 1560 ctgatcgaga aaatcaagct gcacgacatg caagaaggca agtgcctgta ctctctggaa 1620 gctatcccac tcgagatct gctgaataat ccattcatt acgaggtgga ccacatcatc 1680 cctagatccg taagctttga caatccttc aaacaag ttctggttaa acaggaggaa 1740 aattctaaaa aagggaccg gaccccgttc cagtacctga gctccagtga caccagatt 1800 agctacgaga cttttaagaa acatattctg aatctggcca aaggcaagg caggatcagc 1860 agaccaga aggagtacct cctcgaagaa cgcgacatta acgatttag tgtgcagaaa 1920 gatttcatca accgaaacct tgtcgatact cggtacgcca cgagaggcct gatgaatctc 1980 ctcaggagct acttccgcgt caataatctg gacgttaag tcagagcat aaatggggga 2040 ttcaccagct ttctgaggag aaagtggaag tttagaag aacgaaaaa aggatacaag caccatgctg aggatgcttt gatcatcgct aacgcggact ttatctttaa ggaatggaaa 2160 aagctgata aggxaagaa agtgatgaa aaccagatgt tcaggagaa gcaggcagag 2220 tcaatgcctg agatcgagac agagcaggaa tacaggaa ttttcatcac ccctcatcag 2280 attaaacaca taaggactt caagactat aaatactctc atagggtgga caaaaaaccc 2340 aatcgcgagc tcattaatga caccctgtac tcacacgga aggatgataa aggtaatacc 2400 ttgattgtga ataatcttta tggattgtat cagaagata acgacaagct cagaagctg 2460 atcacaagt ctccagagaa gctccttatg tatcaccacg acccacagac ttacagaaa 2520 ttgaaactga tcatggagca atacgggat gagagaacc cactctacaa atttatgag 2580 gaaacaggta attacctgac caagtactcc aagaaggata acggaccagt gathaaaag 2640 aaagtact atggcaaca acttaatgcg catttggaca taactgacga ttaccccaat 2700 tctcgaaaca aggttgtgaa gctctccctg aagccttata gatttgacgt gtacctggat 2760 aatggggttt aataattcgt caccgtgaaa atctggacg tgatcaaaa ggagaactat 2820 tatgaagtaa actcaagtg ctatgaggag gcgaagaagc tgaagagat ctccaatcag 2880 gccgagttca tcgctcctt ctataatac gatctcatca agatcaatgg agagctttat 2940 cgcgtcattg gtgtgaacaa tgacttgctg aacaggatcg aagtcaat gatagacatt 3000 acctaccggg agtatctcga aaacatgaat gataaacggc cgcctcgcat catcaagaca 3060 atcgcatcta aaactcagtc aataaaaaag tactctaccg atatcctggg gaatctctat 3120 gaagtgaagt caaagaagca cccacaaatc attaaaaaag gtggatcccc caagaagaag 3180 aggaaagtct cgagcgacta caaagaccat gacggtgatt ataaagatca tgacatcgat 3240 tacaaggatg acgatgacaa gtaa 3264 <210> 18 <211> 422 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> plasmid BPK1520 <400> 18 tgtacaaaaa agcaggcttt aaaggaacca attcagtcga ctggatccgg taccaaggtc 60 gggcaggaag agggcctatt tcccatgatt ccttcatatt tgcatatacg atacaaggct 120 gttagagaga taattagaat taatttgact gtaaacacaa agatattagt acaaaatacg 180 tgacgtagaa agtaataatt tcttgggtag tttgcagttt taaaattatg ttttaaaatg 240 gactatcata tgcttaccgt aacttgaaag tatttcgatt tcttggcttt atatatcttg 300 tggaaaggac gaaacaccgg agacgattaa tgcgtctccg ttttagagct agaaatagca 360 agttaaaata aggctagtcc gttatcaact tgaaaaagtg gcaccgagtc ggtgcttttt 420 tt 422 <210> 19 <211> 471 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Plasmid BPK2301 <400> 19 tgtacaaaaa agcaggcttt aaaggaacca attcagtcga ctggatccgg taccaaggtc 60 gggcaggaag agggcctatt tcccatgatt ccttcatatt tgcatatacg atacaaggct 120 gttagagaga taattagaat taatttgact gtaaacacaa agatattagt acaaaatacg 180 tgacgtagaa agtaataatt tcttgggtag tttgcagttt taaaattatg ttttaaaatg 240 gactatcata tgcttaccgt aacttgaaag tatttcgatt tcttggcttt atatatcttg 300 tggaaaggac gaaacacccg agacgattaa tgcgtctcgg tttttgtact ctcaagattt 360 aagtaactgt acaacgaaac ttacacagtt acttaaatct tgcagaagct acaaagataa 420 ggcttcatgc cgaaatcaac accctgtcat tttatggcag ggtgtttttt t 471 <210> 20 <211> 466 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> pVVT1 plasmid <400> 20 tgtacaaaaa agcaggcttt aaaggaacca attcagtcga ctggatccgg taccaaggtc 60 gggcaggaag agggcctatt tcccatgatt ccttcatatt tgcatatacg atacaaggct 120 gttagagaga taattagaat taatttgact gtaaacacaa agatattagt acaaaatacg 180 tgacgtagaa agtaataatt tcttgggtag tttgcagttt taaaattatg ttttaaaatg 240 gactatcata tgcttaccgt aacttgaaag tatttcgatt tcttggcttt atatatcttg 300 tggaaaggac gaaacacccg agacgattaa tgcgtctcgg ttttagtact ctgtaatttt 360 aggtatgagg tagacgaaaa ttgtacttat acctaaaatt acagaatcta ctaaaacaag 420 gcaaaatgcc gtgtttatct cgtcaacttg ttggcgagat tttttt 466 <210> 21 <211> 3609 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Plasmid MSP2283 <400> 21 taatacgact cactataggg gaattgtgag cggataacaa ttcccctgta gaaataattt 60 tgtttaactt taataaggag atataccatg ggcaaacgga actacatcct ggggcttgac 120 attgggataa ccagcgttgg ctacggaatt attgattatg agacacgcga tgtgattgac 180 gccggggtta ggctgttcaa agaggccaac gttgaaaaca acgagggaag acggagtaag 240 cgcggagcaa gaagactcaa gcgcagacgg agacatcgga ttcagagggt gaaaaagctg 300 ctcttcgatt acaatctcct gaccgatcat agtgagctga gcggaatcaa cccctacgag 360 gcgcgagtga aagggctttc ccagaagctg tccgaagagg agttctccgc cgcgttgctg 420 cacctggcca aacggagggg ggttcacaat gtaaacgaag tggaggagga cacgggcaat 480 gaacttagta cgaaagaaca gatcagtagg aactctaagg ctctcgaaga gaaatacgtc 540 gctgagttgc agcttgagag actgaaaaaa gacggcgaag tacgcggatc tattaatagg 600 ttcaagactt cagattacgt aaggaagcc aagcagctcc tgaagtaca gaaagcgtac 660 catcagctcg atcagagctt catcgatacc tacatagatt tgctggagac acggaggaca 720 tactacgagg gcccagggga aggatctcct tttggtgga aggacatcaa ggaatggtac 780 gagatgctta tgggacattg tacatatttt ccggaggagc tcaggagcgt caagtacgcc 840 tacaatgccg acctgtacaa tgccctcaat gacctcaata acctcgtgat taccagggac 900 gagaacgaga agctggagta ctatgaaag ttccagatta tcgagaatgt gtttaagcag 960 aagaagaagc cgacacttaa gcagattgca aaggaatcc tcgtgaatga ggaagatatc 1020 aagggataca gagtgacaag tacaggcaag cccgagttca caatctgaa ggtgtaccac 1080 gatattagg acataccgc acgaaggg ataatcgaaa acgctgagct cctcgatcag 1140 atcgcaaaaa 1200 ctgaacagtg agctcaccca agaggaatt gagcagattt siaacctgaa agctcacc 1260 gggacgcaca atctgagcct caagcaatc aacctcattc tggatgact ttggcacaca 1320 aatgacaacc aaattgccat attcaaccc ctgaactgg tgccaaaaaa agtggatctg 1380 tcacagcaaa aggaatccc tacaacttg gttgacgatt ttattctc cccgttgtc 1440 1500 1500 ccaacgata taattatcga gcttgcacga gagagaact ccaaggacgc ccagagatg 1560 attack attacgaaa tgcagaagcg aaccgccag aacgaac gcatagagg attattaag 1620 acaaccggca agagaatgc caagtatctg atcgagaaaa tcaagctgca cgacatgca 1680 gaaggcaagt gcctgtactc tctggaagct atcccactcg aagatctgct gatataatcca 1740 ttcattacg aggtggacca catcatccct agatccgtaa gcttgaca ttccttcaat 1800 aacaagttc tggttaaaca ggaggaaat tctaaaaag ggaccggac cccgttccag 1860 tacctgagct ccagtgacag tacctgactc tacgagactt tgaaca tattctgaat 1920s ctggccaaag gcaaggcag gatcagcaag accagaagg agtaccctcct cgaagaacggc 1980 gabattaaca gatttagtgt gcagaagat ttcatcacc gaaaccttgt cgatactcgg 2040 tacgccacga gaggcctgat gatctccctc aggagctact tccgcgtca taatctggac 2100 gttaaagtca agagcataaa tggggttc accagctttc tgaggagaaa gtggaagtttt 2160 aagaaggaac gaaaaagg atacaagcac catgctgagg atgctttgat catcgctaac 2220 gcggacttta tctttaagga atggaaaag ctggaagg CAagaaagt gatgaaaac 2280 cagatgttcg aggaagca ggcagagtca atgcctgaga tcgagacaga gcaggaatac 2340 aaggaaattt tcatcacccc tcatcagatt aaacacata aggactcaa agacttaaa 2400 tactctcata gggtggacaa aaaacccaat cgcgagctca ttaatgacac cctgtactca 2460 acacggaagg atgataagg taataccttg attgtgaata atcttaatgg attgtatgac 2520 aaagataacg acaagctca gaagctgatc aacaagtctc cagagaagct ccttatgtat 2580 caccacgacc cacagactta tcagaaattg aaactgatca tggagcaata cggggatgag 2640 aagaacccac tctacaaata ttatgaggaa acaggtatt acctgaccaa gtactccaag 2700 aagtaacg gaccagtgat ciaagata aagtacttg gcaaac taatgcgcat 2760 ttggacataa ctgacgatta ccccaattct cgaaacaagg ttgtgaagct ctccctgaag 2820 ccttatagat ttgacgtgta cctggataat ggggtttata aattcgtcac cgtgaaaaat 2880 ctggacgtga tcaaaaagga gaactattat gaagtaaact caaagtgcta tgaggaggcg 2940 aagaagctga agaagatctc caatcaggcc gagttcatcg cttccttcta tataacgat 3000 ctcatcaaga tcaatggaga gctttatcgc gtcattggtg tgaacaatga cttgctgaac 3060 aggatcgaag tcaatatgat agacattacc taccgggagt atctcgaaaa catgaatgat 3120 aaacggccgc ctcgcatcat caagacaatc gcatctaaaa ctcagtcaat aaaaaagtac 3180 tctaccgata tcctggggaa tctctatgaa gtgaagtcaa agaagcaccc acaaatcatt 3240 aaaaaaggtg gatcccccaa gaagaagagg aaagtctcga gcgactacaa agaccatgac 3300 ggtgattata aagatcatga catcgattac aaggatgacg atgacaagta aagcggccgc 3360 ataatgctta agtcgaacag aaagtaatcg tattgtacac ggccgcataa tcgaaattaa 3420 tacgactcac tataggggca cgggcagctt gccgggtttt agtactctgt aatgaaaatt 3480 acagaatcta ctaaaacaag gcaaaatgcc gtgtttatct cgtcaacttg ttggcgagat 3540 tttttttccg ctgagcaata actagcataa ccccttgggg cctctaaacg ggtcttgagg 3600 ggttttttg 3609 <210> 22 <211> 3609 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Plasmid MSP2262 <400> 22 taatacgact cactataggg gaattgtgag cggataacaa ttcccctgta gaaataattt 60 tgtttaactt taataaggag atataccatg ggcaaacgga actacatcct ggggcttgcc 120 attgggataa ccagcgttgg ctacggaatt attgattatg agacacgcga tgtgattgac 180 gccggggtta ggctgttcaa agaggccaac gttgaaaaca acgagggaag acggagtaag 240 cgcggagcaa gaagactcaa gcgcagacgg agacatcgga ttcagagggt gaaaaagctg 300 ctcttcgatt acaatctcct gaccgatcat agtgagctga gcggaatcaa cccctacgag 360 gcgcgagtga aagggctttc ccagaagctg tccgaagagg agttctccgc cgcgttgctg 420 cacctggcca aacggagggg gttcacaat gtaaacgaag tggagga cacggggcaat 480 gaacttagta cgaagaaca gatcagtagg aactctagg ctctcgaaga gaatacgtc 540 gctgagttgc agcttgagag actgaaaaaa gacggcgaag tacgcggatc tattatagg 600 ttcaagactt cagattacgt aaggaagcc aagcagctcc tgaagtaca gaaagcgtac 660 catcagctcg atcagagctt catcgatacc tacatagatt tgctggagac acggaggaca 720 tactacgagg gcccagggga aggatctcct tttggtgga aggacatcaa ggaatggtac 780 gagatgctta tgggacattg tacatatttt ccggaggagc tcaggagcgt caagtacgcc 840 tacaatgccg acctgtacaa tgccctcaat gacctcaata acctcgtgat taccagggac 900 gagaacgaga agctggagta ctatgaaag ttccagatta tcgagaatgt gtttaagcag 960 aagaagaagc cgacacttaa gcagattgca aaggaatcc tcgtgaatga ggaagatatc 1020 aagggataca gagtgacaag tacaggcaag cccgagttca caatctgaa ggtgtaccac 1080 gatattagg acataccgc acgaaggg ataatcgaaa acgctgagct cctcgatcag 1140 atcgcaaaaa 1200 ctgaacagtg agctcaccca agaggaatt gagcagattt siaacctgaa agctcacc 1260 gggacgcaca atctgagcct caagcaatc aacctcattc tggatgact ttggcacaca 1320 aatgacaacc aaattgccat attcaaccc ctgaactgg tgccaaaaaa agtggatctg 1380 tcacagcaaa aggaatccc tacaacttg gttgacgatt ttattctc cccgttgtc 1440 1500 1500 ccaacgata taattatcga gcttgcacga gagagaact ccaaggacgc ccagagatg 1560 attack attacgaaa tgcagaagcg aaccgccag aacgaac gcatagagg attattaag 1620 acaaccggca agagaatgc caagtatctg atcgagaaaa tcaagctgca cgacatgca 1680 gaaggcaagt gcctgtactc tctggaagct atcccactcg aagatctgct gatataatcca 1740 ttcattacg aggtggacgc catcatccct agatccgtaa gctttgaca ttccttcaat 1800 aacaagttc tggttaaaca ggaggaaat tctaaaaag ggaccggac cccgttccag 1860 tacctgagct ccagtgacag tacctgactc tacgagactt tgaaca tattctgaat 1920s ctggccaaag gcaaggcag gatcagcaag accagaagg agtaccctcct cgaagaacggc 1980 gabattaaca gatttagtgt gcagaagat ttcatcacc gaaaccttgt cgatactcgg 2040 tacgccacga gaggcctgat gatctccctc aggagctact tccgcgtca taatctggac 2100 gttaaagtca agagcataaa tggggttc accagctttc tgaggagaaa gtggaagtttt 2160 aagaaggaac gaaaaagg atacaagcac catgctgagg atgctttgat catcgctaac 2220 gcggacttta tctttaagga atggaaaag ctggaagg CAagaaagt gatgaaaac 2280 cagatgttcg aggaagca ggcagagtca atgcctgaga tcgagacaga gcaggaatac 2340 aaggaaattt tcatcacccc tcatcagatt aaacacata aggactcaa agacttaaa 2400 tactctcata gggtggacaa aaaacccaat cgcgagctca ttaatgacac cctgtactca 2460 acacggaagg atgataagg taataccttg attgtgaata atcttaatgg attgtatgac 2520 aaagataacg acaagctca gaagctgatc aacaagtctc cagagaagct ccttatgtat 2580 caccacgacc cacagactta tcagaaattg aaactgatca tggagcaata cggggatgag 2640 aagaacccac tctacaaata ttatgaggaa acaggtatt acctgaccaa gtactccaag 2700 aagtaacg gaccagtgat ciaagata aagtacttg gcaaac taatgcgcat 2760 ttggacataa ctgacgatta ccccaattct cgaaacaagg ttgtgaagct ctccctgaag 2820 ccttatagat ttgacgtgta cctggatat ggggttatata aattcgtcac cgtgaaaaat 2880 ctggacgtga tcaaaagga gaactatt gaagtaact caagtgcta tgaggaggcg 2940 agaagctga agagatctc caatcaggcc gagttcatcg cttccttcta taataacgat 3000 ctcatcaaga tcaatggaga gctttatcgc gtcattggtg tgacaatga ctgctgaac 3060 aggatcgaag tcaatgat agacattacc taccggt atctcgaaaa catgaatgat 3120 aaacggccgc ctcgcatcat caacaatc gcatctaaaa ctcagtcaat aaaaagtac 3180 tctaccgata tcctggggaa tcttatgaa gtgaagtca agaagcaccc acaaatcatt 3240 aaaaaaggtg gatccccca gagagagg aaagtctcga gcgactaca agaccacatgac 3300 ggtgattata aagatcatga catcgattac aaggatgacg atgacaagta aagcggccgc 3360 ataatgctta agtcgaacag aaagtaatcg tattgtacac ggccgcataa tcgaaattaa 3420 tacgactcac tataggggca cgggcagctt gccgggtttt agtactctgt aatgaaaatt 3480 acagaatcta ctaaaacaag gcaaaatgcc gtgtttatct cgtcaacttg ttggcgagat 3540 tttttttccg ctgagcaata actagcataa ccccttgggg cctctaaacg ggtcttgagg 3600 ttttttg 3609 <210> 23 <211> 3609 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> MSP2253 plasmid <400> 23 taatacgact cactataggg gaattgtgag cggataacaa ttcccctgta gaaataattt 60 tgtttaactt taataaggag atataccatg ggcaaacgga actacatcct ggggcttgac 120 attgggataa ccagcgttgg ctacggaatt attgattatg agacacgcga tgtgattgac 180 gccggggtta ggctgttcaa agaggccaac gttgaaaaca acgagggaag acggagtaag 240 cgcggagcaa gaagactcaa gcgcagacgg agacatcgga ttcagagggt gaaaaagctg ctcttcgatt acaatctcct gaccgatcat agtgagctga gcggaatcaa cccctacgag gcgcgagtga aagggctttc ccagaagctg tccgaagagg agttctccgc cgcgttgctg cacctggcca aacggagggg ggttcacaat gtaaacgaag tggaggagga cacgggcaat gaacttagta cgaaagaaca gatcagtagg aactctaagg ctctcgaaga gaatacgtc gctgagttgc agcttgagag actgaaaaaa gacggcgag tacgcggatc tattaatagg ttcaagactt cagattacgt aaaggaagcc aagcagctcc tgaaagtaca gaaagcgtac catcagctcg atcagagctt catcgatacc tacatagatt tgctggagac acggagaca 720 tactacgagg gcccagggga aggatctcct tttgggtgga aggacatcaa ggaatggtac 780 gagatgctta tgggacattg tacatatttt ccggaggagc tcaggagcgt caagtacgcc 840 900. acctgccg acctgtaca tgccctcaat acctcgtgat acctcggac gagaacgaga agctggagta ctatgaaaag ttccagatta tcgagaatgt gtttaagcag aagaagaagc cgacacttaa gcagattgca aaggaatcc tcgtgaatga ggaagatatc 1020 aagggataca gagtgacaag tacaggcaag cccgagttca caatctgaa ggtgtaccac 1080 gatattagg acataccgc acgaaggg ataatcgaaa acgctgagct cctcgatcag 1140 atcgcaaaaa 1200 ctgaacagtg agctcaccca agaggaatt gagcagattt siaacctgaa agctcacc 1260 gggacgcaca atctgagcct caagcaatc aacctcattc tggatgact ttggcacaca 1320 aatgacaacc aaattgccat attcaaccc ctgaactgg tgccaaaaaa agtggatctg 1380 tcacagcaaa aggaatccc tacaacttg gttgacgatt ttattctc cccgttgtc 1440 1500 1500 ccaacgata taattatcga gcttgcacga gagagaact ccaaggacgc ccagagatg 1560 attack attacgaaa tgcagaagcg aaccgccag aacgaac gcatagagg attattaag 1620 acaaccggca agagaatgc caagtatctg atcgagaaaa tcaagctgca cgacatgca 1680 gaaggcaagt gcctgtactc tctggaagct atcccactcg aagatctgct gatataatcca 1740 ttcattacg aggtggacca catcatccct agatccgtaa gcttgaca ttccttcaat 1800 aacaagttc tggttaaaca ggaggaaat tctaaaaag ggaccggac cccgttccag 1860 tacctgagct ccagtgacag tacctgactc tacgagactt tgaaca tattctgaat 1920s ctggccaaag gcaaggcag gatcagcaag accagaagg agtaccctcct cgaagaacggc 1980 gabattaaca gatttagtgt gcagaagat ttcatcacc gaaaccttgt cgatactcgg 2040 tacgccacga gaggcctgat gatctccctc aggagctact tccgcgtca taatctggac 2100 gttaaagtca agagcataaa tggggttc accagctttc tgaggagaaa gtggaagtttt 2160 aagaaggaac gaaaaagg atacaagcac catgctgagg atgctttgat catcgctaac 2220 gcggacttta tctttaagga atggaaaag ctggaagg CAagaaagt gatgaaaac 2280 cagatgttcg aggaagca ggcagagtca atgcctgaga tcgagacaga gcaggaatac 2340 aaggaaattt tcatcacccc tcatcagatt aaacacata aggactcaa agacttaaa 2400 tactctcata gggtggacaa aaaacccaat cgcaagctca ttaatgacac cctgtactca 2460 acacggaagg atgataaagg tataccttg attgtgaata atcttaatgg attgtatgac 2520 aaagataacg acaagctcaa gaagctgatc aacaagtctc cagagaagct ccttatgtat 2580 caccacgacc cacagactta tcagaaattg aaactgatca tggagcaata cggggatgag 2640 aagaacccac tctacaaata ttatgaggaa acaggtaatt acctgaccaa gtactccaag 2700 aaggataacg gaccagtgat caaaaagata aagtactatg gcaacaaact taatgcgcat 2760 ttggacataa ctgacgatta ccccaattct cgaaacaagg ttgtgaagct ctccctgaag 2820 ccttatagat ttgacgtgta cctggataat ggggtttata aattcgtcac cgtgaaaaat 2880 ctggacgtga tcaaaaagga gaactattat gaagtaaact caaagtgcta tgaggaggcg 2940 aagaagctga agaagatctc caatcaggcc gagttcatcg cttccttcta taagaacgat 3000 ctcatcaaga tcaatggaga gctttatcgc gtcattggtg tgaacaatga cttgctgaac 3060 aggatcgaag tcaatatgat agacattacc taccgggagt atctcgaaaa catgaatgat 3120 aaacggccgc ctcacatcat caagacaatc gcatctaaaa ctcagtcaat aaaaaagtac 3180 tctaccgata tcctggggaa tctctatgaa gtgaagtcaa agaagcaccc acaaatcatt 3240 aaaaaaggtg gatcccccaa gaagaagagg aaagtctcga gcgactacaa agaccatgac 3300 ggtgattata aagatcatga catcgattac aaggatgacg atgacaagta aagcggccgc 3360 ataatgctta agtcgaacag aaagtaatcg tattgtacac ggccgcataa tcgaaattaa 3420 tacgactcac tataggggca cgggcagctt gccgggtttt agtactctgt aatgaaaatt 3480 acagaatcta ctaaaacaag gcaaaatgcc gtgtttatct cgtcaacttg ttggcgagat 3540 tttttttccg ctgagcaata actagcataa ccccttgggg cctctaaacg ggtcttgagg 3600 ttttttg 3609 <210> 24 <211> 3609 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Plasmid MSP2266 <400> 24 taatacgact cactataggg gaattgtgag cggataacaa ttcccctgta gaaataattt 60 120. tgtttaactt taataggag fatherccatg ggcaaacgga actacatcct ggggcttgac attgggataa ccagcgttgg ctacggaatt attgattatg agacacgcga tgtgattgac gccggggtta ggctgttcaa agaggccaac gttgaaaaca acgaggag acggagtaag cgcggagcaa gaagactcaa gcgcagacgg agacatcgga ttcagagggt gaaaaagctg ctcttcgatt acaatctcct gaccgatcat agtgagctga gcggaatcaa cccctacgag gcgcgagtga aagggctttc ccagaagctg tccgaagagg agttctccgc cgcgttgctg cacctggcca aacggagggg ggttcacaat gtaaacgaag tggaggagga cacgggcaat gaacttagta cgaaagaaca gatcagtagg aactctaagg ctctcgaaga gaatacgtc gctgagttgc agcttgagag actgaaaaaa gacggcgag tacgcggatc tattaatagg ttcaagactt cagattacgt aaaggaagcc aagcagctcc tgaaagtaca gaaagcgtac catcagctcg atcagagctt catcgatacc tacatagatt tgctggagac acggagaca 720 tactacgagg gcccagggga aggatctcct tttgggtgga aggacatcaa ggaatggtac 780 gagatgctta tgggacattg tacatatttt ccggaggagc tcaggagcgt caagtacgcc 840 tacaatgccg acctgtacaa tgccctcaat gacctcaata acctcgtgat taccagggac 900 gagaacgaga agctggagta ctatgaaag ttccagatta tcgagaatgt gtttaagcag 960 aagaagaagc cgacacttaa gcagattgca aaggaatcc tcgtgaatga ggaagatatc 1020 aagggataca gagtgacaag tacaggcaag cccgagttca caatctgaa ggtgtaccac 1080 gatattagg acataccgc acgaaggg ataatcgaaa acgctgagct cctcgatcag 1140 atcgcaaaaa 1200 ctgaacagtg agctcaccca agaggaatt gagcagattt siaacctgaa agctcacc 1260 gggacgcaca atctgagcct caagcaatc aacctcattc tggatgact ttggcacaca 1320 aatgacaacc aaattgccat attcaaccc ctgaactgg tgccaaaaaa agtggatctg 1380 tcacagcaaa aggaatccc tacaacttg gttgacgatt ttattctc cccgttgtc 1440 1500 1500 ccaacgata taattatcga gcttgcacga gagagaact ccaaggacgc ccagagatg 1560 attack attacgaaa tgcagaagcg aaccgccag aacgaac gcatagagg attattaag 1620 acaaccggca agagaatgc caagtatctg atcgagaaaa tcaagctgca cgacatgca 1680 gaaggcaagt gcctgtactc tctggaagct atcccactcg aagatctgct gatataatcca 1740 ttcattacg aggtggacca catcatccct agatccgtaa gcttgaca ttccttcaat 1800 aacaagttc tggttaaaca ggaggaaat tctaaaaag ggaccggac cccgttccag 1860 tacctgagct ccagtgacag tacctgactc tacgagactt tgaaca tattctgaat 1920s ctggccaaag gcaaggcag gatcagcaag accagaagg agtaccctcct cgaagaacggc 1980 gabattaaca gatttagtgt gcagaagat ttcatcacc gaaaccttgt cgatactcgg 2040 tacgccacga gaggcctgat gatctccctc aggagctact tccgcgtca taatctggac 2100 gttaaagtca agagcataaa tggggttc accagctttc tgaggagaaa gtggaagtttt 2160 aagaaggaac gaaaaagg atacaagcac catgctgagg atgctttgat catcgctaac 2220 gcggacttta tctttaagga atggaaaag ctggaagg CAagaaagt gatgaaaac 2280 cagatgttcg aggaagca ggcagagtca atgcctgaga tcgagacaga gcaggaatac 2340 aaggaaattt tcatcacccc tcatcagatt aaacacata aggactcaa agacttaaa 2400 tactctcata gggtggacaa aaaacccaat cgcgagctca ttaatgacac cctgtactca 2460 acacggaagg atgataagg taataccttg attgtgaata atcttaatgg attgtatgac 2520 aaagataacg acaagctca gaagctgatc aacaagtctc cagagaagct ccttatgtat 2580 caccacgacc cacagactta tcagaaattg aaactgatca tggagcaata cggggatgag 2640 aagaacccac tctacaaata ttatgaggaa acaggtatt acctgaccaa gtactccaag 2700 aagtaacg gaccagtgat ciaagata aagtacttg gcaaac taatgcgcat 2760 ttggacataa ctgacgatta ccccaattct cgaaacaagg ttgtgaagct ctccctgaag 2820 ccttatagat ttgacgtgta cctggatat ggggttatata aattcgtcac cgtgaaaaat 2880 ctggacgtga tcaaaagga gaactatt gaagtaact caagtgcta tgaggaggcg 2940 aagaagctga agaagatctc caatcaggcc gagttcatcg cttccttcta tataacgat 3000 ctcatcaaga tcaatggaga gctttatcgc gtcattggtg tgaacaatga cttgctgaac 3060 aggatcgaag tcaatatgat agacattacc taccgggagt atctcgaaaa catgaatgat 3120 aaacggccgc ctcgcatcat caagacaatc gcatctaaaa ctcagtcaat aaaaaagtac 3180 tctaccgata tcctggggaa tctctatgaa gtgaagtcaa agaagcaccc acaaatcatt 3240 aaaaaaggtg gatcccccaa gaagaagagg aaagtctcga gcgactacaa agaccatgac 3300 ggtgattata aagatcatga catcgattac aaggatgacg atgacaagta aagcggccgc 3360 ataatgctta agtcgaacag aaagtaatcg tattgtacac ggccgcataa tcgaaattaa 3420 tacgactcac tataggtcgc cctcgaactt cacctgtttt agtactctgt aatgaaaatt 3480 acagaatcta ctaaaacaag gcaaaatgcc gtgtttatct cgtcaacttg ttggcgagat 3540 ttttttccg ctgagcaata actagcataa ccccttgggg cctctaaacg ggtcttgagg 3600 ggttttttg 3609 <210> 25 <211> 3609 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> MSP2279 plasmid <400> 25 taatacgact cactataggg gaattgtgag cggataacaa ttcccctgta gaaataattt 60 tgtttaactt taataaggag atataccatg ggcaaacgga actacatcct ggggcttgcc 120 attgggataa ccagcgttgg ctacggaatt attgattatg agacacgcga tgtgattgac 180 gccggggtta ggctgttcaa agaggccaac gttgaaaaca acgagggaag acggagtaag 240 cgcggagcaa gaagactcaa gcgcagacgg agacatcgga ttcagagggt gaaaaagctg 300 ctcttcgatt acaatctcct gaccgatcat agtgagctga gcggaatcaa cccctacgag 360 gcgcgagtga aagggctttc ccagaagctg tccgaagagg agttctccgc cgcgttgctg 420 cacctggcca aacggagggg gttcacaatg taaacgaag tggaggagga cacgggcaat 480 gaacttagta cgaaagaaca gatcagtagg aactctaagg ctctcgaaga gaaatacgtc 540 gctgagttgc agcttgagag actgaaaaaa gacggcgaag tacgcggatc tattaatagg 600 ttcaagactt cagattacgt aaggaagcc aagcagctcc tgaagtaca gaaagcgtac 660 catcagctcg atcagagctt catcgatacc tacatagatt tgctggagac acggaggaca 720 tactacgagg gcccagggga aggatctcct tttggtgga aggacatcaa ggaatggtac 780 gagatgctta tgggacattg tacatatttt ccggaggagc tcaggagcgt caagtacgcc 840 tacaatgccg acctgtacaa tgccctcaat gacctcaata acctcgtgat taccagggac 900 gagaacgaga agctggagta ctatgaaag ttccagatta tcgagaatgt gtttaagcag 960 aagaagaagc cgacacttaa gcagattgca aaggaatcc tcgtgaatga ggaagatatc 1020 aagggataca gagtgacaag tacaggcaag cccgagttca caatctgaa ggtgtaccac 1080 gatattagg acataccgc acgaaggg ataatcgaaa acgctgagct cctcgatcag 1140 atcgcaaaaa 1200 ctgaacagtg agctcaccca agaggaatt gagcagattt siaacctgaa agctcacc 1260 gggacgcaca atctgagcct caagcaatc aacctcattc tggatgact ttggcacaca 1320 aatgacaacc aaattgccat attcaaccc ctgaactgg tgccaaaaaa agtggatctg 1380 tcacagcaaa aggaatccc tacaacttg gttgacgatt ttattctc cccgttgtc 1440 1500 1500 ccaacgata taattatcga gcttgcacga gagagaact ccaaggacgc ccagagatg 1560 attack attacgaaa tgcagaagcg aaccgccag aacgaac gcatagagg attattaag 1620 acaaccggca agagaatgc caagtatctg atcgagaaaa tcaagctgca cgacatgca 1680 gaaggcaagt gcctgtactc tctggaagct atcccactcg aagatctgct gatataatcca 1740 ttcattacg aggtggacgc catcatccct agatccgtaa gctttgaca ttccttcaat 1800 aacaagttc tggttaaaca ggaggaaat tctaaaaag ggaccggac cccgttccag 1860 tacctgagct ccagtgacag tacctgactc tacgagactt tgaaca tattctgaat 1920s ctggccaaag gcaaggcag gatcagcaag accagaagg agtaccctcct cgaagaacggc 1980 gabattaaca gatttagtgt gcagaagat ttcatcacc gaaaccttgt cgatactcgg 2040 tacgccacga gaggcctgat gatctccctc aggagctact tccgcgtca taatctggac 2100 gttaaagtca agagcataaa tggggttc accagctttc tgaggagaaa gtggaagtttt 2160 aagaaggaac gaaaaagg atacaagcac catgctgagg atgctttgat catcgctaac 2220 gcggacttta tctttaagga atggaaaag ctggaagg CAagaaagt gatgaaaac 2280 cagatgttcg aggaagca ggcagagtca atgcctgaga tcgagacaga gcaggaatac 2340 aaggaaattt tcatcacccc tcatcagatt aacacata aggactcaa agacttaa 2400 tactctcata gggtggacaa aaaacccaat cgcgagctca ttaatgacac cctgtactca 2460 acacggaagg atgataagg taataccttg attgtgaata atcttaatgg attgtatgac 2520 aaagataacg acaagctca gaagctgatc aacaagtctc cagagaagct ccttatgtat 2580 caccacgacc cacagactta tcagaaattg aaactgatca tggagcaata cggggatgag 2640 aagaacccac tctacaaata ttatgaggaa acaggtatt acctgaccaa gtactccaag 2700 aagtaacg gaccagtgat ciaagata aagtacttg gcaaac taatgcgcat 2760 ttggacataa ctgacgatta ccccaattct cgaaacaagg ttgtgaagct ctccctgaag 2820 ccttatagat ttgacgtgta cctggataat ggggtttata aattcgtcac cgtgaaaaat 2880 ctggacgtga tcaaaaagga gaactattat gaagtaaact caaagtgcta tgaggaggcg 2940 aagaagctga agaagatctc caatcaggcc gagttcatcg cttccttcta tataacgat 3000 ctcatcaaga tcaatggaga gctttatcgc gtcattggtg tgaacaatga cttgctgaac 3060 aggatcgaag tcaatatgat agacattacc taccgggagt atctcgaaaa catgaatgat 3120 aaacggccgc ctcgcatcat caagacaatc gcatctaaaa ctcagtcaat aaaaaagtac 3180 tctaccgata tcctggggaa tctctatgaa gtgaagtcaa agaagcaccc acaaatcatt 3240 aaaaaaggtg gatcccccaa gaagaagagg aaagtctcga gcgactacaa agaccatgac 3300 ggtgattata aagatcatga catcgattac aaggatgacg atgacaagta aagcggccgc 3360 ataatgctta agtcgaacag aaagtaatcg tattgtacac ggccgcataa tcgaaattaa 3420 tacgactcac tataggtcgc cctcgaactt cacctgtttt agtactctgt aatgaaaatt 3480 acagaatcta ctaaaacaag gcaaaatgcc gtgtttatct cgtcaacttg ttggcgagat 3540 tttttttccg ctgagcaata actagcataa ccccttgggg cctctaaacg ggtcttgagg 3600 ggttttttg 3609 <210> 26 <211> 3609 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> Plasmid MSP2292 <400> 26 taatacgact cactataggg gaattgtgag cggataacaa ttcccctgta gaaataattt 60 tgtttaactt taataaggag atataccatg ggcaaacgga actacatcct ggggcttgac 120 attgggataa ccagcgttgg ctacggaatt attgattatg agacacgcga tgtgattgac 180 gccggggtta ggctgttcaa agaggccaac gttgaaaaca acgagggaag acggagtaag 240 cgcggagcaa gaagactcaa gcgcagacgg agacatcgga ttcagagggt gaaaaagctg 300 ctcttcgatt acaatctcct gaccgatcat agtgagctga gcggaatcaa cccctacgag 360 gcgcgagtga aagggctttc ccagaagctg tccgaagagg agttctccgc cgcgttgctg 420 cacctggcca aacggagggg gttcacaat gtaaacgaag tggagga cacggggcaat 480 gaacttagta cgaagaaca gatcagtagg aactctagg ctctcgaaga gaatacgtc 540 gctgagttgc agcttgagag actgaaaaaa gacggcgaag tacgcggatc tattatagg 600 ttcaagactt cagattacgt aaggaagcc aagcagctcc tgaagtaca gaaagcgtac 660 catcagctcg atcagagctt catcgatacc tacatagatt tgctggagac acggaggaca 720 tactacgagg gcccagggga aggatctcct tttggtgga aggacatcaa ggaatggtac 780 gagatgctta tgggacattg tacatatttt ccggaggagc tcaggagcgt caagtacgcc 840 tacaatgccg acctgtacaa tgccctcaat gacctcaata acctcgtgat taccagggac 900 gagaacgaga agctggagta ctatgaaag ttccagatta tcgagaatgt gtttaagcag 960 aagaagaagc cgacacttaa gcagattgca aaggaatcc tcgtgaatga ggaagatatc 1020 aagggataca gagtgacaag tacaggcaag cccgagttca caatctgaa ggtgtaccac 1080 gatattagg acataccgc acgaaggg ataatcgaaa acgctgagct cctcgatcag 1140 atcgcaaaaa 1200 ctgaacagtg agctcaccca agaggaatt gagcagattt siaacctgaa agctcacc 1260 gggacgcaca atctgagcct caagcaatc aacctcattc tggatgact ttggcacaca 1320 aatgacaacc aaattgccat attcaaccc ctgaactgg tgccaaaaaa agtggatctg 1380 tcacagcaaa aggaatccc tacaacttg gttgacgatt ttattctc cccgttgtc 1440 1500 1500 ccaacgata taattatcga gcttgcacga gagagaact ccaaggacgc ccagagatg 1560 attack attacgaaa tgcagaagcg aaccgccag aacgaac gcatagagg attattaag 1620 acaaccggca agagaatgc caagtatctg atcgagaaaa tcaagctgca cgacatgca 1680 gaaggcaagt gcctgtactc tctggaagct atcccactcg aagatctgct gatataatcca 1740 ttcattacg aggtggacca catcatccct agatccgtaa gcttgaca ttccttcaat 1800 aacaagttc tggttaaaca ggaggaaat tctaaaaag ggaccggac cccgttccag 1860 tacctgagct ccagtgacag tacctgactc tacgagactt tgaaca tattctgaat 1920s ctggccaaag gcaaggcag gatcagcaag accagaagg agtaccctcct cgaagaacggc 1980 gabattaaca gatttagtgt gcagaagat ttcatcacc gaaaccttgt cgatactcgg 2040 tacgccacga gaggcctgat gatctccctc aggagctact tccgcgtca taatctggac 2100 gttaaagtca agagcataaa tggggttc accagctttc tgaggagaaa gtggaagtttt 2160 aagaaggaac gaaaaagg atacaagcac catgctgagg atgctttgat catcgctaac 2220 gcggacttta tctttaagga atggaaaag ctggaagg CAagaaagt gatgaaaac 2280 cagatgttcg aggaagca ggcagagtca atgcctgaga tcgagacaga gcaggaatac 2340 aaggaaattt tcatcacccc tcatcagatt aacacata aggactcaa agacttaa 2400 tactctcata gggtggacaa aaaacccaat cgcaagctca ttaatgacac cctgtactca 2460 acacggaagg atgataagg taataccttg attgtgaata atcttaatgg attgtatgac 2520 aaagataacg acaagctca gaagctgatc aacaagtctc cagagaagct ccttatgtat 2580 caccacgacc cacagactta tcagaaattg aaactgatca tggagcaata cggggatgag 2640 aagaacccac tctacaaata ttatgaggaa acaggtaatt acctgaccaa gtactccaag 2700 areataacg gaccagtgat caaaaagata aagtactatg gcaacaaact taatgcgcat 2760 ttggacataa ctgacgatta ccccaattct cgaaaaagg ttgtgaagct ctccctgaag 2820 ccttatagat ttgacgtgta cctggataat ggggtttata aattcgtcac cgtgaaaaat 2880 2940 aagaagctga agaagatctc caatcaggcc gagttcatcg cttccttcta taagaacgat 3000 ctcatcaaga tcaatggaga gctttatcgc gtcattggtg tgaacaat cttgctgaac 3060 aggatcgaag tcaatatgat agacattacc taccgggagt atctcgaaaa catgaatgat 3120 aaacggccgc ctcacatcat caagacaatc gcatctaaaa ctcagtcaat aaaaaagtac 3180 tctaccgata tcctggggaa tctctatgaa gtgaagtcaa agaagcaccc acaaatcatt 3240 aaaaaaggtg gatcccccaa gaagaagagg aaagtctcga gcgactacaa agaccatgac 3300 ggtgattata aagatcatga catcgattac aaggatgacg atgacaagta aagcggccgc 3360 ataatgctta agtcgaacag aaagtaatcg tattgtacac ggccgcataa tcgaaattaa 3420 tacgactcac tataggtcgc cctcgaactt cacctgtttt agtactctgt aatgaaaatt 3480 acagaatcta ctaaaacaag gcaaaatgcc gtgtttatct cgtcaacttg ttggcgagat 3540 tttttttccg ctgagcaata actagcataa ccccttgggg cctctaaacg ggtcttgagg 3600 ttttttg 3609 <210> 27 <211> 3264 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> MSP1830 plasmid <400> 27 atggggaaac ggaactacat cctggggctt gacattggga taaccagcgt tggctacgga 60 attattgatt atgagacacg cgatgtgatt gacgccgggg ttaggctgtt caaagaggcc 120 aacgttgaaa acaacgaggg aagacggagt aagcgcggag caagaagact caagcgcaga 180 cggagacatc ggattcagag ggtgaaaaag ctgctcttcg attacaatct cctgaccgat 240 catagtgagc tgagcggaat caacccctac gaggcgcgag tgaaagggct ttcccagaag 300 ctgtccgaag aggagttctc cgccgcgttg ctgcacctgg ccaaacggag gggggttcac 360 aatgtaaacg aagtggaga ggacacggc aatgaactta gtacgaaaga acagatcagt 420 aggaactcta aggctctcga agaaatac gtcgctgagt tgcagcttga gagactgaaa 480 aaagacggcg aagtacgcgg atctattaat aggttcaaga cttcagatta cgtaaaggaa 540 gccaagcagc tcctgaaagt agaaagcg taccatcagc tcgatcagag cttcatcgat 600 acctacatag atttgctgga gacacggagg acatactacg agggcccagg ggaaggatct 660 ccttttgggt ggaaggacat caaggaatgg tacgagatgc ttatgggaca ttgtacatat 720 tttccggagg agctcaggag cgtcaagtac gcctacaatg ccgacctgta caatgccctc 780 aatgacctca ataacctcgt gattaccagg gacgagaacg agaagctgga gtactatgaa 840 aagttccaga ttatcgagaa tgtgtttaag cagaaaga agccgacact taagcagatt 900 ccaaaggaaa tcctcgtgaa tgaggaagat atcaagggat acagagtgac aagtacaggc 960 aagcccgagt tcacaaatct gaaggtgtac cacgatatta aggacataac cgcacgaaag 1020 gagataatcg aaaacgctga gctcctcgat cagatcgcaa aaattcttac catctaccag 1080 tctagtgagg acattcagga ggaactgact aatctgaaca gtgagctcac ccaagaggaa 1140 attgagcaga tttcaaacct gaaaggctac accgggacgc acaatctgag cctcaaagca 1200 atcaacctca ttctggatga actttggcac acaaatgaca accaaattgc catattcaac 1260 cgcctgaaac tggtgccaaa aaaagtggat ctgtcacagc aaaaggaaat ccctacaacc 1320 ttggttgacg attttattct gtcccccgtt gtcaagcgga gcttcatcca gtcaatcaag 1380 gtgatcaatg ccatcattaa aaaatacgga ttgccaaacg atataattat cgagcttgca 1440 cgagagaaga actcaaagga cgcccagaag atgattaacg aaatgcagaa gcgcaaccgc 1500 cagacaaacg aacgcataga ggaaattata agaacaaccg gcaaagagaa tgccaagtat 1560 ctgatcgaga aaatcaagct gcacgacatg caagaaggca agtgcctgta ctctctggaa 1620 gctatcccac tcgaagatct gctgaataat ccattcaatt acgaggtgga ccacatcatc 1680 cctagatccg tagctttga caattccttc aatacaaag ttctggttaa acaggagga aattctaaaa aagggaaccg gaccccgttc cagtacctga gctccagtga cagcaagatt agctcgaga cttttaaga acatattctg aatctggcca aaggcaaagg caggatcagc aagaccaaga aggagtacct cctcgaaga cgcgacatta acagatttag tgtgcagaaa gatttcatca accgaaacct tgtcgatact cggtacgcca cgagaggcct gatgaatctc ctcaggagct acttccgcgt caataatctg gacgttaaag tcaagagcat aaatggggga ttcaccagct ttctgaggag aaagtggaag tttaagaagg aacgaaacaa aggatacaag caccatgctg aggatgcttt gatcatcgct aacgcggact ttatctttaa ggaatggaa aagctggata aggcaaagaa agtgatgga aaccagatgt tcgaggaga gcaggcagag tcaatgcctg agatcgagac agagcagga tacaagga ttttcatcac ccctcatcag attaacaca taaaggactt aaatactctc atagggtgga caaaaaaccc aatcgcaagc tcattaatga caccctgtac tcaacacgga aggatgata aggtaatacc ttgattgtga ataatcttaa tggattgtat gacaaagata acgacaagct caagaagctg 2460 atcaacaagt ctccagagaa gctccttatg tatcaccacg acccacagac ttatcagaaa 2520 ttgaaactga tcatggagca atacggggat gagaagaacc cactctacaa atattatgag 2580 gaaacaggta attacctgac caagtactcc aagaaggata acggaccagt gatcaaaaag 2640 ataaagtact atggcaacaa acttaatgcg catttggaca taactgacga ttaccccaat 2700 tctcgaaaca aggttgtgaa gctctccctg aagccttata gatttgacgt gtacctggat 2760 aatggggttt ataaattcgt caccgtgaaa aatctggacg tgatcaaaaa ggagaactat 2820 tatgaagtaa actcaaagtg ctatgaggag gcgaagaagc tgaagaagat ctccaatcag 2880 gccgagttca tcgcttcctt ctataagaac gatctcatca agatcaatgg agagctttat 2940 cgcgtcattg gtgtgaacaa tgacttgctg aacaggatcg aagtcaatat gatagacatt 3000 acctaccggg agtatctcga aaacatgaat gataaacggc cgcctcacat catcaagaca 3060 atcgcatcta aaactcagtc aataaaaaag tactctaccg atacctggg gaatctctat 3120 gaagtgaagt caaagaagca cccacaaatc attaaaaaag gtggatcccc caagaagaag 3180 aggaaagtct cgagcgacta caaagaccat gacggtgatt ataaagatca tgacatcgat 3240 tacaaggatg acgatgacaa gtaa 3264 <210> 28 <211> 430 <212> DNA <213> Artificial Sequence <220> <221> unsure <223> BPK2660 plasmid <400> 28 tgtacaaaaa agcaggcttt aaaggaacca attcagtcga ctggatccgg taccaaggtc 60 gggcaggaag agggcctatt tcccatgatt ccttcatatt tgcatatacg atacaaggct 120 gttagagaga taattagaat taatttgact gtaaacacaa agatattagt acaaaatacg 180 tgacgtagaa agtaataatt tcttgggtag tttgcagttt taaaattatg ttttaaaatg 240 gactatcata tgcttaccgt aacttgaaag tatttcgatt tcttggcttt atatatcttg 300 tggaaaggac gaaacacccg agacgattaa tgcgtctcgg ttttagtact ctgtaatgaa 360 aattacagaa tctactaaaa caaggcaaaa tgccgtgttt atctcgtcaa cttgttggcg 420 agattttttt 430

Claims

1. An isolated Streptococcus pyogenes Cas9 (SpCas9) protein, wherein the amino acid sequence of the SpCas9 protein is as shown in SEQ ID NO. The amino acid sequence shown in NO:1 contains only mutations selected from the following groups: D1135V / R1335Q / T1337R, D1135V / G1218R / R1335Q / T1337R, L1111H / D1135V / R1335Q / T1337R, L1111H / D1135V / G1218R / R1335Q / T1337R, D1135V / N1317K / R1335Q / T1337R, D1135V / G1218R / N1317K / R1335Q / T1337R, G1104R / D1135V / R1335Q / T1337R, G11 04R / D1135V / G1218R / R1335Q / T1337R, S1109T / D1135V / R1335Q / T1337R, S1109T / D1135V / G1218R / R1335Q / T1337R, D1135N / S1136N / R1335Q / T1337R, D1135N / S1136N / G1218R / R1335Q / T1337R, D1135E / R1335Q / T1337R, and D1135V / G1218R / R1335E / T1337R, wherein the SpCas9 protein has altered PAM specificity.

2. A fusion protein comprising the isolated protein of claim 1, fused with an optional interventional linker and a heterologous functional domain, wherein the linker does not interfere with the activity of the fusion protein.

3. The fusion protein of claim 2, wherein the heterologous functional domain is a transcriptional activation domain.

4. The fusion protein of claim 3, wherein the transcriptional activation domain is derived from VP64 or NF-κB p65.

5. The fusion protein of claim 2, wherein the heterologous functional domain is a transcriptional silencer or a transcriptional repressor domain.

6. The fusion protein of claim 5, wherein the transcriptional repression domain is a Kluber-associated box domain, an ERF repression factor domain, or an mSin3A interaction domain, or wherein the transcriptional silencer is heterochromatin protein 1.

7. The fusion protein of claim 2, wherein the heterologous functional domain is an enzyme that modifies the methylation state of DNA.

8. The fusion protein of claim 7, wherein the enzyme that modifies the methylation state of DNA is a DNA methyltransferase or a TET protein.

9. The fusion protein of claim 8, wherein the TET protein is TET1.

10. The fusion protein of claim 2, wherein the heterologous functional domain is an enzyme that modifies histone subunits.

11. The fusion protein of claim 10, wherein the enzyme modifying the histone subunit is a histone acetyltransferase, histone deacetylase, histone methyltransferase, or histone demethylase.

12. The fusion protein of claim 2, wherein the heterologous functional domain is a biological chain.

13. The fusion protein of claim 12, wherein the biological chain is an MS2, Csy4, or λN protein.

14. The fusion protein of claim 2, wherein the heterologous functional domain is FokI.

15. An isolated nucleic acid encoding a protein as described in any one of claims 1-14.

16. A vector comprising the isolated nucleic acid as described in claim 15.

17. The vector of claim 16, wherein the isolated nucleic acid of claim 15 is operatively linked to one or more regulatory domains for expression of isolated Streptococcus pyogenes Cas9 (SpCas9) protein, wherein the amino acid sequence of the SpCas9 protein is as shown in SEQ ID NO. The amino acid sequence shown in NO:1 contains only mutations selected from the following groups: D1135V / R1335Q / T1337R, D1135V / G1218R / R1335Q / T1337R, L1111H / D1135V / R1335Q / T1337R, L1111H / D1135V / G1218R / R1335Q / T1337R, D1135V / N1317K / R1335Q / T1337R, D1135V / G1218R / N1317K / R1335Q / T1337R, G1104R / D1135V / R1335 Q / T1337R, G1104R / D1135V / G1218R / R1335Q / T1337R, S1109T / D1135V / R1335Q / T1337R, S1109T / D1135V / G1218R / R1335Q / T1337R, D1 135N / S1136N / R1335Q / T1337R, D1135N / S1136N / G1218R / R1335Q / T1337R, D1135E / R1335Q / T1337R and D1135V / G1218R / R1335E / T1337R.

18. A host cell comprising the nucleic acid as described in claim 15.

19. The host cell of claim 18, wherein the host cell expresses the protein of any one of claims 1-14.

20. The host cell as claimed in claim 18 or 19, wherein the host cell is a mammalian host cell.

21. A method for altering the genome of a cell for non-therapeutic purposes, the method comprising expressing in the cell a protein as claimed in any one of claims 1-14 and a guide RNA having a region complementary to a selected portion of the genome of the cell, or contacting the cell with the protein and the guide RNA having a region complementary to a selected portion of the genome of the cell.

22. The protein according to any one of claims 1-14, characterized in that... A method for using the protein to alter a cell genome, the method comprising expressing the protein in a cell and a guide RNA having a region complementary to a selected portion of the cell's genome, or contacting a cell with the protein and the guide RNA having a region complementary to a selected portion of the cell's genome.

23. The method of claim 21 or the protein of claim 22, wherein the protein comprises one or more of a nuclear localization sequence, a cell-penetrating peptide sequence, and / or an affinity tag.

24. The method of claim 21 or the protein of claim 22, wherein the cell is a stem cell.

25. The method of claim 24 or the protein of claim 24, wherein the cell is an embryonic stem cell, a mesenchymal stem cell, or an induced pluripotent stem cell.

26. A method for altering a double-stranded DNA (dsDNA) molecule in vitro, the method comprising contacting the dsDNA molecule with a protein as described in any one of claims 1-14 and a guide RNA having a region complementary to a selected portion of the dsDNA molecule.

27. The protein according to any one of claims 1-14, characterized in that... A method for using the protein to alter a double-stranded DNA (dsDNA) molecule, the method comprising contacting the dsDNA molecule with the protein and a guide RNA having a region complementary to a selected portion of the dsDNA molecule.

Citation Information

Patent Citations

  • TARGET DNA INTERFERENCE WITH crRNA

    US20100076057A1

  • PROKARYOTIC RNAi-LIKE SYSTEM AND METHODS OF USE

    US20110189776A1

  • Methods of Generating Nucleic Acid Fragments

    US20110223638A1

  • Endoribonuclease compositions and methods of use thereof

    US20130130248A1

  • Crispr-CAS systems and methods for altering expression of gene products

    US20140170753A1