Improved high-throughput combinatorial gene modification system and optimized Cas9 enzyme variants

The high-throughput genetic engineering system using type IIS restriction enzymes efficiently generates and screens combinatorial mutants of Cas9 enzymes, enhancing on-target activity and minimizing off-target effects.

JP7813045B2Active Publication Date: 2026-02-12THE UNIVERSITY OF HONG KONG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023119639
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-09-19
Filing Date
2023-07-24
Publication Date
2026-02-12
Estimated Expiration
2039-09-17

AI Technical Summary

Technical Problem

Current systems for generating and screening large numbers of genetic variants of recombinant proteins, such as Cas9 enzymes, are tedious and labor-intensive, lacking efficiency in high-throughput combinatorial recombination.

Method used

A high-throughput genetic engineering system using type IIS restriction enzymes for seamless linkage of genetic elements in DNA constructs, enabling the creation of combinatorial mutants without artificial sequences, and the development of SpCas9 variants with improved on-target cleavage and reduced off-target activity.

Benefits of technology

Facilitates the rapid generation and screening of combinatorial mutants with enhanced protein functionality, specifically improving the on-target activity and reducing off-target effects of Cas9 enzymes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007813045000047
    Figure 0007813045000047
  • Figure 0007813045000048
    Figure 0007813045000048
  • Figure 0007813045000049
    Figure 0007813045000049
Patent Text Reader

Abstract

To provide specific polypeptides and methods for cleaving a DNA molecule at a target site using the polypeptides.SOLUTION: The invention provides a polypeptide comprising a specific amino acid sequence where a residue corresponding to residue 1003 of a specific sequence is substituted and a residue corresponding to residue 661 of the specific sequence is substituted. The invention provides such a polypeptide where the residue corresponding to residue 1003 of the specific sequence is substituted with histidine and the residue corresponding to residue 661 of the specific sequence is substituted with alanine. A method for cleaving a DNA molecule at a target site is provided which comprises bringing the DNA molecule comprising the target DNA site into contact with the polypeptide and a short guide RNA (sgRNA) that specifically binds the target DNA site, thereby causing the DNA molecule to be cleaved at the target DNA site.SELECTED DRAWING: Figure 1a
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 733,410, filed September 19, 2018, the entire contents of which are incorporated herein by reference for all purposes. [Background technology]

[0002] background Recombinant proteins are becoming increasingly important in a variety of applications, including industrial and medical applications. Because the functionality of recombinant proteins, particularly enzymes and antibodies, can be improved by genetic mutation, ongoing efforts have been made to generate and select a wide range of genetic variants of potential recombinant proteins in order to identify recombinant proteins with more desirable properties that can improve their effectiveness in those applications.

[0003] Cas9 (CRISPR-associated protein 9) is an RNA-guided DNA endonuclease associated with the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat) adaptive immune system in bacteria, such as Streptococcus pyogenes, a Gram-positive bacterium in the genus Streptococcus. Given the increasing use of CRISPR for gene editing in recent years, Cas9 has attracted the interest of many researchers seeking to improve the performance of genetic engineering. However, currently available systems for systematically generating and screening large numbers of genetic variants of specific proteins are often tedious, labor-intensive, and therefore inefficient.

[0004] Thus, there is a clear need for new high-throughput combinatorial recombination systems / methods and recombination proteins (e.g., Cas9 enzymes) with improved characteristics. The present invention fulfills this and other related needs. Summary of the Invention

[0005] Summary of the Invention Previously, a research group led by the present inventors devised a system for high-throughput functional analysis of high-order barcoded combinatorial gene libraries, which we termed combinatorial gene en masse or CombiGEM. This system has been used, for example, to generate libraries of barcoded dual guide RNA (gRNA) combinations and libraries of two- or three-way barcoded human microRNA (miRNA) precursors, which can then be further screened for desired functionality (see, for example, Wong et al. (Nat. Biotechnol. 2015 September; 33(9):952-961), Wong et al. (Proc. Nat. Acad. Sci., March 1, 2016, 113(9):2544-2549), WO2016 / 070037, and WO2016 / 115033). See also U.S. Patent No. 9,315,806. The present inventors have now further improved the CombiGEM system and developed an improved CombinSEAL platform that provides seamless linkages between any two adjacent genetic elements of each member of a high-order combinatorial mutant library. In other words, this platform does not introduce any artificial or foreign amino acid sequences at each linkage site, making it possible to generate a large collection of protein variants containing combinatorial mutations while maintaining the native amino acid sequence of the wild-type protein.

[0006] Thus, the present invention provides, for the first time, an improved high-throughput genetic engineering system for systematically generating and screening combinatorial mutants. In one aspect, the present invention provides a DNA construct comprising, in the 5' to 3' direction of the DNA strand, a first recognition site for a first type IIS restriction enzyme; a DNA element; first and second recognition sites for a second type IIS restriction enzyme; a barcode uniquely assigned to the DNA element; and a second recognition site for the first type IIS restriction enzyme. In one embodiment, the DNA construct is a linear construct. In another embodiment, the DNA construct is a circular construct or a DNA vector, including a bacterial-based DNA plasmid or DNA virus vector. The DNA construct is preferably isolated, i.e., isolated in the absence of significant amounts of other DNA sequences. In one embodiment, the present invention provides a library comprising at least two, and possibly more, of the DNA constructs described above and herein, each of the library members having a distinct DNA element with a different polynucleotide sequence along with a uniquely assigned barcode.

[0007] In another aspect of the present invention, another DNA construct is provided: the DNA construct comprises, in the 5' to 3' direction of the DNA strand, a recognition site for a first type IIS restriction enzyme; a plurality of DNA elements; a primer binding site; and a recognition site for a second type IIS restriction enzyme, each of a plurality of barcodes being uniquely assigned to one of the plurality of DNA elements, wherein the plurality of DNA elements are connected to each other to form a coding sequence for a protein (e.g., a coding sequence for a natural or wild-type protein) without any extraneous sequence at the junction between any two of the plurality of DNA elements, and wherein each of the plurality of barcodes is arranged in the reverse order of the assigned DNA element. In one embodiment, the DNA construct is linear. In another embodiment, the DNA construct is circular, for example, a DNA vector including a bacterial-based DNA plasmid or a DNA virus vector. A library of such constructs is also provided, comprising at least two, and sometimes more, constructs, each member having a set of DNA elements of a different polynucleotide sequence and a set of uniquely assigned barcodes.

[0008] In some embodiments of any of the DNA constructs described above and herein, the first type IIS restriction enzyme and the second type IIS restriction enzyme cleave the DNA molecule to create compatible ends. In some embodiments, the first type IIS restriction enzyme is BsaI. In some embodiments, the second type IIS restriction enzyme is BbsI.

[0009] In a further aspect, the invention relates to a method for making combinatorial genetic constructs. The method comprises the steps of: (a) cleaving the first DNA vector of claim 2 with a first type IIS restriction enzyme to release a first DNA fragment comprising a first DNA segment, first and second recognition sites for a second type IIS restriction enzyme, and a first barcode adjacent to the first and second ends created by the first type IIS restriction enzyme; (b) cleaving a first expression vector comprising a promoter with a second type IIS restriction enzyme to linearize the first expression vector near the 3' end of the promoter and create two ends compatible with the first and second ends of the DNA fragment of (a); (c) annealing and ligating the first DNA fragment of (a) to the linearized expression vector of (b) to form a 1-way composite expression vector in which the first DNA fragment and the first barcode are operably linked to the promoter at their 3' ends; and (d) (e) cleaving the second DNA vector of claim 2 with a first type IIS restriction enzyme to release a second DNA fragment comprising a second DNA segment, first and second recognition sites for the second type IIS restriction enzyme, and a second barcode adjacent to the first and second ends created by the first type IIS restriction enzyme; (e) cleaving the composite expression vector of (c) with a second type IIS restriction enzyme to linearize the composite expression vector between the first DNA element and the first barcode and create two ends compatible with the first and second ends of the DNA fragment of (d);and (f) annealing and ligating the second DNA fragment of (d) between the first DNA element and the first barcode to the linearized composite expression vector of (e) to form a two-way composite expression vector, wherein the first DNA fragment, the second DNA fragment, the second barcode, and the first barcode are operably linked to a promoter at their 3' ends in this order, wherein the first and second DNA elements encode first and second segments of a preselected protein from their N-termini immediately adjacent to each other, the first and second DNA fragments are joined to each other in the two-way composite expression vector without any exogenous nucleotide sequence resulting in an amino acid residue not found in the preselected protein, and the first and second DNA elements each comprise one or more mutations;

[0010] In one embodiment of this method, steps (d) through (f) are repeated an nth time to incorporate an nth DNA fragment comprising an nth DNA element, a first and second recognition site for a second type IIS restriction enzyme, and an nth barcode into an n-way hybrid expression vector, wherein the nth DNA element encodes the nth or second-to-last segment of a preselected protein from its C-terminus. The method further comprises the steps of: (x) providing a final DNA vector comprising an (n+1)th DNA element, a primer binding site, and an (n+1)th barcode between the first and second recognition sites for the first type IIS restriction enzyme; (y) cleaving the final DNA vector with the first type IIS restriction enzyme to release a final DNA fragment comprising, in 5' to 3' order, the (n+1)th DNA element, the primer binding site, and the (n+1)th barcode adjacent to the first and second ends generated by the first type IIS restriction enzyme; (z) annealing and ligating the final DNA fragment to an n-way composite expression vector generated after repeating steps (d) to (f) n times and linearized with the second type IIS restriction enzyme to form a final composite expression vector. forming a final composite expression vector, wherein the first, second and through nth and (n+1)th DNA elements encode the first, second and through nth and final segments of a preselected protein immediately adjacent to each other from their N-termini, the first, second and through nth and final DNA fragments being joined together in the final composite expression vector free of any extraneous nucleotide sequence that results in any amino acid residues not found in the preselected protein, and wherein each of the DNA elements contains one or more mutations.

[0011] In certain embodiments of the methods described above and herein, the first type IIS restriction enzyme and the second type IIS restriction enzyme cleave the DNA molecule to create compatible ends. In certain embodiments, the first type IIS restriction enzyme is BsaI. In certain embodiments, the second type IIS restriction enzyme is BbsI.

[0012] In a further aspect, the present invention provides a library comprising at least two, and optionally more, final composite expression vectors produced by the methods described above and herein.

[0013] Second, the present invention provides SpCas9 variants with improved on-target cleavage ability and reduced off-target cleavage ability, generated and identified by using the improved high-throughput genetic engineering system described herein. In one aspect, the present invention provides a polynucleotide (preferably an isolated polynucleotide) comprising the amino acid sequence set forth in any one of SEQ ID NOS: 1 and 4-13, where the polypeptide serves as a base sequence, wherein at least one, and possibly more, residues corresponding to residues 661, 695, 848, 923, 924, 926, 1003, or 1060 of SEQ ID NO: 1 have been modified, e.g., by substitution. Some exemplary polypeptides of the present invention are provided in Table 2 herein. In some embodiments, the residue corresponding to residue 1003 of SEQ ID NO: 1 has been substituted, and the residue corresponding to residue 661 of SEQ ID NO: 1 has been substituted. In some embodiments, the polypeptide is further substituted with a residue corresponding to residue 926 of SEQ ID NO: 1. For example, the polypeptide has a residue corresponding to residue 1003 of SEQ ID NO:1 substituted with histidine and a residue corresponding to residue 661 of SEQ ID NO:1 substituted with alanine. In another example, the polypeptide has the basic amino acid sequence set forth in SEQ ID NO:1, in which residue 1003 is substituted with histidine, residue 661 is substituted with alanine, and optionally further includes an alanine substitution at residue 926. In a further example, the polypeptide has the basic amino acid sequence set forth in SEQ ID NO:1, in which residues 695, 848, and 926 are substituted with alanine, residue 923 is substituted with methionine, and residue 924 is substituted with valine. Also provided are compositions comprising: (1) a polypeptide as described above and herein; and (2) a physiologically acceptable excipient.

[0014] In another aspect, the present invention provides nucleic acids (preferably isolated nucleic acids) comprising a polynucleotide sequence encoding a polypeptide described above and herein, and compositions comprising said nucleic acids. The present invention also provides expression cassettes comprising a promoter operably linked to a polynucleotide sequence encoding a polypeptide of the invention, and vectors (e.g., bacterial-based plasmids or viral-based vectors) comprising said expression cassettes, and host cells comprising said expression cassettes or polypeptides of the invention.

[0015] In a further aspect, the present invention provides a method for cleaving a DNA molecule at a target site. The method comprises contacting a DNA molecule containing a target DNA site with a polypeptide described above and herein and a short guide RNA (sgRNA) that specifically binds to the target DNA site, thereby cleaving the DNA molecule at the target DNA site. In some embodiments of this method, the DNA molecule is genomic DNA in a living cell, and the cell is transfected with a polynucleotide sequence encoding the sgRNA and the polypeptide. In some cases, the cell is transfected with a first vector encoding the sgRNA and a second vector encoding the polypeptide. In other cases, the cell is transfected with a vector encoding both the sgRNA and the polypeptide. In some embodiments of this method, each of the first and second vectors is a viral vector, such as a retroviral vector, particularly a lentiviral vector.

[0016] The high-throughput combinatorial genetic recombination systems, methods, and related compositions described above and herein are suitable for use in either prokaryotic or eukaryotic cells, with appropriate modifications. Some equivalents can also be derived from the above and herein. For example, the arrangement of the DNA elements and their corresponding barcodes in each DNA construct can be swapped, i.e., the DNA construct comprises, from 5' to 3', a first recognition site for a first type IIS restriction enzyme, a barcode uniquely assigned to the DNA element, first and second recognition sites for a second type IIS restriction enzyme, the DNA element, and a second recognition site for the first type IIS restriction enzyme. Such DNA constructs and libraries thereof can be used in a manner similar to that described herein to produce intermediate and final vectors similar to those described herein, except that the relative positions of the DNA elements and barcodes in these vectors are swapped as appropriate. [Brief explanation of the drawings]

[0017] [Figure 1]Creation of a high-coverage combinatorial mutant library of SpCas9 and efficient delivery of the library into human cells. a) Strategy for assembling the SpCas9 combinatorial mutant library. The SpCas9 coding sequence was modularized into four configurable parts (i.e., P1–P4), each containing a repertoire of barcoded fragments encoding predetermined amino acid residue mutations at defined positions, as shown in the figure. A library of 952 SpCas9 mutants was assembled by sequential rounds of one-pot seamless ligation of the parts, generating concatenated barcodes uniquely tagging each mutant (see Figure 7 for details). b) Cumulative distribution of sequencing reads for the barcoded combinatorial mutant library within a plasmid pool extracted from E. coli and within an infected OVCAR8-ADR cell pool. High coverage of libraries within the plasmid and infected cell pools (~99.9% and ~99.6%, respectively) was detected from approximately 800,000 reads per sample, with most combinations detected with at least 300 absolute barcoded reads (highlighted with shading).

[0018] [Figure 2]Strategy for profiling the on-target and off-target activity of SpCas9 mutants in human cells. a) The SpCas9 library was delivered via lentivirus at a multiplicity of infection of ~0.3 into the OVCAR8-ADR reporter cell line, which expresses RFP and GFP genes driven by the UBC and CMV promoters, respectively, and a tandem U6 promoter-driven expression cassette (RFPsg5 or RFPsg8) of a gRNA targeting RFP. RFP and GFP expression was analyzed by flow cytometry. SpCas9 on-target activity was measured when the gRNA spacer sequence perfectly matched the RFP target site, and its off-target activity was measured when the RFP target site contained a synonymous mutation. Cells harboring active SpCas9 mutants were expected to lose RFP fluorescence. Cells were sorted into bins containing approximately 5% of the population based on RFP fluorescence, and their genomic DNA was extracted by Illumina HiSeq to quantify barcoded SpCas9 variants. b) Scatter plot comparing the barcode counts of each SpCas9 variant between the sorted bins (i.e., A, B, and C) and the unsorted population. Each dot represents an SpCas9 variant; wild-type (WT) SpCas9 and eSpCas9 (1.1) are labeled within the plot. The solid reference line indicates a 1.5-fold enrichment and a 0.5-fold reduction in barcode counts, while the dotted reference line indicates no change in the barcode counts in the sorted bins compared to the unsorted population.

[0019] [Figure 3]High-throughput profiling reveals broad-spectrum specificity and efficiency of SpCas9 combinatorial variants. a) SpCas9 combinatorial variants are ranked by the log-transformed enrichment ratio (i.e., log2(E)), which represents their relative abundance in sorted RFP-depleted cell populations for each of the target molecule (x-axis) and another molecule (y-axis) reporter cell lines, based on profiling data from two biological replicates (see Table 2 and Methods for details). Each dot in the scatter plot represents an SpCas9 variant; WT SpCas9, eSpCas9(1.1), Opti-SpCas9, and OptiHF-SpCas9 are labeled. More than 99% of the combination mutants had lower log2(E) values ​​than WT in the two off-target reporter lines, RFPsg5-OFF5-2 and RFPsg8-OFF5, while 16.2% and 2.5% of the mutants had higher log2(E) values ​​than WT in the two on-target reporter lines, RFPsg5-ON and RFPsg8-ON, respectively. b) OVCAR8-ADR reporter cells bearing on-target (upper panel) and off-target (lower panel) target sites were infected with individual SpCas9 combination mutants. The editing efficiency of the SpCas9 mutants was measured by the percentage of cells with depleted RFP levels and compared with WT.

[0020] [Figure 4]Heatmap showing editing efficiency and epistasis at on-target and off-target sites. Editing efficiency (upper panel; measured by log2(E)) and epistasis (lower panel; ε) scores were measured for each SpCas9 combination variant as described in the Methods section. For visualization, amino acid residues predicted to contact the target DNA strand or located in the linker region connecting the HNH and RuvC domains of SpCas9 are grouped on the y-axis, while amino acid residues predicted to interact with the non-target DNA strand are shown on the x-axis. P values ​​for log2(E) for each combination were calculated by comparing the log2(E) values ​​with those in the entire population obtained from two independent biological replicates using a two-sample, two-tailed Student's t-test (MATLAB function 'ttest2'). Adjusted P values ​​(i.e., Q values) were calculated based on the distribution of P values ​​(MATLAB function 'mafdr') to correct for multiple hypothesis testing. Log2(E) was considered statistically significant for the entire population based on a Q-value cutoff of <0.1 and is boxed. The full heatmap is shown in Figure 10. Combinations for which no enrichment ratio or epistasis score was measured are shown in gray.

[0021] [Figure 5]Opti-SpCas9 demonstrates robust on-target activity and reduced off-target activity. ab) Evaluation of SpCas9 mutants for efficient on-target editing using gRNAs targeting endogenous loci. The rate of insertion or deletion mutations (indels) was measured using a T7 endonuclease I (T7E1) assay. The ratio of on-target activity of SpCas9 mutants relative to WT (a) and Opti-SpCas9 (b) was determined, and the median and interquartile range of the normalized percentage of indel formation are shown for the 10–16 loci tested. Each locus was measured once or twice, and the complete dataset is shown in Figure 12. c) GUIDE-Seq genome-wide specificity profiles of a panel of SpCas9 mutants, each paired with the indicated gRNA. Mismatched positions at off-target sites are highlighted in color, and the number of GUIDE-Seq reads was used as an indicator of cleavage efficiency at a specific site. The list of gRNA sequences used is shown in Table 5.

[0022] [Figure 6] Example strategies for characterizing combinatorial mutations on protein sequences.

[0023] [Figure 7]Strategy for seamless assembly of barcoded combinatorial mutant library pools. a) To create barcoded DNA segments in storage vectors, gene inserts were generated by PCR or synthesis and cloned into storage vectors with random barcodes (pAWp61 and pAWp62; digested with EcoRI and BamHI) using a Gibson assembly reaction. BsaI digestion was performed to generate barcoded DNA segments (i.e., P1, P2, ... P(n)). A BbsI site and a primer binding site for barcode sequencing were introduced between the insert and barcode in pAWp61 and pAWp62, respectively. b) To generate barcoded combinatorial mutant libraries, the pooled DNA segments and the destination assembly vector were digested with BsaI and BbsI, respectively. One-pot ligation generated a pooled vector library, which was further iteratively digested and ligated with subsequent pooled DNA segments to generate higher-order combinatorial mutants. After digestion with type IIS restriction enzymes (i.e., BsaI and BbsI), the barcoded inserts were ligated via compatible overhangs derived from the protein-coding sequence, thereby preventing the formation of a fusion scar in the ligation reaction. All barcodes were localized in a contiguous stretch of DNA. The final combinatorial mutant library was encoded into lentivirus and delivered to targeted human cells. The integrated barcodes representing each combination were amplified in an unbiased manner from genomic DNA within pooled cell populations and quantified using high-throughput sequencing to identify shifts in presentation under different experimental conditions. c) Highly reproducible presentation was demonstrated between plasmids and infected cell pools, and between biological replicates of the infected cell pools.

[0024] [Figure 8]Fluorescence-activated cell sorting of human cells infected with an SpCas9 library containing on-target and off-target reporters. The OVCAR8-ADR reporter cell line, expressing the RFP and GFP genes driven by the UBC and CMV promoters, respectively, and a tandem U6 promoter-driven expression cassette (RFPsg5 or RFPsg8) of a gRNA targeting the RFP site, was either uninfected or infected with the SpCas9 library. The RFPsg5-ON and RFPsg8-ON lines have a site that perfectly matches the gRNA sequence, while the RFPsg5-OFF5-2 and RFPsg8-OFF5 lines contain synonymous mutations in RFP that do not match the gRNA. Cells were sorted by flow cytometry into bins containing approximately 5% of the population with low RFP fluorescence. These experiments were repeated twice independently with similar results.

[0025] [Figure 9] Positive correlation between enrichment scores determined from pooled screens and individual validation data. Normalized log2(E) for each SpCas9 combinatorial mutant is the average score determined from pooled screens with two biological replicates, and normalized RFP disruption value is the average percentage of cells with depleted RFP levels compared to WT determined from three biological replicates. R is Pearson's r (product-moment correlation coefficient).

[0026] [Figure 10]Heatmap showing the editing efficiency of on-target and off-target sites. Editing efficiency was measured by the log-transformed enrichment factor (log2(E)) determined for each SpCas9 combination mutant. Enriched and depleted mutants have values ​​>0 and <0, respectively. To aid visualization, amino acid residues predicted to contact the target DNA strand or located in the linker region connecting the HNH and RuvC domains of SpCas9 are grouped on the y-axis, while amino acid residues predicted to contact the non-target DNA strand are shown on the x-axis. Non-enriched combinations are shown in gray.

[0027] [Figure 11] Frequency of N20-NGG and G-N19-NGG sites in a control human genome. As an estimate of the target coverage of Opti-SpCas9 and other engineered SpCas9 variants, including eSpCas9(1.1), SpCas9-HF1, HypaCas9, and evoCas9, we used custom Python code to find the occurrence of N20-NGG and G-N19-NGG sites in both strands of the control human genome, hg19. N20-NGG sites are approximately 4.3 times more frequent than G-N19-NGG sites in the human genome.

[0028] [Figure 12] Summary of T7 endonuclease I (T7E1) assay results for DNA mismatch cleavage in OVCAR8-ADR cells. Cells were infected with SpCas9 mutants and the indicated gRNAs, and genomic DNA was collected for T7E1 assays 11–16 days postinfection. Indel quantification of infected samples is displayed as a bar graph.

[0029] [Figure 13]Expression of SpCas9 mutants in OVCAR8-ADR cells. Cells were infected with lentiviruses encoding WT SpCas9, Opti-SpCas9, eSpCas9(1.1), HypaCas9, SpCas9-HF1, Sniper-Cas9, evoCas9, xCas9, or OptiHF-SpCas9. Protein lysates were extracted for Western blot analysis and immunoblotted with anti-SpCas9 antibodies. β-actin was used as a loading control. Expression of SpCas9-HF1 and xCas9 was not detected in OVCAR8-ADR cells, likely due to their non-optimized sequences for expression in mammalian cells. Therefore, SpCas9-HF1 and xCas9 were not included in other activity assays. These experiments were repeated three times independently with similar results.

[0030] [Figure 14] Evaluation of the editing efficiency of SpCas9 mutants with gRNAs containing or lacking an additional mismatched 5' guanine (5'G) using a GFP disruption assay. OVCAR8-ADR cells expressing WT SpCas9, Opti-SpCas9, eSpCas9(1.1), or HypaCas9 were infected with lentiviruses encoding gRNAs containing or lacking an additional mismatched 5'G. Editing efficiency was measured by the percentage of cells with depleted GFP levels using flow cytometry. Values ​​and error bars reflect the mean and standard deviation (sd) of four independent biological replicates.

[0031] [Figure 15]Opti-SpCas9 exhibits reduced off-target activity compared to wild-type SpCas9. Evaluation of SpCas9 mutants for off-target editing delivered by VEGFA site 3 or DNMT1 site 4 gRNAs at eight endogenous loci. The percentage of indels was measured using the T7E1 assay, averaged from three independent experiments. Dashes indicate none were detected. The specificity of WT SpCas9 and its mutants with VEGFA site 3 gRNA at the OFF1 locus was plotted as the ratio of on-target activity to off-target activity (on-target activity data taken from Figure 12).

[0032] [Figure 16] Using a GFP disruption assay, we characterized SpCas9 mutants for editing target sites with sequences that either perfectly matched or contained mismatches with the spacer of the gRNA. OVCAR8-ADR cells expressing WT SpCas9, Opti-SpCas9, eSpCas9(1.1), or HypaCas9 were infected with lentivirus encoding gRNAs with no mismatches or 1-4 base mismatches to the target. Editing efficiency was measured by flow cytometry as the percentage of cells with depleted GFP levels. Values ​​and error bars reflect the mean and standard deviation of three independent biological replicates.

[0033] [Figure 17]On-target editing activity of SpCas9 mutants using truncated gRNAs. a, b) OVCAR8-ADR cells expressing WT SpCas9, Opti-SpCas9, eSpCas9(1.1), or HypaCas9 were infected with lentiviruses encoding GFP sequences (a) and gRNAs of different lengths (17–19 nucleotides) targeting the endogenous locus (b). Editing efficiency was measured by the percentage of cells with depleted GFP levels using flow cytometry (a) and the T7E1 assay (b). A list of the gRNA sequences used is shown in Table 5. For (a), values ​​and error bars reflect the mean and standard deviation of four independent biological replicates.

[0034] [Figure 18] Multiple sequence alignment - comparison of Streptococcus pyogenes Cas9 homologs. Conserved amino acid residues among Cas9 homologs are highlighted, particularly those corresponding to SpCas9 residues 661 and 1003. DETAILED DESCRIPTION OF THE INVENTION

[0035] definition As used herein, "CRISPR-Cas9" or "Cas9" refers to CRISPR-associated protein 9, an RNA-guided DNA endonuclease enzyme associated with the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) adaptive immune system found in several bacterial species, including Streptococcus pyogenes. The Cas9 protein from Streptococcus pyogenes, SpCas9, has the amino acid sequence set forth in SEQ ID NO:1, which is encoded by the polynucleotide sequence set forth in SEQ ID NO:2. For additional Cas9 enzymes with significant sequence homology, including at least some (e.g., at least two, three, four, five, or more, e.g., at least half, but not necessarily all) of known key conserved residues, such as residues 661, 695, 848, 923, 924, 926, 1003, and 1060 of SEQ ID NO:1, see the sequence alignment in Figure 18. As used herein, the term "Cas9 protein" includes any RNA-guided DNA endonuclease enzyme having substantial amino acid sequence identity, e.g., at least 50%, 60%, 70%, 75%, up to 80%, 85% or more overall sequence identity, to SEQ ID NO: 1. Exemplary wild-type Cas9 proteins include those derived from the bacterial species Streptococcus mutans, Streptococcus dysgalactiae, Streptococcus equi, Streptococcus oralis, Streptococcus mitis, Listeria monocytogenes, Enterococcus timonensis, Streptococcus thermophilus, and Streptococcus parasanguinis, which have the amino acid sequences set forth in SEQ ID NOs: 4-13, respectively.

[0036] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single- or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to natural nucleotides. Unless otherwise specified, a particular nucleic acid sequence also encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and relative sequences as well as the explicitly set forth sequence. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res., 19:5081 (1991); Ohtsuka et al., J. Biol. Chem., 260:2605-2608 (1985); and Cassol et al., (1992); Rossolini et al., Mol. Cell. Probes, 8:91-98 (1994)). The terms nucleic acid and polynucleotide are used interchangeably with gene, cDNA, and mRNA encoded by a gene.

[0037] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers in which one or more amino acid residues are artificial chemical mimetics of corresponding naturally occurring amino acids, as well as natural and unnatural amino acid polymers. As used herein, these terms encompass amino acid chains of any length, including full-length proteins (i.e., antigens), in which the amino acid residues are linked by covalent peptide bonds.

[0038] The term "amino acid" refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, such as hydroxyproline, γ-carboxyglutamate, and O-phosphoserine. Amino acid analogs refer to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., compounds that have an α-carbon bonded to a hydrogen, a carboxyl group, an amino group, and an R group, such as homoserine, norleucine, methionine sulfoxide, and methionine methylsulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. An "amino acid mimetic" refers to a compound that has a structure that is different from the general chemical structure of an amino acid, but that functions in a manner similar to a naturally occurring amino acid.

[0039] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission (CBN). Similarly, nucleotides may be referred to by their commonly accepted one-letter symbols.

[0040] An "expression cassette" is a recombinantly or synthetically produced nucleic acid construct that contains a series of designated nucleic acid elements that enable transcription of a particular polynucleotide sequence in a host cell. An expression cassette may be part of a plasmid, a viral genome, or a nucleic acid fragment. Generally, an expression cassette contains a polynucleotide to be transcribed operably linked to a promoter. "Operatively linked" in this context means that two or more genetic elements, such as a polynucleotide coding sequence and a promoter, are positioned relative to one another in a manner that allows for the proper biological function of elements such as the promoter to direct transcription of the coding sequence. Other elements that may be present in an expression cassette include those that enhance transcription (e.g., enhancers) and those that terminate transcription (e.g., terminators), as well as elements that confer specific binding affinity or antigenicity to the recombinant protein produced from the expression cassette.

[0041] A "vector" is a circular nucleic acid construct recombinantly produced from a bacterial (e.g., a plasmid) or viral (e.g., a viral genome)-based construct. Generally, a vector contains an origin of autonomous replication in addition to one or more genetic components of interest (e.g., a polynucleotide sequence encoding one or more proteins). In some cases, a vector may contain an expression cassette, making it an expression vector. In other cases, a vector does not contain an apparatus for expressing a coding sequence but rather may act as a carrier or shuttle for the storage and / or transfer of one or more genetic components of interest (e.g., a coding sequence) from one genetic construct to another. Optionally, a vector may further contain a sequence encoding one or more selectable or identifiable markers, which may encode proteins such as antibiotic resistance proteins (e.g., for detecting bacterial host cells) or fluorescent proteins (e.g., for detecting eukaryotic host cells), to allow for easy detection of transformed or transfected host cells that harbor the vector and allow protein expression from the vector.

[0042] The term "heterologous," when used in the context of describing the relationship between two elements, such as two polynucleotide sequences or two polypeptide sequences, in a recombinant construct describes the two elements as being derived from two different sources and positioned relative to one another in a manner not currently found in nature. For example, a "heterologous" promoter that drives expression of a protein-coding sequence is a promoter not naturally found to drive expression of that coding sequence. As another example, in the case of a peptide fused to a "heterologous" peptide to form a recombinant polypeptide, the two peptide sequences may be derived from two different parent proteins, or may be two separate portions derived from the same protein but that are not immediately adjacent to one another. In other words, the positioning of two elements "heterologous" to one another does not result in a longer polynucleotide or polypeptide sequence than may be found in nature.

[0043] As used herein, the term "barcode" refers to a short stretch of polynucleotide sequence (generally up to 30 nucleotides in length, e.g., between about 4 or 5 nucleotides in length and about 6, 7, 8, 9, 10, 12, 20, or 25 nucleotides) uniquely assigned to another, predetermined polynucleotide sequence (e.g., a segment of the coding sequence of a protein of interest, such as SpCas9) to enable detection / identification of the predetermined polynucleotide sequence or its encoded amino acid sequence based on the presence of the barcode.

[0044] "Type IIS restriction enzymes" are endonucleases that recognize asymmetric DNA sequences and cleave outside (3' or 5') their recognition sequence. These enzymes behave differently from type IIP restriction enzymes, which recognize symmetric or palindromic DNA sequences and cleave within their recognition sequence. Because type IIS restriction enzymes cleave DNA strands outside their recognition sequence, they can generate overhangs of any sequence, regardless of the recognition sequence. Therefore, two different type IIS restriction enzymes can be used to generate overhangs of the same size and orientation (i.e., both overhangs are 3' or 5' overhangs and have the same number of nucleotides), as well as matched overhangs or compatible ends (i.e., the overhangs on the two opposing strands are completely complementary), which allows annealing and ligation between the two ends generated by two different type IIS restriction enzymes.

[0045] As used herein, the term "short guide RNA" or "sgRNA" refers to an RNA molecule approximately 15-50 (e.g., 20, 25, or 30) nucleotides in length that specifically binds to a DNA molecule at a predetermined target site and guides a CRISPR nuclease to cleave the DNA molecule adjacent to the target site.

[0046] A nucleotide sequence "specifically binds" to another when two polynucleotide sequences, particularly two single-stranded DNA or RNA sequences, complex with each other to form a double-stranded structure based on substantial or complete (e.g., at least about 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or up to 100%) Watson-Crick complementarity between the two sequences.

[0047] "Physiologically acceptable excipient / carrier" and "pharmaceutically acceptable excipient / carrier" refer to substances that aid in administration of an active agent to, and often aid in absorption by, a delivery target (cell, tissue, or living organism), and can be included in the compositions of the present invention without significantly affecting the recipient. Non-limiting examples of physiologically / pharmaceutically acceptable excipients include water, NaCl, normal saline, emulsified Ringer's, normal sucrose, normal glucose, binders, fillers, disintegrants, lubricants, coatings, sweeteners, flavorings, and coloring agents. As used herein, the term "physiologically / pharmaceutically acceptable excipient / carrier" is intended to include any and all solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like, that are compatible with the intended use.

[0048] When used in connection with a predetermined value, the term "about" means a range encompassing ±10% of that value.

[0049] Detailed Description I. General The present invention relates to a new and improved high-order genetic engineering and screening platform for the highly efficient production and identification of recombinant proteins with desired biological functionality. The present invention also provides recombinant proteins produced by this platform.

[0050] A. recombinant technology Basic texts disclosing general methods and techniques in the field of recombinant genetics include Sambrook and Russell, Molecular Cloning, A Laboratory Manual (3rd ed. 2001); Kriegler, Gene Transfer and Expression: A Laboratory Manual (1990); and Ausubel et al., eds., Current Protocols in Molecular Biology (1994).

[0051] For nucleic acids, sizes are given in either kilobases (kb) or base pairs (bp). These are estimates derived from agarose or acrylamide gel electrophoresis, sequenced nucleic acids, or published DNA sequences. For proteins, sizes are given in kilodaltons (kDa) or amino acid residue numbers. Protein sizes are estimated from gel electrophoresis, sequenced proteins, the amino acid sequences from which they are derived, or published protein sequences.

[0052] Non-commercially available oligonucleotides can be chemically synthesized, for example, using an automated synthesizer as described by Van Devanter et al., Nucleic Acids Res. 12: 6159-6168 (1984), following the solid-phase phosphoramidite triester method first described by Beaucage & Caruthers, Tetrahedron Lett. 22: 1859-1862 (1981). Purification of oligonucleotides can be performed using any art-recognized strategy, for example, native acrylamide gel electrophoresis or anion-exchange HPLC as described by Pearson & Reanier, J. Chrom. 255: 137-149 (1983).

[0053] Polynucleotide sequences encoding a polypeptide of interest, e.g., an SpCas9 protein or fragment thereof, and synthetic oligonucleotides can be verified after cloning or subcloning using, for example, the chain termination method for sequencing double-stranded templates described in Wallace et al., Gene 16: 21-26 (1981).

[0054] B. Modification of Polynucleotide Coding Sequences Given the known amino acid sequence of a preselected protein of interest (e.g., SpCas9), modifications can be made to achieve desirable characteristics or improved biological functionality of the protein, as can be determined by in vitro or in vivo methods known in the art and described herein. Possible modifications to the amino acid sequence can include substitutions (conservative or non-conservative); deletions or additions of one or more amino acid residues at one or more positions in the amino acid sequence.

[0055] A variety of mutagenesis protocols have been established and described in the art and can be readily used to modify polynucleotide sequences encoding proteins of interest. See, for example, Zhang et al., Proc. Natl. Acad. Sci. USA, 94: 4504-4509 (1997); and Stemmer, Nature, 370: 389-391 (1994). These methods can be used separately or in combination to generate mutant sets of nucleic acids as well as mutants of the encoded proteins.

[0056] Examples of mutagenesis methods for generating diversity include site-directed mutagenesis (Botstein and Shortle, Science, 229: 1193-1201 (1985)), mutagenesis using uracil-containing templates (Kunkel, Proc. Natl. Acad. Sci. USA, 82: 488-492 (1985)), oligonucleotide-directed mutagenesis (Zoller and Smith, Nucl. Acids Res., 10: 6487-6500 (1982)), phosphorothioate-modified DNA mutagenesis (Taylor et al., Nucl. Acids Res., 13: 8749-8764 and 8765-8787 (1985)), and gapped double-stranded DNA mutagenesis (Kramer et al., Nucl. Acids Res., 12: 9441-9456 (1985)). (1984)).

[0057] Other possible mutagenesis methods include point mismatch repair (Kramer et al., Cell, 38: 879-887 (1984)), mutagenesis using repair-deficient host strains (Carter et al., Nucl. Acids Res., 13: 4431-4443 (1985)), deletion mutagenesis (Eghtedarzadeh and Henikoff, Nucl. Acids Res., 14: 5115 (1986)), restriction-selection and restriction-purification (Wells et al., Phil. Trans. R. Soc. Lond. A, 317: 415-423 (1986)), mutagenesis by total gene synthesis (Nambiar et al., Science, 223: 1299-1301 (1984)), and double-strand break repair (Mandecki, Proc. Natl. Acad. Sci. USA, 83: 7177-7181 (1986)), polynucleotide chain termination mutagenesis (U.S. Patent No. 5,965,408), and error-prone PCR (Leung et al., Biotechniques, 1: 11-15 (1989)).

[0058] C. Modification of nucleic acids for preferred codon usage Polynucleotide sequences encoding proteins of interest or fragments thereof can be further modified based on the principle of codon degeneracy to conform to preferred codon usage to enhance recombinant expression in specific host cell types or to facilitate further genetic manipulations, such as allowing the construction of restriction endonuclease recognition sequences at desired sites for potential cleavage / religation. The latter use is particularly important in the present invention, since seamless joining of multiple coding segments of a target protein (e.g., an SpCas9 protein) undergoing combinatorial mutagenesis relies on digesting the coding segments with type IIS restriction enzymes to generate overhangs specifically derived from the coding sequence of the native protein, eliminating any extraneous sequence or so-called scar sequence at the junction between any two of these segments.

[0059] Upon completion of modification, the coding sequence is verified by sequencing and then subcloned into an appropriate vector for further manipulation or for recombinant expression of the protein.

[0060] D. Expression of recombinant polypeptides A recombinant polypeptide of interest (e.g., an improved Cas9 protein) can be expressed using techniques conventional in the field of recombinant genetics, depending on the polynucleotide sequence encoding the polypeptide described herein.

[0061] (I) Expression system To obtain high-level expression of a nucleic acid encoding a polypeptide of interest, a polynucleotide coding sequence is generally subcloned into an expression vector containing a strong promoter to direct transcription, a transcription / translation terminator, and a ribosome binding site for translation initiation. Suitable bacterial promoters are well known in the art and are described, for example, in Sambrook and Russell (supra) and Ausubel et al. (supra). Bacterial expression systems for expressing recombinant polypeptides are available, for example, for Escherichia coli, Bacillus sp., Salmonella, and Caulobacter. Kits for such expression systems are commercially available. Eukaryotic expression systems for mammalian cells, yeast, and insect cells are well known in the art and are commercially available. Some exemplary eukaryotic expression vectors include adenoviral vectors, adeno-associated vectors, and retroviral vectors, such as lentivirus-derived viral vectors.

[0062] The promoter used to direct the expression of a heterologous polynucleotide sequence encoding a protein of interest will vary depending on the particular application. The promoter is optionally positioned approximately the same distance from the heterologous transcription start site as it is from the transcription start site in its natural environment. However, as is known in the art, some variation in this distance is possible without impairing promoter function.

[0063] In addition to a promoter, an expression vector generally contains a transcription unit or expression cassette that contains all additional elements required for the expression of a desired polypeptide in a host cell. Thus, a typical expression cassette contains a promoter operably linked to a nucleic acid sequence encoding a polypeptide, as well as signals required for efficient polyadenylation of the transcript, a ribosome binding site, and translation termination. For recombinant expression of secreted proteins, the polynucleotide sequence encoding the protein is generally linked to a cleavable signal peptide sequence to facilitate secretion of the recombinant polypeptide by transformed cells. On the other hand, if the recombinant polypeptide is intended to be expressed on the host cell surface, an appropriate anchor sequence is used in conjunction with the coding sequence. Additional elements of the cassette may include enhancers and, when genomic DNA is used as the structural gene, introns with functional splicing donor and acceptor sites.

[0064] In addition to a promoter sequence, the expression cassette should also contain a transcription termination region downstream of the coding sequence to provide for efficient termination. The termination region may be obtained from the same gene as the promoter sequence or from a different gene.

[0065] Expression vectors containing regulatory elements derived from eukaryotic viruses are commonly used in eukaryotic expression vectors, such as SV40 vectors, papillomavirus vectors, lentivirus vectors, and vectors derived from Epstein-Barr virus. Other exemplary eukaryotic expression vectors include pMSG, pAV009 / A, and pAV009 / A. + , pMTO10 / A + , pMAMneo-5, baculovirus pDSVE and any other vector that allows expression of proteins under the direction of the SV40 early promoter, SV40 late promoter, metallothionein promoter, mouse mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter or other promoters shown to be effective for expression in eukaryotic cells.

[0066] Elements commonly included in expression vectors may also include a replicon that functions in E. coli, a gene encoding antibiotic resistance to allow selection of bacteria hosting the recombinant plasmid, and a unique restriction site in a non-essential region of the plasmid to allow insertion of eukaryotic sequences. The particular antibiotic resistance gene selected is not critical; any of the many resistance genes known in the art are suitable. If necessary, prokaryotic sequences can be selected so that they do not interfere with DNA replication in eukaryotic cells. Similar to antibiotic resistance selection markers, metabolic selection markers based on known metabolic pathways can also be used as a means to select transformed host cells.

[0067] As noted above, those skilled in the art will recognize that various conservative substitutions can be made to a protein or its coding sequence while retaining the biological activity of the protein. Additionally, modifications of the coding sequence of a polynucleotide can also be made to accommodate preferred codon usage in a particular expression host, or to create restriction enzyme cleavage sites without altering the resulting amino acid sequence.

[0068] (II) Transfection method Standard transfection methods are used to generate bacterial, mammalian, yeast, insect, or plant cell lines that express large amounts of recombinant polypeptides, which are purified using standard techniques (see, e.g., Colley et al., J. Biol. Chem. 264: 17619-17622 (1989); Guide to Protein Purification, in Methods in Enzymology, vol. 182 (Deutscher, ed., 1990)). Transformation of eukaryotic and prokaryotic cells is carried out according to standard techniques (see, e.g., Morrison, J. Bact. 132: 349-351 (1977); Clark-Curtiss & Curtiss, Methods in Enzymology 101: 347-362 (Wu et al., eds, 1983)).

[0069] Any well-known method for introducing foreign nucleotide sequences into host cells can be used. These include the use of calcium phosphate transfection, polybrene, protoplast fusion, electroporation, liposomes, microinjection, plasma vectors, viral vectors, and any other well-known method for introducing cloned genomic DNA, cDNA, synthetic DNA, or other foreign genetic material into host cells (see, e.g., Sambrook and Russell, supra). It is only necessary that the particular genetic engineering method used be capable of successfully introducing at least one gene into the host cell capable of expressing a recombinant polypeptide.

[0070] II. Improved Combinatorial Gene Modification System Based on previously developed high-throughput CombiGEM combinatorial gene engineering systems and others, the inventors further improved these systems by seamlessly joining DNA elements encoding multiple protein segments, each corresponding to a portion of a protein of interest (e.g., SpCas9) and containing at least one, and potentially multiple, mutations in its amino acid sequence, so that the resulting composite protein variants contain no extraneous amino acid residues other than the intentionally introduced mutations. Previous methods utilized type IIP restriction endonucleases to cleave and religate DNA sequences (encoding segments of combinatorial protein variants). However, the nature of this type of endonucleases (which bind to and cleave short palindromic sequences of nucleotide sequences) generally requires the user to design the cleavage sites by introducing extra nucleotides, resulting in the creation of extraneous amino acid residue(s) or "scar" sequences at each junction between two segments of the protein variants generated by the system. These extraneous amino acid residues could further alter the protein sequence and potentially hinder functional screening of the variants.

[0071] To avoid the introduction of these unwanted extra amino acid residues, the present inventors discovered that such undesired "scar" sequences between segments can be completely eliminated when constructing libraries of combinatorial gene variants by instead using type IIS restriction enzymes to assemble and ligate multiple DNA coding sequences encoding protein segments. This method takes advantage of the fact that type IIS endonucleases can cleave DNA strands outside their asymmetric recognition sites, allowing compatible ends or matched overhangs to be generated that contain portions of the native DNA coding sequence of the wild-type protein after DNA cleavage with these enzymes. The use of coding sequences from native proteins in the compatible ends or matched overhangs not only facilitates seamless joining between protein segments but also allows for directional ligation, further improving the efficiency of the process of constructing combinatorial protein variants.

[0072] A. Generation of libraries of DNA segments encoding protein segments The first step in generating a library of combinatorial protein variants is to create a library for each segment of the protein: protein variants can be designed to be generated by joining a predetermined number of protein segments or modules (e.g., 3, 4, 5, 6, or more) end-to-end. As described herein, the predetermined number is represented as n+1, and for the protein of interest, n=5 is designed to consist of 6 segments. First, a library or collection of individual members of DNA elements encoding a first protein segment corresponding to the extreme N-terminal portion of the wild-type protein and containing one or more possible mutations in this portion of the protein may be generated by known methods, such as recombinant production or chemical synthesis, and incorporated into a DNA vector (a so-called storage vector for this purpose) containing appropriate restriction enzyme sites and a barcode sequence uniquely assigned to the DNA element bearing the predetermined mutation (or set of predetermined mutations). When the DNA elements are relatively long, they may first be generated by joining short fragments using known methods, such as Gibson assembly, before being incorporated into the storage vector. As noted above, methods for generating DNA sequence mutations are well known to those of skill in the art and can readily be used to generate sequence variants by modifying the native version or wild-type sequence, for example, by deletion, insertion and / or substitution of one or more nucleotides.

[0073] Figure 7a shows an example of a method for inserting and ligating DNA elements encoding protein segments into a vector to form a DNA construct containing, from 5' to 3', a first recognition site for a first type IIS restriction enzyme (e.g., BsaI), the DNA element, first and second recognition sites for a second type IIS restriction enzyme (e.g., BbsI), a barcode uniquely assigned to the DNA element for the specific mutation(s) it possesses, and a second recognition site for the first type IIS restriction enzyme (e.g., BsaI). For proteins engineered or "deconstructed" to have (n+1) segments or modules for combinatorial mutation testing, libraries of storage vectors containing DNA segments can be constructed in a similar manner for each subsequent DNA element, the second, third, and nth DNA elements (encoding the second, third, and nth protein segments, respectively), and the nth protein segment corresponding to the second to last or most C-terminal portion of the protein.

[0074] For DNA elements encoding the final or most C-terminal segment of a protein, a structurally distinct storage vector is used to construct a library of vectors containing the (n+1)th DNA element. As illustrated in Figure 7a, the final or (n+1)th DNA element is inserted into this storage vector to form a DNA construct containing, in 5' to 3' order, a first recognition site for a first type IIS restriction enzyme (e.g., BsaI), the (n+1)th DNA element, a short nucleotide sequence extension that serves as a primer binding site, a barcode uniquely assigned to the DNA element due to the specific mutation(s) it contains, and a second recognition site for the first type IIS restriction enzyme (e.g., BsaI). The presence and position of the primer binding site allows for rapid sequencing of the combined barcodes using a universal primer (that specifically binds to the primer binding site) after a composite coding sequence for a protein variant (combining all n+1 DNA elements) is generated, allowing for easy identification of the mutations contained in the variants and eliminating the need for the time-consuming task of sequencing the entire composite coding sequence.

[0075] To ensure equal opportunity for each possible combinatorial protein variant in the library, each DNA element with a unique set of mutations is preferably present in the library in equimolar ratios.

[0076] B. Generation of combinatorial protein mutant libraries Once a library of storage vectors containing the first, second, and nth DNA elements, as well as the (n+1)th DNA element, has been constructed, DNA fragments containing DNA elements encoding protein segments or modules are first released by enzymatically digesting the storage vector, for example, by cleaving the vector at two sites using a first type IIS restriction endonuclease (e.g., BsaI). Digestion of the storage vector releases DNA fragments containing the DNA element encoding the protein segment (with mutations) and its uniquely assigned barcode, flanked by two type IIS restriction enzyme (e.g., BbsI) recognition sites. The two ends of the DNA fragments have overhangs resulting from cleavage with the first type IIS restriction enzyme.

[0077] On the other hand, a DNA vector intended to carry and express the final composite DNA element encoding the entire protein variant (for this purpose, a so-called destination vector) is an expression vector that contains all the genetic elements necessary for the expression of a DNA coding sequence. As described in the previous section, one essential element for transcription is a promoter that must be operably linked to the coding sequence to direct the transcription of the sequence. Typically, the promoter is heterologous to the coding sequence.

[0078] To obtain DNA fragments produced from a storage vector library, the destination vector is also linearized by digestion with a type IIS restriction enzyme at a site downstream from the promoter at an appropriate distance to allow insertion / ligation of the DNA fragment and place the DNA element (encoding the protein segment) within the DNA fragment under the control of the promoter for transcription. The type IIS restriction enzyme used to linearize the destination vector is often different from that used to release the DNA fragment from the storage vector. However, they preferably produce overhangs of the same size and matching to allow ligation of the DNA fragment into the destination vector.

[0079] As shown in Figure 7b, a library of storage vectors containing the full variety of first DNA elements encoding the full variety of first protein segments is digested with a first type IIS restriction enzyme, releasing a library of DNA fragments containing the full variety of first DNA elements along with their corresponding barcodes from the storage vector. This library of first DNA fragments, preferably in an equimolar ratio for each sequence species, is then ligated into a linearized destination vector, resulting in a 1-wise library. Each member of the resulting 1-wise library may contain a functional expression cassette in which a promoter is operably linked to the first DNA element and capable of directing expression of the first or most N-terminal protein segment encoded by the first DNA element.

[0080] The uniform library is then digested again with a type IIS restriction enzyme, cutting each member of the library twice between the first DNA element and its barcode, generating two overhangs at each cut site.

[0081] Meanwhile, digestion of a library of storage vectors containing the full species of a second DNA element encoding the full species of a second protein segment with a first type IIS restriction enzyme releases a library of DNA fragments containing the full species of the second DNA element along with the corresponding barcode from the storage vector. This library of second DNA fragments, preferably in an equimolar ratio for each sequence species, is then ligated into a linearized 1-wise expression vector between the first DNA element and its corresponding barcode, resulting in a new library of 2-wise expression vectors. Each member of the resulting 2-wise library may contain a functional expression cassette in which a promoter is operably linked to the first DNA element fused to the second DNA element and capable of directing expression of the fused first and second protein segments encoded by the fusion between the first DNA element and the second DNA element. To remove any extraneous amino acid residues or "scar" sequences at the fusion point between the first and second protein segments, two cleavage sites located between the first DNA element and its barcode must be carefully designed to ensure that (1) there is a perfect match (both in sequence and overhang size / orientation) between the two terminal overhangs of the linearized one-way vector and the two terminal overhangs of the second DNA fragment released from the library of storage vectors containing the complete species of the second DNA element, and (2) upon their ligation, the matched overhang sequence between the tail or 3' end of the first DNA element and the head or 5' end of the second DNA element encodes the amino acid sequence extension found in the wild-type protein of interest at the same position. In other words, the design of the cleavage sites ensures seamless joining of the two adjacent protein segments.

[0082] Upon completion of ligation of the second library of DNA fragments released from the second library of storage vectors into the linearized library of unidirectional expression vectors, a library of bidirectional composite expression vectors is constructed. The cycle of steps outlined in the last two paragraphs can be repeated to incorporate third DNA fragments into the composite expression vector up to the nth and (n+1)th DNA fragments to obtain a final library of composite expression vectors, which contains the complete sequence of DNA coding sequences encoding full-length protein variants containing all possible combinations of mutations, with each variant coding sequence followed by a composite barcode sequence, which has all corresponding barcodes uniquely assigned to the DNA elements and can indicate how the DNA elements are fused in reverse order.

[0083] C. Functional screening of protein variants Because the final library of destination vectors is an expression vector, each of which has a promoter operably linked to a composite DNA coding sequence containing all n+1 DNA elements for encoding full-length protein variants containing a particular set of mutations, these protein variants can be easily expressed, screened, and selected for any particular desired functional characteristic in an appropriate reporting system. For example, viral-based destination vectors can be used to transfect host cells and directly express the protein variants of interest in a cellular environment suitable for functional analysis.

[0084] Figure 2a shows an example of how SpCas9 variants can be screened for their functionality: a cell line stably expressing red fluorescent protein (RFP) and a gRNA targeting the RFP gene sequence is transfected with a lentiviral vector containing sequences encoding SpCas9 variants to demonstrate the on-target activity of each variant, and another cell line stably expressing RFP and a gRNA with a synonymous mutation is transfected to demonstrate the off-target activity of the variant. Because the CombiSEAL platform is designed to generate useful variants of any protein, various functional screening assays can be envisioned depending on the specific functionality of the protein of interest. Once a clone with the desired functional properties (such as on-target and off-target activity profiles, as in the case of Cas9 proteins) is discovered, sequencing of the composite barcode allows immediate identification of the specific mutations in a particular variant.

[0085] III. Optimized CAS9 Enzyme We utilized a newly improved CombiSEAL combinatorial gene recombination system to identify a series of SpCas9 mutants and characterize their functional properties. Among the mutants tested, one particular mutant, designated Opti-SpCas9, was found to possess a highly desirable functional profile: it possesses enhanced gene editing specificity without compromising efficacy and a broad test range. Given its functional properties, this improved Cas9 enzyme is an extremely valuable tool in CRISPR genome editing schemes.

[0086] The wild-type SpCas9 protein has the amino acid sequence set forth in SEQ ID NO:1, and its corresponding DNA coding sequence is set forth in SEQ ID NO:2. Previous studies of this endonuclease have provided insight into the protein's structure, including the regions and amino acid residues that interact with DNA. During testing to develop the CombiSEAL platform, the inventors confirmed that mutations, particularly substitutions, introduced into specific residues in the SpCas9 amino acid sequence previously predicted to interact with target and non-target DNA strands have a direct impact on the endonuclease's performance. Specifically, substitutions at residues such as R661, Q695, K848, Q926, K1003, and K1060 were found to alter the enzyme's on-target and off-target editing activities. Mutant Opti-SpCas9 is a double mutant of wild-type SpCas9: residue 661 of SEQ ID NO:1 is substituted with alanine and residue 1003 is substituted with histidine. Its amino acid sequence is set forth in SEQ ID NO:3. These substitutions are responsible for the increased on-target editing efficiency of the modified endonucleases and reduced off-target activity, a highly desirable phenotype.

[0087] The inventors also identified a triple mutant of R661A, K1003H, and Q926A, which further reduced off-target editing from Opti-SpCas9 by approximately 80% while also substantially reducing its on-target activity. This triple mutant may be valuable in situations where avoiding off-target cleavage is particularly important. Additionally, a second mutant, designated OptiHF-SpCas9, has been generated, which contains five point mutations: Q695A, K848A, E923M, T924V, and Q926A (see mutant 46 in Table 2). The amino acid sequences of Opti-SpCas9 and OptiHF-SpCas9 are set forth in SEQ ID NO: 3 and SEQ ID NO: 13, respectively. Table 2 provides a summary of the SpCas9 mutants analyzed in this study, detailing the point mutation(s) they contain and their on-target and off-target cleavage profiles.

[0088] The SpCas9 mutants described herein are useful tools for genetic manipulation of the genome of living cells. To use these mutants for targeted DNA cleavage using the CRISPR system, an expression vector directing the expression of the mutant (e.g., Opti-SpCas9) and an expression vector encoding an sgRNA with an appropriate sequence for directing the SpCas9 mutant to a preselected target site in the genome of the cell to cleave genomic DNA at the target site are generally introduced into living cells. In some embodiments, the expression vector is a viral vector such as a retroviral vector, particularly a lentiviral vector. The expression vector encoding the SpCas9 mutant and the expression vector encoding the sgRNA are often two separate vectors, but in some embodiments, a single expression vector contains the coding sequences of both the SpCas9 mutant and the sgRNA, and the two coding sequences are operably linked to either the same promoter or two separate promoters. Because promoters are generally heterologous to the coding sequence, it may be further considered to use a promoter suitable for a specific type of recipient cell. [Example]

[0089] Example The following examples are offered by way of illustration only, and not by way of limitation. Those of ordinary skill in the art will readily recognize a variety of noncritical parameters that could be changed or modified to yield essentially the same or similar results.

[0090] Example 1: CombiSEAL as a high-throughput platform for seamlessly assembling barcoded combinatorial genetic units and providing novel methods for protein optimization, such as screening of SpCas9 mutants Because the combined effects of multiple mutations on protein function are difficult to predict, the ability to functionally evaluate vast numbers of protein sequence variants would be practically useful in protein engineering. This invention provides a high-throughput platform that enables the scalable assembly and parallel characterization of barcoded protein variants with combinatorial modifications. This platform, CombiSEAL, is illustrated by systematically analyzing a library of 948 combinatorial variants of the widely used Streptococcus pyogenes Cas9 (SpCas9) nuclease to optimize their genome editing activity in human cells. The ease of simultaneously evaluating the editing activity of SpCas9 variants at multiple on- and off-target sites accelerates the identification of optimized variants and facilitates the study of mutational epistasis. Opti-SpCas9 was successfully identified, which possesses enhanced editing specificity and a broad target range without compromising potency. This platform is broadly applicable to protein engineering through combinatorial bulk modification.

[0091] explanation Protein engineering has proven to be an important method for generating enzymes, antibodies, and genome-edited proteins with new or enhanced properties. 1-7 Combinatorial optimization of protein sequences relies on strategies for generating and screening large numbers of variants, but current methods are limited in their ability to systematically and efficiently construct and test multiple variants in a high-throughput manner. 8-11 While traditional site-directed mutagenesis methods based on structural and biochemical knowledge facilitate the generation of functionally relevant variants, screening combinatorial mutants using such a one-to-one approach lacks throughput and scalability. Gene synthesis techniques can be employed to generate combinatorial mutants in a pooled format, which typically yields 1–10 errors per kilobase synthesized. 12,13However, this approach is prohibitively expensive when the mutations to be introduced are scattered across different regions of the protein. 14,15 and recombination and shuffling 16 While methods such as [1] create combinatorial mutants by fusing multiple mutant sequences together to assemble the entire protein sequence, subsequent genotyping and mutation characterization require the selection of clonal isolates or long-read sequencing, neither of which is feasible for tracking large numbers of mutants. Mutagenesis using error-prone polymerase chain reaction (PCR) and directed evolution (DEV) mutant strains can actively select desired mutants, but suffer from selection bias toward a subset of amino acids due to the rare occurrence of two or more specific nucleotide mutations within a codon. Even if sequence randomization can achieve diversity in protein mutants, the very limited throughput of individually genotyping and analyzing selected hits presents a major obstacle to protein engineering. Furthermore, pinpointing the exact mutations that confer the desired phenotype from the remaining non-target mutations can help accelerate the combinatorial optimization process.

[0092] Here, we present a method, Combinatorial Genetics En Masse (CombiGEM), for the pooled assembly of barcoded combinatorial variants that can be easily tracked by high-throughput short-read sequencing (Figure 1). 17-19We have devised a novel cloning method to combine seamless combinatorial DNA assembly with the barcode ligation strategy used in our platform, called CombiSEAL. CombiSEAL works by modularizing protein sequences into configurable parts, each containing a repertoire of variants tagged with barcodes that specify predetermined mutations at defined positions. Type IIS restriction enzyme sites are used to flank the barcoded parts, creating cleaved overhangs derived from the protein-coding sequence, thereby achieving seamless ligation upon fusion with the part. Unique barcodes are then ligated to each protein-coding sequence variant in the resulting library after repeated pooled cloning of multiple parts. This method is advantageous over other strategies because it avoids the need for long-read sequencing across the entire protein-coding region covering multiple mutations. This provides a cost-effective way to quantitatively track each variant in a pool by high-throughput sequencing of short (e.g., ~50 base pairs) barcodes without the need for clonal isolate selection. Furthermore, characterization of pooled variants allows direct comparison under the same experimental conditions, facilitating the study of mutational epistasis. Unlike CombiGEM, which allows the combinatorial assembly of individual genetic elements, CombiSEAL leaves no fusion scar sequences to seamlessly join consecutive sequences (e.g., different segments of a protein), thus making this new platform of great potential for protein engineering.

[0093] result High-throughput screening of SpCas9 combinatorial mutants. CombiSEAL was used to assemble a combinatorial mutant library of SpCas9, a CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) nuclease widely used in genome engineering, with the aim of identifying optimized mutants with high editing specificity and activity.20-23 So far, eSpCas9(1.1) 3 , SpCas9-HF1 4 , HypaCas9 5 and evoCas9 6 SpCas9 nucleases with specific mutation combinations, including , have been engineered to minimize off-target editing. However, these mutants have fewer targetable sites due to incompatibility with gRNAs that start with a mismatched 5'-guanine (5'G). 3-6,24-27 The number of combinatorial mutants that have been generated and tested to date is limited (Table 1), and therefore, a more systematic search for other SpCas9 mutants that have good compatibility with gRNAs with extra 5' Gs is necessary.

[0094] Using CombiSEAL, we modularized the SpCas9 sequence into four parts, and barcoded inserts containing different random and specific mutations in each part were cloned into a storage vector (Figure 1a; Figure 7a, b; see Methods for details). The combinatorial barcoded library (containing 4 x 2 x 17 x 7 = 952 SpCas9 mutants, wild-type (WT) SpCas9, and eSpCas9(1.1) sequences) was pooled and integrated into a lentiviral vector. Individual parts in the library and the assembled construct were sequenced to confirm the high-precision assembly of the barcoded mutants (see Methods for details). We detected high coverage of the library in both the E. coli-stored plasmid pool (i.e., 951 of 952 variants) and the infected human cell pool (i.e., 948 of 952 variants) (Fig. 1b), and highly reproducible representation between the plasmid and infected cell pools, and between biological replicates of the infected cell pools (Fig. 7c).

[0095] To explore robust and specific SpCas9 variants, we established a reporter system using monoclonal human cell lines stably expressing red fluorescent protein (RFP) and gRNAs targeting the RFP gene sequence (hereafter referred to as RFPsg5-ON and RFPsg8-ON; Figure 2a). 3-6 Unlike previous screens that primarily used 20-nucleotide gRNAs starting from the nucleotide 5', we used gRNAs carrying an additional 5' G in our reporter system to search for compatible SpCas9 mutants that do not sacrifice targeting range. Cells were then infected with the SpCas9 mutant library and sorted into bins based on RFP fluorescence levels at 14 days post-infection. Loss of RFP fluorescence reflects DNA cleavage and indel-mediated disruption of the target site; therefore, cells with active SpCas9 mutants can be enriched in sorted bins with low RFP levels. Using Illumina HiSeq to track barcoded SpCas9 mutants, we found that the mutant subpopulation was more than 1.5-fold enriched in the sorted bin encompassing approximately 5% of the cell population with the lowest RFP levels (i.e., bin A) compared to the unsorted population (Figure 2b; Figure 8). WT SpCas9 was one of those enriched for both the RFPsg5-ON and RFPsg8-ON reporter systems, while eSpCas9(1.1) was enriched for RFPsg8-ON. To facilitate parallel characterization of the on-target and off-target activities of SpCas9 mutants, we further generated cell lines with synonymous mutations in RFPs such that targeting of mismatched sites would reveal the off-target activity of SpCas9 mutants (i.e., RFPsg5-OFF5-2 and RFPsg8-OFF5; Figure 2a). WT SpCas9, but not eSpCas9(1.1), was enriched for both RFPsg5-OFF5-2 and RFPsg8-OFF5 (Figure 2b; Figure 8).

[0096] The on-target and off-target activities of the library of SpCas9 mutants were ranked and plotted based on the enrichment of the sorted bins relative to the unsorted population, revealing that the majority of mutants reduced both the on-target and off-target activities of SpCas9 (Figure 3a). Activity-optimized mutants were defined as those with enrichment ratios of at least 90% of WT for both RFPsg5-ON and RFPsg8-ON, and less than 60% of WT for both RFPsg5-OFF5-2 and RFPsg8-OFF5. The nOne mutant (hereafter referred to as Opti-SpCas9) met these criteria and was evaluated for further characterization (Table 2). We also identified a high-fidelity mutant, designated OptiHF-SpCas9, based on enrichment ratios of at least 50% of WT for both RFPsg5-ON and RFPsg8-ON, and less than 90% of WT for both RFPsg5-OFF5-2 and RFPsg8-OFF5 (Table 2). The efficiency and specificity of Opti-SpCas9 and OptiHF-SpCas9 were verified by individual validation assays to measure their on-target and off-target activities. Using multiple cell lines expressing gRNAs targeting either matched or mismatched RFP sites, we confirmed that Oppi-SpCas9 exhibited comparable on-target activity (i.e., 94.6%; average from three mismatched sites) and substantially reduced off-target activity (i.e., 1.7%; average from three mismatched sites) compared to WT, whereas OptiHF-SpCas9 exhibited reduced activity both on-target (i.e., 63.6%; average from two matched sites) and off-target (i.e., 2.0%; average from two mismatched sites) (Figure 3b).

[0097] Examining mutational epistasis for SpCas9 editing efficiency. Systematic construction of protein mutants using CombiSEAL allows us to classify sets of amino acid substitutions as neutral, beneficial, or deleterious, allowing us to explore their difficult-to-predict epistatic interactions. Using enrichment ratios as an indicator of SpCas9 editing activity (Figure 9), we constructed heat maps showing the on-target and off-target activity resulting from mutation combinations and epistatic interactions (Figure 4; Figure 10). The results revealed that the number and type of substitutions introduced into SpCas9 amino acid residues predicted to interact with target and non-target DNA strands (e.g., R661, Q695, K848, Q926, K1003, K1060, etc.) governs the optimal balance between maximizing on-target efficiency and minimizing off-target activity. The activity-optimized mutant Opti-SpCas9 differs from the WT by two substitution mutations at these DNA contact residues (i.e., R661A and K1003H). Comparison of the three conserved basic residues (i.e., lysine, arginine, and histidine) introduced at amino acid position 1003 of SpCas9 revealed that K1003H exhibited positive epistatic interactions with the R661A mutation and was the favorable substitution that conferred high editing efficiency to Opti-SpCas9 (Figure 4). SpCas9-HF1 4Adding the Q926A substitution to Opti-SpCas9, which has been shown to confer higher specificity for SpCas9, slightly reduced its off-target effect (i.e., from 1.0% for Opti-SpCas9 to 0.2% for Opti-SpCas9+Q926A, averaged from three mismatched target sites), but significantly reduced its on-target activity across the three matched sites tested to 21.6%, 62.4%, and 99.9% (Figure 3b). Furthermore, most SpCas9 mutants with three or more mutations in these DNA-contacting residues demonstrated reduced editing at both on-target and off-target target sites (Figure 4). These results are consistent with previous findings that excessive alanine substitutions in these DNA-contacting residues significantly reduced the editing activity of SpCas9. 25 However, interestingly, the HNH and RuvC nuclease domains of SpCas9, such as the E923M+T924V and E923H+T924L mutations located in the linker region connecting the two domains, 28 Several SpCas9 mutants with three or more mutations in DNA-contacting residues restored on-target editing at the RFPsg5-ON site through additional substitutions introduced into residues involved in conformational control (Figure 4). The high-fidelity mutant OptiHF-SpCas9, which contains the E923M+T924V mutation in addition to the Q695A, K848A, and Q926A substitutions, also exhibited slightly higher on-target activity at the RFPsg8-ON site than the mutant with only the Q695A, K848A, and Q926A triple mutation (Figure 4). These data support the model that the DNA binding and cleavage activities of SpCas9 functionally couple to determine its editing specificity and editing efficiency. 5,29 , highlighting the possibility of programming the editing performance of SpCas9 by modifying linker residues.

[0098] Characterization of optimized SpCas9 mutants. In gRNA design and construction, a 5' G is typically included or added to the beginning of the gRNA sequence to promote efficient transcription under the U6 promoter. WT SpCas9 is compatible with gRNAs that have an additional 5' G mismatched to the protospacer sequence. On the other hand, eSpCas9(1.1), SpCas9-HF1, HypaCas9, and evoCas9 are compatible with gRNAs that have an additional 5' G (i.e., GN 20 ) or has a starting guanine (i.e., HN 19 ) lose their editing efficiency when using 20-nucleotide gRNAs lacking the 4,6,24-26,30 The use of gRNAs with a 5' G matched to the protospacer sequence allows for the 20 - Compared to NGG, GN 19 We were able to significantly reduce the number of editable sites in the human genome based on the availability of -NGG sites by approximately 4.3-fold (Figure 11). The editing activity of Opti-SpCas9 was further characterized using gRNAs with 5' G additions, demonstrating that Opti-SpCas9 is more efficient than the previous assays of endogenous loci tested by the inventors. 3-5,18,31SpCas9, Opti-SpCas9, eSpCas9(1.1), and HypaCas9 showed on-target DNA cleavage activity comparable to that of WT (i.e., 95.1%), whereas eSpCas9(1.1) and HypaCas9 showed significantly reduced activity (i.e., 32.4% and 25.6%, respectively) (Figure 5a; Figure 12). The reduced editing was not due to a decrease in the protein expression levels of the two SpCas9 mutants (Figure 13). These results are consistent with the on-target activity of these mutants observed in our screening system (Figures 2; 3a) in which gRNAs with an additional 5' G were used, as well as with the results based on independent validation experiments using a green fluorescent protein (GFP) disruption assay (Figure 3b; Figure 14). Furthermore, Opti-SpCas9, eSpCas9(1.1), and HypaCas9 showed editing activity comparable to that of WT (i.e., 109.1%, 103.3%, and 106.8%, respectively) when using 20-nucleotide gRNAs starting with the matching 5' G (Figure 5a). Opti-SpCas9 was further classified as OptiHF-SpCas9 and a more recently characterized high-fidelity variant, evoCas9. 6 and Sniper-Cas9 32 Further comparison revealed that OptiHF-SpCas9, evoCas9, and Sniper-Cas9 produced fewer on-target edits than Opti-SpCas9 (i.e., 60.7%, 99.8%, and 51.7%, respectively, when expressed with gRNAs containing an additional 5' G, and 40.1%, 87.7%, and 63.9%, respectively, when expressed with gRNAs starting with a matching 5' G in the 20-nucleotide gRNA sequence) (Figure 5b; Figure 12; Figure 13). That is, the restriction of having a matching 5' G as the first base of the 20-nucleotide gRNA sequence for transcription under U6 limits the utility of other previously engineered SpCas9s with improved specificity but does not apply to Opti-SpCas9, which works interchangeably with gRNAs containing an additional 5' G. These findings emphasize that engineered SpCas9s do not necessarily have to sacrifice target coverage for specificity.

[0099] We further investigated the off-target activity of different SpCas9 mutants. We used gRNAs for VEGFA site 3 and DNMT1 site 4 to amplify eight potential off-target loci edited by WT SpCas9. 3-5,31Genomic indels induced by WT SpCas9 were detected at four of these sites (i.e., VEGFA OFF1, VEGFA OFF2, VEGFA OFF3, and DNMT1 OFF1) in OVCAR8-ADR cells. When Opti-SpCas9, eSpCas9(1.1), and HypaCas9 were used instead of WT, off-target editing was detected only at the VEGFA OFF1 site (Figure 15). Of the four mutants, Opti-SpCas9 showed the greatest on-target and off-target activity at that site (Figure 15). To compare the mismatch tolerance of different SpCas9 mutants, gRNAs containing one to four mismatches to the reporter gene target (i.e., the GFP gene sequence integrated into the genome) were generated. These mismatched bases spanned different positions in the gRNA spacer sequence. The disappearance of GFP fluorescence was measured to reflect DNA cleavage and indel-mediated disruption of the target site. Opti-SpCas9 was found to be largely intolerant of gRNAs with two or more mismatched bases, although a relatively low level of activity (i.e., 3.5% for Opti-SpCas9 and 73.2% for WT) was detected at one of eight sites with two mismatches (Figure 16). eSpCas9(1.1) and HypaCas9 were observed to produce less editing (i.e., more than 60% reduction) at both on-target and off-target sites in our reporter system (Figure 16). With comparable on-target activity between WT and Opti-SpCas9 (i.e., 97.6% of WT), Opti-SpCas9 exhibited higher specificity than WT, as indicated by significantly fewer off-target edits at 13 of 20 sites containing single-base mismatches, although significant amounts of off-target editing were still detected (Figure 16). Others include eSpCas9(1.1), SpCas9-HF1, HypaCas9, evoCas9, and Sniper-Cas9. 3,5,6,32Editing activity at single-base mismatch sites using gRNAs has also been reported. Nevertheless, the majority of off-target sites in genomes predicted in silico contain two or more mismatches to the gRNA sequence. 33 Therefore, tolerance to single-base mismatches should not limit the usefulness of SpCas9 for achieving precise genome editing. Furthermore, GUIDE-Seq was performed to examine the genome-wide cleavage activity produced by Opti-SpCas9 and other engineered SpCas9 variants. The results showed that Opti-SpCas9 produced significantly fewer off-target cleavages than WT, and OptiHF-SpCas9 exhibited an increased on- / off-target ratio comparable to other reported high-fidelity variants, such as eSpCas9(1.1), HypaCas9, evoCas9, and Sniper-Cas9 (Figure 5c, Table 3). Compared to eSpCas9(1.1) and HypaCas9, Opti-SpCas9 exhibited better compatibility with the use of cleaving gRNAs (Figure 17), which may provide a complementary strategy for improving the editing specificity of Opti-SpCas9. 34 .

[0100] Consideration To address the unmet need for rapid and simultaneous profiling of high-order combinatorial mutations for protein engineering, we established a simple yet highly powerful platform called CombiSEAL. This strategy utilizes a pooled assembly approach to avoid the laborious process of constructing individual combinatorial mutants one by one and utilizes barcoding to facilitate protein engineering, enabling parallel testing to identify top performers from a large number of protein mutants. Furthermore, this method can also be applied to mapping epistatic relationships between mutations. Using the CombiSEAL method, we successfully identified Opti-SpCas9 and OptiHF-SpCas9, novel mutants with excellent genome editing efficiency and specificity for a wide range of endogenous targets in human cells (Table 3). We further expanded the CombiSEAL pipeline to include a broader range of protospacer-adjacent motifs with greater flexibility. 7 and those with enhanced compatibility with ribonucleoprotein delivery. 35 CombiSEAL can be easily applied to construct even more Cas9 variants to broaden the search for variants with pleiotropic or other properties. CombiSEAL is a novel method for the precise editing of genomes using CRISPR enzymes (SaCas9). 36 and Cpf1 37 ) and derivatives thereof (e.g., base editors 38-41 It is anticipated that this approach will accelerate the generation of engineered proteins. Moreover, the generalizability of this approach expands the possibilities for systematically designing not only diverse proteins but also other biomolecules and systems, including synthetic DNA and genetic control circuits, with relevance to many biomedical and biotechnological applications.

[0101] method Construction of DNA vectors The vectors used in this study (Table 4) were constructed using standard molecular cloning techniques, including PCR, restriction enzyme digestion, ligation, and Gibson assembly. Custom oligonucleotides were purchased from Integrated DNA Technologies and Genewiz. The vector constructs were transformed into the E. coli strain DH5α, and colonies containing the constructs were isolated using 50 μg / ml carbenicillin / ampicillin. DNA was extracted and purified using the Plasmid Mini (Takara) or Midi (Qiagen) kit. The sequences of the vector constructs were confirmed by Sanger sequencing.

[0102] To generate lentiviral expression vectors encoding eSpCas9(1.1), HypaCas9, or SpCas9-HF1 with Zeocin as a selectable marker, SpCas9 sequences were amplified / mutated from pAWp30 (Addgene #73857), eSpCas9(1.1) (Addgene #71814), and VP12 (Addgene #72247) by PCR using Phusion DNA polymerase (New England Biolabs) and cloned into the pFUGW lentiviral expression vector backbone using Gibson Assembly Master Mix (New England Biolabs). Lentiviral expression vectors encoding evoCas9, Sniper-Cas9, and xCas9(3.7) were generated by amplifying their SpCas9 sequences from Addgene constructs #107550, #113912, and #1803380, respectively, and cloning them into the pFUGW vector backbone. To construct a storage vector containing a U6 promoter-driven expression of gRNA targeting a specific gene, we used a previously reported 18Oligo pairs of gRNA and target sequences were synthesized, annealed, and cloned into BbsI-digested pAWp28 vector (Addgene #73850) using T4 DNA ligase (New England Biolabs) as shown in Figure 1. To explore SpCas9 variants compatible with gRNAs with an additional 5' G at the start of the 20-nucleotide spacer sequence to favor transcription under the U6 promoter, gRNAs with an additional 5' G were used in this study, except for some used in Figures 5 and 14. The spacer sequences of the gRNAs are listed in Table 5. To construct a lentiviral vector for U6-driven expression of gRNA, the storage vector was digested with BglII and MfeI enzymes (ThermoFisher Scientific) to prepare a U6-gRNA expression cassette, which was then inserted into the pAWp12 (Addgene #72732) vector backbone using ligation via compatible cohesive ends generated by digesting the vector with BamHI and EcoRI enzymes (ThermoFisher Scientific).To express gRNA together with dual RFP and GFP fluorescent protein reporters, the U6-driven gRNA expression cassette was inserted into the lentiviral vector backbone, pAWp9 (Addgene #73851), instead of pAWp12, using the same strategy as above.

[0103] Preparation of barcoded DNA parts for SpCas9 Guided by prior knowledge available at the time of initiating this study, we aimed to identify the gRNA-guided genomic site (SpCas9-HF1 4 and eSpCas9(1.1) 3 contacting target and non-target DNA strands in the target and non-target DNA strands (including those identified in [1], [2], [3], [4], [5], [6], [7], [8], [9],

[10] ,

[11] ,

[12] ,

[13] ,

[14] ,

[15] ,

[16] ,

[17] ,

[18] ,

[19] ,

[20] ,

[21] ,

[22] , [ 28We focused on constructing a library of combinatorial mutants at amino acid residues predicted to be ubiquitous. Eight amino acid residues were selected and modified to carry specific or randomly generated substitution mutations (Figure 1a). Basic residues were mutated to alanine to evaluate the role of these charged residues. In addition to the alanine substitution at K1003 previously introduced in eSpCas9 (1.1), this residue was also mutated to other positively charged residues (i.e., arginine and histidine) to minimize its impact on protein stability. We hypothesized that specific combinations of these mutations on SpCas9 could maximize its on-target editing efficiency and enhance gRNA compatibility while minimizing unwanted off-target activity.

[0104] To construct combinatorial mutants, the SpCas9 sequence was modularized into four parts (i.e., P1, P2, P3, and P4), creating four inserts in P1, two inserts in P2, 17 inserts in P3, and seven inserts in P4. Each insert was amplified from pAWp30 (Addgene #73857) or eSpCas9(1.1) (Addgene #71814) and mutated by PCR using Phusion (New England Biolabs) or Kapa HiFi (Kapa Biosystems) DNA polymerase. To generate site-specific mutations at amino acid positions 923, 924, and 926 of SpCas9, the three original codon sequences were replaced with the degenerate codon NNS in the PCR primers. After cloning into the storage vector (pAWp61 or pAWp62), a unique 8-base pair barcode was added to each DNA insert. BsaI restriction enzyme sites were added adjacent to the ends (a BbsI site and a primer binding site for barcode sequencing were introduced between the insert and barcode for pAWp61 and pAWp62, respectively). Thus, the pAWp61 and pAWp62 storage vectors of the present invention were constructed as "BsaI-insert-BbsI-BbsI-barcode-BsaI" and "BsaI-insert-primer-binding site-barcode-BsaI," respectively. Sanger sequencing was performed to confirm the sequence identity of the individual inserts and their barcodes. If the engineered sequence of interest contains a BsaI or BbsI site, other type IIS restriction enzyme sites can be used in place of BsaI and BbsI, or synonymous mutations can be introduced into the protein-coding sequence to remove the restriction sites while still encoding the same amino acid residues.

[0105] Generation of barcoded combinatorial mutation libraries for SpCas9 Storage vectors containing each part of SpCas9 insert were mixed in an equimolar ratio. Pooled inserts were generated by a single-pot digestion of the mixed storage vector with BsaI. The target vector (pAWp60) was digested with BbsI. The digested P1 insert and vector were ligated to create a pooled P1 library into the target vector. This P1 library was again digested with BbsI and ligated with the digested P2 insert to create a two-way combination library (P1 x P2). Sequential ligation reactions were performed to create three-way (P1 x P2 x P3) and four-way (P1 x P2 x P3 x P4) combination libraries. After the pooled assembly step, the protein-encoding portions of the inserts were seamlessly integrated and localized at one end of the vector construct, and their respective barcodes were ligated to the other end. A four-part (4 × 2 × 17 × 7) combinatorial library of 952 SpCas9 variants was constructed to examine their interactions with the target and non-target DNA strands of the gRNA-guided genomic site. 3,4 or altering the conformational dynamics of the SpCas9 nuclease domain. 28Each of the SpCas9 mutants contained one to eight mutations (excluding WT) at predicted amino acid residues (Figure 1a). The complexity of this combination can be expanded by adding barcoded parts, allowing for scaling up to simultaneously test tens of thousands or more combinatorial modifications. Sanger sequencing analysis confirmed that the majority of the assembled barcoded combinatorial mutant constructs contained the expected mutations in the two-way (20 / 20 colonies), three-way (14 / 15 colonies), and four-way (8 / 8 colonies) libraries. Except for one three-way combinatorial mutant construct with an unintended base substitution, no random mutations were detected in the other constructs. The final library was subcloned into the pFUGW lentiviral vector, and SpCas9 mutants were expressed under the EFS promoter with the selectable marker Zeocin. Sanger sequencing of the full-length sequences of the barcoded SpCas9 mutants incorporated into the lentiviral vector (seven of the seven colonies sampled from the library) confirmed the presence of only the expected mutations and no random mutations.

[0106] Generation of SpCas9 variants for individual validation Lentiviral vectors encoding individual SpCas9 mutants, including Opti-SpCas9, were constructed using the same method as used to generate the combinatorial mutant library described above, except that they were assembled one by one with individual inserts and vectors.

[0107] human cell culture HEK293T cells were obtained from the American Type Culture Collection (ATCC). OVCAR8-ADR cells were obtained from Ochiya (National Cancer Center, Japan). 42The cells were kindly donated by the OVCAR8-ADR Research Center. The identity of the OVCAR8-ADR cells was confirmed by cell line authentication test (Genetica DNA Laboratories). The monoclonal stable OVCAR8-ADR cell line was generated by introducing a tandem U6 promoter-driven expression cassette of gRNA targeting the RFP site into cells along with lentivirus encoding the RFP and GFP genes expressed from the UBC and CMV promoters, respectively. The RFPsg5-ON, RFPsg8-ON, and RFP-sg6-ON lines contain target sites on RFP that perfectly match the spacer of the gRNA, whereas the RFPsg5-OFF5-2, RFPsg8-OFF5, and RFPsg5-OFF5 lines contain synonymous mutations and target sites on RFP that mismatch the gRNA spacer (Table 6). HEK293T cells were cultured in DMEM supplemented with 10% heat-inactivated FBS and 1x antibiotic-antimycotic (Life Technologies) at 37°C with 5% CO. OVCAR8-ADR cells were cultured in RPMI supplemented with 10% heat-inactivated FBS and 1x antibiotic-antimycotic (Life Technologies) at 37°C with 5% CO.

[0108] Lentivirus production and transduction Lentivirus was added at 2.5 x 10 cells per well. 5HEK293T cells were transfected in 6-well plates. Cells were transfected for 15 minutes using FuGENE HD Transfection Reagent (Promega) with 0.5 μg of lentiviral vector, 1 μg of pCMV-dR8.2-dvpr vector, and 0.5 μg of pCMV-VSV-G vector mixed in 100 μl of OptiMEM medium (Life Technologies). One day after transfection, the medium was replaced with fresh medium. Viral supernatants were then collected every 24 hours between 48 and 96 hours posttransfection, pooled together, and filtered through a 0.45 μm polyethersulfone membrane. For transductions with individual vector constructs, 500 μl of filtered viral supernatant was used to transduce 2.5 × 10 cells in the presence of 8 μg / ml polybrene (Sigma). 5 Cells were infected overnight. The same test conditions were used to scale up lentivirus production to transduce the pooled library into human cells (OVCAR8-ADR). To ensure high-coverage libraries with sufficient reproducibility across most combinations, infections were performed with a starting cell population containing more than 300 times the size of the library being tested. The lentivirus was adjusted to a multiplicity of infection of approximately 0.3, resulting in an infection efficiency of approximately 30% in the presence of 8 μg / ml polybrene, ensuring the SpCas9 mutant library was delivered at low copy number.

[0109] Cell sorting Cell sorting was performed using a BD Influx cell sorter (BD Biosciences). Drop delay was measured using BD Accudrop beads. Cells were filtered through a 70 μm nylon mesh filter and then sorted through a 100 μm nozzle using the 1.0 Drop Pure sorting mode. Cells were gated by GFP-positive signals and classified into three bins (i.e., A, B, and C) based on RFP fluorescence intensity. Approximately 5% of the population was collected in each bin, encompassing cells with low RFP levels. The proportion of cells sorted into each bin can be adjusted to balance the trade-off between the representation of individual combinations in the sorted population and the sensitivity of detecting mutant enrichment between bins. Approximately 200,000–300,000 cells were collected in each sorted bin for each sample.

[0110] Sample preparation for barcode sequencing For the combinatorial mutant vector library, plasmid DNA was extracted from E. coli transformed with the vector library using the Plasmid Mini kit (Qiagen). For the human cell pool infected with the combinatorial mutant library, genomic DNA was extracted from cells harvested under various test conditions using the DNeasy Blood & Tissue Kit (Qiagen). DNA concentration was measured using the Quant-iT PicoGreen dsDNA Assay Kit (Life Technologies). PCR amplification of a 393-base pair fragment containing a unique barcode representing each combinatorial mutant, an Illumina anchor sequence, and an 8-base pair index barcode for multiplex sequencing was performed using Kapa HiFi Hotstart Ready-mix (Kapa Biosystems). The forward and reverse primers used were 5'-AATGATACGGCGACCACCGAGATCTACACGGAACCGCAACGGTATTC-3' and 5'-CAAGCAGAAGACGGCATACGAGAT NNNNNNNNGGTTGCGTCAGCAAACACAG-3', where NNNNNNNN indicates the specific index barcode assigned to each test sample. To avoid bias in PCR that could distort the population distribution, PCR conditions were optimized to ensure amplification occurred during logarithmic growth. PCR amplicons were purified through two rounds of size selection using Agencourt AMPure XP beads (Beckman Coulter Genomics) at ratios of 1:0.5 and 1:0.95 before real-time PCR quantification using Kapa SYBR Fast qPCR Master Mix (Kapa Biosystems) on a StepOnePlus Real Time PCR system (Applied Biosystems). The forward and reverse primers used for quantitative PCR were 5'-AATGATACGGCGACCACCGA-3' and 5'-CAAGCAGAAGACGGCATACGA-3', respectively. The quantified samples were then pooled at the desired ratios for multiplexing, evaluated on an Agilent 2100 Bioanalyzer using a high-sensitivity DNA chip (Agilent), and analyzed on an Illumina HiSeq using primer (5'-CCACCGAGATCTACGGAACCGCAACGGTATTC-3') and index barcode primer (5'-GTGGCGTGGTGCACTGTTTGCTGACGCAACC-3').

[0111] Barcode sequencing data analysis Barcode reads for each combination variant were processed from the sequencing data. Barcode reads representing each combination were normalized per million reads for each sample, as classified by the index barcode. Profiling was performed in two biological replicates. The frequency of each combination variant between the sorted bin A and the unsorted population was measured, and the enrichment factor (E) between them relative to the remaining population was calculated. Bin A was chosen because the enrichment of variants was most evident in this bin (Figure 2b). The formula used was as follows:

number

[0112] The log-transformed mean score (i.e., log ) detected from replicates comparing sorted bin A with the unsorted population 2 The log2(E) score was used as a measure of target editing activity. To increase the reliability of the data, only barcodes that yielded ≥300 absolute reads in the unsorted population were analyzed. The correlation between the log2(E) scores obtained from the pooled screen and the individual validation data (Figure 9) could be improved by increasing the cell expansion fold for each combination in the pooled screen to reduce experimental noise. 43 Mutants with optimized activity (i.e., Opti-SpCas9 identified in this study) were defined as those whose log2(E) (bin A vs. unsorted population) was at least 90% of WT for both RFPsg5-ON and RFPsg8-ON and less than 60% of WT for both RFPsg5-OFF5-2 and RFPsg8-OFF5. OptiHF-SpCas9 was identified as a mutant with high fidelity based on enrichment ratios at least 50% of WT for both RFPsg5-ON and RFPsg8-ON and less than 90% of WT for both RFPsg5-OFF5-2 and RFPsg8-OFF5. The complete list is shown in Table 2.

[0113] To determine epistasis, published protein compatibility literature was used. 44,45 A similar scoring system was applied to calculate the epistasis (ε) score for each combination in Figure 4. The ε score was determined as follows: observed fitness - expected fitness (where the expected fitness of a combination [X, Y] is the log(E [X] )+log2(E [Y])). Generally, combinations showing better than expected fitness were defined as positive epistasis, and combinations showing lower than expected fitness were defined as negative epistasis. For comparison, the log2(E) values ​​of lethal or near-lethal combination variants were set equal to that of the SpCas9 mutant with eight mutations (i.e., R661A + Q695A + K848A + E923M + T924V + Q926A + K1003A + R1060A) in this study. Our individual validation data confirmed the minimal activity of the mutants in disrupting the target RFP sequence (Figure 3b). The expected fitness was capped at the log2(E) values ​​of lethal or near-lethal combination variants to minimize false epistasis due to meaningless predicted fitness. In the future, it may be beneficial to include nucleolytic mutants of SpCas9 as lethal mutants in pooled screening for comparison.

[0114] Fluorescent protein disruption assay Fluorescent protein disruption assays were performed to assess DNA cleavage and indel-mediated disruption at the target site of a fluorescent protein (i.e., GFP or RFP) resulting from SpCas9 and gRNA expression, resulting in the loss of cellular fluorescence. Cells carrying the GFP or RFP reporter gene along with SpCas9 and gRNA were washed, resuspended in 1x PBS supplemented with 2% heat-inactivated FBS, and assayed on an LSR Fortessa analyzer (Becton Dickinson). Cells were gated by forward and side scatter. At least 1x10 cells per sample were included in each dataset. 4 Cells were recorded.

[0115] Immunoblot analysis Cells were lysed in 2x RIPA buffer supplemented with protease inhibitors (Gold Biotechnology #GB-108-2). Lysates were collected by scraping the culture plate on ice and centrifuged at 15,000 rpm for 15 minutes at 4°C. The supernatant was quantified using a Bradford assay (BioRad). Proteins were denatured at 99°C for 5 minutes and then subjected to gel electrophoresis on a 10% polyacrylamide gel (BioRad). Proteins were transferred to a polyvinylidene difluoride membrane at 110V for 2 hours at 4°C. The primary antibodies used were anti-Cas9 (7A9-3A3) (1:2,000, Cell Signaling #14697) and anti-β-actin (1:10,000, Sigma #A2228). The secondary antibody used was HRP-conjugated anti-mouse IgG (1:20,000, Cell Signaling #7076). The membrane was developed with WesternBright ECL HRP substrate (Advansta #K-12045-D20).

[0116] T7 endonuclease I assay A T7 endonuclease I assay was performed to evaluate DNA mismatch cleavage at the genomic locus targeted by the gRNA. Genomic DNA was extracted from cell cultures using QuickExtract DNA Extraction Solution (Epicentre) or the DNeasy Blood & Tissue Kit (Qiagen). PCR was performed using the primers and PCR conditions listed in Table 7, and amplicons containing the target locus were purified using Agencourt AMPure XP beads (Beckman Coulter Genomics). Approximately 400 ng of PCR amplicon was denatured, self-annealed, and incubated with 4 units of T7 endonuclease I (New England Biolabs) at 37°C for 40 minutes. Reaction products were separated using 2% agarose gel electrophoresis. Quantification was based on relative band intensities measured using ImageJ. The percentage of indels was calculated using the formula 100 × (1-(1-(b+c) / (a+b+c)). 1 / 2) (where a is the integrated intensity of the uncleaved PCR product, and b and c are the integrated intensities of each cleavage product.) 46 .

[0117] Genome-wide off-target GUIDE-Seq detection Genome-wide off-target detection with GUIDE-Seq 47 For each GUIDE-Seq sample, 1.5 million OVCAR8-ADR cells infected with SpCas9 variants and gRNA were electroporated with 1,000 pmol of freshly annealed GUIDE-seq end-protected dsODN using a 100 μl Neon tip (ThermoFisher Scientific) according to the manufacturer's protocol. The sequences of the dsODN oligos used were as follows: 5'-PG*T*TTAATTGAGTTGTCATATGTTAATAACGGT*A*T-3' and 5'-PA*T*ACCGTTATTAACATATGACAACTCAATTAA*A*C-3' Here, P indicates 5' phosphorylation and * indicates a phosphorothioate bond. Genomic DNA was extracted 72 hours after electroporation using the DNeasy Blood and Tissue kit (Qiagen). Genomic DNA concentration was quantified using the Qubit fluorometer dsDNA HS assay (ThermoFisher Scientific), and 400 ng was used for library construction according to the GUIDE-Seq protocol with minor modifications. Briefly, DNA was enzymatically fragmented using the KAPA Frag Kit (KAPA Biosystems), followed by adapter ligation and two rounds of hemi-nested PCR to enrich for dsODN-incorporated sequences. To integrate the Illumina sequencing workflow to obtain dual-index data using a single-index sequencing workflow on various Illumina platforms, we redesigned the semi-functional adapters by placing the sample index (index 2) at the beginning of read 1 according to the unique molecular index (Table 8). The final sequencing library was quantified using the KAPA Library Quantification Kit for Illumina and sequenced on an Illumina NextSeq 500 System. Data de-multiplexing for index 1 was performed with bcl2fq v2.19, followed by de-multiplexing for index 2 and GUIDE-Seq software. 48 Custom scripts were used to format the data for analysis.

[0118] All patents, patent applications and other publications cited herein, including GenBank Accession Numbers or equivalent sequence identification numbers, are hereby incorporated by reference in their entirety for all purposes. Further aspects of the present invention are described below: [Section 1] In order from 5' to 3', a first recognition site for a first type IIS restriction enzyme, DNA elements, a first and second recognition site for a second type IIS restriction enzyme; a barcode uniquely assigned to the DNA element, and A second recognition site for the first type IIS restriction enzyme A DNA construct comprising: [Section 2] Item 1. The DNA construct of item 1, which is a DNA vector. [Section 3] A library comprising two or more of the DNA constructs described in Item 1. [Section 4] In order from 5' to 3', a recognition site for a first type IIS restriction enzyme, Multiple DNA elements, a primer binding site, and a plurality of barcodes and a recognition site for a second type IIS restriction enzyme, each uniquely assigned to one of the plurality of DNA elements; wherein a plurality of DNA elements are linked to one another to form a protein-encoding sequence that does not contain any extraneous sequence at any junction between any two of the plurality of DNA elements, and wherein the plurality of barcodes are arranged in reverse order of the DNA elements to which they are assigned. [Section 5] Item 5. The DNA construct of item 4, which is a DNA vector. [Section 6] 6. The DNA construct of any one of paragraphs 1, 2, 4 and 5, wherein the first type IIS restriction enzyme and the second type IIS restriction enzyme cleave the DNA molecule to create compatible ends. [Section 7] 6. The DNA construct of any one of paragraphs 1, 2, 4 and 5, wherein the first type IIS restriction enzyme is BsaI and the second type IIS restriction enzyme is BbsI. [Section 8] A method for producing a combinatorial gene construct, comprising: (a) cleaving the first DNA vector of paragraph 2 with a first type IIS restriction enzyme to release a first DNA fragment comprising a first DNA segment, first and second recognition sites for a second type IIS restriction enzyme, and a first barcode adjacent to the first and second ends created by the first type IIS restriction enzyme; (b) cutting the promoter-containing first expression vector with a second type IIS restriction enzyme to linearize the first expression vector near the 3' end of the promoter and create two ends compatible with the first and second ends of the DNA fragment of (a); (c) annealing and ligating the first DNA fragment of (a) to the linearized expression vector of (b) to form a unidirectional composite expression vector in which the first DNA fragment and the first barcode are operably linked to a promoter at their 3' ends; (d) cleaving the second DNA vector of paragraph 2 with a first type IIS restriction enzyme to release a second DNA fragment comprising a second DNA segment, first and second recognition sites for the second type IIS restriction enzyme, and a second barcode adjacent to the first and second ends created by the first type IIS restriction enzyme; (e) cutting the composite expression vector of (c) with a second type IIS restriction enzyme to linearize the composite expression vector between the first DNA element and the first barcode and create two ends compatible with the first and second ends of the DNA fragment of (d); (f) annealing and ligating the second DNA fragment of (d) to the linearized hybrid expression vector of (e) between the first DNA element and the first barcode to form a two-way hybrid expression vector, in which the first DNA fragment, the second DNA fragment, the second barcode, and the first barcode are operably linked to a promoter at their 3' ends in this order. wherein first and second DNA elements encode first and second segments of a preselected protein from their N-termini immediately adjacent to each other, the first and second DNA fragments being joined together in a two-way composite expression vector that does not contain any exogenous nucleotide sequence that results in an amino acid residue not found in the preselected protein, and the first and second DNA elements each contain one or more mutations. How to make it. [Section 9] Item 9. The method according to item 8, wherein steps (d) to (f) are repeated n times to incorporate an n-th DNA fragment comprising an n-th DNA element, a first and second recognition site for a second type IIS restriction enzyme, and an n-th barcode into an n-directional hybrid expression vector, wherein the n-th DNA element encodes, from its C-terminus, an n-th or second to last segment of a preselected protein; (x) providing a final DNA vector comprising an (n+1)th DNA element, a primer binding site and an (n+1)th barcode between the first and second recognition sites for a first type IIS restriction enzyme; (y) cleaving the final DNA vector with a first type IIS restriction enzyme to release a final DNA fragment comprising, in 5' to 3' order, the (n+1)th DNA element, a primer binding site, and an (n+1)th barcode adjacent to the first and second ends created by the first type IIS restriction enzyme; (z) annealing and ligating the final DNA fragment to the n-way composite expression vector generated after repeating steps (d) to (f) n times and linearized with a second type IIS restriction enzyme to form a final composite expression vector. wherein the first, second and through (n+1)th DNA elements encode the first, second and through nth and last segments of a preselected protein immediately adjacent to each other from their N-termini, and the first, second and through nth and last DNA fragments are joined together in a final composite expression vector that does not contain any exogenous nucleotide sequence that results in an amino acid residue not found in the preselected protein, and each of the DNA elements contains one or more mutations. How to make it. [Section 10] 10. The method of paragraph 8 or 9, wherein the first type IIS restriction enzyme and the second type IIS restriction enzyme create compatible ends by cleaving the DNA molecule. [Section 11] Item 10. The method according to Item 8 or 9, wherein the first type IIS restriction enzyme is BsaI and the second type IIS restriction enzyme is BbsI. [Section 12] A library comprising two or more final composite expression vectors produced by the method described in Item 9. [Section 13] A polypeptide comprising the amino acid sequence set forth in any one of SEQ ID NOs: 1 and 4-13, in which a residue corresponding to residue 1003 of SEQ ID NO: 1 has been substituted and a residue corresponding to residue 661 of SEQ ID NO: 1 has been substituted. [Section 14] 14. The polypeptide of Paragraph 13, wherein the residue corresponding to residue 1003 of SEQ ID NO:1 is substituted with histidine and the residue corresponding to residue 661 of SEQ ID NO:1 is substituted with alanine. [Section 15] 15. The polypeptide of claim 14, comprising the amino acid sequence set forth in SEQ ID NO:1, wherein residue 1003 is substituted with histidine and residue 661 is substituted with alanine, and optionally further comprising an alanine substitution at residue 926. [Section 16] 14. The polypeptide of paragraph 13, wherein residues corresponding to residues 695, 848, and 926 of SEQ ID NO:1 are substituted with alanine, the residue corresponding to residue 923 of SEQ ID NO:1 is substituted with methionine, and the residue corresponding to residue 924 of SEQ ID NO:1 is substituted with valine. [Section 17] 17. The polypeptide of Paragraph 16, comprising the amino acid sequence set forth in SEQ ID NO:1, wherein residues corresponding to residues 695, 848, and 926 of SEQ ID NO:1 are substituted with alanine, the residue corresponding to residue 923 of SEQ ID NO:1 is substituted with methionine, and the residue corresponding to residue 924 of SEQ ID NO:1 is substituted with valine. [Section 18] Item 14. A composition comprising the polypeptide of Item 13 and a physiologically acceptable excipient. [Section 19] 18. A nucleic acid comprising a polynucleotide sequence encoding the polypeptide of any one of paragraphs 13 to 17. [Section 20] Item 18. A composition comprising the nucleic acid according to Item 17 and a physiologically acceptable excipient. [Section 21] 18. An expression cassette comprising a promoter operably linked to a polynucleotide sequence encoding the polypeptide of any one of paragraphs 13 to 17. [Section 22] A vector comprising the expression cassette according to Item 21. [Section 23] Item 23. The vector according to Item 22, which is a viral vector. [Section 24] A host cell comprising the expression cassette of paragraph 19 or the polypeptide of any one of paragraphs 13 to 17. [Section 25] 18. A method for cleaving a DNA molecule at a target site, comprising contacting a DNA molecule containing the target DNA site with the polypeptide of any one of paragraphs 13 to 17 and a short guide RNA (sgRNA) that specifically binds to the target DNA site, thereby cleaving the DNA molecule at the target DNA site. [Section 26] 26. The method of claim 25, wherein the DNA molecule is genomic DNA in a living cell, and the cell has been transfected with a polynucleotide sequence encoding the sgRNA and the polypeptide. [Section 27] 27. The method of paragraph 26, wherein the cell is transfected with a first vector encoding the sgRNA and a second vector encoding the polypeptide. [Section 28] 27. The method of paragraph 26, wherein the cell is transfected with a vector encoding both the sgRNA and the polypeptide. [Section 29] 28. The method of paragraph 27, wherein the first and second vectors are each viral vectors. [Section 30] 29. The method of paragraph 28, wherein the vector is a viral vector. [Section 31] 31. The method of paragraph 29 or 30, wherein the viral vector is a retroviral vector. [Section 32] 32. The method of claim 31, wherein the retroviral vector is a lentiviral vector.

[0119] [Table 1]

[0120] [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] Table 2-6 Table 2-7 Table 2-8 Table 2-9 Table 2-10 Table 2-11 Table 2-12 Table 2-13 Table 2-14 Table 2-15 Table 2-16 Table 2-17 Table 2-18 Table 2-19 Table 2-20

[0121] Table 3

[0122] Table 4-1 Table 4-2 Table 4-3 Table 4-4 Table 4-5

[0123] Table 5-1 Table 5-2 Table 5-3

[0124] Table 6

[0125] Table 7

[0126] Table 8-1 Table 8-2

[0127] JPEG0007813045000036.jpg234164 JPEG0007813045000037.jpg229164 JPEG0007813045000038.jpg235164 JPEG0007813045000039.jpg242164 JPEG0007813045000040.jpg253164 JPEG0007813045000041.jpg253162 JPEG0007813045000042.jpg237164 JPEG0007813045000043.jpg69164

[0128] literature JPEG0007813045000044.jpg219164 JPEG0007813045000045.jpg227164 JPEG0007813045000046.jpg132164

Claims

Claim 1: A polypeptide having 90% or more sequence identity to the amino acid sequence of SEQ ID NO:1 and having increased or equivalent on-target editing efficiency and decreased off-target activity compared to SpCas9 of SEQ ID NO:1, wherein the polypeptide comprises an amino acid sequence having the following mutations relative to the amino acid sequence set forth in SEQ ID NO:1: (a) a residue corresponding to residue 1003 of SEQ ID NO:1 is substituted with a histidine, and a residue corresponding to residue 661 of SEQ ID NO:1 is substituted with an alanine; (b) the residue corresponding to residue 695 of SEQ ID NO:1 is substituted with alanine, the residue corresponding to residue 848 of SEQ ID NO:1 is substituted with alanine, the residue corresponding to residue 923 of SEQ ID NO:1 is substituted with methionine, the residue corresponding to residue 924 of SEQ ID NO:1 is substituted with valine, and the residue corresponding to residue 926 of SEQ ID NO:1 is substituted with alanine; (c) a residue corresponding to residue 661 of SEQ ID NO:1 is substituted with alanine, a residue corresponding to residue 926 of SEQ ID NO:1 is substituted with alanine, and a residue corresponding to residue 1003 of SEQ ID NO:1 is substituted with histidine; or (d) the residue corresponding to residue 661 of SEQ ID NO:1 is substituted with alanine, the residue corresponding to residue 848 of SEQ ID NO:1 is substituted with alanine, the residue corresponding to residue 923 of SEQ ID NO:1 is substituted with histidine, the residue corresponding to residue 924 of SEQ ID NO:1 is substituted with leucine, and the residue corresponding to residue 1003 of SEQ ID NO:1 is substituted with histidine.

2. A nucleic acid comprising a polynucleotide sequence encoding the polypeptide of claim 1.

3. A composition comprising the polypeptide of claim 1 or the nucleic acid of claim 2 and a physiologically acceptable excipient.

4. An expression cassette comprising a polynucleotide sequence encoding the polypeptide of claim 1 and a promoter operably linked to the polynucleotide sequence.

5. A vector comprising the expression cassette of claim 4.

6. The vector of claim 5 , which is a viral vector.

7. A host cell comprising the expression cassette of claim 4 or the polypeptide of claim 1.

8. 10. An in vitro method for cleaving a DNA molecule at a target site, comprising contacting a DNA molecule containing the target DNA site with the polypeptide of claim 1 and a short guide RNA (sgRNA) that specifically binds to the target DNA site, thereby cleaving the DNA molecule at the target DNA site.

9. 9. The method of claim 8, wherein the DNA molecule is genomic DNA in a living cell, and the cell is transfected with a polynucleotide sequence encoding an sgRNA and a polypeptide of claim 1.

10. 10. The method of claim 9, wherein the cell is transfected with a first vector encoding an sgRNA and a second vector encoding the polypeptide of claim 1.

11. 10. The method of claim 9, wherein the cell is transfected with a vector encoding both the sgRNA and the polypeptide of claim 1.

12. The method of claim 10 , wherein the first and second vectors are each viral vectors.

13. The method of claim 11 , wherein the vector is a viral vector.

14. The method of claim 12 or 13, wherein the viral vector is a retroviral vector.

15. 15. The method of claim 14, wherein the retroviral vector is a lentiviral vector.

16. 4. The composition of claim 3 for cleaving a DNA molecule at a target site, wherein the composition cleaves the DNA molecule at the target DNA site by contacting the DNA molecule containing the target DNA site with a short guide RNA (sgRNA) that specifically binds to the target DNA site.

17. 17. The composition of claim 16, wherein the DNA molecule is genomic DNA in a living cell, and the cell is transfected with a polynucleotide sequence encoding an sgRNA and the polypeptide of claim 1.

18. 18. The composition of claim 17, wherein the cell is transfected with a first vector encoding an sgRNA and a second vector encoding the polypeptide of claim 1.

19. 18. The composition of claim 17, wherein the cell is transfected with a vector encoding both the sgRNA and the polypeptide of claim 1.

20. 20. The composition of claim 18, wherein the first and second vectors are each viral vectors.

21. The composition of claim 19 , wherein the vector is a viral vector.

22. The composition of claim 20 or 21, wherein the viral vector is a retroviral vector.

23. 23. The composition of claim 22, wherein the retroviral vector is a lentiviral vector.

Citation Information

Patent Citations

  • Engineered crispr-cas9 nuclease

    JP2018525019A

  • Cas9 variants and methods of use thereof

    WO2016196655A1

  • High-fidelity CAS9 variants and applications thereof

    WO2018149888A1