Stable genome editing complex with few side effects and nucleic acid encoding the same
A genome editing complex with a proteolytic tag and nucleic acid modifying enzymes stabilizes and enhances bacterial gene editing efficiency, addressing toxicity and mutation issues in conventional methods.
Patent Information
- Application Number
- JP2025121334
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-11-22
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-07
AI Technical Summary
Conventional genome editing technologies are highly toxic to bacterial hosts, leading to vector instability and nonspecific mutations, and current methods like CRISPR/Cas9 induce cell death in bacteria lacking NHEJ pathways, while deaminase-mediated base editing has insufficient efficiency for multiple-site editing.
A genome editing complex with a nucleic acid sequence recognition module and a proteolytic tag, such as LVA, to stabilize the complex in bacteria, reducing nonspecific mutations and enhancing editing efficiency by using nucleic acid modifying enzymes like deaminases, which do not rely on host-dependent factors like RecA.
The complex allows stable amplification and efficient gene modification in bacteria with reduced toxicity and nonspecific mutations, applicable to a wide range of bacterial species.
Smart Images

Figure 2025148567000009 
Figure 2025148567000010 
Figure 2025148567000011
Abstract
Description
[Technical Field]
[0001] The present invention relates to a genome editing complex that is stable and has few side effects, a nucleic acid encoding the same, and a genome editing method using the complex. [Background technology]
[0002] Genome editing, which does not require the integration of selectable marker genes and can minimize the impact on the expression of downstream genes in the same operon, is particularly advantageous in prokaryotes. Phage-derived RecET and λ-Red recombinases have been used as recombination technologies to facilitate homology-dependent integration / replacement of donor DNA or oligonucleotides (e.g., Non-Patent Document 1). By combining this with strains deficient in methyl-directed mismatch repair (MMR), highly efficient recombination can be achieved without the integration of selectable markers (Non-Patent Document 2). This technology has been utilized in multiple automated genome engineering (MAGE) to generate genetic diversity at multiple target loci within a few days. However, this recombination technology relies on host-dependent factors such as MMR deficiency and RecA, a central component of the recombination DNA repair system, which harms most Escherichia coli used as hosts for cloning, making it difficult to transfer to bacterial species with different backgrounds (Non-Patent Document 3).
[0003] CRISPR (clustered regularly interspaced short palindromic repeats) and CRISPR The associated (Cas) proteins encode a single guide RNA (sgRNA) and a protospacer-adjacent motif. It is known to function as a bacterial adaptive immune system by cleaving target DNA in a PAM-dependent manner. The Cas9 nuclease from Streptococcus pyogenes has been widely used as a powerful genome editing tool in eukaryotes that possess DNA double-strand break (DSB) repair pathways (e.g., Non-Patent Documents 4 and 5). During DSB repair via the non-homologous end joining (NHEJ) pathway, small insertions and / or deletions (indels) are introduced into the target DNA, resulting in site-specific mutations or gene disruption. For more precise editing, homology-directed repair (HDR) can be promoted by providing donor DNA containing homologous arms to the target region, although the efficiency depends on the host cell.
[0004] However, current genome editing technologies rely on the host's DNA repair system. Further ingenuity is required for application to prokaryotes. In most bacteria, DNA cleavage by artificial nucleases results in cell death due to the lack of the NHEJ pathway (Non-Patent Documents 6 and 7). Therefore, CRISPR / Cas9 is a promising tool for gene modification in vivo, as it allows for the efficient translation of genes into other target genes through other methods, such as the λ-Red recombination system. It has only been used as a counter-selector for transformed cells (eg, Non-Patent Documents 8 and 9).
[0005] Recently, it has been reported that target genes can be cloned without using donor DNA containing homology arms to the target region. Deaminase-mediated targeted base editing, which directly edits nucleotides at a locus, has been demonstrated (e.g., Patent Document 1, Non-Patent Documents 10 to 12). This technique uses DNA deamination instead of nuclease-mediated DNA cleavage, and therefore does not induce bacterial cell death and can be applied to bacterial genome editing. However, its mutation efficiency, especially the efficiency of simultaneous editing at multiple sites, is not sufficient. [Prior art documents] [Patent documents]
[0006]
Patent Document 1
Non-licensed literature
[0007] [Non-licensed document 1] Datsenko, KA & Wanner, BL, Proc. Natl. Acad. Sci. USA 97, 6640-5 (2000). [Non-licensed document 2] Costantino, N. & Court, DL, Proc. Natl. Acad. Sci. USA 100, 15748-53 (2003). [Non-licensed document 3] Wang, J. et al., Mol. Biotechnol. 32, 43-53 (2006).
Non-licensed Document 4
Non-licensed Document 5
Non-licensed Document 6
Non-licensed Document 7
Non-licensed Document 8
Non-licensed literature 9
Non-licensed literature 10
[0008] Conventional genome editing vectors are expressed from the vector and act on the host's genomic DNA. The high toxicity of genome editing complexes places a heavy burden on hosts, particularly bacteria, and can lead to vector instability within the host. Genome editing can cause side effects such as nonspecific and off-target mutations. In particular, when mutation efficiency is increased using uracil DNA glycosylase inhibitors (UGIs), the tradeoff is strong toxicity to the host, causing cell death and increased rates of nonspecific mutations. Therefore, the present invention aims to provide nucleic acids such as low-toxicity vectors that can be stably amplified within a host, and genome editing complexes encoded by such nucleic acids. It also aims to provide a genome editing method that uses such vectors, and if necessary, nucleic acid modifying enzymes, to modify bacterial DNA while suppressing nonspecific mutations and is applicable to a wide range of bacteria without relying on host-dependent factors such as RecA. [Means for solving the problem]
[0009] The inventors have demonstrated that by suppressing the abundance of genome editing complexes, which are highly toxic to bacteria as hosts, in bacteria, it is possible to stabilize vectors in bacteria and prevent non-specific mutations in bacterial DNA. Therefore, in order to suppress the abundance of genome editing complexes, we focused on the LVA tag, a proteolytic tag known to promote protein degradation in bacteria and shorten its half-life, and conducted research. As a result, we were able to It was demonstrated that adding the proteolytic tag to the target complex reduces nonspecific mutations while maintaining the mutation efficiency at the target site, and that even when UGI is combined, nonspecific mutations can be reduced and the target sequence can be modified with high efficiency (Figures 9 and 10). Based on these findings, the present inventors conducted further research and have completed the present invention.
[0010] That is, the present invention is as follows. [1] A nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in double-stranded DNA. and a protein comprising (i) a peptide containing three hydrophobic amino acid residues at the C-terminus, or (ii) a peptide containing three amino acid residues at the C-terminus, in which at least some of the amino acid residues are substituted with serine. A complex in which the protein degradation tag is bound. [2] The complex further comprises a nucleic acid modifying enzyme that converts one or more nucleotides at the targeted site into one or more other nucleotides or deletes them; or The complex described in [1], which inserts one or more nucleotides into the targeted site. [3] The complex according to [1] or [2], wherein the three amino acid residues are leucine-valine-alanine, leucine-alanine-alanine, alanine-alanine-valine, or alanine-serine-valine. [4] The method according to any one of [1] to [3], wherein the nucleic acid sequence recognition module is a CRISPR-Cas system in which only one or both of the two DNA cleavage abilities of Cas are inactivated. A complex of [5] The complex according to any one of [1] to [3], wherein the complex is a complex in which a CRISPR-Cas system is bound to a proteolytic tag. [6] The nucleic acid modifying enzyme is a nucleic acid base conversion enzyme or a DNA glycosylase. [4] The complex according to any one of [4]. [7] The complex according to [6], wherein the nucleic acid base conversion enzyme is a deaminase. [8] The complex according to [6] or [7], further comprising an inhibitor of base excision repair bound thereto. [9] A nucleic acid encoding the complex according to any one of [1] to [8].
[10] A method for modifying a targeted site in bacterial double-stranded DNA or regulating the expression of a gene encoded in double-stranded DNA near the site, comprising: a nucleic acid sequence recognition module that specifically binds to a nucleotide sequence; and (i) a hydrophobic amino acid sequence of 3 or (ii) a peptide containing at least some of the amino acid residues at the C-terminus thereof, wherein at least some of the amino acid residues are substituted with serine. A proteolytic tag consisting of a peptide containing the three selected amino acid residues at its C-terminus was attached to the contacting the complex with said double-stranded DNA.
[11] The method described in
[10] , wherein the complex further comprises a nucleic acid modifying enzyme bound thereto, and the method comprises converting or deleting one or more nucleotides at the targeted site into one or more other nucleotides, or inserting one or more nucleotides at the targeted site.
[12] The method according to
[10] or
[11] , wherein the three amino acid residues are leucine-valine-alanine, leucine-alanine-alanine, alanine-alanine-valine, or alanine-serine-valine.
[13] Any of
[10] to
[12] , wherein the nucleic acid sequence recognition module is a CRISPR-Cas system in which only one or both of the two DNA cleavage abilities of Cas are inactivated. The method described above.
[14] The method according to any one of
[10] to
[12] , wherein the complex is a complex in which a CRISPR-Cas system is bound to a proteolytic tag.
[15] The method according to any one of
[10] to
[14] , characterized in that two or more types of nucleic acid sequence recognition modules, each specifically binding to a different target nucleotide sequence, are used.
[16] The method according to
[15] , wherein the different target nucleotide sequences are present in different genes.
[17] The method according to any one of
[0010] to
[13] ,
[15] and
[16] , wherein the nucleic acid modifying enzyme is a nucleic acid base conversion enzyme or a DNA glycosylase.
[18] The method according to
[17] , wherein the nucleic acid base conversion enzyme is a deaminase.
[19] The method according to
[17] or
[18] , wherein the complex further has an inhibitor of base excision repair bound thereto.
[20] The method according to any one of
[10] to
[19] , wherein the contact between the double-stranded DNA and the complex is achieved by introducing a nucleic acid encoding the complex into a bacterium having the double-stranded DNA. [Effects of the Invention]
[0011] The present invention provides a low-toxicity nucleic acid (e.g., a vector) that can be stably amplified even in a host bacterium, and a genome editing complex encoded by the nucleic acid. Genome editing techniques using the nucleic acid and nucleic acid-modifying enzyme of the present invention make it possible to modify genes in host bacteria while suppressing nonspecific mutations, or to regulate the expression of genes encoded in double-stranded DNA. Because this technique does not depend on host-dependent factors such as RecA, it can be applied to a wide range of bacteria. [Brief explanation of the drawings]
[0012] [Figure 1]Figure 1 shows an overview of the Target-AID system in bacteria. (a) A schematic model of Target-AID (dCas9-PmCDA1 / sgRNA) base editing is shown. The dCas9-PmCDA1 / sgRNA complex binds to double-stranded DNA and forms an R-loop in an sgRNA- and PAM-dependent manner. PmCDA1 catalyzes the deamination of cytosines located on the upper (non-complementary) strand within 15–20 bases upstream of the PAM, resulting in C-to-T mutagenesis. (b) A single bacterial Target-AID plasmid is shown. This plasmid contains a chloramphenicol resistance (CmR) gene, a temperature-sensitive (ts) λ cI repressor, a pSC101 origin of replication (ori), and RepA101(ts). The lambda operator drives expression of the dCas9-PmCDA1 fusion at high temperatures (>37°C) as the cI repressor (ts) is inactivated. The sgRNA is driven by the constitutive promoter J23119. dCas9 represents a nuclease-deficient Cas9 with the D10A H840A mutation, and PmCDA1 represents the P. marinus (lamprey) cytosine deaminase. [Figure 2] Figure 2 shows the transformation efficiency of Cas9 and Target AID vectors in E. coli. Plasmids expressing each modifying protein (Cas9, dCas, Cas9-CDA, nCas-CDA, or dCas-CDA) along with an sgRNA targeting the galK gene were transformed into E. coli DH5α strain and selected with a chloramphenicol resistance marker. Viable cells were counted and calculated as colony-forming units (CFU) per amount of transformed plasmid DNA. Dots represent three independent experiments, and boxes indicate the 95% confidence interval of the geometric mean from a t-test analysis. [Figure 3]Figure 3 shows mutations induced at specific sites in the galK9 gene by dCas-CDA. DH5α cells expressing dCas-CDA with an sgRNA targeting galK_9 were spotted onto LB agar plates, and single colonies were isolated. Eight randomly selected clones were sequenced and aligned. The translated amino acid sequence is shown at the bottom of each nucleotide sequence. The frequency of each sequence is shown as the number of clones. Boxes and inverted boxes indicate the target sequence and PAM sequence, respectively. ORF numbers are shown at the top. The mutated site is highlighted with black shading, and the mutated base is highlighted in bold. The mutated codon is underlined. [Figure 4] Figure 4 shows mutation frequencies assessed by drug resistance. (a) shows the frequency of galK mutagenesis and 2-DOG resistance. DH5α cells expressing dCas-CDA with a non-targeting sgRNA (vector) or galK_9-targeting ssRNA were serially diluted and spotted onto agar plates of M63 medium with or without 2-DOG, followed by colony counts. (b) shows the frequency of rpoB mutagenesis and rifampicin resistance. Cells expressing dCas-CDA with a non-targeting sgRNA (vector) or rpoB_1-targeting ssRNA were serially diluted and spotted onto LB agar plates with or without rifampicin, followed by colony counts. Drug resistance frequencies were calculated as the number of drug-resistant colonies relative to the number of unselected colonies. Dots represent four independent experiments, and boxes indicate the 95% confidence intervals of the geometric mean from t-test analysis. [Figure 5]Figure 5 shows gain-of-function mutagenesis of the rpoB gene. (a) Sequence alignment of rpoB mutations induced by dCas-CDA. DH5α cells expressing dCas-CDA with sgRNA targeting rpoB_1 were spotted onto LB agar plates to isolate single colonies. Eight randomly selected clones were sequenced and the sequences aligned. The translated amino acid sequence is shown at the bottom of each nucleotide sequence. The frequency of each sequence is shown as the number of clones. Boxes and inverted boxes indicate the target and PAM sequences. ORF numbers are shown at the top. Mutated sites are highlighted with black shading, and mutated bases and amino acids are highlighted in bold. Mutated codons are underlined. (b) shows the results of whole-genome sequencing of cells subjected to mutagenesis of rpoB. Three independent clones selected with rifampicin were subjected to whole-genome sequencing. Sequence coverage was calculated as the total base pairs of sequence mapped across 4,631 Mbp of the E. coli BW25113 genome sequence. Parental / variable mutations are shown as the number of variants detected at frequencies greater than 50%, including insertions, deletions, single nucleotide variants (SNVs), and multiple nucleotide variants (MNVs), minus the common parental mutation. Detected mutations are shown as the number of variants (count), genomic locus (region / gene), reference genome sequence, and mutant allele. Variant calling was performed as described in the Examples section. (c) shows the sequences surrounding the detected mutations listed in (b). Mutated sites are highlighted in gray, and mutated bases and amino acids are highlighted in bold. [Figure 6]Figure 6 shows the mutation locations and frequencies with and without UGI-LVA and with sgRNAs of different lengths. Target sequences longer than 20 nt (galK_8, 9, 11, and 13) were tested using dCas-CDA (white bars on the left) or dCas-CDA-UGI-LVA (black bars on the right) and analyzed by deep sequencing. Averages from three independent experiments are plotted. Gray-shaded and inverted boxes indicate the galK target sequence and PAM, respectively. Mutated bases are underlined. [Figure 7] Figure 7 shows the effect of target sequence characteristics on the location and frequency of mutations induced by Target-AID. Cells expressing dCas-CDA and each targeting sgRNA were analyzed by deep sequencing. The target sequence (20 nt in length or as indicated) was located on the upper (+) or lower (-) DNA strand of the galK ORF, and missense (M) or nonsense (N) mutations were introduced as expected. The corresponding ORF number is indicated (Position). The mutation frequency of the peak base position (highlighted in gray in the sequence) was obtained as the average of three independent experiments. Mutation frequencies of >50%, 10-50%, or <10% were distinguished by gray shading. [Figure 8] Figure 8 shows the effect of target length on mutational spectra. (a) Mutation frequencies using target sequences of various lengths in gsiA. Target sequences containing a polyC at the distal site were edited using dCas-CDA-UGI-LVA and analyzed by deep sequencing. Mutational spectra for sgRNAs with lengths of 18 nt, 20 nt, 22 nt, or 24 nt are differentiated by gray shades. Averages from three independent experiments are shown. Inverted boxes indicate PAMs. Mutated bases are underlined. (b) Mutation frequencies for targets in ycbF and yfiH. Targets were set on the lower strand. Mutational spectra for sgRNAs with lengths of 18 nt, 20 nt, or 22 nt are shown in the same manner as in (a). (c) Averaged mutational spectra for each sgRNA length in (a) and (b). Peak positions are numbered. [Figure 9]Figure 9 shows multiple mutagenesis in the galK gene. (a) Nonspecific mutagenesis effects were assessed by rifampicin resistance. Cells expressing each protein (vector, dCas, dCas-CDA, dCas-CDA-LVA, or dCas-CDA1-UGI-LVA) along with the tandem-sgRNA unit against the galK_10-galK_11-galK_13 target were spotted onto LB agar plates with or without rifampicin to assess the frequency of nonspecific mutations. Dots represent at least three independent experiments, and boxes indicate the 95% confidence intervals of the geometric means from t-test analysis. (b) The frequency of on-target multiple mutations induced in the target region. Eight randomly selected clones from (a) were sequenced at the three targeted loci, and the frequencies of single-, double-, or triple-mutant clones were shown. (c and d) show sequence alignments of the mutants. dCas-CDA-UGI-LVA was used to mutate single targets (galK_10, galK_11, or galK_13) (c) or triple targets (d). Eight randomly selected clones were sequenced and the sequences aligned. The boxes and inverted boxes indicate the target and PAM sequences, respectively. The mutation sites are highlighted in black shading and bold. [Figure 10] Figure 10 shows multiple mutagenesis. (a) Schematic diagram of two plasmids for multiple mutagenesis (a modification vector expressing dCas-CDA-UGI-LVA and the plasmid pSBP80608 containing two tandem repeat sgRNA-units, each containing three targeting sgRNAs). (b) Sequence alignment of the target regions. Eight randomly selected clones were sequenced and aligned for each target region. Clone numbers are indicated to the left of the sequences. Boxes and inverted boxes indicate the target sequence and PAM. Mutated sites and bases are highlighted in black shading and bold. [Figure 11]Figure 11 shows the simultaneous disruption of multiple copies of the transposase gene. IS1, 2, 3, and 5 were simultaneously targeted using dCas-CDA-UGI-LVA. sgRNAs were designed to introduce a stop codon into the consensus sequence for the same type of transposase. All sequences were aligned except for those that could not be amplified from the DH10B reference genome. The translated amino acid sequence is shown above each consensus sequence. The genomic region of each sequence is indicated on the left. All target sequences were designed on the complementary strand, and corresponding regions were aligned with the complementary PAM sequence (inverted). Mutated bases are highlighted in black. [Figure 12] Figure 12 shows the method for isolation and confirmation of IS-edited cells. Clonal isolation and sequence confirmation were performed stepwise. Isolated clones were numbered as shown in the top row of each table and sequenced at the IS site indicated in the left column. Genotypes were determined based on Sanger sequencing spectra as confirmed (mutated), unmutated (wt), or heterologous (mutated and wt) with the targeted mutation. [Figure 13] Figure 13 shows a schematic diagram of the yeast expression vector (background: pRS315 vector) used in Example 5. In the figure, Gal1p represents the GAL1-10 promoter. DETAILED DESCRIPTION OF THE INVENTION
[0013] 1. Genome editing complex and nucleic acid encoding same The present invention relates to a nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in double-stranded DNA. In one embodiment of the genome editing complex of the present invention, a complex is provided to which a nucleic acid modification enzyme is further bound (i.e., a complex to which a nucleic acid sequence recognition module, a nucleic acid modification enzyme, and a proteolytic tag are bound), which is capable of modifying nucleic acid at a targeted site. In one embodiment, the complex contains a base excision molecule to improve the efficiency of double-stranded DNA modification. In another embodiment of the genome editing complex of the present invention, a complex is provided in which at least a nucleic acid sequence recognition module and a proteolytic tag are bound, and the complex inhibits the expression of a gene encoded by double-stranded DNA in the vicinity of the targeted site. The present invention provides a complex capable of regulating the expression of a nucleic acid. In one embodiment, a transcriptional regulatory factor may further be bound to the complex. Hereinafter, a complex bound to at least one of a nucleic acid modifying enzyme, a base excision repair inhibitor, and a transcriptional regulatory factor, and a complex bound to none of these may be collectively referred to as a "complex of the present invention" or a "genome editing complex," and a complex bound to a nucleic acid modifying enzyme may particularly be referred to as a "nucleic acid modifying enzyme complex." Furthermore, nucleic acids encoding these complexes may be collectively referred to as "nucleic acids of the present invention."
[0014] The nucleic acid of the present invention is introduced into a host bacterium (e.g., a bacterium containing a bacterium or a bacterium) for the purpose of replication, not for the purpose of modifying DNA. When the nucleic acid of the present invention is introduced into a host bacterium (e.g., a bacterium) and cultured, even if a complex is unintentionally expressed from the nucleic acid, the complex is quickly degraded by the proteolytic tag, thereby reducing toxicity to the host bacterium. In fact, the Examples below demonstrate that when the nucleic acid of the present invention is introduced into a host bacterium for the purpose of replicating the nucleic acid, the transformation efficiency of the host bacterium is higher than when a nucleic acid not encoding a proteolytic tag is introduced. Therefore, the nucleic acid of the present invention containing a sequence encoding a proteolytic tag can be stably replicated in bacteria as a nucleic acid for genome editing in a host other than bacteria (e.g., a eukaryote). Therefore, it is also useful to add a sequence encoding the proteolytic tag of the present invention to a vector intended for genome editing in a host other than a bacterium.
[0015] In the present invention, "modification" of double-stranded DNA refers to the conversion of a nucleotide (e.g., dC) on a DNA strand to another nucleotide (e.g., dT, dA, or dG) or deletion, or the insertion of a nucleotide or nucleotide sequence between nucleotides on a DNA strand. The double-stranded DNA to be modified is not particularly limited as long as it is double-stranded DNA present in a host cell, but is preferably genomic DNA. Furthermore, the "targeted site" of double-stranded DNA refers to all or a portion of the "target nucleotide sequence" that a nucleic acid sequence recognition module specifically recognizes and binds to, or the vicinity of the target nucleotide sequence (either 5' upstream or 3' downstream, or both). Furthermore, the "target nucleotide sequence" refers to the sequence in double-stranded DNA to which a nucleic acid sequence recognition module binds. In the present invention, the term "genome editing" is used to refer not only to the modification of double-stranded DNA, but also to the promotion or suppression of expression of a gene encoded by double-stranded DNA near the targeted site.
[0016] In the present invention, a "nucleic acid sequence recognition module" refers to a specific nucleotide sequence on a DNA strand. The term "nucleic acid modification enzyme complex" refers to a molecule or molecular complex that has the ability to specifically recognize and bind to a target nucleotide sequence (i.e., a target nucleotide sequence). When a nucleic acid modification enzyme complex is used, the nucleic acid sequence recognition module binds to the target nucleotide sequence, allowing the nucleic acid modification enzyme and / or base excision repair inhibitor linked to the module to act specifically on the targeted site of double-stranded DNA. This makes it possible to:
[0017] In the present invention, the term "nucleic acid modifying enzyme" refers to an enzyme that modifies nucleic acids, thereby directly or indirectly modifying DNA, and as long as it has catalytic activity, it can be used to generate peptide fragments. Such DNA modification reactions include those catalyzed by nucleases. , a reaction that cleaves a DNA strand (hereinafter also referred to as "DNA strand cleavage reaction"), or a reaction that does not directly involve DNA strand cleavage but is catalyzed by a nucleic acid base conversion enzyme, which is a reaction that cleaves a purine or pyrimidine of a nucleic acid base. Examples of such reactions include a reaction that converts a substituent on a ring to another group or atom (hereinafter also referred to as a "nucleobase conversion reaction") (e.g., base deamination reaction), and a reaction that hydrolyzes an N-glycosidic bond in DNA catalyzed by DNA glycosylase (hereinafter also referred to as an "abasic reaction"). As shown in the examples below, the toxicity of a nucleic acid modification enzyme complex containing a nucleobase conversion enzyme to a host bacterium can be reduced by adding a proteolytic tag to the complex. Therefore, the technology of the present invention can be applied not only to nucleobase conversion enzymes, but also to genome editing using nucleases, which have traditionally been difficult to apply to bacteria due to their strong toxicity. Therefore, the nucleic acid modification enzymes used in the present invention include nucleases, nucleobase conversion enzymes, and DNA From the viewpoint of reducing cytotoxicity, nucleic acid base conversion enzymes and DNA glycosylases are preferred, and the use of these enzymes allows for the formation of a target site. and modifying the targeted site without cleaving at least one strand of the double-stranded DNA. It is possible.
[0018] In the present invention, the term "proteolytic tag" refers to a peptide that is primarily composed of three or more hydrophobic amino acid residues, and that, when added to a genome editing complex, shortens the half-life of the protein compared to a protein that does not have the amino acid residues. Examples of such amino acids include glycine, alanine, valine, leucine, isoleucine, methionine, proline, phenylalanine, and tryptophan. The proteolytic tag of the present invention is not particularly limited as long as it contains any three of these amino acid residues at the C-terminus. The proteolytic tag of the present invention may be a peptide consisting of these three amino acid residues. Peptides in which some or all of the hydrophobic amino acid residues are substituted with serine or threonine are also encompassed within the scope of the proteolytic tag of the present invention. Preferred examples of the three amino acid residues include, but are not limited to, leucine-valine-alanine (LVA), leucine-alanine-alanine (LAA), alanine-alanine-valine (AAV), and the like, which have been shown to be highly effective in Escherichia coli and Pseudomonas putida (Andersen JB et al., Apr. Environ. Microbiol., 64:2240-2246 (1998)). Examples of the three amino acid residues that contain serine include alanine-serine-valine (ASV). Furthermore, proteolytic tags containing these three amino acid residues can be found in databases of tmRNA tag peptides (e.g., tmRDB, http: / / www.ag.auburn.edu / mirror / tmRDB / peptide / peptidephylolist.html). Specifically, the tmRNA tag peptides YAASV (SEQ ID NO: 324), YALAA (SEQ ID NO: 325), ANDENYALAA (SEQ ID NO: 181), and AANDENYALAA (SEQ ID NO: 182) known as tmRNA tag peptides of Escherichia coli, and the tmRNA tag peptides of Bacillus spp. GKQNNLSLAA (SEQ ID NO: 183), GKSNNNFALAA (SEQ ID NO: 184), GKENNNFALAA (SEQ ID NO: 185), GKTNSFNQNVALAA (SEQ ID NO: 186), GKSNQNLALAA (SEQ ID NO: 187), and GKQNYALAA (SEQ ID NO: 188), known as tmRNA tag peptides of Pseudomonas spp., and ANDDNYALAA (SEQ ID NO: 189), ANDDQYGAALAA (SEQ ID NO: 190), ANDENYGQEFALAA (SEQ ID NO: 191), ANDETYGDYALAA (SEQ ID NO: 192), ANDETYGEYALAA (SEQ ID NO: 193), ANDETYGEETYALAA (SEQ ID NO: 194), ANDENYGAEYKLAA (SEQ ID NO: 195), and ANDENYGAQLAA (SEQ ID NO: 196), known as tmRNA tag peptides of Streptococcus spp. Examples of known tag peptides include, but are not limited to, AKNTNSYALAA (SEQ ID NO: 197), AKNTNSYAVAA (SEQ ID NO: 198), AKNNTTYALAA (SEQ ID NO: 199), AKNTNTYALAA (SEQ ID NO: 200), and AKNNTSYALAA (SEQ ID NO: 201). Proteolytic tags typically consist of 3 to 15 amino acid residues, but are not limited to this range. In one embodiment, the proteolytic tag consists of 3 to 5 amino acid residues. Those skilled in the art can select an appropriate proteolytic tag depending on the type of host bacterium, etc. Unless otherwise specified, capital letters in the present specification indicate the single-letter codes of amino acids, and amino acid sequences are written from left to right, from the N-terminus to the C-terminus.
[0019] In the present invention, the term "genome editing complex" refers to a molecular complex having nucleic acid modification activity or expression regulation activity, which is endowed with the ability to recognize a specific nucleotide sequence and comprises a complex in which the nucleic acid sequence recognition module is linked to a proteolytic tag. The term "nucleic acid modification enzyme complex" refers to a molecular complex having nucleic acid modification activity, which is endowed with the ability to recognize a specific nucleotide sequence and comprises a complex in which the nucleic acid sequence recognition module is linked to a nucleic acid modification enzyme and a proteolytic tag. The complex may further be linked to a base excision repair inhibitor. Here, the term "complex" includes not only those composed of multiple molecules, but also those containing the molecules constituting the complex of the present invention within a single molecule, such as fusion proteins. Furthermore, the complex of the present invention also includes molecules or molecular complexes in which a nucleic acid sequence recognition module and a nucleic acid modification enzyme function together, such as restriction enzymes and CRISPR / Cas systems, to which a proteolytic tag is attached. Furthermore, "encoding a complex" includes both encoding each of the molecules constituting the complex and encoding a fusion protein containing the constituent molecules within a single molecule.
[0020] The nuclease used in the present invention is not particularly limited as long as it can catalyze the above reaction, and examples thereof include nucleases (e.g., Cas effector proteins (e.g., Cas9, Cpf1), Endonucleases (e.g., restriction enzymes), exonucleases, recombinases, DNA Examples include gyrase, DNA topoisomerase, and transposase.
[0021] The nucleic acid base conversion enzyme used in the present invention is not particularly limited as long as it can catalyze the above reaction, and examples include deaminases belonging to the nucleic acid / nucleotide deaminase superfamily that catalyze the deamination reaction that converts an amino group to a carbonyl group. Preferred examples include cytidine deaminase, which can convert cytosine or 5-methylcytosine to uracil or thymine, respectively, adenosine deaminase, which can convert adenine to hypoxanthine, and guanosine deaminase, which can convert guanine to xanthine. More preferred examples of cytidine deaminase include activation-induced cytidine deaminase (hereinafter also referred to as AID), an enzyme that introduces mutations into immunoglobulin genes in the adaptive immunity of vertebrates.
[0022] The origin of the nucleobase conversion enzyme is not particularly limited, and examples include lamprey-derived PmCDA1 (Petromyzon marinus cytosine deaminase 1) and mammalian (e.g., human, pig, cow, horse, monkey, etc.)-derived AID (Activation-induced cytidine deaminase; AICDA). For example, the nucleotide sequence and amino acid sequence of PmCDA1 cDNA can be found in GenBank accession numbers EF094822 and ABO15149, and the nucleotide sequence and amino acid sequence of human AID cDNA can be found in GenBank accession numbers NM_020661 and NP_065712, respectively. From the viewpoint of enzymatic activity, PmCDA1 is preferred.
[0023] The DNA glycosylase used in the present invention may be any enzyme capable of catalyzing the above reaction. There are no particular limitations, and thymine DNA glycosylase, oxoguanine glycosylase, alkyl Adenine DNA glycosylase (e.g., yeast 3-methyladenine-DNA glycosylase (MAG1)) The present inventors have previously demonstrated that DNA glycosylases can be used to synthesize unstrained duplexes. Use a DNA glycosylase that has sufficiently low reactivity to unrelaxed DNA. It has been reported that this method can reduce cytotoxicity and efficiently modify target sequences (WO 2016 / 072399). Therefore, it is preferable to use a DNA glycosylase that has sufficiently low reactivity with DNA in an unstrained double helix structure. Examples of such DNA glycosylases include the UNG (uracil-DNA glycosylase) mutants having cytosine-DNA glycosylase (CDG) activity and / or thymine-DNA glycosylase (TDG) activity described in WO 2016 / 072399, and vaccinia virus-derived UDG mutants.
[0024] Specific examples of the UNG mutant include the N222D / L304A double mutant of yeast UNG1 and the N222D / R308E double mutant, N222D / R308C double mutant, Y164A / L304A double mutant, Y164A / R308E double mutant, Y164A / R308C double mutant, Y164G / L304A double mutant, Y164G / R308E double mutant, Y164G / R308C double mutant, N222D / Y164A / L304A triple mutant, N222D / Y164A / R308E triple mutant, N222D / Y164A / R308C triple mutant, N222D / Y164G / L304A triple mutant, N222D / Y164G / R308E triple mutant, N222D / Y164G / R308C triple mutant, etc. When another UNG is used instead of yeast UNG1, a mutant having a similar mutation introduced into the amino acid corresponding to each of the above mutants may be used. Examples of vaccinia virus-derived UDG mutants include the N120D mutant, Y70G mutant, Y70A mutant, N120D / Y70G double mutant, and N120D / Y70A double mutant. Alternatively, the enzyme may be a split enzyme designed such that a DNA glycosylase is split into two fragments, each of which binds to one of the two split nucleic acid sequence recognition modules to form two complexes, and when both complexes are refolded, the nucleic acid sequence recognition module can specifically bind to a target nucleotide sequence, and this specific binding enables the DNA glycosylase to catalyze an abasic reaction. Split enzymes can be designed and produced with reference to, for example, the descriptions in International Publication No. 2016 / 072399, Nat Biotechnol. 33(2): 139-142 (2015), and PNAS 112(10): 2984-2989 (2015).
[0025] In the present invention, "base excision repair" refers to one of the DNA repair mechanisms possessed by living organisms, which repairs damaged bases by enzymatically excising the damaged bases and reconnecting them. The removal of damaged bases is achieved by the enzyme hydrolyzing the N-glycosidic bond of DNA. This is carried out by DNA glycosylase, which is a base debasing enzyme. Apurinic / apyrimidic (AP) sites are processed by downstream enzymes in the base excision repair (BER) pathway, such as AP endonucleases, DNA polymerases, and DNA ligases. Genes or proteins involved in the BER pathway include UNG (NM_003362), SMUG1 (NM_014311), MBD4 (NM_003925), TDG (NM_003211), OGG1 (NM_002542), MYH (NM_012222), NTHL1 (NM_002528), MPG (NM_002434), NEIL1 (NM_024608), NEIL2 (NM_145043), and NEIL3. (NM_018248), APE1 (NM_001641), APE2 (NM_014481), LIG3 (NM_013975), XRCC1 (NM_006297), ADPRT (PARP1) (NM_0016718), ADPRTL2 (PARP2) (NM_005484) and others (the numbers in parentheses indicate the refseq numbers where the base sequence information of each gene (cDNA) has been registered), but are not limited to these.
[0026] In the present invention, the term "base excision repair inhibitor" refers to an inhibitor of any step in the BER pathway. By inhibiting the expression of the molecules involved in the BER pathway, The term "base excision repair inhibitor" refers to a protein that specifically inhibits BER. The inhibitor is not particularly limited as long as it ultimately inhibits BER, but from the viewpoint of efficiency, an inhibitor of DNA glycosylase located upstream of the BER pathway is preferred. Examples of the DNA glycosylase inhibitor used in the present invention include inhibitors of thymine DNA glycosylase, inhibitors of uracil DNA glycosylase, inhibitors of oxoguanine DNA glycosylase, and inhibitors of alkylguanine DNA glycosylase. For example, when cytidine deaminase is used as the nucleic acid modifying enzyme, the U:G or G:U mutation in DNA caused by mutation can be inhibited. To prevent the repair of the match, inhibitors of uracil DNA glycosylase can be used. Suitable.
[0027] Examples of such uracil DNA glycosylase inhibitors include the uracil DNA glycosylase inhibitor (UGI) derived from PBS1, a Bacillus subtilis bacteriophage, and the uracil DNA glycosylase inhibitor (UGI) derived from PBS2, a Bacillus subtilis bacteriophage (Wang, Z., and Mosbaugh, DW (1988) J. Bacteriol. 170, 1082-1091). However, without being limited thereto, any inhibitor of DNA mismatch repair can be used in the present invention. In particular, UGI derived from PBS2 is known to have the effect of making it difficult for mutations, cleavage, and recombination to occur at positions other than C to T in DNA, so it is suitable to use UGI derived from PBS2.
[0028] As mentioned above, in the base excision repair (BER) mechanism, when a base is removed by DNA glycosylase, AP endonuclease nicks the abasic site (AP site), and then exonuclease completely removes the AP site. DNA ligase creates new bases using the bases on the opposite strand as a template, and finally DNA ligase fills in the nick. Repair is completed by the addition of the AP endonuclease. Mutant AP endonucleases that have lost their enzymatic activity but retain the ability to bind to AP sites are known to competitively inhibit BER. Therefore, these mutant AP endonucleases can also be used as inhibitors of base excision repair in the present invention. The origin of the mutant AP endonuclease is not particularly limited, and AP endonucleases derived from Escherichia coli, yeast, mammals (e.g., human, mouse, pig, cow, horse, monkey, etc.), etc., can be used. For example, the amino acid sequence of human Ape1 can be found under UniprotKB No. P27695. Examples of mutant AP endonucleases that have lost enzymatic activity but retain the ability to bind to AP sites include proteins with mutations in the active site or the Mg-binding site (cofactor). For example, in the case of human Ape1, mutations include E96Q, Y171A, Y171F, Y171H, D210N, D210A, and N212A.
[0029] In the present invention, the term "transcriptional regulatory factor" refers to a protein or domain thereof that has the activity of promoting or suppressing the transcription of a target gene. Hereinafter, a factor that has the activity of promoting transcription may be referred to as a "transcriptional activator," and a factor that has the activity of suppressing transcription may be referred to as a "transcriptional repressor."
[0030] The transcription activator used in the present invention is not particularly limited as long as it can promote the transcription of a target gene, and examples thereof include the activation domain of HSV (Herpes simplex virus) VP16, the p65 subunit of NFκB, VP64, VP160, HSF, P300, and Epstein-Barr virus (EB virus) RTA, as well as fusion proteins thereof. The transcription repressor used in the present invention is not particularly limited as long as it can repress the transcription of a target gene, and examples thereof include KRAB, MBD2B, v-ErbA, SID (including a concatemer of SID (SID4X)), MBD2, MBD3, the DNMT family (e.g., DNMT1, DNMT3A, DNMT3B), Rb, MeCP2, ROM2, AtHD2A, and fusion proteins thereof.
[0031] The target nucleotide in double-stranded DNA recognized by the nucleic acid sequence recognition module of the complex of the present invention The peptide sequence is not particularly limited as long as the module can specifically bind to it. The length of the target nucleotide sequence may be any sequence as long as it is sufficient for the nucleic acid sequence recognition module to specifically bind to it, for example, a specific site in the genomic DNA of a mammal. When a mutation is introduced into a gene, the length is 12 nucleotides or more, preferably 15 nucleotides or more, more preferably 17 nucleotides or more, depending on the genome size. There is no particular upper limit to the length, but it is preferably 25 nucleotides or less.
[0032] Examples of the nucleic acid sequence recognition module of the complex of the present invention include a Cas effector protein. Examples of suitable proteins that can be used include, but are not limited to, CRISPR-Cas systems in which at least one DNA cleavage ability of the protein has been inactivated (hereinafter also referred to as "CRISPR-mutant Cas"), zinc finger motifs, TAL effectors, and PPR motifs, as well as fragments containing the DNA-binding domain of proteins capable of specifically binding to DNA, such as restriction enzymes, transcriptional regulators, and RNA polymerases. When a nucleic acid modifying enzyme is used, a CRISPR-Cas system in which a nucleic acid sequence recognition module and a nucleic acid modifying enzyme are integrated may be used (the Cas effector protein of this system maintains both DNA cleavage activities). Preferred examples include CRISPR-mutant Cas, zinc finger motifs, TAL effectors, and PPR motifs.
[0033] The zinc finger motif consists of different zinc finger units (1 fin) of the Cys2His2 type. Zinc finger motifs are composed of three to six linked zinc finger motifs (each of which recognizes approximately three bases) and can recognize a target nucleotide sequence of 9 to 18 bases. Zinc finger motifs can be prepared by known methods such as the modular assembly method (Nat Biotechnol (2002) 20: 135-141), the OPEN method (Mol Cell (2008) 31: 294-301), the CoDA method (Nat Methods (2011) 8: 67-69), or the E. coli one-hybrid method (Nat Biotechnol (2008) 26: 695-701). For details on the preparation of zinc finger motifs, see Japanese Patent No. 4968498.
[0034] TAL effectors have a modular repeat structure consisting of approximately 34 amino acids. The binding stability and base specificity are determined by the 12th and 13th amino acid residues (called RVD) of each module. Since each module is highly independent, it is possible to create a TAL effector specific to a target nucleotide sequence simply by connecting modules. Methods for producing TAL effectors using open resources (such as the REAL method (Curr Protoc Mol Biol (2012) Chapter 12: Unit 12.15), the FLASH method (Nat Biotechnol (2012) 30: 460-465), and the Golden Gate method (Nucleic Acids Res (2011) 39: e82)) have been established, making it relatively easy to design TAL effectors for target nucleotide sequences. For details on the production of TAL effectors, see JP-A-2013-513389.
[0035] The PPR motif consists of 35 amino acids and is composed of a series of PPR motifs that recognize one nucleic acid base. Each motif is configured to recognize a specific nucleotide sequence, and only the 1st, 4th, and ii(-2)th amino acids of each motif recognize the target base. Since there is no dependency on the motif structure and no interference from the flanking motifs, it is possible to create a PPR protein specific to a target nucleotide sequence simply by linking PPR motifs, similar to TAL effectors. For details on the creation of PPR motifs, see JP 2013-128413 A.
[0036] In addition, when fragments of restriction enzymes, transcriptional regulatory factors, RNA polymerases, etc. are used, The DNA binding domains of these proteins are well known, and therefore, for example, Furthermore, fragments that do not have the ability to cleave DNA double strands can be easily designed and constructed.
[0037] When a nucleic acid modifying enzyme is used, any of the above nucleic acid sequence recognition modules can be provided as a fusion protein with the above nucleic acid modifying enzyme and / or base excision repair inhibitor, or a protein binding domain such as an SH3 domain, PDZ domain, GK domain, or GB domain and its binding partner can be fused to the nucleic acid sequence recognition module and the nucleic acid modifying enzyme and / or base excision repair inhibitor, respectively, and provided as a protein complex via the interaction between the domain and its binding partner. Alternatively, an intein can be fused to each of the nucleic acid sequence recognition module and the nucleic acid modifying enzyme and / or base excision repair inhibitor, and the two can be linked by ligation after synthesis of each protein. The proteolytic tag can be attached to a component molecule ( The proteolytic tag may be bound to any of the components (nucleic acid sequence recognition module, nucleic acid modification enzyme, and base excision repair inhibitor), or may be bound to multiple component molecules. Similarly to the above, when a transcriptional regulatory factor is used, the transcriptional regulatory factor may be provided as a fusion protein with the nucleic acid sequence recognition module, or may be bound to the nucleic acid recognition module via the above-mentioned protein binding domain and its binding partner. Similarly to the above, the proteolytic tag may be bound as a fusion protein, or may be bound to the genome editing complex or its component molecules via the above-mentioned protein binding domain and its binding partner. Furthermore, it is preferable that the proteolytic tag be bound to the C-terminus of the genome editing complex or its component molecules.
[0038] The nucleic acid of the present invention can be prepared as a nucleic acid encoding a fusion protein of a nucleic acid sequence recognition module, a proteolytic tag, and, if necessary, a nucleic acid modifying enzyme and / or a base excision repair inhibitor, or a transcriptional regulator, or as a nucleic acid encoding each of these in a form that can form a complex in a host cell after translation into a protein using a binding domain, intein, etc. Here, the nucleic acid may be DNA or RNA. In the case of DNA, it is preferably double-stranded DNA and is provided in the form of an expression vector placed under the control of a promoter functional in the host cell. In the case of RNA, it is preferably single-stranded RNA.
[0039] DNA encoding a nucleic acid sequence recognition module such as a zinc finger motif, a TAL effector, or a PPR motif can be obtained by any of the methods described above for each module. It is possible to encode sequence recognition modules such as restriction enzymes, transcriptional regulators, and RNA polymerases. The DNA to be encoded is selected based on, for example, the cDNA sequence information of the desired portion of the protein. Cloning can be performed by synthesizing an oligo DNA primer that covers the region encoding the protein (the portion containing the DNA-binding domain), and amplifying it by RT-PCR using total RNA or an mRNA fraction prepared from cells that produce the protein as a template. DNA encoding nucleic acid modifying enzymes and inhibitors of base excision repair are also useful. Based on the cDNA sequence information of the enzyme, an oligo DNA primer is synthesized and the cDNA is cloned from the cell that produces the enzyme. The total RNA or mRNA fraction prepared by this method was used as a template and amplified by RT-PCR. For example, the DNA encoding UGI derived from PBS2 can be cloned by RT-PCR from PBS2-derived mRNA using appropriate primers designed upstream and downstream of the CDS based on the DNA sequence registered in the NCBI / GenBank database (accession no. J04434). The cloned DNA may be used as is, or optionally digested with restriction enzymes or digested with appropriate restriction enzymes. A linker (e.g., GS linker, GGGAR linker, etc.), a spacer (e.g., FLAG sequence, etc.), and / or a nuclear localization signal (NLS) (if the target double-stranded DNA is mitochondrial or chloroplast DNA, By adding a gene encoding a protein, DNA encoding the protein can be prepared. Furthermore, it is ligated with DNA encoding a nucleic acid sequence recognition module to form a fusion protein. DNA encoding the protein can be prepared.
[0040] The DNA encoding the genome editing complex of the present invention can be prepared by chemically synthesizing a DNA strand, or by synthesizing partially overlapping short oligo DNA strands using PCR or Gibson Assembly. By connecting the fragments, it is possible to construct DNA encoding the entire fragment. The advantages of constructing full-length DNA by synthesis or in combination with PCR or Gibson Assembly are The advantage of this method is that the codons used can be designed over the entire length of the CDS to suit the host into which the DNA is introduced. When expressing heterologous DNA, converting the DNA sequence to codons frequently used in the host organism is expected to increase the amount of protein expressed. Data on codon usage in the host can be obtained from, for example, the genetic code usage database published on the website of the Kazusa DNA Research Institute (http: / / www.kazusa.or.jp / codon / index.html), or literature listing codon usage in each host can be referenced. By referring to the obtained data and the DNA sequence to be introduced, codons used in the DNA sequence that are less frequently used in the host can be converted to codons that code for the same amino acid and are frequently used.
[0041] An expression vector containing DNA encoding the conjugate of the present invention can be produced, for example, by ligating the DNA downstream of a promoter in an appropriate expression vector. Expression vectors that can be used include plasmids derived from Escherichia coli (e.g., pBR322, pBR325, pUC12, pUC13); plasmids derived from Bacillus subtilis (e.g., pUB110, pTP5, pC194); yeast-derived plasmids (e.g., pSH19, pSH15); insect cell expression plasmids (e.g., pFast-Bac); animal cell expression plasmids (e.g., pA1-11, pXT1, pRc / CMV, pRc / RSV, pcDNAI / Neo); bacteriophages such as λ phage; insect virus vectors such as baculovirus (e.g., BmNPV, AcNPV); and animal virus vectors such as retrovirus, vaccinia virus, and adenovirus. Any promoter may be used as long as it is appropriate for the host used to express the gene. When a nuclease is used as the nucleic acid modifying enzyme, the survival rate of the host cells may be significantly reduced due to toxicity, so it is desirable to use an inducible promoter to increase the number of cells before the start of induction. On the other hand, when a nucleobase converting enzyme and a DNA glycosylase are used as the nucleic acid modifying enzyme, or when no nucleic acid modifying enzyme is used, In this case, sufficient cell growth can be achieved even when the complex of the present invention is expressed, and therefore, a constitutive promoter can also be used without any restrictions. For example, when the host is an animal cell, the SRα promoter, the SV40 promoter, the LTR promoter, Motor, CMV (cytomegalovirus) promoter, RSV (Rous sarcoma virus) promoter, MoMuLV (Moloney murine leukemia virus) LTR, HSV-TK (herpes simplex virus) Thymidine kinase) promoters are used. Among them, CMV promoter, SR promoter, etc. The α promoter is preferred. When the host is E. coli, J23 series promoters (e.g., J23119 promoter), trp promoter, lac promoter, recA promoter, λP L promoter, lpp promoter, T7 promoter, etc. are preferred. When the host is a bacterium of the genus Bacillus, the SPO1 promoter, SPO2 promoter, penP promoter, etc. are preferred. When the host is yeast, the Gal1 / 10 promoter, PHO5 promoter, PGK promoter, GAP promoter, ADH promoter, etc. are preferred. When the host is an insect cell, the polyhedrin promoter, P10 promoter, etc. are preferred. It's nice. When the host is a plant cell, the CaMV35S promoter, CaMV19S promoter, NOS promoter, A rotor or the like is preferred.
[0042] In addition to the above, expression vectors that contain, if desired, an enhancer, a splicing signal, a terminator, a polyA addition signal, a selection marker such as a drug resistance gene or an auxotrophy complementing gene, a replication origin, etc. can be used.
[0043] RNA encoding the complex of the present invention can be prepared, for example, by using a vector containing DNA encoding each protein as a template and transcribing it into mRNA in a known in vitro transcription system.
[0044] The host bacterium used to replicate the nucleic acid of the present invention is not particularly limited as long as it has a proteolytic system using tmRNA (ssrA). Examples of such bacteria include bacteria of the genus Escherichia, Bacillus, Pseudomonas (e.g., Pseudomonas putida), Streptococcus (e.g., Streptococcus), Streptomyces, Staphylococcus, Yersinia, Acinetobacter, Klebsiella, Bordetella, Lactococcus, Neisseria, Bacteria of the genus Aeromonas, Franciecella, Corynebacterium, Citrobacter, Chlamydia, Haemophilus, Brucella, Mycobacterium, Legionella, Rhodococcus, Pseudomonas, Helicobacter, Salmonella, Staphylococcus, Vibrio, and Erysipelothrix can be used. Examples of Escherichia bacteria that can be used include Escherichia coli K12 DH1 [Proc. Natl. Acad. Sci. USA, 60, 160 (1968)], Escherichia coli JM103 [Nucleic Acids Research, 9, 309 (1981)], Escherichia coli JA221 [Journal of Molecular Biology, 120, 517 (1978)], Escherichia coli HB101 [Journal of Molecular Biology, 41, 459 (1969)], Escherichia coli C600 [Genetics, 39, 440 (1954)], Escherichia coli DH5α, and Escherichia coli BW25113. Examples of Bacillus bacteria that can be used include Bacillus subtilis MI114 [Gene, 24, 255 (1983)] and Bacillus subtilis 207-21 [Journal of Biochemistry, 95, 87 (1984)].
[0045] When a nucleic acid base conversion enzyme or a DNA glycosylase is used as the nucleic acid modification enzyme, the nucleic acid modification The enzyme and / or base excision repair inhibitor is provided as a complex with the mutant Cas by a method similar to the linkage mode with the zinc finger or the like. And / or the base excision repair inhibitor and the mutant Cas can be bound using an RNA scaffold formed by RNA aptamers such as MS2F6 and PP7 and their binding proteins. The guide RNA forms a complementary strand to the target nucleotide sequence, and the mutant Cas is recruited to the following tracrRNA, which recognizes the DNA cleavage site recognition sequence PAM (protospacer adjacent motif). (In the case of SpCas9, the PAM is a trinucleotide NGG (where N is any base), theoretically allowing targeting anywhere in the genome.) However, it is unable to cleave one or both DNA strands. Instead, the action of a nucleobase conversion enzyme or DNA glycosylase linked to the mutant Cas results in nucleobase conversion or abasic site at the targeted site (which can be adjusted to any range of several hundred bases, including all or part of the target nucleotide sequence). This creates a mismatch (e.g., when cytidine deaminases such as PmCDA1 or AID are used as nucleobase conversion enzymes, cytosine on the sense or antisense strand at the targeted site is converted to uracil, resulting in a U:G or G:U mismatch) or an abasic site (AP site). The cell's BER system attempts to repair this by introducing various mutations. For example, if a mismatch or a base is not repaired correctly, the base on the opposite strand may be repaired to pair with the base on the converted strand (TA or AT in the above example), or the repair may result in further substitution with another nucleotide (e.g., U → A, G), or the deletion or insertion of one to several dozen bases, resulting in the introduction of various mutations. The combined use of a base excision repair inhibitor inhibits the intracellular BER mechanism, increasing the frequency of repair errors and improving the efficiency of mutagenesis.
[0046] The efficiency of producing zinc finger motifs that specifically bind to target nucleotide sequences is not high, and the selection of zinc fingers with high binding specificity is complicated, so it is not easy to produce a large number of zinc finger motifs that actually function. Effector and PPR motifs have a higher degree of autonomy in target nucleic acid sequence recognition than zinc finger motifs. Although this method offers a high degree of flexibility, it requires the design and construction of a large protein each time according to the target nucleotide sequence, which leaves problems in terms of efficiency. In contrast, the CRISPR-Cas system recognizes the sequence of a desired double-stranded DNA using a guide RNA that is complementary to the target nucleotide sequence, so any sequence can be targeted simply by synthesizing an oligo DNA that can specifically hybridize with the target nucleotide sequence. Therefore, in a more preferred embodiment of the present invention, a CRISPR-Cas system in which both DNA cleavage activities are maintained, or a CRISPR-Cas system in which the DNA cleavage activities of only one or both Cass are inactivated (CRISPR-mutant Cas) is used as the nucleic acid sequence recognition module.
[0047] The nucleic acid sequence recognition module of the present invention using CRISPR-mutant Cas is provided as a complex of CRISPR-RNA (crRNA) containing a sequence complementary to the target nucleotide sequence, and optionally a trans-activating RNA (tracrRNA) required for recruiting the mutant Cas effector protein (if tracrRNA is required, it can be provided as a chimeric RNA with crRNA), and the mutant Cas effector protein. The RNA molecule consisting of crRNA alone or a chimeric RNA of crRNA and tracrRNA, which is combined with the mutant Cas effector protein to form the nucleic acid sequence recognition module, is collectively referred to as a "guide RNA." The same applies when using a CRISPR / Cas system without mutations.
[0048] The Cas effector protein used in the present invention forms a complex with a guide RNA and binds to the target nucleotide sequence in the target gene and its adjacent protospacer adjacent motif (PAM). There are no particular limitations as long as it can recognize and bind to the Cas9 gene, but Cas9 or Cpf1 is preferred. Examples of Cas9 include Cas9 derived from Streptococcus pyogenes (SpCas9; PAM sequence NGG (N is A, G, T, or C; the same applies below)), Cas9 derived from Streptococcus thermophilus ... Examples of Cas9 include, but are not limited to, Cas9 derived from Streptococcus thermophilus (StCas9; PAM sequence: NNAGAAW) and Cas9 derived from Neisseria meningitidis (NmCas9; PAM sequence: NNNNGATT). SpCas9, which has fewer PAM constraints (it is essentially two bases long and can theoretically target almost anywhere in the genome), is preferred. Examples of Cpf1 include, but are not limited to, Cpf1 derived from Francisella novicida (FnCpf1; PAM sequence: NTT), Cpf1 derived from Acidaminococcus sp. (AsCpf1; PAM sequence: NTTT), and Cpf1 derived from Lachnospiraceae bacteria (LbCpf1; PAM sequence: NTTT). The mutant Cas effector proteins (sometimes abbreviated as "mutant Cas") used in the present invention can be either Cas effector proteins that have lost their ability to cleave both strands of double-stranded DNA, or Cas effector proteins that have lost their ability to cleave only one strand but retain nickase activity. For example, in the case of SpCas9, the D10A mutant, in which the Asp residue at position 10 is replaced by an Ala residue and the effector protein lacks the ability to cleave the strand opposite the strand that is complementary to the guide RNA (and therefore retains nickase activity against the strand that is complementary to the guide RNA), or the H840A mutant, in which the His residue at position 840 is replaced by an Ala residue and the effector protein lacks the ability to cleave the strand that is complementary to the guide RNA (and therefore retains nickase activity against the strand that is complementary to the guide RNA), or a double mutant thereof (dCas9) can be used. In the case of FnCpf1, mutants lacking the ability to cleave both strands can be used, such as those in which the Asp residue at position 917 is replaced by an Ala residue (D917A) or the Glu residue at position 1006 is replaced by an Ala residue (E1006A). Other mutant Cass can also be used as long as they lack the ability to cleave at least one strand of double-stranded DNA.
[0049] DNA encoding Cas effector proteins (including mutant Cas, the same applies below) is The enzyme can be isolated by a method similar to that described above for DNA encoding an inhibitor of excision repair. Mutant Cas can be cloned from cells that produce the mutant Cas. The DNA encoding the Cas was subjected to a site-directed mutagenesis method known per se to generate a mutant that has a significant effect on DNA cleavage activity. amino acid residues at key sites (e.g., in the case of SpCas9, the 10th Asp residue and the 840th His residue) In the case of FnCpf1, examples include the 917th Asp residue and the 1006th Glu residue, but are not limited to these. The amino acid sequence can be obtained by introducing a mutation to replace the amino acid sequence (not specified) with another amino acid. Alternatively, DNA encoding a Cas effector protein can be constructed as DNA with codon usage suitable for expression in the host cell to be used by chemical synthesis or in combination with PCR or Gibson assembly using methods similar to those described above for DNA encoding a nucleic acid sequence recognition module or DNA encoding a DNA glycosylase.
[0050] The resulting Cas effector proteins, nucleic acid modifying enzymes, and base excision repair inhibitors The DNA encoding the transcriptional regulatory factor may be expressed in the same manner as above depending on the target cell. It can be inserted downstream of the promoter of the target gene.
[0051] On the other hand, the DNA encoding the guide RNA contains a crRNA sequence (for example, in the case of recruiting FnCpf1 as a Cas effector protein, SEQ ID NO: 19; AAUU) that contains a nucleotide sequence complementary to the target nucleotide sequence (also referred to as a "targeting sequence" herein). UCUAC UGUU GUAGAU-containing crRNA can be used, and the underlined sequences form base pairs to form a stem-loop structure. code sequence, or the crRNA coding sequence and, if necessary, a known tracrRNA coding sequence (e.g., For example, the tracrRNA coding sequence for recruiting Cas9 as a Cas effector protein. The oligo DNA sequence can be designed by linking the above sequence to the above sequence (gttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgcttttttt; SEQ ID NO: 18), and chemically synthesized using a DNA / RNA synthesizer. Here, "target strand" refers to the strand of the target nucleotide sequence that hybridizes with crRNA. The opposite strand that becomes single-stranded upon hybridization between the target strand and crRNA is The nucleotide base conversion reaction is generally It is assumed that this usually occurs on the single-stranded non-target strand. Therefore, when expressing the target nucleotide sequence on one strand (for example, when expressing the PAM sequence or the target nucleotide sequence), When expressing the positional relationship between the sequence and PAM, the sequence should be represented by the sequence of the non-target strand.
[0052] The length of the targeting sequence is not particularly limited as long as it can specifically bind to the target nucleotide sequence, but it can be, for example, 15 to 30 nucleotides, preferably 18 to 25 nucleotides. The selection of the target nucleotide sequence is limited by the presence of a PAM adjacent to the 3' (in the case of Cas9) or 5' (in the case of Cpf1) side of the sequence. However, as demonstrated in the Examples below, in the system of the present invention combining CRISPR-mutagenized Cas9 and cytidine deaminase, as the target nucleotide sequence becomes longer, the easily substitutable C shifts toward the 5' end. Therefore, by appropriately selecting the length of the target nucleotide sequence (its complementary strand, the targeting sequence), the position of the base at which mutations can be introduced can be shifted. This at least partially relieves the constraints imposed by the PAM (NGG in the case of SpCas9), further increasing the flexibility of mutagenesis.
[0053] The design of the targeting sequence can be performed, for example, using Cas9 as the Cas effector protein. If available, public guide RNA design websites (e.g., CRISPR Design Tool, CRISPRdirect) This can be done by using a list of 20-mer sequences adjacent to the PAM (e.g., NGG in the case of SpCas9) on the 3' side from the CDS sequence of the target gene, and selecting a sequence that will cause an amino acid change in the protein encoded by the target gene when C is converted to T within 7 nucleotides from the 5' end to the 3' end. In addition, it is also possible to select a target sequence of a length other than 20-mer. When using a target sequence, an appropriate sequence can be selected. From these candidates, a candidate sequence with a small number of off-target sites in the target host genome can be used as the target sequence. If there is no function to search for target sites, for example, a Blast search can be performed against the host genome for 8 to 12 nucleotides on the 3' side of the candidate sequence (a seed sequence with high discriminatory power for the target nucleotide sequence). By applying this, off-target sites can be searched for.
[0054] The DNA encoding the guide RNA can also be inserted into the same expression vector as above, but as the promoter, a pol III promoter (e.g., SNR6, SNR52, SCR1, RPR1, U3, U6, H1 promoter, etc.) and a terminator (e.g., poly T sequence (T6 sequence, etc.)) can be used. It is preferable that:
[0055] The DNA encoding the guide RNA (crRNA or crRNA-tracrRNA chimera) is An oligo-RNA sequence can be designed that combines a sequence complementary to the target strand of the sequence with a known tracrRNA sequence (to recruit Cas9) or a direct repeat sequence of crRNA (to recruit Cpf1), and then chemically synthesized using a DNA / RNA synthesizer.
[0056] 2. Method for modifying targeted sites in double-stranded DNA of a host bacterium The complex or nucleic acid of the present invention described in 1. is introduced into a host, particularly a bacterium, and the host is cultured to modify the targeted site of the double-stranded DNA of the host, or to modify the targeted site. The expression of genes encoded in double-stranded DNA in the vicinity of the site can be regulated. In another embodiment, the nucleic acid modifying enzyme complex is contacted with double-stranded DNA of the host bacterium, and the target A method for modifying a targeted site in bacterial double-stranded DNA (hereinafter also referred to as the "modification method of the present invention"), comprising the steps of converting or deleting one or more nucleotides at the targeted site into one or more other nucleotides, or inserting one or more nucleotides at the targeted site. ) is provided. A nucleic acid base conversion enzyme or a DNA glycosylase is used as the nucleic acid modification enzyme. This allows cleavage of at least one strand of double-stranded DNA at the targeted site. In yet another embodiment, the complex of the present invention is contacted with double-stranded DNA of a host bacterium, and the target site is modified by the complex. Methods for regulating the transcription of a gene are provided.
[0057] The contact of the complex of the present invention with double-stranded DNA results in the formation of a target double-stranded DNA (e.g., genomic DNA). This is carried out by introducing the complex or a nucleic acid encoding it into a bacterium. Considering the efficiency of introduction and expression, it is preferable to introduce the genome editing complex into the bacterium in the form of a nucleic acid encoding it, rather than the complex itself, and express the complex in the bacterium.
[0058] Bacteria used in the modification method of the present invention include the same bacteria as those used for nucleic acid replication in 1.
[0059] Introduction of the expression vector can be carried out according to known methods (e.g., lysozyme method, competent method, PEG method, CaCl2 co-precipitation method, electroporation method, microinjection method, particle gun method, lipofection method, Agrobacterium method, etc.) depending on the type of bacterium. E. coli can be transformed according to the methods described in, for example, Proc. Natl. Acad. Sci. USA, 69, 2110 (1972) or Gene, 17, 107 (1982). A vector can be introduced into a bacterium of the genus Bacillus according to the method described in, for example, Molecular & General Genetics, 168, 111 (1979).
[0060] Bacteria into which a vector has been introduced can be cultured according to known methods depending on the type of bacteria.
[0061] For example, when culturing Escherichia coli or Bacillus bacteria, a liquid medium is preferred. The medium preferably contains a carbon source, a nitrogen source, inorganic substances, and the like necessary for the growth of the transformant. Examples of carbon sources include glucose, dextrin, soluble starch, and sucrose; examples of nitrogen sources include inorganic or organic substances such as ammonium salts, nitrates, corn steep liquor, peptone, casein, meat extract, soybean meal, and potato extract; and examples of inorganic substances include calcium chloride, sodium dihydrogen phosphate, and magnesium chloride. The medium may also contain yeast extract, vitamins, growth-promoting factors, and the like. The pH of the medium is preferably about 5 to about 8.
[0062] As a medium for culturing E. coli, for example, M9 medium containing glucose and casamino acids (Journal of Experiments in Molecular Genetics, 431-433, Cold Spring Harbor Laboratory) is used. Laboratory, New York 1972) is preferred. For this purpose, an agent such as 3β-indolylacrylic acid may be added to the medium. E. coli is usually cultured at about 15 to about 43° C. If necessary, aeration or stirring may be performed. Bacteria of the genus Bacillus are usually cultured at about 30 to about 40° C. If necessary, aeration or stirring may be performed. Furthermore, the inventors have confirmed that when PmCDA1 is used as a nucleic acid modification enzyme, the efficiency of mutagenesis can be increased by culturing animal or plant cells at a lower temperature than usual (e.g., 20 to 26°C, preferably about 25°C), and it is also preferable to culture bacteria at the above-mentioned low temperatures.
[0063] The RNA encoding the complex of the present invention is introduced into the host bacterium by microinjection. This can be done by lipofection, etc. The RNA introduction can be done once or repeatedly multiple times (for example, 2 to 5 times) at appropriate intervals.
[0064] The inventors have also confirmed using budding yeast that by creating sequence recognition modules for multiple adjacent target nucleotide sequences and using them simultaneously, the efficiency of mutagenesis is significantly increased compared to targeting a single nucleotide sequence, and a similar effect can be expected in bacteria. This effect is achieved even when both target nucleotide sequences partially overlap, and even when the two are about 600 bp apart. The nucleotide sequences may be in the same direction (the target strand is the same strand) or in opposite directions (both strands of double-stranded DNA). This can occur in either case (the target strand being the target strand).
[0065] In a preferred embodiment, the method of the present invention detects six positions in the genomic DNA of a bacterium. It has been demonstrated that simultaneous mutations can be introduced (Figure 10), resulting in extremely high mutation introduction efficiency. Therefore, the method for modifying a genome sequence or the method for regulating expression of a target gene of the present invention allows for target modification of multiple DNA regions at completely different locations, or for regulating the expression of multiple target genes. Therefore, in a preferred embodiment of the present invention, two or more nucleic acid sequence recognition modules can be used, each specifically binding to a different target nucleotide sequence (which may be within a single target gene or within two or more different target genes. These target genes may be located on the same chromosome or plasmid, or on separate chromosomes or plasmids). In this case, each of these nucleic acid sequence recognition modules forms a complex with a proteolytic tag attached, together with a nucleic acid modification enzyme and / or a base excision repair inhibitor, or a transcriptional regulator. Here, the nucleic acid modification enzyme, base excision repair inhibitor, and transcriptional regulator may be the same. For example, when the CRISPR-Cas system is used as the nucleic acid sequence recognition module, a common complex (including a fusion protein) of the Cas effector protein with a nucleic acid modification enzyme and / or a base excision repair inhibitor, or a transcriptional regulator can be used, and two or more chimeric RNAs of two or more guide RNAs that form complementary strands with different target nucleotide sequences and the tracrRNA can be prepared and used as the guide RNA-tracrRNA. On the other hand, when a zinc finger motif or TAL effector is used as the nucleic acid sequence recognition module, for example, a nucleic acid modification enzyme and / or a base excision repair inhibitor, or a transcriptional regulator can be fused to each nucleic acid sequence recognition module that specifically binds to a different target nucleotide.
[0066] In order to express the complex of the present invention in a host bacterium, an expression vector containing DNA encoding the complex is introduced into the host bacterium as described above. To adequately regulate the expression of the target gene, it is desirable to maintain the expression of the genome editing complex at a certain level for a certain period of time or longer. From this perspective, it is certain that the expression vector is incorporated into the host genome. However, since continuous expression of the genome editing complex increases the risk of off-target cleavage, it is necessary to remove it promptly after successful mutagenesis. It is preferable to remove the DNA that has been integrated into the host genome. Examples of such methods include a method using the Cre-loxP system and a method using transposons.
[0067] Alternatively, by transiently expressing the complex of the present invention in a host bacterium for the period required for a nucleic acid reaction to occur at the desired time and for the modification at the targeted site to be fixed, host genome editing can be efficiently achieved while avoiding the risk of off-target cleavage. The period required for the nucleic acid modification reaction to occur and for the modification at the targeted site to be fixed varies depending on the type of host bacterium, culture conditions, etc., but is thought to be approximately 2-3 days, as at least several generations of cell division are required. Those skilled in the art can appropriately determine an appropriate expression induction period based on the culture conditions used, etc. The expression induction period of a nucleic acid encoding the complex of the present invention may be extended beyond the above-mentioned "period required for the modification at the targeted site to be fixed," as long as it does not cause side effects in the host bacterium.
[0068] As a means for transiently expressing the complex of the present invention at a desired time for a desired period, nucleic acids encoding the complex (in the mutant CRISPR-Cas system, DNA encoding a guide RNA, DNA encoding a Cas effector protein, and, if necessary, DNA encoding a nucleic acid modifying enzyme and / or a base excision repair inhibitor, or a transcriptional regulator) can be introduced into the cells at a desired time for a desired period. One example is a method of preparing a construct (expression vector) containing the complex of the present invention in a controllable form and introducing it into a host. A specific example of a "controllable expression period" is one in which the nucleic acid encoding the complex of the present invention is placed under the control of an inducible regulatory region. The "inducible regulatory region" is not particularly limited, but examples include an operon consisting of a temperature-sensitive (ts) mutant repressor and an operator controlled by the repressor. Examples of ts mutant repressors include, but are not limited to, the ts mutant of the cI repressor derived from λ phage. In the case of the λ phage cI repressor (ts), at temperatures below 30°C (e.g., 28°C), it binds to the operator and represses downstream gene expression, but at temperatures above 37°C (e.g., 42°C), it dissociates from the operator, thereby inducing gene expression. Therefore, host bacteria into which nucleic acid encoding the complex of the present invention has been introduced are typically cultured at 30°C or below, and then the temperature is raised to 37°C or above at an appropriate time and cultured for a certain period of time to allow the nucleic acid conversion reaction to occur.After a mutation has been introduced into the target gene, the temperature is quickly returned to 30°C or below, thereby minimizing the period during which target gene expression is suppressed.Even when targeting a gene essential for the host cell, efficient editing can be achieved while minimizing side effects. When a temperature-sensitive mutation is used, for example, a temperature-sensitive mutant of a protein required for the autonomous replication of a vector is incorporated into a vector containing DNA encoding the complex of the present invention. After expression of the complex, autonomous replication is quickly disabled, and the vector is naturally lost with cell division. Examples of such temperature-sensitive mutant proteins include, but are not limited to, temperature-sensitive mutants of Rep101 ori, which is necessary for replication of pSC101 ori. (ts) acts on the pSC101 ori at temperatures below 30°C (e.g., 28°C) to enable autonomous replication of the plasmid. However, at temperatures above 37°C (e.g., 42°C), the function is lost and the plasmid becomes unable to autonomously replicate. Therefore, by using the complex of the present invention in combination with the cI repressor (ts) of the λ phage, transient expression of the complex of the present invention and removal of the plasmid can be achieved simultaneously.
[0069] Alternatively, transient expression of the complex can be achieved by introducing DNA encoding the complex of the present invention into a host bacterium under the control of an inducible promoter (e.g., lac promoter (induced by IPTG), cspA promoter (induced by cold shock), araBAD promoter (induced by arabinose), etc.), adding an inducer to the culture medium (or removing it from the culture medium) at an appropriate time to induce expression of the complex, culturing the bacterium for a certain period of time to carry out a nucleic acid modification reaction, etc., and stopping the induction of expression after a mutation has been introduced into the target gene.
[0070] The present invention will be described below with reference to examples, although the present invention is not limited to these examples. [Example]
[0071] In the examples described below, experiments were carried out as follows. <Strains, Plasmids, Primers, and Targeting gRNA Design> E. coli strain DH5α ((F - endA1 supE44 thi-1 recA1 relA1 gyrA96 deoR phoA Φ80dlacZ ΔM15 Δ(lacZYA-argF)U169, hsdR17 (rK - , mK + ), λ - ) (TaKaRa-Bio), BW25113(lacI + rrnB T14 ΔlacZ WJ16 hsdR514 ΔaraBAD AH33 ΔrhaBAD LD78 rph-1 Δ(araB-D)567 Δ(rhaD-B)568ΔlacZ4787(::rrnB-3) hsdR514 rph-1) and Top10(F- mcrA Δ(mrr-hsdRMS-mcrBC) φ80lacZΔM15 ΔlacX74 nupG recA1 araD139 Δ(ara-leu)7697 galE15 galK16 rpsL(Str R ) endA1 λ -) (Invitrogen) was used. The plasmids and primers used in the examples are listed in Tables 1 and 2, respectively. The oligo DNA pair for constructing the targeting gRNA vector was designed as follows: 5'-tagc-(target sequence)-3' and 5'-aaac-(reverse complementary sequence of the target sequence)-3'.
[0072] [Table 1-1]
[0073] [Table 1-2]
[0074] [Table 2-1]
[0075] [Table 2-2]
[0076] <Plasmid construction> pCas9 and pCRISPR plasmids were obtained via Addgene from the Marraffini laboratory (8). The nuclease-deficient Cas9 (nCas9) (D10A or H840A) and nuclease-deficient Cas9 (dCas9) (D10A and H840A) (SEQ ID NOs: 1 and 2) (Jinek, M. et al., Science 337, 816-822 (2012)) were generated by PCR. PmCDA1 (SEQ ID NOs: 3 and 4) was fused to the C-terminus of nCas9 or dCas9 using a 121-amino acid peptide linker (SEQ ID NOs: 5 and 6) (Figure 1).
[0077] The plasmid pScI_dCas9-PmCDA1_J23119-sgRNA, which contains an sgRNA unit (SEQ ID NO: 15) driven by the artificial constitutive promoter J23119 (BBa_J23119 in the registry for standard biological parts) (http: / / parts.igem.org / Part:BBa_J23119) (SEQ ID NO: 16), was amplified by PCR using primers p346 / p426. The sgRNA expression unit contains two BsaI restriction enzyme sites for insertion of target sequences. A pair of oligo DNAs containing the target sgRNA sequence was annealed and ligated into BsaI-digested pScI_dCas9-PmCDA1_sgRNA.
[0078] pScI and pScI_dCas9 carry only the lambda operator and operator-dCas9, respectively. pScI_dCas9-PmCDA1 carries the dCas9-PmCDA1 gene. C-terminus of the dCas9-PmCDA1 gene A degradation tag (LVA tag) and UGI gene were added to the plasmid pScI_dCas9-PmCDA1-LVA and pScI_dCas9-PmCDA1-UGI-LVA, respectively.
[0079] The vector plasmid pTAKN-2 contains the pMB1 replication origin corresponding to pSC101. The sgRNA unit containing -J23119 was excised from the synthetic oligonucleotides using EcoRI-HindIII and ligated into the cloning vector pTAKN2. Plasmids carrying three tandem target sequences (pSBP804, galK_10-galK_11-gal_13; pSBP806, galK_2-xylB_1-manA_1; pSBP808, pta_1-adhE_3-tpiA_2) were constructed by Golden Gate assembly of PCR products using the nucleotide sequence library (Engler, C. et al., PLoS One 4, (2009).). A plasmid carrying six different target sequences (pSBP80608) was constructed using Gibson assembly of PCR products amplified from pSBP808 (pta_1-adhE_3-tpiA_2 tandem sequence) with primers p597 / 598 and from pSBP806 (vector and galK_2-xylB_1-manA_1 tandem sequence) with primers p599 / p600. For the IS-editing plasmid, the sgRNA expression units were arranged in tandem in the order IS1, IS2, IS3, and IS5.
[0080] Mutation induction assay DH5α or BW25113 cells chemically transformed with the desired plasmid were pre-cultured in 1 mL of SOC medium (2% Bacto Tryptone, 0.5% yeast extract, 10 mM NaCl, 2.5 mM KCl, 1 mM MgSO4, and 20 mM glucose). After 2–3 h of incubation at 28°C, the cell culture was diluted 1:10 into 1 mL of Luria-Bertani (LB) medium or terrific broth (TB), supplemented with antibiotics (chloramphenicol (25 μg / mL) and / or kanamycin (30 μg / mL)) as needed. The cells were grown overnight at 28°C and 100 rpm using a maximizer (TAITEC). The next day, the cell culture was again diluted 1:10 into 1 mL of medium and cultured at 37°C for 6 h for induction, followed by overnight incubation at 28°C. The cell cultures were then serially diluted and spotted onto LB or TB agar plates supplemented with the appropriate antibiotic and incubated overnight at 28°C to allow the formation of single colonies.
[0081] For positive selection of the galK gene disruption, cells were grown in M63 minimal medium (2 g / L (NH4)2SO4, 13.6 g / L KH2PO4, 0.5 mg / L FeSO4-7H2O, 1 mM MgSO4, 0.1 mM CaCl2, and 10 μg / ml thiamine) containing 0.2% glycerol and 2-deoxy-galctose (2-DOG) (Warming, S. et al., Nucleic Acids Res. 33, 1-12 (2005).). rifampicin-resistant mutations in the rpoB gene For selection of β-actin, cells were grown in LB medium containing 50 μg / ml rifampicin. For sequence analysis, colonies were randomly picked and amplified by PCR using appropriate primers. Direct amplification was performed and the DNA was analyzed by Sanger sequencing using a 3130XL Genetic Analyzer (Applied Biosystems). Statistical analysis was performed using the t-test software (Microsoft).
[0082] <Whole genome sequencing> Each expression construct (dCas9, dCas9-PmCDA1, dCas9-PmCDA1-LVA-UGI, and rpoB_1 target) Pre-culture BW25113 cells carrying the target (dCas9-PmCDA1) overnight and dilute them at 1:10 in 1 mL of LB medium. The cells were diluted and grown for 6 hours at 37°C for induction, followed by overnight incubation at 28°C. The cells were plated on rifampicin-containing plates to isolate single colonies. Three independent colonies were inoculated onto TB medium. Genomic DNA was extracted using the Wizard Genomic DNA Purification Kit (Promega) and then sonicated using the Bioruptor UCD-200 TS Sonication System (Diagenote) to obtain fragments with a size distribution of 500–1000 bp. A genomic DNA library was prepared using the NEBNext Ultra DNA Library Prep Kit from Illumina (New England Biolabs) and labeled with Dual Index Primers. Size selection of the reads was performed using an Agencourt AMPure XP (Beckman Coulter) to obtain labeled fragments ranging in length from 600 to 800 bp. Size distribution was assessed using an Agilent 2100 Bioanalyzer system (Agilent Technologies). DNA was quantified using a Qibit HS dsDNA HS Assay Kit and a fluorometer (Thermo Fisher Scientific). Sequencing was performed using a MiSeq sequencing system (Illumina) and MiSeq Reagent Kit v3 to obtain a read length of 2 × 300 bp, which is expected to provide approximately 20-fold coverage of the genome size. Data analysis was performed using CLC Genomic Workbench 9.0 (CLC bio). Sequencing reads were paired, and overlapping reads within a read pair were merged and trimmed based on a quality limit of 0.01 with a maximum ambiguity of 2. The following settings were used: Masking mode = no masking, Mismatch cost = 2, Reads were mapped to the E. coli BW25113 reference genome with the following settings: Insertion cost = 3, Deletion cost = 3, Length fraction = 0.5, Similarity fraction = 0.8, Global alignment = No, Auto-detect paired distances = Yes, Nonspecific match handling = ignore. Local realignment was performed with default settings (Realign unaligned ends = Yes, Multi-pass realignment = 2). Variant calling was performed with the following settings: Ignore positions with coverage = 1,000,000, Ignore broken pairs = Yes, Ignore Nonspecific matches = Reads, Minimum coverage = 5, Minimum count = 2, Minimum frequency = 50%, Base quality filter = No, Read detection filter = No, Relative read direction filter = Yes, Significance = 1%, Read position filter = No, Remove pyro-error variants = No). Output files were sorted using Excel (Microsoft).
[0083] Deep sequencing DH5α cells expressing dCas9-PmCDA1 or dCas9-PmCDA1-UGI-LVA with gRNAs targeting the galK, gsiA, ycbF, or yfiH genes were incubated overnight, diluted 1:10 in 1 mL of LB medium, and grown at 37°C for 6 hours for induction. The cell culture was harvested and genomic DNA extracted. A fragment (~0.3 kb) containing the target region was directly amplified from the extracted genomic DNA using a primer pair (p685–p696). The amplicon was labeled with Dual Index Primer. An average of over 30,000 reads per sample were analyzed using a MiSeq sequencing system. Sequencing reads were paired and trimmed with a maximum ambiguity of 2 based on a quality limit of 0.01, and overlapping reads within a read pair were merged. Each read was mapped to a reference sequence using the following settings: Masking mode = no masking, Mismatch cost = 2, Insertion cost = 3, Deletion cost = 3, Length fraction = 0.5, Similarity fraction = 0.8, Global alignment = No, Auto-detect paired distances = Yes, Nonspecific match handling = Map randomly. The output file was sorted using Excel.
[0084] Example 1 Deaminase-mediated targeted mutagenesis in E. coli To assess whether deaminase-mediated targeted mutagenesis can be applied to bacteria, we constructed a bacterial targeting AID (Target-AID) vector expressing catalytically inactive Cas9 (dCas: D10A and H840A mutations) fused to the cytosine deaminase PmCDA1 from P. marinus (sea lamprey) (NPL 11) under a temperature-inducible λ operator system (Wang, Y. et al., Nucleic Acids Res. 40, (2012)). The vector also expresses CDA under a 20-nucleotide (nt) target sequence-gRNA scaffold hybrid (sgRNA) under the artificial constitutive promoter J23119 (Figure 1(b)). In eukaryotes, nickase Cas9 (nCas:D10A mutation) can be used in combination with a deaminase to achieve higher mutation efficiency (Non-Patent Documents 10, 11). However, a plasmid expressing nCas(D10A)-CDA was shown to have low transformation efficiency. This suggests that, like the intact Cas9 nuclease, nCas(D10A)-CDA induces severe cell proliferation and / or cell death in E. coli (Figure 2). On the other hand, nCas(H840)-CDA showed high transformation efficiency, similar to dCas and dCas-CDA, and was shown to be advantageous for cell proliferation and survival. Next, to quantitatively evaluate the efficiency of targeted mutagenesis, we used the galactose analog 2-deoxyglucose (GAL) as a target. The galK gene, which can be positively selected for loss of function by oxy-D-galactose (2-DOG), was used as a target. 2-DOG is catalyzed by the galK gene product, galactokinase, to produce a toxic compound (Warming, S. et al., Nucleic Acids Res. 33, 1-12 (2005)). Target-AID is known to induce mutations at cytosine nucleotides (C) located approximately 15-20 bases upstream of the protospacer adjacent motif (PAM) sequence (Non-Patent Document 11) (Figure 1(a)). The target sequence (Figure 3) was selected to introduce a stop codon into the galK gene. 2-DOG induced a nearly 100% survival rate, suggesting highly efficient mutagenesis (Figure 4(a)). When sequencing analysis was performed on cells grown in medium without 2-DOG, 6 of 8 colonies were found to be mutated as expected, with C to T substitutions at positions -17 and / or 20.
[0085] Next, we targeted rpoB, an essential gene encoding the β-subunit of RNA polymerase. Disruption of rpoB gene function leads to cell growth inhibition and cell death, and specific point mutations in the rpoB gene are known to confer rifampicin resistance (Jin, DJ et al., J. Mol. Biol. 202, 245-253 (1988)). The target sequence was designed to induce point mutations conferring rifampicin resistance (Figure 5(a)). There was no obvious growth inhibition, and transformants acquired rifampicin resistance at a frequency of nearly 100% (Figure 4(b)). Sequencing analysis of clones selected in rifampicin-free medium confirmed the expected C-to-T substitutions at positions -16 and / or 17 from the PAM sequence (positions 1545 and 1546 of the rpoB gene) (Figure 5(a)). We performed whole-genome sequencing to assess the potential for nonspecific mutagenic effects of Target-AID in E. coli. Three independent clones expressing dCas-CDA and sgRNA targeting rpoB_1 were analyzed and found to contain zero to two unique single nucleotide variants (SNVs) at apparently unrelated genomic locations (Figure 5(b)). The flanking sequences of the detected SNVs showed no similarity to the rpoB target sequence (Figure 5(c)).
[0086] Example 2 Effect of sgRNA length and uracil DNA glycosylase inhibitors on mutation frequency and location To comprehensively analyze the mutation efficiency and location, we performed deep sequencing analysis using 18 target sequences in the galK gene (Figures 6 and 7). Seven targets showed high mutagenesis efficiencies (61.7–95.1%), while five showed low mutagenesis efficiencies (1.4–9.2%). The most effective mutation location was 17–20 bases upstream of the PAM, consistent with previous studies in higher organisms. Mutation frequency also varied depending on the length of the target sequence, as can be seen from the fact that sgRNAs with longer target sequences showed higher mutagenesis efficiencies for galK_8 and galK_13 and lower mutagenesis efficiencies for galK_9 and galK_11 (Figure 6, left bar).
[0087] To improve the mutagenesis efficiency, uracil DNA glycosylators derived from bacteriophage PBS2 were used. UGI (Ultraligandase inhibitor) (Zhigang, W. et al., Gene 99, 31-37 (1991).) and proteolytic enzyme inhibitors The LVA tag (Andersen, JB et al., Appl. Environ. Microbiol. 64, 2240-2246 (1998)) was introduced by fusing it to the C-terminus of dCas-CDA. UGI inhibits the removal of uracil (the direct product of cytosine deamination) from DNA (Non-Patent Documents 10, 11), promoting mutagenesis by cytidine deamination. The use of the LVA tag is expected to protect cells from damage and suppress the development of escaper cells by reducing the half-life of the dCas-CDA-UGI protein, which can be potentially harmful if overexpressed. To evaluate the nonspecific mutagenesis effect, whole-genome sequence analysis was performed on cells expressing dCas, dCas-CDA, and dCas-CDA-UGI-LVA. While dCas-CDA induced 0–2 SNV mutations, dCas-CDA-UGI-LVA induced 21–30 mutations without positional bias across the genome (Tables 3 and 4).
[0088] [Table 3]
[0089] Each construct (dCas, dCas-CDA, or dCas-CDA-LVA-UGI) without sgRNA was expressed. The rifampicin-selected clones were subjected to whole-genome sequencing. Biological triplicates of dCas-CDA and dCas-CDA-LVA-UGI are shown. Sequence coverage was calculated as the total base pairs mapped to the 4,631 Mbp of E. coli BW25113 genome sequence. A list of unique mutations is shown in Table 4.
[0090] [Table 4]
[0091] dCas-CDA-UGI-LVA showed strong mutagenesis at all target sites, regardless of the length and position of the target sequence (Figure 6, right bar), and the mutation spectra using sgRNAs of different lengths were compared. The results showed that the mutation spectrum of galK_9 and galK_11 extended toward the 5' end (Figure 6). We further investigated the effect of the length of the sgRNA target sequence. To characterize the mutations, C-rich target sequences of 18 nt, 20 nt, 22 nt, and 24 nt in length were tested (Figures 8(a) and (b)). The mutation spectra for each of the five target sites consistently showed a peak shift toward the 5' end and an expansion of the window as the target sequence became longer (Figure 8(c)).
[0092] Example 3 Multiple Mutagenesis For multiplex editing, tandem repeats of sgRNA expression units are inserted into the modifying plasmid and was constructed on a separate plasmid. Three sites of the galK gene (galK_10, galK_11 and gal A plasmid targeting dCas (K_13) was constructed and co-transfected into cells carrying the engineered vectors expressing dCas, dCas-CDA, dCas-CDA-LVA, or dCas-CDA1-UGI-LVA. The effect of nonspecific mutagenesis was evaluated by analyzing the occurrence of resistance mutations (Figure 9). dCas-CDA showed an approximately 10-fold increase over background mutation frequency, while dCas-CDA-UGI-LVA showed an even 10-fold increase over that of dCas-CDA. Although dCas-CDA and dCas-CDA-LVA were not efficient enough to simultaneously generate triple mutants, single mutations occurred in both cases, and at least the target mutation rate did not differ significantly between the presence and absence of LVA. Therefore, the addition of LVA demonstrated the ability to suppress nonspecific mutagenesis while maintaining mutagenesis efficiency. Furthermore, although the mutation frequency was lower with dCas-CDA-UGI-LVA compared to the results for each single target, which yielded 100% (8 / 8) for each target (Figure 9(c) and (d)), dCas-CDA-UGI-LVA successfully induced triple mutations in 5 of the 8 clones analyzed (Figure 9(b) and (d)). Therefore, it was demonstrated that the combination of UGI and LVA can achieve high mutation efficiency while suppressing nonspecific mutagenesis.
[0093] Then, six different genes (galK, xylB (xylulokinase), manA (mannose-6- We targeted dCas-CDA-UGI-LVA with sgRNAs targeting six different genes, and found that seven of eight clones contained mutations at all target loci (Figure 10).
[0094] Example 4. Multicopy gene editing with Target-AID Unlike other methods involving recombination or genome cleavage, Target-AID allows multiple copies of the same sgRNA sequence to be replicated without inducing genome instability. To prove this concept, we simultaneously targeted four major transposable elements (TEs: IS1, 2, 3, and 5) in the E. coli genome using four sgRNAs. The 10, 12, 5, and 14 loci for IS1, 2, 3, and 5, respectively, were The locus could be specifically amplified using unique PCR primers. The sgRNAs were designed to contain the consensus sequence of the transposase gene of each TE and introduce a stop codon (Figure 11). E. coli Top10 cells were transformed with two plasmids expressing dCas-CDA-UGI-LVA and four target sgRNAs, respectively. The procedure for isolation and verification of IS-edited cells is shown in Figure 12 and described in detail below. After double transformation and selection, colonies were amplified by PCR and sequenced first at IS5-1, IS5-2, IS5-11, and IS5-12. The IS5 target proved to be inefficient. Of the four colonies analyzed, one contained three mutation sites and one heterologous site (IS5-1). The cells were then suspended in liquid medium and plated for re-isolation. Three of the eight colonies contained mutations at IS5-1, and two of these were further sequenced at the remaining 24 IS loci, showing that they contained all mutated sites but one incomplete, heterologous site (IS5-5). The cells were then suspended and expanded, yielding four of the six reisolated clones containing mutations at IS5-5. One of these clones was sequenced at the IS5 site and found to contain one heterologous site (IS5-2). Eight clones were reisolated, six of which contained mutations at IS5-2. Two clones were expanded onto non-selective medium to obtain cells that had lost the plasmid. The cells were then genome extracted and sequenced to confirm mutations at all IS sites (Figure 11). Further genome sequencing was performed to assess genome-wide off-target effects. Of 34 potential off-target sites from the reference genome, including sequences with 1-8 base pairs adjacent to the PAM, two sites were found to be mutated (Table 5).
[0095] [Table 5]
[0096] Region indicates the target site in the DH10B database. Strand indicates the orientation of the target sequence. Expected off-target sequences were determined as described herein. Mismatches indicate the number of mismatches between the on-target and off-target sequences. Mismatched nucleotides are highlighted in bold. The frequency of C to T mutations in each sequence is indicated by a gray box.
[0097] Example 5 Comparison of transformation efficiency of E. coli using yeast expression vectors LbCpf1 (SEQ ID NOs: 326 and 327) was used as the Cas effector protein. Yeast expression vectors encoding YAASV and YALAA as proteolytic tags (Vector 3685: Cpf1-NLS-3xFlag-YAASV (SEQ ID NO: 328) and Vector 3687: Cpf1-NLS-3xFlag-YALAA (SEQ ID NO: 329)), as well as a control vector lacking a nucleic acid encoding a proteolytic tag (Vector 3687: Cpf1-NLS-3xFlag (SEQ ID NO: 330)), were constructed based on the pRS315 vector. The transformation efficiency of E. coli was evaluated using these vectors. Figure 13 shows a schematic diagram of each vector. As shown in Table 6 below, a DNA solution containing each vector was adjusted to 2 ng / μl. 20 μl of E. coli Top10 competent cells were transformed with 1 μl (2 ng) of the DNA solution. Subsequently, 200 μl of SOC was added, and the cells were recovered at 37°C for 1 hour. Growth was stopped by placing the cells on ice for 5 minutes, after which 1 μl of 50 mg / ml Amp was added. A portion of the culture medium (1 μl and 10 μl) was diluted with TE, spread onto an LB+Amp plate, and cultured overnight at 37° C. The number of colonies was counted. The results are shown in Table 6.
[0098] [Table 6]
[0099] It was shown that the transformation efficiency of E. coli was higher when vectors 3685 and 3686, which contain nucleic acids encoding proteolytic tags, were used than when the control vector 3687 was used. Therefore, the use of proteolytic tags is expected to improve the replication efficiency of vectors used for the expression of heterologous organisms, even when the vectors are replicated in bacteria such as E. coli.
[0100] This application is based on patent application No. 2017-225221 filed in Japan (filing date: November 22, 2017), the contents of which are incorporated in their entirety herein. [Industrial Applicability]
[0101] The present invention provides a low-toxicity vector that can be stably amplified even in a host bacterium, and a genome editing complex encoded by the vector. Genome editing techniques using the vector and nucleic acid modifying enzyme of the present invention make it possible to modify the genes of a host bacterium while suppressing nonspecific mutations, etc. Because this technique does not depend on host-dependent factors such as RecA, it can be applied to a wide range of bacteria and is extremely useful.
Claims
[Claim 1] The invention as shown in the drawings.
Citation Information
Patent Citations
Genomic sequence modification method for specifically converting nucleic acid bases of targeted DNA sequence, and molecular complex for use in same
WO2015133554A1