Protein expression
Patent Information
- Application Number
- JP2024533803
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-06
- Filing Date
- 2022-12-06
- Publication Date
- 2025-12-17
AI Technical Summary
Existing codon optimization methods struggle to achieve high and sustained expression of target nucleic acid sequences in host cells, particularly in human induced pluripotent stem cells (iPSCs) and their differentiated cell types, limiting applications such as CRISPR-Cas9 genome-wide genetic screening.
Codon optimization of target nucleic acid sequences based on the codon bias of highly expressed proteins in the host cell, replacing non-preferred codons with preferred synonymous codons, as exemplified by tubulin III in neuronal cells, to enhance expression efficiency.
Significantly improves expression levels and maintains expression throughout differentiation, enabling effective CRISPR-Cas9 genome-wide screening in various iPSC-derived cell types.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to methods for codon-optimizing a target nucleic acid sequence for expression in a host cell. The present invention also relates to codon-optimized nucleic acids for improved expression in a host cell, and to vectors and host cells comprising the codon-optimized nucleic acids. [Background technology]
[0002] A codon is a trinucleotide sequence of DNA or RNA that either codes for a specific amino acid or signals the end of translation (a "stop" or "termination" codon). There is degeneracy in the genetic code because there are more codon sequences than there are amino acids or termination codons. In fact, 18 of the 20 common amino acids are coded for by multiple "synonymous codons" (i.e., different codons that code for the same amino acid). Codon usage can vary greatly between species: different species typically show a "bias" for certain codons, with some species using certain codons very rarely or not at all. If a gene of interest contains a codon that is rarely used by the host, the gene will encounter translational arrest in cells from that host, thereby reducing expression efficiency or preventing expression altogether. Codon optimization approaches are designed to improve the codon composition of a target nucleic acid sequence by taking into account differences in codon bias between species and replacing codons that are rarely used by the host with synonymous codons that are more frequently used by the host and are therefore "preferred" by the host.
[0003] Codon usage has recently attracted attention as a key determinant of translation elongation rate and co-translational protein folding, with codons preferred by the host enhancing translation efficiency and folding fidelity. The unequal use of synonymous codons, termed "codon bias," and the universal nature of this bias from yeast to humans, suggest the existence of a secondary code within the more general genetic code. This secondary code has emerged as a major regulator of translation rate and co-translational protein folding, and thereby as a key determinant of the cellular levels of specific proteins.
[0004] To identify the codon bias of a particular host, the frequency of codon usage is typically determined across hundreds or thousands of coding DNA sequences (CDS). To optimize the codons of a gene of interest, codons in the gene that are less frequent (or absent altogether) in the host (referred to as "non-preferred codons") are replaced with synonymous codons that are more commonly used by the host (referred to as "preferred codons"). Codon optimization aims to improve the efficiency of expression of a gene of interest without changing the sequence of the encoded protein.
[0005] Although codon optimization methods are well established in the art, some genes are still difficult to express, and in some cases, codon-optimized genes do not achieve high enough expression levels or cannot maintain sufficient expression levels for long periods of time.
[0006] For example, human induced pluripotent stem cells (hiPSCs / iPSCs) are powerful research tools with the potential to differentiate into multiple cell types. However, the application of these cells to genome-wide genetic screens using the CRISPR (clustered regularly interspaced short palindromic repeats)-Cas (CRISPR-associated protein) gene editing system is hindered by the inability of these cells to efficiently express Cas proteins, e.g., Cas9, even though the genes encoding these proteins have been codon-optimized for expression in human cell lines. The mechanism by which Cas genes are silenced in iPSC-derived differentiated cell types is currently unknown. Summary of the Invention [Problem to be solved by the invention]
[0007] There is an urgent and unmet need for improved codon optimization methods that allow efficient expression of target nucleic acid sequences in host cells. [Means for solving the problem]
[0008] The present inventors have developed a novel method for optimizing the codons of a target nucleic acid sequence for expression in a host cell. According to the present invention, codon optimization utilizes the codon usage of genes encoding proteins that are highly expressed by the host cell, or genes encoding proteins that are highly expressed in cells of the same species as the host cell. Codons in the target nucleic acid that are used at low frequency by genes encoding highly expressed proteins are replaced with synonymous codons that are used at high frequency by genes encoding highly expressed proteins.
[0009] The current "gold standard" for optimizing codons of a target nucleic acid is based on species-level codon bias derived from hundreds or thousands of coding sequences. Surprisingly, the present inventors have discovered that codon-optimizing a target nucleic acid sequence based on the codon bias of genes encoding highly expressed proteins significantly improves expression efficiency compared to the corresponding nucleic acid optimized using the current gold standard. Codon optimization according to the present invention achieves high levels and sustained expression even in cell types that typically do not express the gene containing the nucleic acid sequence on which the codon optimization was based.
[0010] Importantly, codon optimization according to the present invention allows for high-level and sustained protein expression to be achieved in iPSCs and iPSC-derived differentiated cell lines, which greatly improves the potential applications of these cells in research.
[0011] The present invention provides methods for optimizing the codons of a target nucleic acid sequence for expression in a host cell comprising altering the codon usage of the target nucleic acid sequence based on the codon usage of genes encoding proteins that are highly expressed in the host cell, or genes encoding proteins that are highly expressed in a cell of the same species as the host cell.
[0012] In some embodiments, the methods involve replacing one or more non-preferred codons in a target nucleic acid sequence with preferred synonymous codons, where: (a) the non-preferred codons are codons used infrequently by genes encoding highly expressed proteins, and (b) the preferred codons are codons used frequently by genes encoding highly expressed proteins.
[0013] In some embodiments, a non-preferred codon is a codon that is used less frequently by genes encoding highly expressed proteins than would be expected if each synonymous codon was used randomly.
[0014] In some embodiments, non-preferred codons are used less than 50%, less than 45%, less than 40%, less than 35%, less than 33%, less than 30%, less than 25%, less than 20%, less than 16%, less than 15%, less than 10%, less than 5%, or 0% frequently by genes encoding highly expressed proteins.
[0015] In some embodiments, a preferred codon is a codon that is used more frequently by genes encoding highly expressed proteins than would be expected if each synonymous codon was used randomly.
[0016] In some embodiments, the preferred codons are used at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% frequently by genes encoding highly expressed proteins.
[0017] In some embodiments, the methods include replacing at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% of the non-preferred codons in the target nucleic acid with preferred synonymous codons.
[0018] In some embodiments, the methods involve replacing all non-preferred codons in a target nucleic acid that are used at 0% frequency by genes encoding highly expressed proteins with preferred synonymous codons.
[0019] In some embodiments, the method comprises replacing all non-preferred codons in a region of a target nucleic acid that encodes an N-terminal region of a protein with preferred synonymous codons.
[0020] In some embodiments, the method includes replacing all non-preferred codons in the 5' region of the target nucleic acid with preferred synonymous codons, and optionally replacing the first at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 codons beginning from the 5' end of the target nucleic acid.
[0021] In some embodiments, the highly expressed protein is a housekeeping protein or a cell marker protein. In some embodiments, the highly expressed protein is selected from GAPDH, β-tubulin, β-actin, and tubulin III. In some embodiments, the highly expressed protein is tubulin III.
[0022] In some embodiments, the one or more non-preferred codons are selected from the alanine codons GCA, GCG, and GCT; the arginine codons AGA and CGT; the cysteine codon TGT; the glutamine codon CAA; the isoleucine codon ATA; the leucine codons CTA and TTA; the lysine codon AAA; the proline codon CCG; the serine codon TCC; the threonine codons ACA, ACG, and ACT; the tyrosine codon TAT; the valine codons GTA and GTT; and the stop codons TAA and TAG.
[0023] In some embodiments, the one or more non-preferred codons are selected from the asparagine codon AAT; the aspartic acid codon GAT; the glutamic acid codon GAA; the glycine codons GGA, GGG, and GGT; the histidine codon CAC; the isoleucine codon ATT; the leucine codons CTC, CTT, and TTG; the phenylalanine codon TTT; the proline codon CCA; the serine codons TCA and TCG; and the valine codon GTC.
[0024] In some embodiments, the preferred codons are selected from the alanine codon GCC; the cysteine codon TGC; the glutamine codon CAG; the lysine codon AAG; the threonine codon ACC; the tyrosine codon TAC; and the stop codon TGA.
[0025] In some embodiments, the preferred codons are selected from the arginine codons AGG, CGA, CGC, and CGG; the asparagine codon AAC; the aspartic acid codon GAC; the glutamic acid codon GAG; the glycine codon GGC; the histidine codon CAT; the isoleucine codon ATC; the leucine codon CTG; the phenylalanine codon TTC; the proline codons CCC and CCT; the serine codons AGC, AGT, and TCT; and the valine codon GTG.
[0026] In some embodiments, the host cell is selected from a human cell, a bacterial cell, a yeast cell, and a fungal cell. In some embodiments, the host cell is a human cell. In some embodiments, the host cell is a HEK293 cell. In some embodiments, the host cell is a human induced pluripotent stem cell (iPSC). In some embodiments, the host cell is a differentiated cell derived from an iPSC, optionally the host cell is selected from an iPSC-derived neuron, such as a cortical neuron, a dopaminergic neuron, or a motor neuron; an iPSC-derived macrophage; an iPSC-derived cardiomyocyte; and an iPSC-derived hepatocyte.
[0027] In some embodiments, the target nucleic acid encodes a Cas protein, optionally wherein the Cas protein is selected from Cas9, Cas12a, and Cas13Rx.
[0028] The present invention also provides a nucleic acid comprising a nucleic acid sequence that has been codon optimized by the methods of the present invention.
[0029] The present invention also provides codon-optimized nucleic acids for improved expression in a host cell, where the codon usage of the nucleic acid corresponds to the codon usage of a gene encoding a protein that is highly expressed by the host cell, or to the codon usage of a gene encoding a protein that is highly expressed in a cell of the same species as the host cell.
[0030] In some embodiments, a codon-optimized nucleic acid contains less frequent non-preferred codons than a non-optimized nucleic acid sequence encoding the same amino acid sequence.
[0031] In some embodiments, a codon-optimized nucleic acid contains a higher frequency of preferred codons than a non-optimized nucleic acid sequence encoding the same amino acid sequence.
[0032] The present invention also provides a nucleic acid encoding Cas9 and comprising a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or 100% sequence identity to SEQ ID NO:1 or SEQ ID NO:3.
[0033] The present invention also provides a nucleic acid encoding Cas12a and comprising a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or 100% sequence identity to SEQ ID NO:4.
[0034] The present invention also provides a nucleic acid encoding Cas13Rx and comprising a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or 100% sequence identity to SEQ ID NO:5.
[0035] The present invention also provides a vector comprising a nucleic acid according to the present invention.
[0036] The present invention also provides a host cell comprising a nucleic acid according to the invention or a vector according to the invention.
[0037] In some embodiments, the host cell is selected from a human cell, a bacterial cell, a yeast cell, and a fungal cell. In some embodiments, the host cell is a human cell. In some embodiments, the host cell is a HEK293 cell. In some embodiments, the host cell is a human induced pluripotent stem cell (iPSC). In some embodiments, the host cell is a differentiated cell derived from an iPSC, optionally the host cell is selected from an iPSC-derived neuron, such as a cortical neuron, a dopaminergic neuron, or a motor neuron; an iPSC-derived macrophage; an iPSC-derived cardiomyocyte; and an iPSC-derived hepatocyte. [Brief description of the drawings]
[0038] [Figure 1] (A) Schematic diagram of the PiggyBac (PB) transposon plasmid containing the start Cas9 (old-Cas9) sequence. (B) Cas9 and (C) GAPDH (glyceraldehyde-3-phosphate dehydrogenase) mRNA levels in iPSC Cas9 cells at days 0 (iPSC), 10, and 20 during differentiation into dopaminergic neurons. Cas9 mRNA is reduced by approximately 60% by day 20. [Diagram 2] (A) Schematic showing homologous recombination used to knock-in Cas9 upstream of GAPDH. (B) Cas9 and GAPDH mRNA levels during iPSC to neuron differentiation in GapdhCas9-iNgn2 iPSCs. (C) GapdhCas9 and WT (no Cas9) GAPDH mRNA levels during iPSC to neuron differentiation in iNgn2 cells. [Diagram 3](A) Western blot showing Cas9 and GAPDH protein expression during iPSC-to-neuron differentiation. (B) Densitometric quantification of Cas9 protein levels in (A) during neuronal differentiation in Bob-iNgn2 Gapdh-Cas9 cells. (C) Schematic of the fluorescent reporter construct used. (D) Flow cytometry plot showing dual fluorescence of the reporter construct (left) and loss of GFP fluorescence in the presence of Cas9 (right). (E) Cas9 cleavage efficiency is quantified by loss of GFP fluorescent reporter 4 days after reporter transduction. [Figure 4] (A) Cas9 protein expression in the presence of MG132 inhibitor, which inhibits proteasomal degradation, or bafilomycin A1 (BafA1), which inhibits the autophagy-lysosomal pathway. (B) Cas9 protein expression in different media compositions. DMEM = Dulbecco's modified Eagle's medium; NEAA = non-essential amino acids. [Diagram 5] Comparison of bacterial (E. coli) codon usage (black bars) with general codon usage in humans (grey bars). Dashed boxes indicate codons with significant differences in usage between humans and bacteria. [Figure 6] Comparison of old-Cas9 codon usage (black bars) with common human codon usage (grey bars). Solid boxes indicate optimized codons, and dashed boxes indicate codons with random distribution. [Figure 7] Comparison of tubulin III codon usage (black bars) with common codon usage in humans (grey bars). Dashed boxes indicate codons that are not used in tubulin III; triangles indicate codons that are more frequently used in tubulin III than common in humans. [Figure 8] Comparison of codon usage of codon-optimized Cas9 (CodOpt-Cas9; SEQ ID NO: 1) (black bars) with common codon usage in humans (grey bars). Dashed boxes indicate codons not used in CodOpt-Cas9; triangles indicate codons that are frequently used in CodOpt-Cas9. [Figure 9] Schematic diagram of Old-Cas9 (left) and CodOpt-Cas9 (right) expression constructs. [Figure 10] (A) Old-Cas9 and CodOpt-Cas9 mRNA levels in HEK293-Cas9 lines. (B) and (C) Old-Cas9 and CodOpt-Cas9 protein levels (B) and quantification (C) in HEK293 cells. (D) Cas9 editing efficiency in HEK293 cells using a reporter plasmid. [Figure 11] Indel formation by Cas9 induced to edit a non-essential gene (ST6GALNAC6) in HEK-Cas9 and control strains. Graph shows TIDE profile obtained by tracking indel formation at http: / / shinyapps.datacurators.nl / tide / . Y-axis = % of sequence. [Figure 12] (A) Old-Cas9 and CodOpt-Cas9 mRNA levels in Bob-iNgn2 Cas9 iPSC cells generated using PiggyBac (B) and (C) Western blot and quantification of Old-Cas9 and CodOpt-Cas9 protein levels in Bob-iNgn2 iPSC cells (C). [Figure 13] 1 is a Western blot showing the levels and relative quantification of Cas9 protein in Old-Cas9 and CodOpt-Cas9 Bob-iNgn2 iPSC lines at different time points during neuronal differentiation. [Figure 14] Schematic diagram of Old-Cas9 (left), NOpt-Cas9 (center), and CodOpt-Cas9 (right) expression constructs. [Figure 15] Flow cytometry plots showing the disappearance of GFP reporter fluorescence in Bob-iNgn2 iPSCs carrying either Old-Cas9, NOpt-Cas9, or CodOpt-Cas9. Editing efficiency of the Cas9 variants was assessed during differentiation into neurons (iPSCs, days 4 and 10). [Figure 16] Quantification of Cas9 cleavage efficiency with Old-Cas9, NOpt-Cas9, CodOpt-Cas9, or no Cas9 (WT) in Bob-iNgn2 iPSCs and at various time points of neuronal differentiation. Highlighted boxes highlight the differences that occur in Cas9 editing as neurons differentiate through the protocol. [Figure 17] Cas9 mRNA (A) and protein (B) levels produced by Old-Cas9, CodOpt-Cas9, or NOpt-Cas9 in Bob-iNgn2 iPSCs and during differentiation into neurons. Quantification of protein levels is shown in (C), with the highlighted grey / black boxes highlighting the differences in Cas9 levels as neurons differentiate and mature. [Figure 18] (A) Cas9 protein levels in iPSC-derived hepatocytes containing either Old-Cas9 or CodOpt-Cas9. (B) Quantification shows elevated levels of CodOpt Cas9 in differentiated hepatoblastoma cells at day 10 (grey / black box). [Figure 19] Comparison of Cas12a codon usage (black bars) with general human codon usage (grey bars). Dashed boxes indicate amino acids that are biased towards certain codons. [Figure 20] Comparison of codon-optimized Cas12a codon usage (black bars) with common human codon usage (grey bars). [Figure 21] Comparison of Cas13Rx codon usage (black bars) with general human codon usage (grey bars). [Figure 22] Comparison of codon-optimized Cas13Rx codon usage (black bars) with common human codon usage (grey bars). [Figure 23]FIG. 1 is a schematic diagram of a plasmid containing LlDr optimized using the existing gold standard method based on human codon usage (denoted "normal codon optimization") and the codon bias of tubulin III described herein (denoted "novel codon optimization"). [Figure 24] (A) Transfection efficiency of normal and novel optimized plasmids in HEK293 and iPSC cells. (B) Western blot showing LlDr(c-Myc) and GAPDH protein expression in Bob-iNgn2 iPSC and HEK293 cells 5 days after transfection. (C) Densitometric quantification of LlDr(c-Myc) levels compared to Gapdh levels in Bob-iNgn2 iPSC and HEK293 cells (B). The existing gold standard codon optimization method is referred to as "normal optimization" and the optimization method using the codon bias of Tubulin III is referred to as "novel optimization." DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0039] The present invention is based on the surprising discovery that expression of a target nucleic acid may be significantly improved by codon-optimizing the target nucleic acid sequence for expression in a host cell (or a cell of the same species as the host cell) based on the codon usage of genes encoding proteins highly expressed in the host cell (or the codon usage of genes encoding proteins highly expressed in cells of the same species as the host cell). In particular, the inventors have discovered that this approach achieves efficient and sustained expression of previously difficult to express target nucleic acids, even when codons are optimized using the current gold standard of codon optimization based on species-level codon bias.
[0040] In many host cell types, transgene expression is difficult. One example of a host cell in which transgene expression may be difficult is cells derived from human induced pluripotent stem cells (hiPSCs / iPSCs). As iPSCs are a powerful research tool with the potential to differentiate into multiple cell types, improvements in transgene expression in such cells are highly desirable. In recent years, numerous hiPSC-based cell lines have been generated that allow controlled and rapid differentiation into various cell types, including macrophages (immune cells), cardiomyocytes (muscle cells), and neurons (nerve cells). These cell lines have been applied in a wide range of research fields. For example, hiPSC-derived neurons provide a powerful alternative to immortalized human cell lines and non-human primary neuronal cells for use in in vitro studies to understand neurodegenerative disorders, as they can be differentiated into specific neuronal subtypes that have been found to affect these disorders. Several differentiation protocols have been optimized to generate specific neuronal subtypes, such as cortical neurons, dopaminergic neurons, and even motor neurons, and have been robustly utilized to model Alzheimer's disease, Parkinson's disease, or motor neuron disease, respectively.
[0041] Another powerful research tool is the CRISPR-Cas gene editing system, which has revolutionized molecular approaches to help elucidate cellular mechanisms, for example, those of neuronal degeneration. CRISPR-Cas9 genetic screens in multiple cell types are essential to identify novel cellular pathways and gene targets that may inform translational research. Many of these studies rely on starting genome-wide screens at the iPSC / progenitor cell stage and extrapolating the findings to iPSC-derived cell types, but performing CRISPR-Cas9 screens in differentiated cells has been challenging, primarily due to the inability to efficiently express Cas9 in iPSC-derived cell lines. The mechanism by which Cas9 is inactivated in iPSC-derived differentiated cell types, including neurons, is currently unknown. This inability to efficiently express key components of the CRISPR-Cas9 system dramatically limits the potential for studying iPSC-derived cell types.
[0042] Multiple approaches have been investigated to overcome Cas9 silencing during iPSC differentiation. These approaches include integrating multiple copies of Cas9 into the genome using lentivirus / transposon; testing Cas9 expression under various mammalian expression promoters; and targeting Cas9 to specific locations in the genome, such as genomic safe harbor sites. Despite these efforts, the levels of Cas9 protein are dramatically reduced in differentiated cells compared to the levels observed in iPSCs. We attempted to circumvent Cas9 silencing and achieve continued expression by inserting Cas9 into the site of a housekeeping gene (glyceraldehyde-3-phosphate dehydrogenase (GAPDH)). Despite successful knock-in of the endogenous GAPDH gene, the promoter of the housekeeping gene failed to maintain constitutive expression of Cas9 protein during differentiation into neuronal cell types. Interestingly, despite the reduced protein levels, the mRNA levels of Cas9 remained detectable during differentiation, suggesting that transcription and translation were uncoupled.
[0043] The Cas9 gene typically used in experimental studies (hereafter "old-Cas9") is derived from Streptococcus pyogenes and has been codon-optimized for expression in humans using common human codon usage (representing the current gold standard codon optimization approach). Based on the uncoupling of Cas9 transcription and translation observed in iPSC-derived cell lines, we hypothesized that Cas9 may require further codon optimization to function in differentiated cell types.
[0044] We attempted to identify whether genes highly expressed in iPSC-derived differentiated cells exhibited specific codon bias by comparing the codon usage of tubulin III (Ensembl transcript:TUBB3-208 ENST00000555576.5; SEQ ID NO:8), a marker gene highly expressed in neuronal cells, with the common codon usage in humans (Figure 7). The common codon usage in humans is typically derived from tens of thousands of human coding DNA sequences (CDS). The common codon usage in humans used herein was obtained from the codon usage database provided by the Kazusa DNA Research Institute and is based on the codon usage of 93,487 human CDSs (Nakamura, Y. et al. Nucleic acids research 2000 28(1):292). Unless otherwise specified, the term "human usage" or "common codon usage in humans" in this specification refers to the codon usage as set out in the Human (Homo sapiens) Codon Usage Table of the Kazusa Codon Usage Database.
[0045] The inventors have found that tubulin III exhibits a different codon bias for some codons compared to the common codon usage in humans. Surprisingly, tubulin III does not use some codons commonly used in humans, such as the cysteine codon TGT (46% human usage); the lysine residue AAA (43% human usage); and the tyrosine residue TAT (44% human usage). Furthermore, tubulin III exhibits a strict preference for the alanine codon GCC (40% human usage) and the threonine codon ACC (36% human usage), which are used exclusively, despite the availability of three additional synonymous codons for each of these amino acids. Tubulin III also exhibits a high preference for certain codons, for example, tubulin III uses the histidine residue CAT more frequently (60% usage) than CAC (40% usage), whereas the common codon usage in humans prefers CAC (58% usage). We propose that the high expression levels achieved by tubulin III in neuronal cells are due to these codon biases contributing to efficient expression in these cell types.
[0046] The inventors have exploited the codon bias of tubulin III to generate a codon-optimized version of Cas9 (CodOpt-Cas9) that more closely reflects the codon usage of tubulin III. The resulting CodOpt-Cas9 sequence after these modifications is represented in SEQ ID NO:1: [ka] [ka] [ka]
[0047] The starting Cas9 (old-Cas9) sequence is represented in SEQ ID NO:2: [ka] [ka] [ka]
[0048] Old-Cas9 is a commercially available Cas9 sequence that has been codon-optimized using common human codon usage. To codon-optimize this sequence using tubulin III codon usage, we replaced 33% of the codons with preferred synonymous codons of tubulin III (463 codons of the original Cas9 sequence were replaced). For example, 100% of the alanine, glutamine, lysine, and tyrosine codons of CodOpt-Cas9 are provided by the preferred codons of tubulin III, GCC, CAG, AAG, and TAC, respectively (these codons are used 93%, 96%, 77%, and 87% in Old-Cas9, respectively). In addition, codons of Old-Cas9 that are not used by tubulin III are replaced with preferred codons of tubulin III, for example, the lysine AAA codon is replaced by AAG, and the tyrosine TAT codon is replaced by TAC.
[0049] To test whether CodOpt-Cas9 could be expressed, both Old-Cas9 and CodOpt-Cas9 were expressed in human embryonic kidney 293 (HEK293) cells. Surprisingly, CodOpt-Cas9 showed higher expression than old-Cas9 at both the mRNA and protein levels in HEK293 cells. With these higher levels of Cas9, CodOpt-Cas9 HEK293 cells showed higher nuclease activity and faster cleavage efficiency compared to HEK293 cells containing old-Cas9. However, it should be noted that since tubulin III is not highly expressed in HEK293 cells, these results suggest that the codon bias of tubulin III is not intrinsic to neurons.
[0050] Next, we tested whether CodOpt-Cas9 could be easily expressed in iPSCs. Similar to the results in HEK293 cells, CodOpt-Cas9 achieved higher expression levels in iPSCs than old-Cas9. As mentioned above, the expression of old-Cas9 dramatically declines during the differentiation of iPSCs into neuronal cell types, so we next attempted to differentiate iPSCs expressing CodOpt-Cas9. Advantageously, we found that CodOpt-Cas9 is expressed throughout the differentiation of iPSCs, and CodOpt-Cas9 remains detectable in differentiated neuronal cells, whereas old-Cas9 expression levels decline rapidly and eventually become undetectable as the cells adopt a more neuronal phenotype. These results support that optimizing Cas9 based on the codon bias of tubulin III achieves efficient and sustained expression of Cas9 in iPSC-derived differentiated neurons.
[0051] To test whether the above advantageous results are limited to iPSC-derived neuronal cells, we attempted to express CodOpt-Cas9 in iPSC-derived hepatocytes (which typically do not express tubulin III). Similar to iPSC-derived neuronal cells, old-Cas9 expression levels declined sharply during hepatocyte differentiation, whereas CodOpt-Cas9 achieved and maintained significantly higher expression levels throughout differentiation. Advantageously, these results, together with the increased expression observed in HEK293 cells, demonstrate that by codon-optimizing sequences based on the codon bias of tubulin III, increased and sustained expression can be achieved in many different cell types, including cells that typically do not express tubulin III.
[0052] We next attempted to determine whether the expression of Cas9 could be "tuned" through partial codon optimization. A Cas9 variant was created in which the first 606 N-terminal amino acids were codon-optimized using tubulin III preferred codons, with the remaining sequence unchanged. This N-terminal codon-optimized Cas9 variant is referred to herein as NOpt-Cas9 and is represented by SEQ ID NO:3. [ka] [ka] [ka]
[0053] Similar to CodOpt-Cas9, NOpt-Cas9 showed improved cleavage efficiency compared to old-Cas9 as iPSC cells progressed to more neuronal cell types. At later stages of differentiation, e.g., days 10 and 14, NOpt-Cas9 showed reduced expression, and therefore reduced cleavage efficiency, compared to CodOpt-Cas9, suggesting that the degree of codon optimization directly impacts protein production levels as neurons mature. Advantageously, these results indicate that it is not necessary to codon-optimize the complete Cas9 nucleic acid sequence to achieve increased expression, and that Cas9 activity can be tuned by adjusting the level of codon optimization, with fully codon-optimized Cas9 exhibiting higher activity than partially codon-optimized variants.
[0054] The results described herein demonstrate that a codon-optimized target nucleic acid sequence (e.g., a gene of interest) achieves higher levels and more sustained expression based on the codon bias exhibited by endogenous genes encoding proteins highly expressed by the host cell, or based on the codon bias exhibited by endogenous genes encoding proteins highly expressed in cells of the same species as the host cell. Advantageously, sequences codon-optimized according to the present invention achieve higher expression than sequences codon-optimized using current gold standard methods that typically rely on species-level codon bias. Furthermore, the inventors have shown that gene expression can be tuned by altering the degree to which a sequence is codon-optimized using the methods described herein.
[0055] The present invention provides methods for optimizing the codons of a target nucleic acid sequence for expression in a host cell, comprising altering the codon usage of the target nucleic acid sequence based on the codon usage of a gene encoding a protein highly expressed in the host cell or the codon usage of a gene encoding a protein highly expressed in a cell of the same species as the host cell. In some embodiments, the present invention provides methods for optimizing the codons of a target nucleic acid sequence to improve expression in a host cell. In some embodiments, the present invention provides methods for optimizing the codons of a target nucleic acid sequence to increase expression in a host cell. In some embodiments, the gene encoding the protein highly expressed in the host cell is an endogenous gene. In some embodiments, the gene encoding the protein highly expressed in the cell of the same species as the host cell is an endogenous gene.
[0056] As used herein, "codon usage" (also referred to herein as "codon frequency" or "usage") refers to the proportion of each synonymous codon (each codon that codes for the same amino acid) present in a sequence or group of sequences. A codon usage of 100% indicates that the codon is used exclusively for a particular amino acid. Methionine (met) and tryptophan (trp) are each encoded by a single codon, so the usage of these codons will always be 100%. A codon usage of 0% indicates that the codon is not used by the sequence or group of sequences. A codon usage of 25% for a particular codon indicates that the codon accounts for 25% of all synonymous codons present in the sequence or group of sequences that code for the encoded amino acid (other synonymous codons account for the remaining 75%). In some embodiments, the method includes determining the codon usage frequency of genes encoding proteins that are highly expressed by the host cell or genes encoding proteins that are highly expressed by cells of the same species as the host cell.
[0057] Hereinafter, a "gene encoding a highly expressed protein" refers to a gene encoding a protein that is highly expressed in a host cell or in a cell of the same species as the host cell.
[0058] In some embodiments, a non-preferred codon is a codon that is used less frequently than would be expected if each synonymous codon were used randomly by a gene encoding a highly expressed protein. The random usage frequency depends on the number of synonymous codons available for a given amino acid. For example, for an amino acid encoded by two synonymous codons, each of these synonymous codons would have a random usage frequency of 50%. In this scenario, a codon usage frequency of less than 50% indicates that the codon is non-preferred. Similarly, for an amino acid encoded by six synonymous codons, each of these synonymous codons would have a random usage frequency of 16.67%, and a codon usage frequency of less than 16.67% indicates that the codon is not preferred. In some embodiments, a non-preferred codon is a codon that is used less frequently than other synonymous codons encoding the same amino acid by a gene encoding a highly expressed protein.
[0059] In some embodiments, a non-preferred codon is a codon that is used less than 50%, less than 45%, less than 40%, less than 35%, less than 33%, less than 30%, less than 25%, less than 20%, less than 16%, less than 15%, less than 10%, less than 5%, or less than 0% by genes encoding highly expressed proteins. In some embodiments, a non-preferred codon is used less than 10% by genes encoding highly expressed proteins. In some embodiments, a non-preferred codon is used 0% by genes encoding highly expressed proteins.
[0060] In some embodiments, a preferred codon refers to a codon that is used more frequently by a gene encoding a highly expressed protein than would be expected if each synonymous codon were used randomly. As mentioned above, the random usage frequency depends on the number of synonymous codons available for a given amino acid. For example, for an amino acid encoded by two synonymous codons, each of these codons will have a random usage frequency of 50%, and thus a codon usage frequency of more than 50% indicates that the codon is preferred. For an amino acid encoded by six synonymous codons, each of these synonymous codons will have a random usage frequency of 16.67%, and thus a codon usage frequency of more than 16.67% indicates that the codon is preferred. In some embodiments, a preferred codon is a codon that is used more frequently by a gene encoding a highly expressed protein than other synonymous codons that encode the same amino acid. In some embodiments, a preferred codon is a codon that is used exclusively by a gene encoding a highly expressed protein.
[0061] In some embodiments, the preferred codon is a codon that is used by genes encoding highly expressed proteins at least 17%, at least 20%, at least 25%, at least 30%, at least 34%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% frequently. In some embodiments, the preferred codon is used by genes encoding highly expressed proteins at least 50% frequently. In some embodiments, the preferred codon is used by genes encoding highly expressed proteins at least 75% frequently.
[0062] In some embodiments, at least 50% of the non-preferred codons in the target nucleic acid sequence are replaced with preferred synonymous codons. In some embodiments, the method comprises replacing at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% of the non-preferred codons in the target nucleic acid sequence with preferred synonymous codons. In some embodiments, the method comprises replacing all non-preferred codons in the target nucleic acid sequence that are used with a frequency of 0% by genes encoding highly expressed proteins with preferred synonymous codons.
[0063] In some embodiments, the methods involve replacing all non-preferred codons in a target nucleic acid sequence with preferred synonymous codons in a particular region of the target nucleic acid, for example, the 5' end of the target nucleic acid (encoding the N-terminal region of a protein). In some embodiments, the methods involve replacing all non-preferred codons with preferred synonymous codons in at least the first 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, or at least 2000 codons beginning from the 5' end of the target nucleic acid.
[0064] As used herein, a highly expressed protein is a protein that is constitutively expressed by a host cell or a related cell of the same species as the host cell. Typically, a highly expressed protein can be easily detected using methods known in the art, such as, for example, Western blot and enzyme-linked immunosorbent assay (ELISA). Preferably, a highly expressed protein is one of the most highly and / or stably expressed proteins produced by a host cell or a related cell. In some embodiments, a highly expressed protein is among the top 10% of the most highly expressed proteins in a host cell or in a cell of the same species as the host cell. A person skilled in the art can easily identify a highly expressed protein using methods known in the art, such as, for example, proteomic approaches including gel electrophoresis and mass spectrometry. A highly expressed protein can also be identified using online protein expression databases, such as, for example, the Human Protein Atlas.
[0065] In some embodiments, the highly expressed protein is a housekeeping protein or marker protein. As used herein, a "housekeeping protein" is a constitutively expressed protein necessary for the maintenance of basic cellular functions in a host cell or a cell of the same species as the host cell, e.g., in humans, GAPDH, β-tubulin, and β-actin are considered housekeeping genes. In some embodiments, the gene encoding the highly expressed protein is the GAPDH gene. In some embodiments, the gene encoding the highly expressed protein is the β-actin gene. In some embodiments, the gene encoding the highly expressed protein is the β-tubulin gene. As used herein, a "cell marker protein" is a protein that is expressed by a particular cell type and can be used to identify that cell type, such as tubulin III (also referred to as β-tubulin III, class III β-tubulin, or βIII-tubulin), a neuronal cell marker, myosin, a muscle cell marker, and α-fetoprotein, a hepatic stem cell marker. In some embodiments, the gene encoding the highly expressed protein is the tubulin III gene. In some embodiments, the gene encoding the highly expressed protein is a tubulin III gene transcript, such as the tubulin III transcript represented by SEQ ID NO:8. In some embodiments, the gene encoding the highly expressed protein is a myosin gene. In some embodiments, the gene encoding the highly expressed protein is an alpha-fetoprotein gene. The highly expressed protein may be a lymphocyte marker protein, such as a T cell marker protein, such as CD4. In some embodiments, the gene encoding the highly expressed protein is a CD4 gene. In some embodiments, the non-preferred codon comprises a codon used with 0% frequency by the gene encoding the highly expressed protein.For example, if the gene encoding the highly expressed protein is a tubulin III gene, the non-preferred codons may include the alanine codons GCA, GCG, and GCT; the arginine codons AGA and CGT; the cysteine codon TGT; the glutamine codon CAA; the isoleucine codon ATA; the leucine codons CTA and TTA; the lysine codon AAA; the proline codon CCG; the serine codon TCC; the threonine codons ACA, ACG, and ACT; the tyrosine codon TAT; the valine codons GTA and GTT; and the stop or end codons TAA and TAG. In some embodiments, the non-preferred codons include codons that are used less frequently by the gene encoding the highly expressed protein than would be expected if each synonymous codon was used randomly. For example, if the gene encoding the highly expressed protein is a tubulin III gene, the non-preferred codons may also include the asparagine codon AAT; the aspartic acid codon GAT; the glutamic acid codon GAA; the glycine codons GGA, GGG, and GGT; the histidine codon CAC; the isoleucine codon ATT; the leucine codons CTC, CTT, and TTG; the phenylalanine codon TTT; the proline codon CCA; the serine codons TCA and TCG; and the valine codon GTC.
[0066] In some embodiments, the preferred codons include those codons that are used 100% frequently by genes encoding highly expressed proteins. For example, if the gene encoding the highly expressed protein is a tubulin III gene, the preferred codons may include the alanine codon GCC; the cysteine codon TGC; the glutamine codon CAG; the lysine codon AAG; the threonine codon ACC; the tyrosine codon TAC; and the stop codon TGA. In some embodiments, the preferred codons include those codons that are used more frequently by genes encoding highly expressed proteins than other synonymous codons that encode the same amino acid. For example, if the gene is a tubulin III gene, the preferred codons may also include the arginine codons AGG, CGA, CGC, and CGG; the asparagine codon AAC; the aspartic acid codon GAC; the glutamic acid codon GAG; the glycine codon GGC; the histidine codon CAT; the isoleucine codon ATC; the leucine codon CTG; the phenylalanine codon TTC; the proline codons CCC and CCT; the serine codons AGC, AGT, and TCT; and the valine codon GTG.
[0067] In some embodiments, the host cell is a human cell. In some embodiments, the host cell is an iPSC cell, or an iPSC-derived differentiated cell. In some embodiments, the host cell is an iPSC-derived neuron. In some embodiments, the host cell is a cortical neuron, a dopaminergic neuron, or a motor neuron. In some embodiments, the host cell is an iPSC-derived macrophage. In some embodiments, the host cell is an iPSC-derived cardiomyocyte. In some embodiments, the host cell is an iPSC-derived hepatocyte. In some embodiments, the host cell is a HEK293 cell. For each of these embodiments, in some embodiments, the gene encoding a highly expressed protein is a tubulin III gene.
[0068] In some embodiments, the host cell is a bacterial cell. In some embodiments, the host cell is a bacterial cell, such as Escherichia coli, Pseudomonas (e.g., P. aeruginosa, P. putida, P. fluorescens), Lactobacillus (e.g., L. lactis), Streptomyces (e.g., S. coelicolor), Bacillus (e.g., B. subtilis), or other species of bacteria. tilis), Acinetobacter, Agrobacterium, Cupriavidus, Clostridium, Rhodobacter, Marinobacter, Klebsiella, Ralstonia, and Rhodococcus.
[0069] In some embodiments, the host cell is a yeast cell. In some embodiments, the host cell is selected from the genera Saccharomyces (e.g., S. cerevisiae), Schizosaccharomyces (e.g., S. pombe), Candida (e.g., C. albicans), Pichia, Hansenula, Klockera, Schwanniomyces, Rhodosporidium, Yarrowia, and Rhodotorula.
[0070] In some embodiments, the host cell is a fungal cell, hi some embodiments, the host cell is selected from the genera Aspergillus (e.g., A. niger), Penicillium, Rhizopus, Chrysosporium, Myceliophthora, Trichoderma (e.g., T. reesei), Humicola, Acremonium, and Fusarium.
[0071] In some embodiments, the target nucleic acid is a heterologous nucleic acid. In some embodiments, the target nucleic acid is an endogenous nucleic acid.
[0072] In some embodiments, the target nucleic acid encodes a Cas enzyme. In some embodiments, the target nucleic acid encodes Cas9. In some embodiments, the target nucleic acid encodes Cas12a. In some embodiments, the target nucleic acid encodes Cas13Rx.
[0073] The invention provides nucleic acid sequences that are codon optimized by the methods of the invention. In some embodiments, the invention provides a codon optimized nucleic acid encoding Cas9. In some embodiments, the codon optimized nucleic acid encoding Cas9 comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:1.
[0074] In some embodiments, the invention provides a nucleic acid encoding Cas9, wherein the 5' region of the nucleic acid has been codon optimized by the methods of the invention. In some embodiments, the nucleic acid comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:3.
[0075] In some embodiments, the invention provides a codon-optimized nucleic acid encoding Cas12a. In some embodiments, the codon-optimized nucleic acid encoding Cas12a comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:4.
[0076] In some embodiments, the invention provides a codon-optimized nucleic acid encoding Cas13Rx. In some embodiments, the codon-optimized nucleic acid encoding Cas13Rx comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:5.
[0077] The present invention also provides a vector comprising a nucleic acid codon-optimized by the method of the present invention. In some embodiments, the vector comprises a nucleic acid of the present invention. The appropriate vector depends on the host cell used and can be easily identified by a person skilled in the art. In some embodiments, the vector is selected from an adeno-associated virus (AAV) vector, an HIV-based lentivirus vector, an equine immunodeficiency virus (EIV) vector, a feline immunodeficiency virus (FIV) vector, and a herpes simplex virus vector.
[0078] The vector may include an origin of replication, a promoter sequence operably linked to the nucleic acid of the present invention, and one or more of a reporter gene or a selectable marker. The promoter may be homologous or heterologous. The promoter may be constitutive or inducible. In some embodiments, the promoter is inducible and is activated in the presence of an inducer. Inducers include, but are not limited to, sugars, metal salts, and antibiotics. Typically, the promoter is operable in the host cell of interest.
[0079] In some embodiments, the vector comprises a codon-optimized nucleic acid encoding a Cas enzyme. In some embodiments, the vector comprises a codon-optimized nucleic acid encoding Cas9, Cas12a, or Cas13Rx. In some embodiments, the vector comprises a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NOs: 1, 3, 4, or 5.
[0080] The invention also provides host cells comprising a nucleic acid sequence codon-optimized by the methods of the invention. In some embodiments, the host cell comprises a nucleic acid of the invention. In some embodiments, the host cell comprises a vector of the invention.
[0081] In some embodiments, the host cell is a human cell. In some embodiments, the host cell is an iPSC cell, or an iPSC-derived differentiated cell. In some embodiments, the host cell is an iPSC-derived neuron. In some embodiments, the host cell is an iPSC-derived macrophage. In some embodiments, the host cell is an iPSC-derived cardiomyocyte. In some embodiments, the host cell is an iPSC-derived hepatocyte. In some embodiments, the host cell is a HEK293 cell.
[0082] In some embodiments, the host cell is a bacterial cell. In some embodiments, the host cell is a bacterial cell, such as Escherichia coli, Pseudomonas (e.g., P. aeruginosa, P. putida, P. fluorescens), Lactobacillus (e.g., L. lactis), Streptomyces (e.g., S. coelicolor), Bacillus (e.g., B. subtilis), or other species of bacteria. tilis, Acinetobacter, Agrobacterium, Cupriavidus, Clostridium, Rhodobacter, Marinobacter, Klebsiella, Ralstonia, and Rhodococcus. In some embodiments, the host cell is a yeast cell. In some embodiments, the host cell is selected from the genera Saccharomyces (e.g., S. cerevisiae), Schizosaccharomyces (e.g., S. pombe), Candida (e.g., C. albicans), Pichia, Hansenula, Klockera, Schwanniomyces, Rhodosporidium, Yarrowia, and Rhodotorula. In some embodiments, the host cell is a fungal cell.In some embodiments, the host cell is selected from the genera Aspergillus (e.g., A. niger), Penicillium, Rhizopus, Chrysosporium, Myceliophthora, Trichoderma (e.g., T. reesei), Humicola, Acremonium, and Fusarium.
[0083] In some embodiments, the host cell comprises a codon-optimized nucleic acid encoding a Cas enzyme. In some embodiments, the host cell comprises a codon-optimized nucleic acid encoding Cas9, Cas12a, or Cas13Rx. In some embodiments, the host cell comprises a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NOs: 1, 3, 4, or 5.
[0084] The invention provides iPSCs comprising a nucleic acid sequence encoding Cas9, wherein the nucleic acid sequence has been codon optimized by the methods of the invention. In some embodiments, the iPSCs comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:1.
[0085] The invention provides iPSC-derived neuronal cells, wherein the neuronal cells comprise a nucleic acid sequence encoding Cas9, wherein the nucleic acid sequence has been codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived neuronal cells comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:1.
[0086] The invention provides iPSC-derived hepatocytes, wherein the hepatocytes comprise a nucleic acid sequence encoding Cas9, wherein the nucleic acid sequence has been codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived hepatocytes comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:1.
[0087] The invention provides iPSC-derived macrophages, wherein the macrophages comprise a nucleic acid sequence encoding Cas9, wherein the nucleic acid sequence has been codon optimized by the methods of the invention. In some embodiments, the iPSC-derived macrophages comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:1.
[0088] The invention provides iPSC-derived cardiomyocytes, wherein the cardiomyocytes comprise a nucleic acid sequence encoding Cas9, wherein the nucleic acid sequence has been codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived cardiomyocytes comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:1.
[0089] The invention provides iPSCs comprising a nucleic acid sequence encoding Cas9, wherein the nucleic acid sequence has been codon optimized by the methods of the invention. In some embodiments, the iPSCs comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:3.
[0090] The invention provides iPSC-derived neuronal cells, wherein the neuronal cells comprise a nucleic acid sequence encoding Cas9, wherein the nucleic acid sequence has been codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived neuronal cells comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:3.
[0091] The invention provides iPSC-derived hepatocytes, wherein the hepatocytes comprise a nucleic acid sequence encoding Cas9, wherein the nucleic acid sequence has been codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived hepatocytes comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:3.
[0092] The invention provides iPSC-derived macrophages, wherein the macrophages comprise a nucleic acid sequence encoding Cas9, wherein the nucleic acid sequence has been codon optimized by the methods of the invention. In some embodiments, the iPSC-derived macrophages comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:3.
[0093] The invention provides iPSC-derived cardiomyocytes, wherein the cardiomyocytes comprise a nucleic acid sequence encoding Cas9, wherein the nucleic acid sequence has been codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived cardiomyocytes comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:3.
[0094] The invention provides iPSCs comprising a nucleic acid sequence encoding Cas12a, wherein the nucleic acid sequence has been codon-optimized by the methods of the invention. In some embodiments, the iPSCs comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:4.
[0095] The invention provides iPSC-derived neuronal cells, wherein the neuronal cells comprise a nucleic acid sequence encoding Cas12a, wherein the nucleic acid sequence is codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived neuronal cells comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:4.
[0096] The invention provides iPSC-derived hepatocytes, wherein the hepatocytes comprise a nucleic acid sequence encoding Cas12a, wherein the nucleic acid sequence is codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived hepatocytes comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:4.
[0097] The invention provides iPSC-derived macrophages, wherein the macrophages comprise a nucleic acid sequence encoding Cas12a, wherein the nucleic acid sequence is codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived macrophages comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:4.
[0098] The present invention provides iPSC-derived cardiomyocytes, wherein the cardiomyocytes comprise a nucleic acid sequence encoding Cas12a, wherein the nucleic acid sequence is codon-optimized by the methods of the present invention. In some embodiments, the iPSC-derived cardiomyocytes comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:4.
[0099] The invention provides iPSCs comprising a nucleic acid sequence encoding Cas13Rx, wherein the nucleic acid sequence has been codon-optimized by the methods of the invention. In some embodiments, the iPSCs comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:5.
[0100] The invention provides iPSC-derived neuronal cells, wherein the neuronal cells comprise a nucleic acid sequence encoding Cas13Rx, wherein the nucleic acid sequence is codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived neuronal cells comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:5.
[0101] The present invention provides iPSC-derived hepatocytes, wherein the hepatocytes comprise a nucleic acid sequence encoding Cas13Rx, wherein the nucleic acid sequence is codon-optimized by the methods of the present invention. In some embodiments, the iPSC-derived hepatocytes comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:5.
[0102] The invention provides iPSC-derived macrophages, wherein the macrophages comprise a nucleic acid sequence encoding Cas13Rx, wherein the nucleic acid sequence is codon-optimized by the methods of the invention. In some embodiments, the iPSC-derived macrophages comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:5.
[0103] The present invention provides iPSC-derived cardiomyocytes, wherein the cardiomyocytes comprise a nucleic acid sequence encoding Cas13Rx, wherein the nucleic acid sequence is codon-optimized by the methods of the present invention. In some embodiments, the iPSC-derived cardiomyocytes comprise a nucleic acid having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any of SEQ ID NO:5. EXAMPLES
[0104] The present invention is further defined by the following non-limiting examples.
[0105] Example 1 Cas9 silencing in differentiated cells In an attempt to overcome Cas9 silencing in iPSCs, multiple approaches have been used. These include integrating multiple copies of Cas9 into the host cell genome using lentivirus / transposon; testing Cas9 expression under various mammalian expression promoters; and targeting Cas9 to safe harbor sites in the genome. Despite these efforts, researchers have observed that the levels of Cas9 protein in differentiated cells are dramatically reduced compared to those observed in iPSCs. Interestingly, Cas9 mRNA levels remain detectable during differentiation, despite the reduced protein levels.
[0106] We first confirmed that Cas9 expression decreased during differentiation of iPSCs (Figure 1). Using PiggyBac transposase (plasmid schematic 1A), we generated stable Cas9-expressing iPSCs, which were then differentiated into dopaminergic neurons. RNA collected on days 10 and 20 was analyzed for Cas9 and GAPDH expression during differentiation. Cas9 mRNA decreased by approximately 60% by day 20 (Figure 1B).
[0107] In an attempt to circumvent Cas9 silencing, we generated Bob-iNgn2 GAPDH-Cas9 iPSC line. To ensure continuous transcription, Cas9 was inserted into the site of the housekeeping gene GAPDH. GAPDH levels were shown to gradually increase during iPSC-derived neuron differentiation protocols, so we selected GAPDH as a good housekeeping gene to knock-in Cas9 (Figure 1C). Using homologous directed recombination, we knocked in Cas9 just upstream of GAPDH (Figure 2A).
[0108] Cas9 and GAPDH mRNA and protein levels were assessed during rapid (14 days) cortical neuron differentiation driven using an inducible Ngn2 transgene. RT-qPCR results for mRNA levels demonstrated that Cas9 expression was comparable to GAPDH expression during iPSC-to-neuron differentiation (Figure 2B), which was encouraging compared to results previously obtained with randomly integrated Cas9 using transposons.
[0109] Cas9 protein levels were determined using Western blotting (Figures 3A and 3B) and Cas9 nuclease activity was determined using a fluorescent reporter construct (Figures 3C-3E). Despite promising evidence at the mRNA level, cell lines showed loss of Cas9 protein and therefore loss of Cas9 activity after 4 days in differentiation medium. Thus, despite successful knock-in of the endogenous GAPDH gene, the promoter of a housekeeping gene was unable to sustain constitutive expression of Cas9 protein during differentiation of iPSCs into neuronal cell types.
[0110] To determine whether Cas9 silencing at the protein level is the result of protein degradation by either the proteasome or autophagy pathway in differentiated neurons, we used MG132 inhibitor to inhibit proteasomal degradation and bafilomycin A1 (BafA1) to inhibit the autophagy-lysosomal pathway. Experiments in day 7 neurons showed that Cas9 levels were reduced by 60% compared to day 4, and blocking these protein degradation pathways could not restore Cas9 levels (Figure 4A).
[0111] Given that Cas9 levels appeared to drop dramatically between days 4 and 7 of cortical neuron differentiation, consistent with changing the differentiation medium following the cortical neuron generation protocol, we further tested alternative protocols during differentiation to determine whether Cas9 levels were reduced in the presence of medium or could be restored by maintaining the supplements used from days 0 to 4 of the protocol. These experiments demonstrated that Cas9 silencing could not be restored by changing the medium composition (Figure 4B).
[0112] Cas9 codon optimization The presence of Cas9 mRNA but the absence of Cas9 protein suggests that Cas9 transcription is uncoupled from Cas9 translation. Cas9 silencing was particularly prominent from day 4 to day 7 of neuronal differentiation, so we suspected that the neuronal phenotype of cells prevented Cas9 expression. There is considerable evidence that synonymous codon selection in natural mRNAs has evolved in response to diverse selective pressures at both the RNA and protein levels. Therefore, we hypothesized that further codon optimization may be required for Cas9 to function in differentiated cell types.
[0113] The general codon usage of Escherichia coli (E. coli) and humans (taken from the codon usage database by Kazusa (Nakamura, Y. et al. Nucleic acid research 2000 28(1):292) - available at https: / / www.kazusa.or.jp / codon / ) was compared (Figure 5). The comparison showed that the codon usage between the two organisms is almost comparable. However, amino acids such as Asp, Arg, His, and Val, which are highlighted by dashed lines, show significant differences in usage in humans compared to bacteria. These small differences result in considerable changes in protein translation, and the impact is even greater for large proteins such as Cas9, which have more than 1000 codons.
[0114] The existing Cas9 (old-Cas9) sequence (SEQ ID NO: 2) has been optimized for expression in human cells based on the existing gold standard method based on human codon usage (Figure 6).
[0115] We sought to determine whether essentially postmitotic differentiated neurons exhibit a different codon bias from the general codon bias of humans by determining the codon usage of a highly expressed neuronal marker, Tuj1. We analyzed the codon distribution of the established protein-coding transcript of Tubulin III (Ensembl transcript:TUBB3-208ENST00000555576.5; SEQ ID NO:8) using the codon calculator tool available at https: / / www.biologicscorp.com / tools / CodonUsageCalculator / . We compared the codon usage of Tubulin III to the general codon usage of humans (Figure 7).
[0116] The codon usage of tubulin III showed that the codon preference of tubulin III differs from the general codon usage of humans. The main differences are highlighted by dashed boxes (no usage) and triangles (high usage) (Figure 7).
[0117] The codon usage of tubulin III was used to generate a novel codon-optimized Cas9 variant with modified codons (CodOpt-Cas9). The DNA sequence of the codon-optimized Cas9 is shown below, with the modified codons highlighted in bold. [ka] [ka]
[0118] To confirm that only the codons of Cas9 and not the amino acid (protein) sequence were altered, we verified the protein sequences obtained from both variants of Cas9 using the ClustalW protein alignment tool.
[0119] The codon usage of CodOpt-Cas9 was compared to the common codon usage in humans (Figure 8). This comparison demonstrated that optimizing the codons of Cas9 with the codon bias of tubulin III resulted in sequences with substantially different codon usage compared to the common codon usage in humans.
[0120] CodOpt-Cas9 expression and activity We cloned the codon-optimized Cas9 into an expression construct that can be directly compared to the old-Cas9 expression construct (Figure 9). First, we tested and compared the expression of old-Cas9 and CodOpt-Cas9 in HEK293 cells. Advantageously, these experiments demonstrated that CodOpt-Cas9 had increased expression at both the mRNA and protein levels compared to old-Cas9 (Figure 10).
[0121] Using HEK293 cells carrying each of these two mutant forms of Cas9, we performed Cas9 activity assays using a fluorescent reporter plasmid. The results of these Cas9 cleavage assays demonstrate that CodOpt-Cas9 exhibits higher nuclease activity and initiates editing much more quickly than old-Cas9 (Figure 10D). This more rapid cleavage efficiency was also observed when Cas9 was induced to edit a non-essential gene (ST6GALNAC6) in the genome, resulting in the formation of indels (Figure 11). These results suggest that CodOpt-Cas9 achieves higher expression and exhibits more rapid and efficient cleavage than old-Cas9.
[0122] Next, we attempted to express CodOpt-Cas9 in Bob-iNgn2 iPSCs. CodOpt-Cas9 strains were generated using PiggyBac transposase. Cas9 expression was confirmed at both the mRNA and protein levels. Similar to HEK293 cells, iPSCs carrying CodOpt-Cas9 showed high levels of Cas9 mRNA and protein (Figure 12).
[0123] Next, Bob-iNgn2 iPSCs containing either CodOpt-Cas9 or old-Cas9 were differentiated into cortical neurons. Western blotting of Cas9 protein levels showed that Cas9 was readily detectable in differentiated neuronal cells expressing CodOpt-Cas9 (Figure 13). However, as observed in previous experiments, cells expressing old-Cas9 showed a rapid decrease in Cas9 levels as the cells became more neuronal phenotype (Figure 13).
[0124] These results indicate that optimizing the codon usage of Cas9 to reflect that of tubulin III, a highly expressed neuronal marker protein, significantly improves Cas9 expression in iPSCs and iPSC-derived neurons. Advantageously, Cas9 expression persists throughout neuronal differentiation, which greatly enhances the potential research applications of both iPSC-derived cell lines and the CRISPR-Cas9 system (Figure 13).
[0125] Codon optimization as a tool to control expression levels Codon usage has recently attracted attention as a key determinant of translation elongation rate and co-translational protein folding, with preferred codons enhancing translation efficiency and folding fidelity. The unequal use of synonymous codons, termed codon bias, and the universal nature of this bias from yeast to humans suggest the existence of a secondary code within the more general genetic code. This secondary code has emerged as a major regulator of translation rate and co-translational protein folding, and thereby as a key determinant of the cellular levels of specific proteins.
[0126] Based on the observation that CodOpt-Cas9 achieved better expression than old-Cas9 at both the mRNA and protein levels in HEK293 cells and iPSCs, we tested whether the level of Cas9 could be tuned through partial codon optimization. A Cas9 variant was generated in which the first 606 amino acid codons were optimized based on the codon usage of tubulin III, while the remaining codons were left unchanged. This version of Cas9, encoding a protein with a codon-optimized N-terminal region, is represented by SEQ ID NO:3 and is referred to herein as NOpt-Cas9.
[0127] Bob iPSC cell lines containing old-Cas9, CodOpt-Cas9, and NOpt-Cas9 were generated with PiggyBac integration (Figure 14) and then differentiated into neurons. Cas9 cleavage efficiency was determined throughout the iPSC stage and various stages of neuronal differentiation.
[0128] NOpt-Cas9 and CodOpt-Cas9 were found to have better cleavage efficiency than Old-Cas9 as cells progressed toward a neuronal fate (Figure 15). Interestingly, experiments performed on more differentiated cells (day 10 and day 14 neurons) demonstrated that the cleavage efficiency of NOpt-Cas9 was slightly reduced compared to CodOpt-Cas9. Despite the reduced editing efficiency, NOpt-Cas9 showed higher cleavage than old-Cas9 at these time points (Figures 15 and 16). Cleavage efficiency was assessed at days 4 and 7 after cell transduction.
[0129] We also evaluated Cas9 expression at the mRNA and protein levels to determine how partial optimization affects transcription and translation. Both full and partial codon optimization of Cas9 results in increased mRNA levels and sustained expression during differentiation (Figure 17A). Interestingly, in cell lines containing NOpt-Cas9, Cas9 levels in neurons were significantly reduced after day 7 (Figures 17B and C, boxes highlight comparison). While this is reflected in a decrease in editing efficiency by NOpt-Cas9, it suggests that modified codons contribute to sustained protein expression as neurons mature in vitro.
[0130] Cas9 expression in non-neuronal iPSC-derived cell types Similar to iPSC-derived neurons, robust and sustained expression of Cas9 has not been achieved in other iPSC-derived cell types, such as hepatocytes and macrophages. Thus, the inability to perform CRISPR-Cas9 genome-wide screening limits the use of these cell lines to a progenitor state, similar to the limitations observed when performing Cas9 screening in differentiated neurons. Therefore, we decided to determine whether Cas9 expression could be achieved when CodOpt-Cas9 iPSC lines were differentiated into other cell types.
[0131] Hepatocytes were derived from iPSCs based on the protocol established by (Hannan et al. Nature protocols. 8, 430-437 (2013)). Similar to differentiating neurons, it has been observed that Cas9 levels decline rapidly after day 7 of differentiation as the cells undergo multiple morphological changes before being committed to an epithelial lineage.
[0132] Bob iPSC cells carrying either old-Cas9 or CodOpt-Cas9 were differentiated into hepatocytes and cell pellets were collected at days 0, 4, and 10 of differentiation. Western blotting revealed that CodOpt-Cas9 levels were significantly higher than old-Cas9 levels in iPSC-derived hepatocyte-like cells, especially from day 7 onwards (Figure 18).
[0133] These results demonstrate that CodOpt-Cas9 can achieve and maintain high expression levels in non-neuronal iPSC-derived cells. Advantageously, these results demonstrate that significant improvements in expression of target nucleic acids can be achieved across a range of cell types, including cells that do not normally express genes encoding highly expressed proteins on which codon optimization is based.
[0134] summary These results suggest that optimizing the codons of a target nucleic acid using the codon bias of a gene encoding a highly expressed protein can significantly improve the expression of that nucleic acid in a range of cell types, even in cells that do not express the highly expressed gene.Furthermore, the target nucleic acid can be partially codon-optimized to regulate the expression level.Therefore, the method described herein can be used as a solution to overcome Cas9 silencing and enable CRISPR-Cas9 genome-wide screening to be performed in various cell lines, including differentiated cell types.
[0135] Example 2 Codon optimization of Cas12a and Cas13Rx In addition to Cas9, Cas12a and Cas13Rx have also emerged as promising tools for gene editing. These CRISPR Cas proteins are used to edit DNA and RNA, respectively, thereby greatly enhancing the potential of gene editing technology. We analyzed existing variants of Cas12a and Cas13Rx to determine whether codon optimization was properly performed for human mammalian cells.
[0136] Codon-optimized Cas12a The codon usage of existing variants of Cas12a is based on the existing gold standard with a similar optimization pattern as observed for old-Cas9 (Figure 19). The starting Cas12a sequence was obtained from addgene plasmid IDs 160573 and 78744 and is represented by SEQ ID NO:6. [ka] [ka]
[0137] As described above, we codon-optimized the Cas12a sequence to match the codon usage of tubulin III (Figure 7). The codon-optimized Cas12a (CodOpt-Cas12a) DNA sequence is represented in SEQ ID NO: 4, where the modified codons are highlighted in bold. [ka] [ka] [ka] The codon usage of CodOpt-Cas12a (Figure 20) is similar to that of tubulin III (Figure 7).
[0138] Codon-optimized Cas13Rx A similar approach was taken for Cas13Rx. The codon usage of existing variants of Cas13Rx is based on the existing gold standard with a similar optimization pattern as observed for old-Cas9 (Figure 21). The starting Cas13Rx sequence was obtained from addgene plasmid ID141320 and is represented by SEQ ID NO:7. [ka] [ka] [ka]
[0139] The Cas13Rx sequence was codon-optimized using the codon bias of tubulin III (Figure 7). The DNA sequence of the codon-optimized Cas13Rx (CodOpt-Cas13Rx) is represented in SEQ ID NO:5, where the modified codons are highlighted in bold. [ka] [ka] The codon usage of CodOpt-Cas13Rx (Figure 22) is similar to that of tubulin III (Figure 7).
[0140] Based on the encouraging results presented herein with respect to Cas9, the inventors hypothesize that other codon-optimized genes, such as the codon-optimized variants of Cas12a and Cas13Rx described herein, will be beneficial for performing genome editing in various iPSC-derived cell types, such as, for example, neurons and hepatocytes.
[0141] Example 3 Codon optimization of L-lactate dehydrogenase We next attempted to confirm whether this novel codon optimization technique could be applied to other bacterial genes. Lldr from E. coli (constituting an element of the L-lactate dehydrogenase operon) was codon-optimized using the existing gold standard method based on human codon usage (denoted as "normal optimization" in Figures 23 and 24) and the codon bias of tubulin III described herein (denoted as "novel optimization" in Figures 23 and 24) and used to construct two plasmids. Both plasmids also contained an eGFP fluorescent reporter, allowing for the evaluation of transfection efficiency.
[0142] HEK293 cells and iPSCs were transfected with a plasmid carrying the gold standard (regular) optimized gene or a plasmid carrying the tubulin III (novel) optimized gene. Transfection efficiency was measured using flow cytometry 3 days after transfection (CytoFLEX, Beckman Coulter Life Sciences, Indianapolis US). Cell pellets were collected 5 days after transfection for Western blotting to determine the expression level of the Lldr gene using a c-myc tagged antibody.
[0143] The starting E. coli LlDr sequence is represented in SEQ ID NO:9. [ka]
[0144] The LlDr sequence codon optimized based on human codon usage is represented in SEQ ID NO:10. [ka]
[0145] The LlDr sequence, codon-optimized using the codon bias of tubulin III, is represented in SEQ ID NO: 11, with the modified codons highlighted in bold. [ka]
[0146] result No significant difference was observed in the transfection efficiency of iPSCs and HEK293 cells with the two plasmids (Figure 24A). Western blotting demonstrated that the novel optimization approach based on the codon bias of tubulin III resulted in increased expression of LlDr in both iPSCs and HEK293 cells compared to the conventional gold standard optimization approach (Figures 24B and 24C).
[0147] The expression of the Lldr gene was significantly increased by a (novel) codon optimization method based on tubulin III codon bias in both HEK293 cells and iPSCs. These experiments demonstrate that this novel method of codon optimization is beneficial to promote and protect the expression of target genes in iPSC-derived cell types and is ideal for modulating gene expression in target cell types.
[0148] These results demonstrate that the codon optimization approach described herein avoids gene silencing upon iPSC differentiation and also promotes the transcription and translation of target genes in the desired cell type.
[0149] Materials and Methods construct All constructs were designed on the backbone created by Metzakopian et al. Sci Rep. 2017 22;7(1):2244. These constructs carry both PiggyBac inverted terminal repeats (PB-transposon) to allow transposase-mediated genome integration, and HIV-1 long terminal repeats (pKLV-PB-backbone) to allow lentiviral genome integration. Every novel construct created was generated using pre-synthesized gene blocks (IDT) integrated into the backbone using Gibson assembly. The three Cas9 variants used were driven by the EF1A promoter and carried blasticidin antibiotic resistance. Genome-targeted constructs were generated using Gibson assembly via PCR fragments amplified from existing plasmids / extracted genomic DNA. A schematic of each construct created and used in this study is shown in the figure. When stable integration by gene transposition was required, a plasmid encoding PiggyBac transposase (HyPBase (Yusa et al. PNAS 2011 108(4):1531-1536)) was co-transfected.
[0150] cell culture All materials and plasticware used for routine cell culture purposes were obtained from Sigma unless otherwise stated.
[0151] HEK293 cells HEK293 cells were cultured in Dulbecco's modified essential medium (Gibco) supplemented with penicillin (100 U / ml), streptomycin (100 μg / ml), L-glutamine (2 mM), and 15% fetal bovine serum. Cells were periodically split when they reached 70% confluence using trypsin-EDTA solution (Sigma), and 1 / 10 of the population was reseeded into a new dish.
[0152] Bob-iNgn2-opti-ox IPS cells TRE-inducible Ngn2-driven Bob iPSCs were kindly provided by Dr. Mark Kotter. Bob-iNgn2-iPSCs were cultured and maintained according to established protocols (Pawlowski et al.Stem Cell Reports 2017 8(4):803-812). Briefly, iPSCs were maintained in TeSR E8 complete medium with supplements (Stem Cell) on vitronectin-coated plates. Upon reaching 70% confluence, iPSCs were detached using 0.5 mM EDTA solution in PBS. After 5 min of incubation, cells were triturated and re-seeded (1 / 4-1 / 6). If gene targeting / transfection needed to be performed, cells were brought into a single cell suspension using Accutase (Stem Cell) for 5 min. Suspended cells were spun down, counted, and then re-seeded in the required number in E8 medium (with Rock inhibitor) on vitronectin-coated plates.
[0153] Differentiation of Bob-iNgn2-opti-ox into cortical neurons To induce differentiation of cortical neurons, iPSCs were made into a single cell suspension and plated on Geltrex-coated plates at 25k cells / cm. 2The following day, cells were fed with differentiation medium containing DMEM / F12 (Gibco), N2 supplement (1x), L-glutamine (1x), non-essential amino acids (1x), 2-mercaptoethanol (5 μM), Pen-Strep (1x), and doxycycline (1 μg / ml) for two consecutive days. From the third day, cells were fed with differentiation medium containing Neurobasal (Gibco), B27 supplement (1x), L-glutamine (1x), 2-mercaptoethanol (5 μM), Pen-Strep (1x), doxycycline (1 μg / ml), NT3 (4 pg / ml), and BDNF (100 pg / ml). Medium was changed daily until the sixth day of differentiation, and then every other day until the end of the experiment.
[0154] Lentivirus production Lentiviruses were generated in the HEK293 FT cell line using the ViraPower Lentiviral Expression system (Invitrogen) according to the manufacturer's instructions as described (Dull et.al. J Virol 1998, Cribbs et.al. BMC Biotechnol 2013) or using the lentiviral packaging plasmid psPAX2 (Addgene, plasmid no. 12260) and the pMD2.G envelope plasmid containing VSV-G (Addgene, plasmid no. 12259). HEK293 FT cells were cultured in DMEM supplemented with 10% FBS (Gibco) and grown on plates coated with 0.02% gelatin (Sigma). Virus production was performed in Opti-Mem (Gibco) using established protocols. Virus was harvested from the medium 3 days after transfection. The supernatant was passed through a 45 μM PVDF filter and then the virus was pelleted by spinning at 6000 g for 18 hours at 4° C. The next day, the virus pellet was dissolved in PBS, aliquoted and stored at −80° C.
[0155] Plasmid transfection HEK293 cells and Bob-iNgn2-iPSCs were grown in 6-well plates to 70% internal confluence. Cells were dissociated with trypsin / EDTA or Accutase, respectively, and resuspended in medium for reverse transfection (approximately 1 × 10 cells in 250 μl per transfection). 6 cells). All cells were transfected with 200ng of PiggyBac transposase and 1000ng of Cas9 construct. Transfection was performed using Lipofectamine LTX (Invitrogen) for HEK293 cells and Lipofectamine-STEM (Invitrogen) for Bob-iNgn2-iPSCs according to the manufacturer's instructions. The medium was changed after 24 hours. Stably transfected cell lines were generated by selection with blasticidin (10μg / ml) for at least 10 days after transfection. Selection against gRNA plasmids or reporter plasmids was omitted when these were introduced.
[0156] Lentiviral transduction All transductions were performed on single cell suspension cells in medium containing lentivirus and polybrene (4 μg / ml) (Sigma) at 37° C. Cells were cultured overnight at 37° C. and the medium was changed the next day.
[0157] Flow cytometry analysis All cells, including nontransfected controls, were harvested periodically (mainly on days 4 and 7 posttransfection) and analyzed for BFP / GFP fluorescence in a flow cytometer (CytoFLEX, Beckman Coulter Life Sciences, Indianapolis US).
[0158] Codon Optimization Codon optimization of Cas9 was performed to reflect the codon usage of tubulin III, a neuronal pan-marker. Codon usage analysis of Cas9, Cas12a, and Cas13Rx was performed using the tool available at https: / / www.biologicscorp.com / tools / CodonUsageCalculator / .
[0159] Codons in the target nucleic acid sequence (Cas9 / Cas12a / Cas13Rx) were manually inspected and, if necessary, changed to preferred codons in the reference nucleic acid sequence (Tubulin III). Codons were preferentially changed with the codons with higher priority for each amino acid. When changing multiple codons within a 60-base sequence, we attempted to achieve a distribution that reflected the codons in the reference sequence. Because sequences with a GC content of more than 60% can be difficult to synthesize, the distribution of nucleotides A, T, G, and C every 300 bases was also considered. Therefore, for amino acids encoded by three or more synonymous codons, A- and T-rich codons were introduced where necessary and applicable.
[0160] Western blotting Cell lysates from HEK293 cells or Bob-iNgn2-iPSCs and neurons were collected after PBS washing at various time points of the experiment. Total cell proteins were extracted using RIPA buffer (SIGMA) supplemented with 1x PIC. Protein amounts were determined using the Bradford assay and 30 μg of lysates were subjected to electrophoresis on 4–15% Mini-PROTEAN® TGX™ Precast Protein Gels (Biorad). Proteins were transferred onto PVDF membranes (Millipore) using the Turboblot system (Biorad). Transferred proteins were immunoblotted for Cas9 ((7A9-3A3) mouse mAb #14697, dilution 1:800) and Gapdh (Sigma, #G8795, dilution 1:4000).
[0161] Quantitative RT-PCR (RT-qPCR) Total RNA was extracted using the RNeasy Mini kit (Qiagen) according to the manufacturer's instructions. First-strand cDNA was synthesized using qScript cDNA Supermix (Quantabio) according to the manufacturer's protocol. All qPCR tests were performed using Sybr green primers designed to amplify the CDS of the genes of interest. qPCR runs were performed on a QuantStudio real-time PCR system (Applied Biosystems). Samples were run in triplicate from three independent experiments for both genes of interest and a housekeeping gene (18S RNA). Expression levels were normalized to 18s RNA.
[0162] Graphical display All graphical representations were generated using GraphPad Prism 7 software. SEQ ID NO:8 - Ensembl transcript TUBB3-208ENST00000555576.5. [ka]
Claims
1. 1. A method for optimizing codons of a target nucleic acid sequence for expression in a host cell, comprising altering the codon usage of the target nucleic acid sequence based on the codon usage of genes encoding proteins that are highly expressed in the host cell or genes encoding proteins that are highly expressed in cells derived from the same species as the host cell.
2. 10. The method of claim 1, comprising substituting one or more non-preferred codons in the target nucleic acid sequence with preferred synonymous codons: (a) the non-preferred codon is a codon used infrequently by a gene encoding said highly expressed protein; and (b) The method, wherein the preferred codon is a codon that is frequently used by a gene that encodes the highly expressed protein.
3. 3. The method of claim 2, wherein non-preferred codons are used by genes encoding the highly expressed proteins at a frequency of less than 50%, less than 45%, less than 40%, less than 35%, less than 33%, less than 30%, less than 25%, less than 20%, less than 16%, less than 15%, less than 10%, less than 5%, or 0%.
4. 3. The method of claim 2, wherein preferred codons are used at a frequency of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% by genes encoding said highly expressed proteins.
5. 3. The method of claim 2, comprising replacing at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% of non-preferred codons in the target nucleic acid sequence with preferred synonymous codons.
6. 3. The method of claim 2, comprising replacing all non-preferred codons used at a frequency of 0% by genes encoding the highly expressed proteins in the target nucleic acid sequence with preferred synonymous codons.
7. 3. The method of claim 2, comprising replacing all non-preferred codons in the region of said target nucleic acid sequence encoding the N-terminal region of a protein with preferred synonymous codons.
8. 8. The method of claim 7, comprising replacing all non-preferred codons in a 5' region of the target nucleic acid with preferred synonymous codons, optionally comprising replacing all non-preferred codons with preferred synonymous codons in at least the first 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 codons starting from the 5' end of the target nucleic acid.
9. The method of claim 1 , wherein the highly expressed protein is a housekeeping protein or a cell marker protein.
10. 10. The method of claim 9, wherein the highly expressed protein is selected from GAPDH, β-tubulin, β-actin, and tubulin III.
11. The method of claim 10, wherein the highly expressed protein is tubulin III.
12. 12. The method of claim 11, wherein the one or more non-preferred codons are selected from the alanine codons GCA, GCG, and GCT; the arginine codons AGA and CGT; the cysteine codon TGT; the glutamine codon CAA; the isoleucine codon ATA; the leucine codons CTA and TTA; the lysine codon AAA; the proline codon CCG; the serine codon TCC; the threonine codons ACA, ACG, and ACT; the tyrosine codon TAT; the valine codons GTA and GTT; and the stop codons TAA and TAG.
13. 12. The method of claim 11, wherein the one or more non-preferred codons are selected from the asparagine codon AAT; the aspartic acid codon GAT; the glutamic acid codon GAA; the glycine codons GGA, GGG, and GGT; the histidine codon CAC; the isoleucine codon ATT; the leucine codons CTC, CTT, and TTG; the phenylalanine codon TTT; the proline codon CCA; the serine codons TCA and TCG; and the valine codon GTC.
14. 12. The method of claim 11, wherein the preferred codons are selected from the alanine codon GCC; the cysteine codon TGC; the glutamine codon CAG; the lysine codon AAG; the threonine codon ACC; the tyrosine codon TAC; and the stop codon TGA.
15. 12. The method of claim 11, wherein the preferred codons are selected from the arginine codons AGG, CGA, CGC, and CGG; the asparagine codon AAC; the aspartic acid codon GAC; the glutamic acid codon GAG; the glycine codon GGC; the histidine codon CAT; the isoleucine codon ATC; the leucine codon CTG; the phenylalanine codon TTC; the proline codons CCC and CCT; the serine codons AGC, AGT, and TCT; and the valine codon GTG.
16. 10. The method of claim 1, wherein the host cell is selected from a human cell, a bacterial cell, a yeast cell, and a fungal cell.
17. 17. The method of claim 16, wherein the host cell is a HEK293 cell or a human induced pluripotent stem cell (iPSC).
18. 17. The method of claim 16, wherein the host cell is an iPSC-derived differentiated cell, optionally selected from an iPSC-derived neuron (e.g., a cortical neuron, a dopaminergic neuron, or a motor neuron); an iPSC-derived macrophage; an iPSC-derived cardiomyocyte; and an iPSC-derived hepatocyte.
19. 2. The method of claim 1, wherein the target nucleic acid sequence encodes a Cas protein, optionally wherein the Cas protein is selected from Cas9, Cas12a, and Cas13Rx.
20. A nucleic acid comprising a nucleic acid sequence that has been codon-optimized by the method of any one of claims 1 to 19.
21. A nucleic acid comprising a nucleic acid sequence that has been codon-optimized for improved expression in a host cell, wherein the codon usage frequency of the nucleic acid sequence corresponds to the codon usage frequency of a gene encoding a protein that is highly expressed by the host cell or to the codon usage frequency of a gene encoding a protein that is highly expressed in a cell derived from the same species as the host cell.
22. 22. The nucleic acid of claim 21, wherein the codon-optimized nucleic acid sequence contains a lower frequency of non-preferred codons than a non-optimized nucleic acid sequence encoding the same amino acid sequence.
23. 22. The nucleic acid of claim 21, wherein the codon-optimized nucleic acid sequence contains a higher frequency of preferred codons than a non-optimized nucleic acid sequence encoding the same amino acid sequence.
24. A nucleic acid comprising a nucleic acid sequence encoding Cas9 and having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:1 or SEQ ID NO:
3.
25. A nucleic acid encoding Cas12a and comprising a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:
4.
26. A nucleic acid encoding Cas13Rx and comprising a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:
5.
27. A vector comprising the nucleic acid according to any one of claims 21 to 26.
28. A host cell comprising the vector of claim 27.
29. 29. The host cell of claim 28, wherein the host cell is selected from a human cell, a bacterial cell, a yeast cell, and a fungal cell.
30. 30. The host cell of claim 29, wherein the host cell is a HEK293 cell or a human induced pluripotent stem cell (iPSC).
31. 30. The host cell of claim 29, wherein the host cell is an iPSC-derived differentiated cell, optionally selected from an iPSC-derived neuron (e.g., a cortical neuron, a dopaminergic neuron, or a motor neuron); an iPSC-derived macrophage; an iPSC-derived cardiomyocyte; and an iPSC-derived hepatocyte.