Chimeric polynucleotides and methods of using same

The chimeric DNA molecule with optimized codon usage and a linked endogenous sequence enhances polypeptide expression and stability, addressing the challenge of optimizing heterologous gene expression for improved biomanufacturing efficiency.

WO2025163647A1PCT designated stage Publication Date: 2025-08-07RAMOT AT TEL AVIV UNIVERSITY LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IL2025/050118
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-04
Filing Date
2025-02-04
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing methods struggle to optimize the expression levels of heterologous genes encoding polypeptides of interest, such as insulin, in a generalized and orthogonal manner, which is crucial for reducing biomanufacturing costs and enhancing biotechnological feasibility.

Method used

A chimeric DNA molecule is designed with a first nucleic acid sequence encoding a polypeptide of interest and a second nucleic acid sequence encoding an endogenous polypeptide, linked operably, optionally with a third sequence, and optimized for codon usage in the target cell, incorporating a stop codon to enhance expression and stability.

Benefits of technology

The chimeric DNA molecule significantly increases the expression and stability of the polypeptide of interest, outperforming conventional methods by maintaining higher expression levels and genetic stability over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050118_07082025_PF_FP_ABST
    Figure IL2025050118_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a chimeric DNA molecule including: (a) a first nucleic acid sequence encoding a polypeptide of interest, wherein a 3' end the first nucleic includes a stop codon; and (b) a second nucleic acid sequence encoding an endogenous polypeptide of a target cell, wherein the first nucleic acid sequence is located upstream to the second nucleic acid sequence in the chimeric DNA molecule, and wherein the first and second nucleic acid sequences are operably linked. Further provided is a method for increasing expression of a polypeptide of interest in a target cell, including culturing a cell including the chimeric DNA molecule, such that a polypeptide of interest encoded by the first nucleic acid sequence of the chimeric DNA molecule is expressed.
Need to check novelty before this filing date? Find Prior Art

Description

CHIMERIC POLYNUCLEOTIDES AND METHODS OF USING SAMEREFERENCE TO AN ELECTRONIC SEQUENCE LISTING

[0001] The contents of the electronic sequence listing (RMT-P-O33-PCT.xml; size: 8,461 bytes; and date of creation: February 3, 2025) is herein incorporated by reference in its entirety.CROSS REFERENCE TO RELATED APPLICATIONS

[0002] This Application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 549,519, titled “CHIMERIC POLYNUCLEOTIDES AND METHODS OF USING SAME” filed 4 February 2024, the contents of which are incorporated herein by reference in their entirety.FIELD OF INVENTION

[0003] The present invention is in the field of synthetic biology, including improvement of expression of recombinant proteins.BACKGROUND

[0004] Heterologous gene expression is fundamental for biotechnology and as such, optimizing the expression level of heterologous genes hold immense importance. Reaching higher level of expression with heterologous genes encoding polypeptide(s) of interest, such as insulin, will reduce biomanufacturing levels and also allow many biotechnologies to be economically feasible.

[0005] Therefore, there is still a great need for a method allowing for the optimization of expression levels of such a polypeptide of interest in an orthogonal and generalized manner which holds great economical and scientific value.SUMMARY

[0006] According to the first aspect, there is provided a chimeric DNA molecule comprising: (a) a first nucleic acid sequence encoding a polypeptide of interest, wherein a 3' end of the first nucleic comprises a stop codon; and (b) a second nucleic acid sequence encoding an endogenous polypeptide of a target cell, wherein the first nucleic acid sequence is located upstream to the second nucleic acid sequence in the chimeric DNA molecule, and wherein the first and second nucleic acid sequences are operably linked.

[0007] According to another aspect, there is provided an expression vector or a plasmid comprising the chimeric DNA molecule of the invention.

[0008] According to another aspect, there is provided a cell comprising: (a) the chimeric DNA molecule of the invention; (b) the expression vector or plasmid of the invention; or (c) both (a) and (b).

[0009] According to another aspect, there is provided an extract, lysate, homogenate, and any fraction thereof, of the cell of the invention.

[0010] According to another aspect, there is provided a composition comprising any one of: (a) the chimeric DNA molecule of the invention; (b) the expression vector or plasmid of the invention; (c) the cell of the invention; (d) the extract, lysate, homogenate, and any fraction thereof of the invention; or (e) any combination of (a) to (d), and an acceptable carrier.

[0011] According to another aspect, there is provided a method for increasing expression of a polypeptide of interest in a target cell, the method comprising culturing the cell of the invention such that the polypeptide of interest encoded by the first nucleic acid sequence is expressed.

[0012] According to another aspect, there is provided a method for increasing genetic stability of a cell producing a polypeptide of interest, the method comprising culturing the cell of the invention such that the polypeptide of interest encoded by the first nucleic acid sequence is expressed.

[0013] In some embodiments, the chimeric DNA molecule further comprises a third nucleic acid sequence located between the first nucleic acid sequence and the second nucleic acid sequence.

[0014] In some embodiments, the chimeric DNA molecule is codon optimized for expression in the target cell.

[0015] In some embodiments, the endogenous gene is a gene highly expressed in the target cell, a gene being essential for viability / fitness of the target cell, or both.

[0016] In some embodiments, the first nucleic acid sequence and the second nucleic acid sequence are in-frame or not in-frame.

[0017] In some embodiments, the cell is a prokaryote cell or a eukaryote cell.

[0018] In some embodiments, the cell is a recombinant cell, a transgenic cell, a transfected cell, or a transduced cell.

[0019] In some embodiments, the cell is characterized by increased or over expression of the polypeptide of interest compared to a control being devoid of the chimeric DNA molecule.

[0020] In some embodiments, the cell is characterized by expression level of the endogenous polypeptide being essentially similar or increased compared to a control cell being devoid of the chimeric DNA molecule; the expression vector or plasmid; or both.

[0021] In some embodiments, the extract, lysate, homogenate, and any fraction thereof, comprises any one of: the polypeptide of interest, the endogenous polypeptide, an mRNA transcribed from the first nucleic acid sequence, an mRNA transcribed from the second nucleic acid sequence, and any combination thereof.

[0022] In some embodiments, the extract, lysate, homogenate, and any fraction thereof, consists essentially of the polypeptide of interest.

[0023] In some embodiments, the composition consists essentially of the polypeptide of interest and the acceptable carrier.

[0024] In some embodiments, increased is compared to a control being devoid of the chimeric DNA molecule.

[0025] In some embodiments, increasing expression comprises increasing any one of: abundance and / or secretion of mRNA transcribed from the first nucleic acid sequence, abundance and / or secretion of the polypeptide of interest, and both.

[0026] In some embodiments, increased abundance of mRNA comprises increased: mRNA transcription, mRNA stability, or both.

[0027] In some embodiments, the mRNA transcribed from the second nucleic acid sequence, the third nucleic acid sequence, or both, is not translated.

[0028] In some embodiments, increasing is compared to a control cell.

[0029] In some embodiments, culturing is for a period of 25 to 75 days.

[0030] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention pertains. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the invention, exemplary methods and / or materials are described below. In case of conflict, the patent specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be necessarily limiting.

[0031] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description given hereinafter. However, it should be understood that the detailed description and specific examples, while indicating preferredembodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description.BRIEF DESCRIPTION OF THE FIGURES

[0032] Fig. 1 includes a vertical bar graph showing expression level of insulin as measured by an enzyme linked immunosorbent assay (ELISA) kit for the detection of insulin. Expression levels are normalized to a construct containing insulin alone. The data shown are an average of three repeats. The P values of the Z-test are 0.003645 and 0.000641 in accordance.

[0033] Fig. 2 includes a scheme of a non-limiting representation of the basic overlay of the current algorithm. Using a set of relevant features generated by the inventors or collected from the literature, the inventors were able to use data of GFP expression levels when attached to different open reading frames (ORFs) in yeast in order to predict beneficiary fusing counterpart for insulin.

[0034] Fig. 3 includes a schematic illustration of the construct of insulin expression cassette with and without the linking to the endogenous gene, as used herein. The linker is rationally designed so as to allow minimum spatial disturbance to both genes (e.g., insulin, and the endogenous essential gene (EG)).

[0035] Figs. 4A-4F include illustrations showing comprehensive approach to enhancing evolutionary stability through gene fusion. (4A) Standard production of a heterological gene - The Gene of Interest (GOI) is inserted into the genome, in place of an existing, non-essential gene. Any mutations that reduce expression or induce misfolding would prove advantageous due to the lower metabolic burden. These mutations proliferate, and a batch must be replaced. (4B) Replacement of endogenous gene with fusion gene - An endogenous gene is removed and replaced by a fusion gene composed of the GOI and EG. Following, mutations leading to loss of expression or misfolding, would be deleterious to the EG, resulting in host death. This limits the spread of many mutations. The EG is selected by an ML model trained on experimental data. (4C) Selection of linker - Different linkers may lead to interaction between the fused proteins, and misfolding. Using biophysical models and a database of fusion linkers, a linker is selected to minimize structural changes between the fused and unfused state. (4D) Sequence optimization - By optimizing the sequence of the GOI and linker, hypermutable sites are avoided, codon usage bias is maximized, and weak mRNA folding is enforced at the start of gene. This further improves stability and expression. (4E) Leaky stop codon enables translation of both GOI and fusion gene - A leaky stop codon is placed between the GOI and linker. Due to partial read-through, both theprotein of interest and the fusion protein are generated. By informed selection of stop codon, large quantities of the protein of interest and just viable quantities of the fusion protein are produced. The GOI’s mutational stability is further enhanced, as more mutations would prove deleterious. (4F) Production of the heterological gene, aided by the current design - The new process has much higher mutational stability, reducing the need to replace batches. The optimized cells exhibit much higher expression.

[0036] Figs. 5A-5D include a block diagram and graphs. (5A) Gene selection model pipeline - the input experimental data comprised measurements for 6,685 EGs and 2 GOI. Biologically meaningful features were extracted to describe the EG, the GOI, and their interactions. For each randomized split of the training and validation data, feature processing and hyperparameter tuning were performed to optimize model performance. After analyzing results across all splits, an ensemble approach combining K-Nearest Neighbors (KNN) and XGBoost (XGB) models was selected. (5B) Performance among top 3 recommendations - The different models were compared on their top recommendations. The test data was resampled with the bootstrap method, and the top recommendations were converted to quantile within this sample. The best performance among the top 3 candidates was presented (x-axis), indicating the expected performance when testing 3 fusion genes, and the height (y-axis) of each bin represents what percentile of samples fall within. Near-optimal performance is attained. (5C) Performance for top recommendation - A similar comparison was performed, taking only the top candidate. There is a very low likelihood of failure, and good performance can be expected. (5D) Distribution of Top 20 Shapley Values Across 20 XGB Models - The graph shows the distribution of the top 20 Shapley Values, averaged across 20 XGB models, to assess feature importance in predicting fusion gene performance. The tRNA Adaptation Index (tAI) emerges as the most predictive feature. Following, features of importance include those related to GC content, codon usage bias, alternative ORF lengths, mRNA folding energy, and amino acid composition similarity between EG and GOI.

[0037] Figs. 6A-6E include graphs, schemes, a table, and a photograph showing real-world applicability of the approach. (6A) Gene fusion affects expression levels at time 0 - The expression level of proinsulin either unfused or fused to Caf20 or ARC15, with or without leader sequence. As previously demonstrated for the unfused case, the leader sequence is vital for expressing proinsulin. For the fused genes, it proved to be of high importance. This might be due to increased RNA stability or better localization of the protein in the cell. (6B) Estimation of proinsulin production, exponential decay estimation - the inventors took proinsulin measurements every 5 days, 3 repeats, for each of 4 variants. The baseline variant is the originalsequence (by Novo Nordisk). The gene is placed instead of canl. The second variant is identical in structure, but with the proinsulin and linker optimized by ESO. Third and fourth are fusion genes, where proinsulin is fused with Caf20 and Arcl5 respectively. The fusion genes were synthesized according to the current design, including selection of EG, selection linker, addition of leader sequence, and sequence optimization. The inventors assume exponential decay, supported by the high R2on this linear estimation in log scale (P< 10’6). The curves have different decay rates (ANOVA F-test, P<1019). The variants of the current design demonstrate a large improvement in stability. (6C) Mutation accumulated in an in-lab evolution experiment - (i) Nano-pore sequencing of unfused proinsulin shows that after the in-lab evolution experiment, most of the promotor and the reading frame of insulin were erased from the plasmid, (ii) Nanopore sequencing of proinsulin fused to Arc 15 reveals a significant deletion in the promoter, and a few small deletions in the promotor region and the proinsulin coding region, (iii) Nano-pore sequencing of proinsulin fused to Caf20 shows a few mutations in the proinsulin coding region, the promotor and in Caf20. (6D) Normalized proinsulin generation across variants and timespans - based on the exponential decay model, the inventors can integrate the expression over time, giving an estimate of the true value of interest - cumulative expression over time. The inventors normalized all values by the estimated cumulative expression for the baseline variant over 10 days. This table presents the expected amount of proinsulin generated by each variant at different times. The variants of the current design have significantly higher cumulative expression. (6E) Western blot validation - western blot analysis validates the current findings - the fused variants expressed proinsulin for much longer than the unfused variant. Actin measurements were taken for control.

[0038] Figs. 7A-7D include graphs and a vector map. (7A) Ratio between the final and initial fluorescence, 15 days - Each bar represents either GFP fused to a different endogenous gene, or the unfused GFP (2ndfrom left). The value represents the mean ratio between final and initial fluorescence, for an experiment running for 15 day, 3 repeats each. Fused genes have a higher ratio of fluorescence than GFP alone at the end of the experiment (student t - test, P ~ 0.048), thus more stable. The different experiments have different ratios throughout the experiment (Kruskal - H, P ~ 0.035), demonstrating the need for gene selection. The inventors have shown that SEC2 is more stable than GFP alone (student t - test, P ~ 0.047). The inventors successfully predicted the top performing gene with the current model. (7B) Eeaky stop codon fluorescence experiment - The design of the stop codon experiment is displayed for clarity. (7C) Determining the termination efficiency of different leaky stop codons - BFP and mCherry were fused using different leaky stop codons in the linker. The fluorescence of each was measured as proxy forexpression level, normalized by the non-fused expression level. The red bar is mCherry expression, which was fused to the C terminal of BFP, analogous to EG in the current design. The blue bar is the expression of BFP, analogous to the GOI. Eighteen (18) sequences were tested in total. In the graph, 3 representatives of the LI -3 stop codon design are presented, as are 3 representatives of codons not selected, due to read-through rate which was either too low or high. (7D) Proinsulin concentration over time for different leaky stop codons - Proinsulin was fused with Arcl5 four times, each with a different stop codon. This graph illustrates the measured proinsulin concentration over time, for 50 days. Using any leaky stop codon improved the stability of the gene fusion (student t-test at t=40, P<0.00001). Different codons lead to different expression profiles (Kruskal H-test, P < 0.003), emphasizing the need for informed selection. The leaky stop codons were selected such that the last amino acid of proinsulin is maintained, the inventors utilize the design principles outlined by Mangkalaphiban, K. et al., (2021) (“Transcriptome-wide investigation of stop codon readthrough in Saccharomyces cerevisiae”. PLoS Genet 17) and select stop codons which displayed a read-through of 0.1-0.25 in the previous experiment. The sequences used for the stop codon and the 3 nt afterwards were: LI - TAGGCG; L2 - TGAGCG; and L3 - TGACAA.

[0039] Fig. 8 includes a graph showing the distribution and relationship of fluorescence values for GFP and mCherry. This graph illustrates the distribution of fluorescence values in the dataset for both univariate GFP and mCherry measurements, along with their relationship. Due to the wide range and skewed nature of the data, values are displayed on a logarithmic scale. A clear mutual trend is observed between GFP and mCherry fluorescence, as expected, indicating their correlated behavior in the dataset.

[0040] Fig. 9 includes a graph showing a top quantile of actual fluorescence among top 3 candidates across model architectures. The graph shows the quantile of the top actual fluorescence (y-axis) among the top 3 predicted candidates for various model architectures. Each column on the x-axis represents results obtained by calculating the mean prediction for all models within a given architecture and applying the bootstrap method, where 6,685 genes were sampled with replacement (matching the original dataset size). The results demonstrate that model architectures incorporating K-Nearest Neighbors (KNN) exhibit a better capacity for catching highly performing EG.

[0041] Fig. 10 includes a heatmap showing quantile distribution of actual fluorescence among top 3 candidates across model architectures. This heatmap displays the binned distribution of the quantile of top performing candidate among top 3 predicted. On one hand, ensembles involving KNN are more likely to find genes in top percentile. On the other hand, ensembles involvingXGBoost are more robust, not getting quantiles under 0.9. The best results, both in likelihood to get top percentile and in robustness to failure, is an XGN and KNN ensemble.

[0042] Fig. 11 includes a heat map showing quantile distribution of actual fluorescence of top candidate across model architectures. This heatmap displays the binned distribution of the quantile of predicted best candidate. The same pattern can be seen, where KNN ensembles are more likely to find top-performers, and XGBoost ensembles are robust to large errors. Thus, the XGBoost KNN ensemble displays the best performance, taking the best of both worlds.

[0043] Fig. 12 includes graphs showing ROC-AUC performance for top5% gene classification across model architectures. The graph depicts the ROC-AUC (y-axis) for the classification task: "Is this gene in the top 5%?". Each column on the x-axis represents results obtained by calculating the mean prediction for all models within a given architecture and applying the bootstrap method, where 6,685 genes were sampled with replacement (matching the original dataset size). Among model architectures incorporating KNearest Neighbors (KNN), the combined KNN+XGB architecture demonstrates the most robust performance.

[0044] Fig. 13 includes a heatmap showing feature selection results from greedy forward selection acros20 KNN models. This graph presents the results of greedy forward selection for the top 5 features across 20 KNN models, where features were selected to maximize AUC, a robust performance metric. Each row represents a feature or family of features, and each column sums to 100% of selections for a specific selection rank. Features related to Codon Usage Bias (CUB) are the most frequently selected, with global CUB emerging as the most important, followed by CUB at initiation and downstream regions. Features related to amino acid k-mer composition are consistently selected in subsequent ranks. Other features contribute less significantly to the selection process.

[0045] Fig. 14 includes a graph showing insulin concentration over time for different constructs. This graph illustrates the measured insulin concentration (y-axis) over time (x-axis, in days) for various constructs. (1) The blue plot represents the concentration of the original insulin, as patented by Novo Nordisk. (2) The orange plot shows the concentration of insulin after optimization using the ESO tool, which demonstrates significantly higher expression levels compared to the original insulin. (3) The gray and (4) yellow plots represent insulin constructs generated using the current method with the CAF20 and ARC 15 genes, respectively. These constructs exhibited significantly higher expression levels and greater stability over time compared to the insulin optimized by the ESO tool.DETAILED DESCRIPTIONChimeric DNA

[0046] According to the first aspect, there is provided a chimeric DNA molecule.

[0047] In some embodiments, the chimeric DNA molecule comprises: (a) a first nucleic acid sequence encoding a polypeptide of interest; and (b) a second nucleic acid sequence encoding an endogenous polypeptide of a target cell.

[0048] In some embodiments, the 3’ end of the first nucleic comprises a stop codon. In some embodiments, the first nucleic acid sequence comprises a stop codon in the 3’ end.

[0049] In some embodiments, the stop codon comprises the nucleic acid sequence TAGGCG (also termed herein “LI”). In some embodiments, the stop codon comprises the nucleic acid sequence TGAGCG (also termed herein “L2”). In some embodiments, the stop codon comprises the nucleic acid sequence TGACAA (also termed herein “L3”).

[0050] In some embodiments, the stop codon comprises the nucleic acid sequence GGATAGGCG (also termed herein “LI”). In some embodiments, the stop codon comprises the nucleic acid sequence CTGTGAGCG (also termed herein “L2”). In some embodiments, the stop codon comprises the nucleic acid sequence GGATGACAA (also termed herein “L3”).

[0051] In some embodiments, the first nucleic acid sequence is located upstream to the second nucleic acid sequence in the chimeric DNA molecule.

[0052] In some embodiments, the first and second nucleic acid sequences are operably linked.

[0053] The term “operably linked” is intended to mean that any combination of the first nucleic acid sequence, the second nucleic acid sequence, and the third nucleic acid sequence, is linked to the regulatory element or elements in a manner that allows for expression of the nucleotide sequences (e.g., in an in vitro system or in a host cell). In some embodiments, operably linked refers to any combination of the first nucleic acid sequence, the second nucleic acid sequence, and the third nucleic acid sequence, being transcribed together or sequentially, (e.g., mRNA(s) production). In some embodiments, operably linked refers to any combination of the first nucleic acid sequence, the second nucleic acid sequence, and the third nucleic acid sequence, being transcribed together or sequentially, and translated (e.g., mRNA(s) production, and polypeptide / protein synthesis). In some embodiments, operably linked needs not to comprise protein translation.

[0054] In some embodiments, a chimeric DNA molecule further comprises a third nucleic acid sequence. In some embodiments, the third nucleic acid sequence is located between the firstnucleic acid sequence and the second nucleic acid sequence. In some embodiments, the third nucleic acid sequence is located downstream to or at the 3’ end of the first nucleic acid sequence. In some embodiments, the third nucleic acid sequence is located upstream to or at the 5’ end of the second nucleic acid sequence. In some embodiments, the chimeric DNA molecule is organized from 5’ end to 3’ as follows: the first nucleic acid sequence, the third nucleic acid sequence, and the second nucleic acid sequence. In some embodiments, the chimeric DNA molecule further comprises additional sequence located between any one of the first nucleic acid sequence, the third nucleic acid sequence, and the second nucleic acid sequence. In some embodiments, the first nucleic acid sequence is conjugated or linked to the third nucleic acid sequence. In some embodiments, the third nucleic acid sequence is conjugated or linked to the second nucleic acid sequence. In some embodiments, the first nucleic acid sequence is conjugated or linked to the second nucleic acid sequence.

[0055] In some embodiments, conjugated or linked is via a covalent bond. In some embodiments, a covalent bond is or comprises a phosphodiester bond.

[0056] In some embodiments, chimeric DNA molecule is codon optimized for expression in a target cell. In some embodiments, the chimeric DNA molecule comprises a sequence comprising codons based on the codon preference of a target cell.

[0057] In some embodiments, the chimeric DNA molecule further comprises at least one start codon. In some embodiments, the chimeric DNA molecule comprises a plurality of start codons. In some embodiments, a start codon is located at the 5’ end of the first nucleic acid sequence. In some embodiments, a start codon is located downstream to the stop codon of the first nucleic acid sequence. In some embodiments, a start codon is located at the 5’ end of the second nucleic acid sequence. In some embodiments, a start codon is located at the 5’ end of the third nucleic acid sequence.

[0058] In some embodiments, the at least one start codon is operably linked to any combination of: the first nucleic acid sequence, the second nucleic acid sequence, and the third nucleic acid sequence.

[0059] In some embodiments, an endogenous gene comprises any gene being highly expressed in a target cell, a gene being essential for viability / fitness of a target cell, or both.

[0060] As used herein, the terms "essential gene" or "essential protein" refer to any gene or protein which a cell cannot maintain or properly maintain life in its absence or lack of functionality. In some embodiments, the removal of expression of the essential protein from the target cell induces death of the target cell, replication arrest of the target cell or both.

[0061] In some embodiments, a target cell comprises an intact, native, naive, wildtype, functional, active, or any combination thereof, at least one copy of a gene encoding the endogenous gene in the genome. In some embodiments, a genome of the target cell comprises at least one intact, native, naive, wildtype, functional, active, or any combination thereof, copy of a gene encoding the endogenous gene.

[0062] In some embodiments, the first nucleic acid sequence and: the second nucleic acid sequence, the third nucleic acid sequence, or both, are in-frame or not in-frame.

[0063] As used herein, the term “in-frame” refers to the ‘reading frame’ of a polynucleotide encoding a polypeptide.

[0064] In some embodiments, the chimeric DNA molecule further comprises an expression regulating sequence. In some embodiments, an expression regulating sequence comprises any nucleic acid sequence which is capable of inducing, promoting, enhancing, reducing, inhibiting, and / or terminating, RNA (including mRNA) transcription from the chimeric DNA molecule of the invention.

[0065] In some embodiments, an expression regulating sequence is located or is further located in an expression vector or a plasmid comprising the chimeric DNA molecule of the invention.

[0066] In some embodiments, the expression regulating sequence comprises or is a promoter. In some embodiments, the expression regulating sequence comprises or is a terminator.

[0067] Types and sequences of expression regulating sequences, including, but not limited to promoters, terminators, etc., would be apparent to one of ordinary skill in the art. Such sequences can be easily retrieved from NCBI.

[0068] In some embodiments, a promoter comprises the endogenous or natural promoter controlling the expression of a gene encoding the polypeptide of interest.

[0069] In some embodiments, the promoter is a constitutive promoter or an inductive promoter.Expression vectors and plasmids

[0070] According to another aspect, there is provided an expression vector or a plasmid comprising the chimeric DNA molecule of the invention.

[0071] The terms “DNA molecule”, "polynucleotide", "polynucleotide sequence", "nucleic acid sequence", and "nucleic acid molecule" are used interchangeably herein. These terms encompass nucleotide sequences and the like. A polynucleotide may be a polymer of RNA, DNA, or a hybrid thereof, that is single- or double- stranded, that optionally contains synthetic, non-natural or altered nucleotide bases.

[0072] In some embodiments, increased expression is compared to a DNA molecule comprising a nucleic acid sequence encoding the polypeptide of interest unlinked to the endogenous polypeptide (e.g., not in a chimera encoding the polypeptide of interest and the endogenous polypeptide). In some embodiments, increased is compared to a DNA molecule comprising a nucleic acid sequence encoding the polypeptide of interest alone. In some embodiments, increased is compared to a DNA molecule encoding the polypeptide of interest not as part of a chimeric polypeptide. In some embodiments, increased is compared to a DNA molecule encoding the polypeptide of interest not as part of a chimeric polypeptide of the invention.

[0073] In some embodiments, the DNA molecule further comprises at least one regulatory sequence. In some embodiments, the regulatory sequence is operably linked to the sequence encoding the chimeric polypeptide.

[0074] As used herein, the term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to the regulatory element or elements in a manner that allows for expression of the nucleotide sequence (e.g., in an in-vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0075] In some embodiments, the regulatory sequence is a promoter sequence.

[0076] The term "promoter" as used herein refers to a group of transcriptional control modules that are clustered around the initiation site for an RNA polymerase i.e., RNA polymerase II. Promoters are composed of discrete functional modules, each consisting of approximately 7-20 bp of DNA, and containing one or more recognition sites for transcriptional activator or repressor proteins.

[0077] In some embodiments, the promoter sequence comprises an endogenous promoter sequence of the gene encoding the essential protein. As used herein, the term "endogenous" and "endogenously" refers to that the essential gene is under the regulation of the promoter in a target cell devoid of the chimeric polypeptide of the invention and / or a polynucleotide encoding the chimeric polypeptide. In some embodiments, the endogenous promoter and the essential gene (e.g., encoding the essential protein) are parts of the same gene and / or are located in the same genomic DNA region.

[0078] In some embodiments, the promoter sequence is a heterologous promoter sequence. As used herein, the terms "heterologous" or "heterologous expression" refers to that polypeptide of interest originates from a different cell type or a different species from the target cell (e.g., configured to expression).

[0079] In some embodiments, a chimeric protein being the product of expression of the chimeric DNA molecule of the invention is encoded by a single reading frame. In some embodiments, a single promoter regulates expression of the single reading frame.

[0080] In some embodiments, there is provided an expression vector comprising the herein disclosed chimeric DNA molecule. In some embodiments, the chimeric DNA molecule is an or ligated in an expression vector. In some embodiments, the expression vector is configured to express in a target cell. In some embodiments, the target cell comprises a genome comprising at least one active copy of the endogenous gene. In some embodiments, the vector is a DNA vector. In some embodiments, the vector is an RNA vector. In some embodiments, the vector further comprises any elements required for expression of the chimeric DNA molecule in a target cell.

[0081] In some embodiments, the expression vector is a prokaryotic expression vector. In some embodiments, the prokaryotic expression vector comprises any sequences necessary for expression of the protein encoded by the nucleic acid molecule of the invention in a prokaryotic cell. In some embodiments, the expression vector is a eukaryotic expression vector.

[0082] In some embodiments, the expression vector is a mammalian expression vector. Mammalian expression vectors include, but are not limited to, pcDNA3, pcDNA3.1 (±), pGL3, pZeoSV2(±), pSecTag2, pDisplay, pEF / myc / cyto, pCMV / myc / cyto, pCR3.1, pSinRep5, DH26S, DHBB, pNMTl, pNMT41, pNMT81, which are available from Invitrogen, pCI which is available from Promega, pMbac, pPbac, pBK-RSV and pBK-CMV which are available from Strategene, pTRES which is available from Clontech, and their derivatives.

[0083] In some embodiments, the expression vector contains regulatory elements from eukaryotic viruses such as retroviruses are used by the present invention. SV40 vectors include pSVT7 and pMT2. In some embodiments, vectors derived from bovine papilloma virus include pBV-lMTHA, and vectors derived from Epstein Bar virus include pHEBO, and p2O5. Other exemplary vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo-5, baculovirus pDSVE, and any other vector allowing expression of proteins under the direction of the SV-40 early promoter, SV-40 later promoter, metallo thionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown effective for expression in eukaryotic cells.

[0084] In some embodiments, a recombinant viral vector, which offers advantages such as lateral infection and targeting specificity, is used for in vivo expression. In one embodiment, lateral infection is inherent in the life cycle of, for example, retrovirus and is the process by which a single infected cell produces many progeny virions that bud off and infect neighboring cells. Inone embodiment, the result is that a large area becomes rapidly infected, most of which was not initially infected by the original viral particles. In one embodiment, viral vectors are produced that are unable to spread laterally. In one embodiment, this characteristic can be useful if the desired purpose is to introduce a specified gene into only a localized number of targeted cells.

[0085] Various methods can be used to introduce the expression vector of the present invention into cells. Such methods are generally described in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Springs Harbor Laboratory, New York (1989, 1992), in Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), Chang et al., Somatic Gene Therapy, CRC Press, Ann Arbor, Mich. (1995), Vega et al., Gene Targeting, CRC Press, Ann Arbor Mich. (1995), Vectors: A Survey of Molecular Cloning Vectors and Their Uses, Butterworths, Boston Mass. (1988) and Gilboa et at. [Biotechniques 4 (6): 504-512, 1986] and include, for example, stable or transient transfection, lipofection, electroporation and infection with recombinant viral vectors. In addition, see U.S. Pat. Nos. 5,464,764 and 5,487,992 for positive-negative selection methods.

[0086] General methods in molecular and cellular biochemistry, such as methods useful for carrying out DNA and protein recombination, as well as other techniques described herein, can be found in such standard textbooks as Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., HaRBor Laboratory Press 2001); Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999); Protein Methods (Bollag et al., John Wiley & Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999); Viral Vectors (Kaplift & Loewy eds., Academic Press 1995); Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998).

[0087] In one embodiment, the expression vector is a plant expression vector. In one embodiment, the expression of a polypeptide coding sequence is driven by a number of promoters. In some embodiments, viral promoters such as the 35S RNA and 19S RNA promoters of CaMV [Brisson et al., Nature 310:511-514 (1984)], or the coat protein promoter to TMV [Takamatsu et al., EMBO J. 6:307-311 (1987)] are used. In another embodiment, plant promoters are used such as, for example, the small subunit of RUBISCO [Coruzzi et al., EMBO J. 3:1671- 1680 (1984); and Brogli et al., Science 224:838-843 (1984)] or heat shock promoters, e.g., soybean hspl7.5-E or hspl7.3-B [Gurley et al., Mol. Cell. Biol. 6:559-565 (1986)]. In one embodiment, constructs are introduced into plant cells using Ti plasmid, Ri plasmid, plant viral vectors, direct DNA transformation, microinjection, electroporation and other techniques well known to the skilled artisan. See, for example, Weissbach & Weissbach [Methods for PlantMolecular Biology, Academic Press, NY, Section VIII, pp 421-463 (1988)]. Other expression systems such as insects and mammalian host cell systems, which are well known in the art, can also be used by the present invention.

[0088] It will be appreciated that other than containing the necessary elements for the transcription and translation of the inserted coding sequence (encoding the chimeric polypeptide), the expression construct of the present invention can also include sequences engineered to optimize stability, production, purification, yield or activity of the expressed polypeptide.

[0089] The term "expression" as used herein refers to the biosynthesis of a gene product, including the transcription and / or translation of the gene product. Thus, expression of a nucleic acid molecule may refer to transcription of the nucleic acid fragment (e.g., transcription resulting in mRNA or other functional RNA) and / or translation of RNA into a precursor or mature protein (polypeptide).

[0090] Expressing of a gene within a cell is well known to one skilled in the art. It can be carried out by, among many methods, transfection, transformation, viral infection, or direct alteration of the cell’s genome. In some embodiments, the gene is in an expression vector such as plasmid or viral vector.

[0091] Recombinant expression vectors generally contains at least an origin of replication for propagation in a cell and optionally additional elements, such as a heterologous polynucleotide sequence, expression control element (e.g., a promoter, enhancer), selectable marker (e.g., antibiotic resistance), poly-Adenine sequence that allows for expression of the nucleotide sequence (e.g. in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0092] As used herein the term "in vitro" refers to any process that occurs outside a living organism. As used herein the term "in-vivo" refers to any process that occurs inside a living organism. In one embodiment, “in-vivo” as used herein is a cell within an intact tissue or an intact organ.

[0093] In some embodiments, the chimeric DNA molecule comprises a nucleic acid(s) which is modified so as to improve transcription efficacy, transcript stability, transcript secretion, translation efficacy, translated protein secretion, or any combination thereof. In some embodiments, modification so as to provide improved transcription efficacy, transcript stability, transcript secretion, translation efficacy, translated protein secretion, or any combination thereof, comprises codon optimization.

[0094] The term “codon optimization” refers to a process directed to improving heterologous gene expression and increase the translational efficiency of a recombinant gene of interest by modifying the nucleic acid sequence of the recombinant gene of interest so as to accommodate codon bias according to the host cell or organism.

[0095] In some embodiments, codon optimization does not alter the amino acid sequence of the polypeptide of interest. In some embodiments the chimeric DNA molecule is modified by means of mutagenesis. In some embodiments, the chimeric DNA molecule comprises at least one mutation compared to the wildtype sequence from which it is derived. In some embodiments, the mutation is a silent mutation. In some embodiments, the mutation is a missense mutation. In some embodiments, the mutation is not a nonsense mutation. In some embodiments, the mutation is any mutation which improves polypeptide production rates, yields, stability, secretion, or any combination thereof, wherein the produced polypeptide is a functional polypeptide of interest of the invention. In some embodiments, the functional polypeptide of interest is in a chimeric polypeptide or a fusion protein with a polypeptide encoded by the endogenous gene.Cells and Extracts

[0096] According to another aspect, there is provided a cell comprising: (a) the chimeric DNA molecule of the invention; (b) the expression vector or plasmid of the invention; or both (a) and (b).

[0097] In some embodiments, the cell is a prokaryote cell or a eukaryote cell.

[0098] In some embodiments, the cell is a recombinant cell, a transgenic cell, a transfected cell, or a transduced cell.

[0099] In some embodiments, the cell is a cell is characterized by increased or over expression of the polypeptide of interest compared to a control cell. In some embodiments, a control cell is devoid of the chimeric DNA molecule of the invention.

[0100] In some embodiments, the cell is a characterized by expression level of the endogenous polypeptide being essentially similar or increased compared to a control cell being devoid of the chimeric DNA molecule of the invention; the expression vector or plasmid of the invention; or both.

[0101] In some embodiments, “essentially similar” comprises being at least 80%, 85%, 90%, 95%, or 99% identical to a control cell being devoid of the chimeric DNA molecule of the invention; the expression vector or plasmid of the invention; or both, or any value and range therebetween. Each possibility represents a separate embodiment of the invention. In some embodiments, “essentially similar” comprises being 80-100%, 85-100%, 90-100%, 95-100%, or99-100% identical to a control cell being devoid of the chimeric DNA molecule of the invention; the expression vector or plasmid of the invention; or both. Each possibility represents a separate embodiment of the invention.

[0102] According to another aspect, there is provided an extract, lysate, homogenate, or any fraction thereof, of the cell of the invention.

[0103] In some embodiments, the extract, lysate, homogenate, or any fraction thereof comprises: the polypeptide of interest, the endogenous polypeptide, an mRNA transcribed from the first nucleic acid sequence, an mRNA transcribed from the second nucleic acid sequence, or any combination thereof.

[0104] In some embodiments, the extract, lysate, homogenate, or any fraction thereof consists essentially of the polypeptide of interest.

[0105] In some embodiments, a polypeptide fraction, portion, or content of the extract, lysate, homogenate, or any fraction thereof, consists essentially of the polypeptide of the invention.

[0106] As used herein, the term “consists essentially of’ denotes that the polypeptide of interest encoded by the chimeric DNA molecule of the invention constitutes the vast majority of polypeptides portion, content, or fraction of the extract, lysate, homogenate, or composition, as disclosed herein.

[0107] In some embodiments, consists essentially of means that: the polypeptide of interest constitute at least 95%, at least 98%, at least 99%, or at least 99.9% by weight, of the polypeptides of the extract, lysate, homogenate, composition, as disclosed herein, or any value and range therebetween. Each possibility represents a separate embodiment of the invention.

[0108] Methods and means for obtaining extract, lysate, homogenate, including any fraction thereof, e.g., polypeptide fraction, portion, or content, are common and would be apparent to one of ordinary skill in the art.Compositions

[0109] According to another aspect, there is provided a composition comprising: (a) the chimeric DNA molecule of the invention; (b) the expression vector or plasmid of the invention; (c) the cell of the invention; (d) the extract, lysate, homogenate, or any fraction thereof of the invention; or (e) any combination of (a) to (d), and an acceptable carrier.

[0110] In some embodiments, the composition consists essentially of the polypeptide of interest and the acceptable carrier.[01 1 1] In some embodiments, the carrier is a pharmaceutically acceptable carrier. In some embodiments, the polypeptide of interest comprises a polypeptide of therapeutic activity, e.g., beneficiary for health purposes of a subject administered therewith.

[0112] As used herein, the term “carrier,” “excipient,” or “adjuvant” refers to any component of a pharmaceutical composition that is not the active agent. As used herein, the term “pharmaceutically acceptable carrier” refers to non-toxic, inert solid, semi-solid liquid filler, diluent, encapsulating material, formulation auxiliary of any type, or simply a sterile aqueous medium, such as saline. Some examples of the materials that can serve as pharmaceutically acceptable carriers are sugars, such as lactose, glucose and sucrose, starches such as com starch and potato starch, cellulose and its derivatives such as sodium carboxymethyl cellulose, ethyl cellulose and cellulose acetate; powdered tragacanth; malt, gelatin, talc; excipients such as cocoa butter and suppository waxes; oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, com oil and soybean oil; glycols, such as propylene glycol, polyols such as glycerin, sorbitol, mannitol and polyethylene glycol; esters such as ethyl oleate and ethyl laurate, agar; buffering agents such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen- free water; isotonic saline, Ringer's solution; ethyl alcohol and phosphate buffer solutions, as well as other non-toxic compatible substances used in pharmaceutical formulations. Some nonlimiting examples of substances which can serve as a carrier herein include sugar, starch, cellulose and its derivatives, powered tragacanth, malt, gelatin, talc, stearic acid, magnesium stearate, calcium sulfate, vegetable oils, polyols, alginic acid, pyrogen-free water, isotonic saline, phosphate buffer solutions, cocoa butter (suppository base), emulsifier as well as other non-toxic pharmaceutically compatible substances used in other pharmaceutical formulations. Wetting agents and lubricants such as sodium lauryl sulfate, as well as coloring agents, flavoring agents, excipients, stabilizers, antioxidants, and preservatives may also be present. Any non-toxic, inert, and effective carrier may be used to formulate the compositions contemplated herein. Suitable pharmaceutically acceptable carriers, excipients, and diluents in this regard are well known to those of skill in the art, such as those described in The Merck Index, Thirteenth Edition, Budavari et al., Eds., Merck & Co., Inc., Rahway, N.J. (2001); the CTFA (Cosmetic, Toiletry, and Fragrance Association) International Cosmetic Ingredient Dictionary and Handbook, Tenth Edition (2004); and the “Inactive Ingredient Guide,” U.S. Food and Drug Administration (FDA) Centre for Drug Evaluation and Research (CDER) Office of Management, the contents of all of which are hereby incorporated by reference in their entirety. Examples of pharmaceutically acceptable excipients, carriers and diluents useful in the present compositions include distilled water, physiological saline, Ringer's solution, dextrose solution, Hank's solution, and DMSO.These additional inactive components, as well as effective formulations and administration procedures, are well known in the art and are described in standard textbooks, such as Goodman and Gillman’s: The Pharmacological Bases of Therapeutics, 8th Ed., Gilman et al. Eds. Pergamon Press (1990); Remington’s Pharmaceutical Sciences, 18th Ed., Mack Publishing Co., Easton, Pa. (1990); and Remington: The Science and Practice of Pharmacy, 21st Ed., Lippincott Williams & Wilkins, Philadelphia, Pa., (2005), each of which is incorporated by reference herein in its entirety. The presently described composition may also be contained in artificially created structures such as liposomes, ISCOMS, slow -releasing particles, and other vehicles which increase the half-life of the peptides or polypeptides in serum. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers and the like. Liposomes for use with the presently described peptides are formed from standard vesicle-forming lipids which generally include neutral and negatively charged phospholipids and a sterol, such as cholesterol. The selection of lipids is generally determined by considerations such as liposome size and stability in the blood. A variety of methods are available for preparing liposomes as reviewed, for example, by Coligan, J. E. et al, Current Protocols in Protein Science, 1999, John Wiley & Sons, Inc., New York, and see also U.S. Pat. Nos. 4,235,871, 4,501,728, 4,837,028, and 5,019,369.

[0113] The carrier may comprise, in total, from about 0.1% to about 99.99999% by weight of the pharmaceutical compositions presented herein.Methods of use

[0114] According to another aspect, there is provided a method for increasing expression of a polypeptide of interest in a target cell.

[0115] According to another aspect, there is provided a method for increasing genetic stability of a cell producing a polypeptide of interest.

[0116] In some embodiments, genetic stability is or comprises evolutionary stability.

[0117] The terms “evolutionary stability” or “genetic stability” are interchangeable, and used herein to refer to genetic stability of a cell or organism over time, particularly in the context of synthetic biology. It involves maintaining the integrity and functionality of a genetically engineered construct, such as a chimeric DNA molecule, across multiple generations or extended periods of culture. This stability is crucial for ensuring consistent expression of the desired polypeptide of interest and minimizing the accumulation of deleterious mutations that could compromise the performance or viability of the engineered cells.

[0118] In some embodiments, increasing genetic stability, evolutionary stability, or both, or a polypeptide of interest and / or of a cell comprising a nucleic acid sequence encoding thereof, comprises reducing the abundance, rate, or both, of accumulated mutation. In some embodiments, the mutation compromises and / or reduces performance, viability, fitness, or any combination thereof, of a cell of the invention. In some embodiments, the cell is an engineered cell. In some embodiments, the cell is a genetically engineered cell. In some embodiments, the mutation is a missense mutation. In some embodiments, the mutation is a nonsense mutation. In some embodiments, the mutation is a deleterious mutation.

[0119] In some embodiments, the method comprises culturing a cell of the invention such that the polypeptide of interest encoded by the first nucleic acid sequence is expressed.

[0120] In some embodiments, increased is compared to a control cell being devoid of the chimeric DNA molecule. In some embodiments, a control cell is not the cell of the invention.

[0121] In some embodiments, increasing comprises at least 5%, 10%, 25%, 50%, 100%, 250%, 500%, 750%, or 1,000% increase, or any value and range therebetween. Each possibility represents a separate embodiment of the invention. In some embodiments, increasing comprises at least 5-100%, 10-200%, 25-300%, 50-150%, 100-400%, 250-550%, 500%, 1-750%, or 100- 1,000% increase. Each possibility represents a separate embodiment of the invention.

[0122] In some embodiments, culturing is for a period of 25 to 75 days, 15 to 65 days, 20 to 70 days, 30 to 60 days, 35 to 70 days, 40 to 60 days, or 45 to 65 days. In some embodiments, culturing is for a period of at least 25 days, at least 30 days, at least 35 days, at least 40 days, at least 45 days, at least 50 days, at least 55 days, at least 60 days, at least 65 days, at least 75 days, or any value and range therebetween. Each possibility represents a separate embodiment of the invention.

[0123] In some embodiments, increasing expression comprises increasing: abundance and / or secretion of mRNA transcribed from the first nucleic acid sequence, abundance and / or secretion of the polypeptide of interest, and both, in or from the cell.

[0124] In some embodiments, increased abundance of mRNA comprises increased: mRNA transcription, mRNA stability, or both.

[0125] In some embodiments, an mRNA transcribed from the second nucleic acid sequence, the third nucleic acid sequence, or both, is not translated into polypeptide(s).

[0126] In some embodiments, an mRNA transcribed from the first nucleic acid sequence and the second nucleic acid sequence is translated, thereby providing a chimeric polypeptide comprising the polypeptide of interest and a polypeptide product of the endogenous gene.

[0127] In some embodiments, an mRNA transcribed from the first nucleic acid sequence and the third nucleic acid sequence is translated, thereby providing a chimeric polypeptide comprising the polypeptide of interest and a linker polypeptide.

[0128] In some embodiments, an mRNA transcribed from the first nucleic acid sequence, the third nucleic acid sequence, and the second nucleic acid sequence is translated, thereby providing a chimeric polypeptide comprising the polypeptide of interest, a linker polypeptide, and a polypeptide product of the endogenous gene.

[0129] The chimeric polypeptides can be further processed in order to separate the polypeptide of interest from the linker polypeptide, the polypeptide product of the endogenous gene, or both.

[0130] Methods and means for separating different polypeptides of a chimeric polypeptide are common and would be apparent to one of skill in the art. Non-limiting examples for such methods, include, but are not limited to proteolytic cleavage, affinity chromatography, or a combination thereof.

[0131] In some embodiments, the linker provides, increases, enables, enhances, any equivalent thereof, or any combination thereof, optimal folding of the polypeptide of interest and the polypeptide product of the endogenous gene of the chimeric polypeptide.

[0132] In some embodiments, the linker as disclosed herein reduces folding disturbance, increases spatial separation, restores folding, or any combination thereof, of the polypeptide of interest and the polypeptide product of the endogenous gene of the chimeric polypeptide.

[0133] In some embodiments, the linker is chosen based on its suitability to increase folding accuracy, proficiency, efficiency, thermodynamic stability of both the polypeptide of interest and the polypeptide product of the endogenous gene of the chimeric polypeptide.

[0134] As used herein, the term “linker” refers to a molecule or macromolecule serving to connect the different moieties of the chimeric polypeptide of the invention, e.g., the polypeptide of interest and the polypeptide product of the endogenous gene. In one embodiment, the linker may also facilitate other functions, including, but not limited to, preserving biological activity, maintaining sub-units and / or domains interactions, and others. In some embodiments, the linker is a flexible linker or a rigid linker.

[0135] In some embodiments, the amino acid sequence of the linker and / or the amino acid sequence of any one of the polypeptide of interest, the polypeptide product of the endogenous gene, and the chimeric polypeptide of the invention, are co-modified so to improve: protein folding, expression, function, or any combination thereof. In some embodiments, amino acid sequence co-modification comprises preventing the common or native folding of the polypeptide of interest, the polypeptide product of the endogenous gene, or both. In some embodiments, any amino acid sequence modification is applicable as long as the provided polypeptide of interest is functional. In some embodiments, the linker is a flexible linker. In some embodiments, the linker is a rigid linker.

[0136] In some embodiments, the linker comprises at least 2, 5, 7, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50 amino acids, or any value and range therebetween. Each possibility represents a separate embodiment of the invention. According to some embodiments, the linker comprises 2-20, 5-50, 1-15, 3-28, 4-44, or 2-50 amino acids. Each possibility represents a separate embodiment of the invention. In some embodiments, the linker and / or its length and / or amino acid content is configured to allow separate folding of the polypeptide of interests and the polypeptide product of the endogenous gene.

[0137] As used herein, the terms “peptide”, "polypeptide" and "protein" are used interchangeably to refer to a polymer of amino acid residues. In another embodiment, the terms "peptide", "polypeptide" and "protein" as used herein encompass native peptides, peptidomimetics (typically including non-peptide bonds or other synthetic modifications) and the peptide analogues peptoids and semipeptoids or any combination thereof. In another embodiment, the peptides polypeptides and proteins described have modifications rendering them more stable while in the body or more capable of penetrating into cells. In one embodiment, the terms “peptide”, "polypeptide" and "protein" apply to naturally occurring amino acid polymers. In another embodiment, the terms “peptide”, "polypeptide" and "protein" apply to amino acid polymers in which one or more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid.

[0138] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes oneor both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0139] As used herein, the term "about" when combined with a value refers to plus and minus 10% of the reference value. For example, a length of about 1,000 nanometers (nm) refers to a length of 1,000 nm ± 100 nm.

[0140] It is noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polynucleotide" includes a plurality of such polynucleotides and reference to "the polypeptide" includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only" and the like in connection with the recitation of claim elements or use of a "negative" limitation.

[0141] In those instances where a convention analogous to "at least one of A, B, and C, etc." is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."

[0142] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub- combination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.

[0143] Additional objects, advantages, and novel features of the present invention will become apparent to one ordinarily skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below finds experimental support in the following examples.

[0144] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.EXAMPLES

[0145] Generally, the nomenclature used herein, and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are thoroughly explained in the literature. See, for example, "Molecular Cloning: A laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, R. M., ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York (1998); methodologies as set forth in U.S. Pat. Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I-III Cellis, J. E., ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley-Liss, N. Y. (1994), Third Edition; "Current Protocols in Immunology" Volumes I-III Coligan J. E., ed. (1994); Stites et al. (eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual" CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this document.Materials and MethodsOverview of Gene Fusion Design

[0146] The inventors developed a multi-step workflow to design and optimize GOI-EG fusion genes for enhanced evolutionary stability and expression. This process included the selection of EGs using an ML model, the identification of optimal linkers to minimize misfolding, sequence optimization of the synthetic gene to maximize expression and stability, and addition of a low- readthrough leaky stop codon.Machine Learning Model for EG Selection

[0147] The ML model was trained on fluorescence datasets from yeast libraries. The dataset included GOLEG fusion genes for 5,185 EG (-78% of all EG in yeast), fused with either GFP or mCherry. Fluorescence was measured across multiple time points.Features and Model Architecture

[0018] The model utilized a diverse set of bioinformatic features derived from the sequences of each GOI and EG. Key feature families included:

[0149] Codon Usage Bias:• tRNA Adaptation Index (tAI): Calculated for the full sequence, first 17 amino acids, and sliding windows across the GOI and EG.• Relative Codon Adaptation (RCA): Evaluated for the GOI, EG, and sliding windows.• Codon Adaptation Index (CAI): Assessed for the full sequence and sliding windows.• Effective Number of Codons (ENC): Measures diversity in codon usage.

[0150] Sequence Composition:• GC Content: Calculated globally and in sliding windows for both GOI and EG.• k-mer Frequencies: Counts of nucleotide or amino acid substrings of lengths 3-5.• Amino Acid Frequency Correlation: Spearman correlation between GOI and EG amino acid compositions.

[0151] Thermodynamic Properties:• Local Folding Energy: Calculated using ViennaRNA in sliding windows (50 bp).

[0152] Chemical Properties:• Molecular Weight and Hydrophobicity: Computed for both GOI and EG sequences.• Isoelectric Point (pl): Assessed at physiological pH.

[0153] Translation Context:1. Start Codon Context: Features describing the nucleotide flanking regions of start codons.2. Shifted Open Reading Frame (sORF) Length: Alternative reading frame lengths.

[0154] ChimeraARS Score:• The ChimeraARS score quantifies sequence similarity to a reference set of genes with high codon usage bias. This metric enhances predictions of sequence stability beyond standard codon usage features.

[0155] The ML pipeline employed an ensemble model combining k-nearest neighbors (KNN) and XGBoost architectures (Fig. 5). The KNN model emphasized features of the EG andprovided high median performance - but was prone to failures. The XGB model emphasized features of the GOI, EG, and interaction features, and provided much higher robustness to failure. Hyperparameter tuning and cross-validation were conducted using Optuna to ensure robust performance.Feature Importance in Machine Learning Predictions

[0156] Feature importance was analyzed using SHAP (SHapley Additive exPlanations) for XGBoost and forward feature selection for KNN. SHAP provided a global ranking of feature contributions, while forward selection iteratively identified the most impactful features for KNN predictions. The XGBoost results, which were more robust, provided the primary conclusions, while KNN analysis offered additional insights, were consistent or complementary.XGBoost Feature Importance

[0157] The SHAP analysis (Fig. 5) revealed tRNA Adaptation Index (tAI) as the dominant feature, significantly outpacing all others. Variants with high tAI scores consistently showed enhanced stability and expression, underscoring its central role in optimizing translation.

[0158] Following tAI, the next most important features were GC content and Relative Codon Adaptation (RCA), both reflecting translational optimization and sequence stability. Additional influential features included:• Amino Acid Frequency Correlation between the GOI and EG, suggesting reduced metabolic burden for similar protein compositions.• Focal Folding Energy downstream in the GOI and near the initiation site, emphasizing the importance of avoiding unstable mRNA secondary structures.• Shifted ORF Length , Ribosomal Coverage Bias Score (RCBS), and tAI in the First 17 Amino Acids, which further highlighted the significance of efficient translation, and successful initiation.KNN Feature Importance

[0159] The KNN forward selection supported the importance of codon usage-related features, particularly in the EG. Key features included RCA, tAI, and CAI metrics within the EG and in specific regions, such as the first 17 amino acids. These findings complemented the more GOL focused insights from XGBoost by emphasizing the role of the EG in overall gene fusion stability.Linkers for Protein Fusion

[0160] Protein misfolding remains a challenge in gene fusion strategies. While tools such as AlphaFold offer accurate prediction of protein structure, they are too computationally intensive to enable a reasonable comparison between many linkers. Rather, the inventors used tools suchas IUPred2A, a biophysical model, and MoreRONN, a machine learning model, to predict protein disorder profiles and assess the impact of linkers. As protein disorder profiles have a profound effect of protein folding, it is taken as a proxy variable - linkers with smaller effect on the disorder profile are less likely to influence the protein folding, and thus less likely to cause misfolding. Linkers were selected to minimize disruptions to the disorder profile upon fusion, likely preserving native folding patterns of both the GOI and the EG. This minimization was applied by calculating the Euclidean distance between the disorder profiles before and after fusion, for the 1,280 linkers in a linker database. The linker inducing the smallest distance was selected.

[0161] Although experimental validation of the linker selection step remains pending, these predictions provide an essential foundation for mitigating misfolding risks in future applications.Scoring and Selection

[0162] Linkers were scored based on the L2 distance between disorder profiles of the unfused and fused states. The optimal linker was selected to minimize disruptions to native folding patterns, preserving the functionality of both the GOI and EG.Optimizing DNA Sequences for Stability and Expression

[0163] Sequence optimization was performed to enhance the stability and expression of the GOL EG fusion genes using the Evolutionary Stability Optimizer (ESO) pipeline. ESO integrates DNAChisel, a tool that optimizes genetic sequences while maintaining biological constraints. ESO has been previously validated for its ability to improve the evolutionary stability of synthetic genes. The ESO pipeline focuses on optimizing codon usage, minimizing mRNA instability, and ensuring the long-term expression of engineered genes by accounting for evolutionary pressures.Optimization Objectives1. mRNA Folding Optimization: It has been demonstrated that minimizing the secondary structure formation in mRNA near the translation initiation site, maximizes the expression. The local folding energy of the first 15 codons in the mRNA sequence was calculated using ViennaRNA and codons were selected to optimize weak folding energy.2. Codon Usage Optimization: For the rest of the sequence, codon usage was optimized to match the tRNA abundances of S. cerevisiae. This was achieved by calculating and optimizing the tRNA Adaptation Index (tAI) for the GOI, linker, and (depending on the application and relevancy) the EG, ensuring that codon usage patterns align with the host’s translational machinery.3. Stability Enhancements: Sequences were optimized to avoid hypermutable regions.

[0164] These optimizations ensure synthetic genes are tailored for host-specific transcription and translation efficiencies, enhancing long-term performance in diverse applications.Experimental ValidationFusion Stability in S. cerevisiae

[0165] The inventors validated the gene fusions by testing their ability to stabilize GOI expression in S. cerevisiae. Taking the Schuldiner lab library (N' SWAp Tag (SWAT)-GFP, derived from S. cerevisiae BY4741 background strain), the inventors selected variants from it such that the EG are varied and representative. This was performed by performing k-means for k = 2 ... 130, 5 iterations each, and selecting the 10 genes most consistently in different clusters, while closest to cluster center. As a control, they were compared with unfused GFP, integrated into the genome by replacing Canl . The canl gene encodes an arginine permease, which is not essential for yeast survival under laboratory conditions where arginine is supplemented in the medium - this is common practice in the field.

[0166] Their yield was measured over 15 days of in-lab evolution, with 3 repeats each. The evolution experiment took place in 96-Deepwell Plates (Starlab group Cat.No S 1896-2110). Fluorescence levels were proxy for protein expression.

[0167] The inventors first normalized this value by the OD, to control for yeast density. Furthermore, we performed this analysis for the wild type (WT) of yeast, to find the inherent fluorescence of yeast. Following, we detracted this value from other measurements, to isolate the fluorescence due to GFP.Proinsulin Production in Yeast

[0168] To test a biotechnologically relevant application, the inventors applied the workflow to human proinsulin. The proinsulin sequence is of Novo Nordisk (patent WO2014195452A1). EGs were selected using the following criteria:1. EG was predicted to be in top 5% for XGBoost architecture.2. EG was predicted to be in top 10% for KNN architecture, due to less robustness.3. EG length of less than 500 nt, to ensure reasonable synthesis of the combined fusion protein.4. EG has no internal repeats of length > 10, to ensure greater stability.5. EG is not ribosomal, due to many interactions of ribosomal proteins, with many points of failure.6. EG does not have induced function - they are more sensitive to changes in expression.7. The inventors chose genes from multiple cellular processes to decrease redundancy.

[0169] Based on these criteria, Caf20 and Arc 15 were selected.

[0170] Four (4) variants were compared:1. The original proinsulin sequence, expressed on a high copy number plasmid (Prs426)2. A similar variant, where the proinsulin sequence was optimized by the ESO.3. A fusion gene composed of optimized proinsulin and Caf20 expressed on a high copy number plasmid in a strain that Caf20 was deleted from the genome.4. A similar variant for Arc 15.

[0171] Each variant included an alpha leader.

[0172] The sequences of all the construct (including - promotor (upper case letters only), proinsulin including alpha leader sequence (lower case letters only), linker (bolded upper case letters) endogenous gene sequence (if applicable; underlined upper case letters); and a terminator sequence (bolded lower case letters)) are as follows:InsulinATATATGGGGCCGTATACTTACATATAGTAGATGTCAAGCGTAGGCGCTTCCCCTGC CGGCTGTGAGGGCGCCATAACCAAGGTATCTATAGACCGCCAATCAGCAAACTACC TCCGTACATTCATGTTGCACCCACACATTTATACACCCAGACCGCGACAAATTACCC ATAAGGTTGTTTGTGACGGCGTCGTACAAGAGAACGTGGGAACTTTTTAGGCTCACC AAAAAAGAAAGAAAAAATACGAGTTGCTGACAGAAGCCTCAAGAAAAGAAAAATT CTTCTTCGACTATGCTGGAGGCAGAGATGATCGAGCCGGTAGTTAACTATATATAGC TAAATTGGTTCCATCACCTTCTTTTCTGGTGTCGCTCCTTCTAGTGCTATTTCTGGCTT TTCCTATTCTTTTCTTTCCATTTTTCTTTCTCTCTTTCTAATATATAAATTCTCTTGCAT TTTCTATTTTTCTCTCTATCTATTCTACTTGTTTATTCCCTTCAAGGTTTTTTTTTAAGG AGTACTTGTTTTTAGAATATACGGTCAACGAACTATAATTAACTAAACACTAGTACC atgaaattgaaaactgttagatctgctgttttgtcttctttgtttgcttctcaagttttgggtcaaccaattgatgatactgaatctcaaactacttctgt taatttgatggctgatgatactgaatctgcttttgctactcaaactaattctggtggtttggatgttgttggtttgatttctatggctgaagaaggtga accaaaaaaaagatttgttaatcaacatttgtgtggttctcatttggttgaagctttgtatttggtttgtggtgaaagaggtttcttttacactccaaa ggaatggaagggtatcgttgaacaatgttgtacttctatctgttctttgtaccaattggaaaattattgtaattcatgtaattagttatgtcacgct tacattcacgccctccccccacatccgctctaaccgaaaaggaaggagttagacaacctgaagtctaggtccctatttatttttttata gttatgttagtattaagaacgttatttatatttcaaatttttcttttttttctgtacagacgcgtgtacgcatgtaacattatactgaaaac cttgcttgagaaggttttgggacgctcgaaggctttaatttgc (SEQ ID NO: 1).Insulin: :CAF20ATATATGGGGCCGTATACTTACATATAGTAGATGTCAAGCGTAGGCGCTTCCCCTGC CGGCTGTGAGGGCGCCATAACCAAGGTATCTATAGACCGCCAATCAGCAAACTACC TCCGTACATTCATGTTGCACCCACACATTTATACACCCAGACCGCGACAAATTACCC ATAAGGTTGTTTGTGACGGCGTCGTACAAGAGAACGTGGGAACTTTTTAGGCTCACC AAAAAAGAAAGAAAAAATACGAGTTGCTGACAGAAGCCTCAAGAAAAGAAAAATT CTTCTTCGACTATGCTGGAGGCAGAGATGATCGAGCCGGTAGTTAACTATATATAGC TAAATTGGTTCCATCACCTTCTTTTCTGGTGTCGCTCCTTCTAGTGCTATTTCTGGCTTTTCCTATTCTTTTCTTTCCATTTTTCTTTCTCTCTTTCTAATATATAAATTCTCTTGCATTTTCTATTTTTCTCTCTATCTATTCTACTTGTTTATTCCCTTCAAGGTTTTTTTTTAAGGAGTACTTGTTTTTAGAATATACGGTCAACGAACTATAATTAACTAAACACTTAAAAAATGAAATTAAAAACTGTTCGATCTGCCGTCCTTTCTTCTTTGTTTGCTTCTCAAGTTTTGGGTcaaccaattgatgatactgaatctcaaactacttctgttaatttgatggctgatgatactgaatctgcttttgctactcaaactaattctgg tggtttggatgttgttggtttgatttctatggctgaagaaggtgaaccaaaaaaaagatttgttaatcaacatttgtgtggttctcatttggttgaag ctttgtatttggtttgtggtgaaagaggtttcttttacactccaaaggaatggaagggtatcgttgaacaatgttgtacttctatctgttctttgtacc aattggaaaattattgtaatATATTGACACACGACTCATCTATCAGATACTTACAAGAAATAT ATAATAGTAACAACCAAAAAATTGTAAATCTAAAGGAAAAGGTCGCTCAACTTGAGGCACAGTGCCAAGAACCATGTAAAGATACTGTTCAGATTCATGATATTACCGGTATGATCAAGTATACTATCGATGAGCTTTTTCAACTGAAGCCAAGTTTAACTTTGGAAGTTAATTTCGATGCGGTGGAATTTAGAGCCATCATTGAAAAAGTTAAGCAATTGCAACACTTGAAAGAGGAAGAGTTTAACAGTCATCATGTTGGTCATTTCGGTCGTAGAAGATCTTCCCACCATCATGGTAGACCAAAGATTAAGCACAACAAGCCTAAGGTTACAACCGATTCAGATGGTTGGTGCACATTTGAAGCCAAGAAGAAGGGTAGTGGAGAAGATGATGAAGAAGAAACAGAAACCACACCAACTTCTACTGTGCCAGTTGCTACCATTGCCCAAGAAACTTTAAAAGTCAAGCCAAATAACAAAAATATTTCTTCCAACAGACCTGCTGATACCAGAGATATTGTTGCGGACAAGCCAATTCTTGGTTTCAACGCATTTGCTGCTTTGGAAAGTGAAGACGAAGACGACGAAGCATAAtcatgtaattagttatgtcacgcttacattca cgccctccccccacatccgctctaaccgaaaaggaaggagttagacaacctgaagtctaggtccctatttatttttttatagttatgtt agtattaagaacgttatttatatttcaaatttttcttttttttctgtacagacgcgtgtacgcatgtaacattatactgaaaaccttgcttg agaaggttttgggacgctcgaaggctttaatttgc (SEQ ID NO: 2).Insulin: :ARC 15ATATATGGGGCCGTATACTTACATATAGTAGATGTCAAGCGTAGGCGCTTCCCCTGCCGGCTGTGAGGGCGCCATAACCAAGGTATCTATAGACCGCCAATCAGCAAACTACCTCCGTACATTCATGTTGCACCCACACATTTATACACCCAGACCGCGACAAATTACCCATAAGGTTGTTTGTGACGGCGTCGTACAAGAGAACGTGGGAACTTTTTAGGCTCACCAAAAAAGAAAGAAAAAATACGAGTTGCTGACAGAAGCCTCAAGAAAAGAAAAATTCTTCTTCGACTATGCTGGAGGCAGAGATGATCGAGCCGGTAGTTAACTATATATAGCTAAATTGGTTCCATCACCTTCTTTTCTGGTGTCGCTCCTTCTAGTGCTATTTCTGGCTTTTCCTATTCTTTTCTTTCCATTTTTCTTTCTCTCTTTCTAATATATAAATTCTCTTGCATTTTCTATTTTTCTCTCTATCTATTCTACTTGTTTATTCCCTTCAAGGTTTTTTTTTAAGGAGTACTTGTTTTTAGAATATACGGTCAACGAACTATAATTAACTAAACACTTAAAAAATGAAATTAAAAACTGTTCGATCTGCCGTCCTTTCTTCTTTGTTTGCTTCTCAAGTTTTGGGTcaaccaattgatgatactgaatctcaaactacttctgttaatttgatggctgatgatactgaatctgcttt tgctactcaaactaattctggtggtttggatgttgttggtttgatttctatggctgaagaaggtgaaccaaaaaaaagatttgttaatcaacattt gtgtggttctcatttggttgaagctttgtatttggtttgtggtgaaagaggtttcttttacactccaaaggaatggaagggtatcgttgaacaatg ttgtacttctatctgttctttgtaccaattggaaaattattgtaatATATTGACACACGACTCATCTATCAGATA CTTACAAGAAATATATAATAGTAACAACCAAAAAATTGTAAATCTAAAGGAAA AGGTCGCTCAACTTGAGGCACAGTGCCAAGAACCATGTAAAGATACTGTTCAG ATTCATGATATTACCGGTATGGAAGCCGATTGGAGGAGAATTGACATCGATGCAT TTGATCCAGAGAGTGGCAGACTAACCGCTGCCGATCTGGTACCACCATACGAAACT ACTGTCACATTACAAGAATTACAACCTCGAATGAATCAATTGCGCTCGCTTGCCAC AAGTGGTGACTCTTTGGGAGCCGTTCAATTACTCACAACCGATCCTCCATACAGTG CAGATGCTCCAACAAAGGAGCAATATTTTAAGAGCGTCCTTGAAGCATTGACACA AGTCAGGCAAGCCGATATTGGTAATGTAATCAAAAATTTGAGTGATTCTCAGAGGG ACGTGCTGGTAAAGTATCTCTACAAAGGAATGTCCGTACCTCAGGGCCAGAAACA AGGGGGTGTCTTGCTTGCGTGGCTGGAAAGAATTACTCAAGTCAGTGGTGTCACAC CTATCGTTCATTATATATCGGATAGAAGAACTGTATGAtcatgtaattagttatgtcacgcttacatt cacgccctccccccacatccgctctaaccgaaaaggaaggagttagacaacctgaagtctaggtccctatttatttttttatagtta tgttagtattaagaacgttatttatatttcaaatttttcttttttttctgtacagacgcgtgtacgcatgtaacattatactgaaaacctt gcttgagaaggttttgggacgctcgaaggctttaatttgc (SEQ ID NO: 3).

[0173] They were tested for yield over 35 days of in-lab evolution, with 3 repeats each. Proinsulin levels were quantified using ELISA plates according to manufacturer instructions (product number RAB0327) for proinsulin detection and verified by western blot using proinsulin specific antibody (sigma i2018).

[0174] The constructs were analyzed through Nanopore sequencing before and after the evolution experiment, for deeper understanding of the accumulated mutations. Nanopore reads were first processed using chopper, trimming the first lOnt of each read and keeping reads of length above 500. The reads were mapped with minimap2 to a reference sequence, including the yeast genome and the relevant construct. Quantification of optimized variants’ mRNA levels was conducted using Salmon, and variant calling was conducted by DeepVariant.Leaky Stop Codon Rate of Read-through

[0175] Selecting and implementing a leaky stop codon is required for the current design. The inventors require a rate of read-through such that:1. The expression of the fusion protein is high enough to generate viable quantities of the fusion protein, required for the cell’s growth.2. It is the lowest value reasonable, to prove more mutations deleterious.3. High quantities of the GOI alone, as this maximizes the output of the system.

[0176] To enable an informed selection for the stop codons, the inventors generated different stop codons using the design principles outlined in (Mangkalaphiban et al., 2021). In this paper, Riboseq is conducted on all genes in S. cerevisiae. They created a predictive model for rate of read-through and found the most influential feature to be the context around the stop codon - from 3 nts before to 6 after, 12 nts total. Further refining the model, they created a ranking of leakiness based on nt at each position. Using these rankings, the inventors generated sequences predicted to be with high read-through rates.

[0177] The inventors conducted an experiment where a fusion gene on a plasmid was implemented. This fusion gene was composed of BFP connected to the C’ terminus of mCherry, with the different leaky stop codons generated. The fluorescence of BFP is proxy for the expression of both the GOI and the fusion protein, and the fluorescence of mCherry is proxy for the expression of only the fusion protein.

[0178] The inventors selected 3 stop codon designs such that:1. The rate of read-through, as measured in the mCherry relative fluorescence, is in the range of 0.1-0.25.2. The relative fluorescence of BFP, indicative of the expression of isolated GOI, is greater than 0.9.3. The last amino acid in proinsulin is unchanged.

[0179] The inventors conducted another evolution experiment. Four (4) variants were generated. In each, they followed the design of the Arc 15 variant from the previous experiment. In one there was no stop codon, the other 3 followed the stop codon designs selected.

[0180] They were tested for yield over 50 days of in-lab evolution, with 3 repeats each. Proinsulin levels were quantified using ELISA plates according to manufacturer instructions (product number RAB0327) for proinsulin detection.Statistical AnalysisComparison of unfused to fused GFP

[0181] In the 10-gene experiment (Fig. 6), the inventors compared the stability of the 10 fused genes to the unfused GFP baseline. The error analysis process was as follows:1. Data Collection:For the 10 fused genes, the inventors calculated the mean (p) and standard deviation (c) of their fluorescence measurements at t = 15, the last day of the experiment. The fluorescence measurements were normalized by the fluorescence at t = 1, the experiment initiation. This was done in order be able to assume normal distribution, whilemaximizing the effect of decay. For the unfused GFP, the inventors used the standard error (SE) of its fluorescence measurements at this time.2. Calculation of T-Score:To determine if the fused genes were significantly more stable than the unfused GFP, the inventors calculated a t-score for the mean fluorescence of the 10 fused genes using the formula:Where:• infused isthe mean fluorescence for each of the fused genes.• / hinfused isthe mean fluorescence of the unfused GFP.• fused ’s l'lcstandard deviation of fluorescence measurements for the fused genes.•nfused is the sample size for the fused genes.• SEunfusied is the standard error of the unfused GFP.3. One-Tailed Test:A one-tailed t-test was used to assess whether the mean fluorescence of the fused genes was significantly higher than the unfused GFP. This approach focused on detecting whether the fused genes exhibited higher stability (i.e., greater fluorescence) than the unfused GFP.4. Calculation of P- Value:The t-score was translated into a p-value using the standard t-distribution. The p-value was used to assess the statistical significance of the difference in stability between the fused genes and the unfused GFP. The inventors calculated a p-value of 0.048. A p-value of less than 0.05 was considered statistically significant, indicating that the fused genes were significantly more stable than the unfused GFP.5. Interpretation of Results :A significant p-value indicated that the fused genes had higher stability than the unfused GFP, supporting the hypothesis that gene coupling contributes to greater stability.6. Error Considerations:Variability due to biological noise and experimental measurement errors was accounted for by incorporating the standard deviation of the fused genes and the standard error of the unfused GFP in the t-test calculation.Comparison between experiments

[0182] The inventors compared the stability of all experiments, to show that they do not have the same decay. The error analysis process was as follows:1. Data Collection:For the 10 fused genes and unfused GFP, the inventors collected all their fluorescencemeasurements along the experiment and normalized them by their fluorescence at time t = 1.2. Statistical test:Under these assumptions, normality is no longer assumed. Rather, the inventors performed a Kruskal-Wallis H test. This test does not assume normality, and has as a null hypothesis that all tests are equal. The inventors derived a P-value of 0.035.3. Interpretation of Results :A significant p-value indicated that the different fused genes and GFP come from different distributions, and do not have the same decay rate.4. Pairwise testing:Following this analysis, the inventors conducted pairwise tests between the different fused genes and GFP, receiving significant results only for SEC2, with a p-value of 0.047. This indicates that top performing EG have a significant advantage over the baseline, while others may not - emphasizing the need for informed selection.

[0183] The inventors further note that similar tests were conducted for other analyses:• Conducting linear regression analysis on the log-values of the expression in the insulin experiment, the inventors get that they are linear, indicating exponential decay (P<10-6). The inventors also found that they have different decay rates (ANOVA F - test, P<1019).• By using the student’s t-test for the leaky codon insulin experiment at time 40, the inventors show that using a leaky stop codon provides slower decay than not using one (P<10“5).• By using the Kruskal-Wallis H test for the leaky codon insulin experiment, the inventors show that each codon leads to a different expression profile (P<0.003).Dataset Construction

[0184] The dataset was taken from the SWAT library as previously published (Zhu et al., “Engineering the robustness of industrial microbes through synthetic biology”. Trends in Microbiology vol. 20 Preprint at https: / / doi.Org / 10.1016 / j.tim.2011.12.003 (2012); and Parker and Kunjapur, “Deployment of Engineered Microbes: Contributions to the Bioeconomy and Considerations for Biosecurity”. Health Secur 18, (2020)). It contained fluorescence measurements of 6,685 endogenous genes in S. cerevisiae. They were engineered as fused genes, fused with either a NOP1 promoter and GFP or TEF with mCherry at the N’ terminus of the gene. GFP and mCherry datasets were analyzed separately to account for any potential variation between labels. A correlation analysis revealed moderate alignment between GFP and RFP fluorescence. All data were complete, and no imputation was required, as the features were derived directly from the genetic sequence. Fluorescence values were used as raw data without normalization to maintain their original scale. The data exhibited highly skewed labels, with correlations between the two target genes, GFP and RFP. To address this skewness, the inventors employed a sub-sampling strategy that introduced variation across training sets, enabling thecreation of an ensemble of models trained on different data subsets. This was repeated 20 times (see Ensemble Approach below).Model Training and EvaluationEnsemble Approach

[0185] An ensemble approach was used, where 85% of the data was sampled for training in each iteration, and 15% held out for validation. This was repeated 20 times with different samples, to ensure robustness. Predictions were averaged among the ensemble models, for the final prediction.Models and Hyperparameters

[0186] Four (4) models were trained: XGBoost, KNN, SVR, and ElasticNet.• XGBoost: Hyperparameters included learning rate (0.01-0.3), tree depth (3-10), subsample fraction (0.5-1), and regularization terms (lambda and alpha).• KNN: Number of neighbors (3-50) and distance metrics (Euclidean, Manhattan).• SVR: Kernel types (linear, RBF), penalty term (C = 0.1-10), and epsilon margin (0.001- 0.1).• ElasticNet: Regularization strength (alpha = 0.01-1) and LI ratio (0.1-1).• Hyperparameters were tuned using Optuna, employing 5-fold cross-validation on training subsets.Performance Metrics

[0187] The primary performance metric was the median model result, expressed as quantile of the top fluorescence within test set in the top 3 candidates. Another performance metric measured was likelihood of failure, defined as probability of top fluorescence being below median performance within test set. These were also measured for performance for only top candidates to check whether the model selected is consistent. Additionally, the models’ performance was measured using the Spearman correlation, and AUC for correctly classifying the top 5%. These metrics tested robustness (rather than performance on top candidates) and agree with the conclusion that ensembling KNN with XGBoost improved the model’s performance.Feature Importance Analysis

[0188] Feature importance was derived for XGBoost using SHAP (SHapley Additive exPlanations), which ranks the contribution of each feature to model predictions. The top 20 features by SHAP values were used to assess the dominant predictors. Forward feature selectionwas applied to the KNN architecture, iteratively adding features based on their ability to maximize the AUC metric defined above.EXAMPLE 1Rationally designed gene fusing increases expression level of heterologous genes

[0189] The inventors report that fusing in-frame heterologous gene to endogenous genes of a host organism, under the same promotor and while using the same sequence, elevates the expression level of the heterologous gene in a statistically significant manner. This was demonstrated for insulin expression in Saccharomyces cerevisiae, either expressed solely (with a leader sequence) or fused to two different endogenous genes of S. cerevisiae (also with leader sequence), namely Arcl5 and Caf20 (Figs. 1 and 3).

[0190] The endogenous gene was chosen by an artificial intelligence (Al) algorithm, that was developed by the inventors for this purpose, and trained on data of fluorescent protein(s) fused with all the ORF in yeast. The inventors constructed a list of features, corresponding to the target gene (insulin in the herein presented exemplification) and to the endogenous gene, and by utilizing XGBoost the inventors build a model that is able to predict beneficiary fusion proteins. The inventors hypothesized that fusing insulin to the algorithmically chosen endogenous gene would elevate insulin’s expression levels due to the presence of numerous regulatory elements, either on the DNA, RNA, or even at the protein level of the endogenous gene, that in turn would lead to possible optimization in transcription, translation, mRNA stability, or any combination thereof (Fig. 2). Future research will focus on pinpointing the regulatory elements offering this increased expression and optimizing synthetic genes with it.

[0191] Protein expression levels were quantified using ELISA kit and analyzed using a plate reader. These results demonstrate a novel method to increase expression level of heterologous genes in a completely orthogonal and easily generalized manner (Fig. 1). The specific genes were chosen by an Al model developed by the inventors, that was devised to optimize the expression levels of any gene of interest by linking it to a rationally picked endogenous gene with a specifically designed linker for each fusing. As seen from the results, different genes have different effects on the expression level and thus, fine tuning the current Al model will allow the inventors to achieve even a greater improvement of the expression levels.EXAMPLE 2Fusion Strategy Improves Evolutionary Stability

[0192] To assess the impact of the fusion strategy, the inventors evaluated the stability of GFP fused to various endogenous genes (EGs) by examining 10 strains from a previously described library of N-terminally GFP-tagged genes in S. cerevisiae. Fluorescence was used as a proxy for expression for 15 days.

[0193] For a baseline comparison, GFP was substituted for canl. The canl gene encodes an arginine permease, which is not essential for yeast survival under laboratory conditions where arginine is supplemented in the medium. This makes it a suitable target for replacement without affecting the viability of the yeast cells.

[0194] This experiment was conducted prior to the creation of the EG selection mechanism - it was designed to give an indication of the need for such a model and its potential impact. For this purpose, meaningful bioinformatic features were generated for all EG, and they were clustered in many clustering configurations. Ten (10) strains were selected such that they were consistently classified to different clusters and were near the centroids of these clusters. This generated a set of strains that were highly varied between them, while being representative of many similar genes (see more details in the Methods section).

[0195] The experiment (Fig. 7) validated the following conclusions, emphasizing the need for an EG selector:

[0196] All strains demonstrated significant decline over the course of the experiment, emphasizing the existence of mutational instability, and the need to improve it.

[0197] GOI-EG fusions exhibited significantly slower declines in fluorescence compared to unfused GFP, confirming that fusing genes enhance stability.

[0198] Different EGs yielded varying degrees of stability, demonstrating that the stability of a fused gene depends on the endogenous gene selected.

[0199] One gene displayed a statistically significant advantage over an unfused GFP, emphasizing the need for more informed gene selection.EXAMPLE 3Machine Learning Predicts Optimal EG-GOI Combinations

[0200] The variability in stability observed across different EGs highlighted the importance of systematic EG selection. To address this, the inventors developed a machine learning model to predict EG-GOI fusions that maximize expression and stability (more details in the Methods section). The model was trained on fluorescence data collected from GOI-EG fusion librariesunder various conditions in S. cerevisiae. As the fluorescence was measured after the variants had time to mutate, this is assumed to capture a combination of both expression and stability.

[0201] Using features such as codon usage bias (tRNA adaptation index, codon adaptation index), GC content, mRNA folding energy, ChimeraARS scores, and other meaningful bioinformatic features, the model successfully ranked potential fusion pairs.

[0202] In a reasonable use case, a user would generate and utilize very few genetic designs. The model would recommend 1-3 EG, which the user would validate experimentally. To test the different architectures in a method that aligns with this task, the models’ performance was estimated on the top prediction, and the best performance in the top 3 recommendations. They were compared based on their median performance (predicting a high quantile of expression within the test set), and their robustness (likelihood of failure, or accidentally recommending a low-performance EG). For performance measurement with more common and less relevant metrics, see supplementary.

[0203] An ensemble model combining k-nearest neighbours (KNN) and XGBoost (XGB) was selected and trained, providing both high median performance and high robustness. Within the top 3 candidates, the median model performance found the 0.995 (CI [0.935, 0.998]) quantile for expression. For the top candidate, the median model performance found the 0.939 (CI [0.927, 0.997]) quantile for expression. For comparison with other architectures, see methods.

[0204] These results underscore the predictive power of the ME model and its ability to systematically identify highly performing gene fusions, enhancing the efficiency of fusion design.EXAMPLE 4Validation with Proinsulin Production

[0205] To demonstrate the real-world applicability of the current approach, the inventors applied it to stabilize human proinsulin expression in yeast, a biotechnologically relevant system. EGs were selected for high performance in both the XGB and KNN models, and additional engineering needs (see Methods). Thus, the current model identified two EGs — Caf20 and Arc 15 — as suitable fusion partners for proinsulin.

[0206] A 30-day in-lab evolution experiment was conducted, where expression level measurements were taken every 5 days, using the ELISA protocol. The strains tested were the original proinsulin (by Novo Nordisk), the proinsulin as optimized by the Evolutionary StabilityOptimizer (ESO), and the optimized sequence fused to Caf20 and Arcl5 as fusion genes. As in the 10-gene experiment, the unfused proinsulin replaced Canl.

[0207] The need for a leader secretion peptide in proinsulin production in yeast has been well documented. However, it was not clear whether it may disrupt the efficacy of the current design. The inventors measured the expression at initiation for the original proinsulin and two fusion genes, with and without a leader sequence (Fig. 6). All variants exhibited much higher expression in the presence of a leader sequence, and thus all further experiments were conducted as such.

[0208] The expression patterns approximately followed an exponential decay pattern, where the fused genes displayed a much slower decay rate (Fig. 6). This enabled the inventors to estimate the cumulative expression of proinsulin over time. These quantities were normalized by the expression estimated for 10 days, for the original proinsulin. The inventors note, the gene fusions showed a fivefold increase in total proinsulin yield over the experimental period (Fig. 6). The inventors note that Arc 15 demonstrated better performance for shorter durations (higher initial expression) while Caf20 demonstrated better performance for longer durations (slower decay) - the selection of the optimal gene will thus depend on the industrial use case.

[0209] As the graph (in Fig. 6) is in log scale, the wish to note two significant takeaways:1. For the fusion gene with Arcl5, the inventors measured 98.5 mg / L at initiation, where for the original sequence, the inventors measured 72.9 mg / L - a -35% increase at initial expression.2. For the fusion gene with Arcl5, the inventors measured 68.7 mg / L after 30 days, meaning -70% of expression is maintained. For the original sequence, the inventors measured 4.8 mg / L, meaning only -7% of expression is maintained.

[0210] Nanopore sequencing was conducted on the final sequences. It reveals that following the experiment, the fused proinsulin suffered few mutations, while the unfused proinsulin sequence was lost completely. This has been further validated by western blot (Fig. 6).EXAMPLE 5Leaky Stop Codon Enables GOI Production and Overexpression

[0211] To enable generation of both the fusion protein and the GOI alone, the inventors utilized leaky stop codons. These stop codons are characterized by their leakiness - some of the time, they function as regular stop codons, ending transcription. The rest of the time, they allow read- through, leading (in the current case) to the generation of the fusion protein.

[0212] Different stop codons lead to different rates of read-through (Fig. 7). The ability to enforce low but non-zero read-through rates is a key factor in improving the stability andindustrial viability of the current design. If the fusion protein’s abundance is reduced to barely viable, then it is highly likely that mutations affecting the GOI’s expression would lead to host lethality, promoting mutational stability. This is in tandem with the enforcement of much higher expression of the isolated GOI, which is the necessary and desired product.

[0213] Thus, informed selection of the stop codon is necessary. To enable this informed decision, the inventors conducted an experiment, creating a fusion protein of BFP and mCherry, with different leaky stop codons between. This was expressed on a plasmid (Fig. 7). The red fluorescence is expressed only in the fusion protein, while the blue fluorescence is expressed in both the fusion protein and BFP alone. Thus, the ratio between the fluorescence measurements can be used to derive the rate of read-through.

[0214] The inventors selected the 3 stop codon designs expressing the lowest mCherry fluorescence (to minimize read-through) which was still detectable (to increase cell viability). The inventors conducted another evolution expression, attaching the proinsulin gene to Arcl5, either without a stop codon or with one of the 3 selected. Using a similar protocol, expression of proinsulin was measured over 50 days, this time capturing the presence of both the fusion protein and the isolated proinsulin. In addition to the secretion of isolated proinsulin (which is a requirement in and of itself), the inventors observed a much higher expression rate, and cumulative expression accordingly (Fig. 7).

[0215] To explicitly state two significant takeaways:

[0216] For the fusion gene without a stop codon, the inventors measured 102 mg / L at initiation, where for the best candidate L3, the inventors measured 98 mg / L- a -4% decrease at initial expression.

[0217] For the best candidate L3, the inventors measured 83 mg / L after 50 days, meaning -85% of expression is maintained. For the fusion gene without a stop codon, the inventors measured 15 mg / L, meaning only -15% of expression is maintained.DiscussionKey Insights and Advantages

[0218] This study provides an organism-agnostic strategy to address evolutionary instability in synthetic biology. By fusing the GOI to an EG under a shared promoter, many deleterious mutations would lead to host lethality, leading to higher GOI mutational stability. By incorporating a leaky stop codon, the inventors enable these gene fusions to generate high expression of the GOI alone, while increasing their mutational stability. These, together withinformed selection of the linker, sequence optimization and use of a leader sequence, are key in the improved stability and performance observed in the current experiments.

[0219] The inventors found that following all current design principles, a 15% decline in expression of proinsulin over 50 days was observed. This contrasts with the original design, which declined by 93% over 30 days. This is in addition to a -30% expected increase in initial expression.

[0220] The current machine learning tool enhances the utility of this approach by selecting better GOI-EG pairings based on biologically meaningful features. Translational efficiency metrics, such as tRNA Adaptation Index (tAI), emerged as dominant predictors, underscoring the central role of translation efficiency in stabilizing gene fusions. Other features, including GC content, RNA folding energy, alternative shifted ORFs, and amino acid composition similarity, reflect the importance of optimizing sequence stability and minimizing regulatory conflicts.

[0221] Protein disorder predictions were employed to guide linker selection, increasing the likelihood that the gene fusions maintain native protein folding and functionality. Combined with avoidance of mutational hotspots and sequence optimization for expression and stability, these elements ensure the practical applicability of the approach across a broad range of use cases.Current Limitations and Future Directions

[0222] Using the current machine learning model, the inventors have already seen empirical evidence for good performance in predicting high-performing GOI-EG pairs for different GOI, as shown in the current experimental validation. Its reliance on sequence-based features enables effective application to non-model organisms, even in the absence of extensive empirical data. However, expanding the dataset used for training the model remains a valuable avenue for further enhancement. Systematic testing for libraries from more host organisms and a broader range of target genes would improve the model’s generalizability and ensure its utility across an even wider variety of synthetic biology applications. Furthermore, in many organisms, epigenetic silencing must be considered when optimizing expression and stability. While the inventors have designed the current model with capabilities to avoid epigenetically silencing motifs, this application should be tested in vivo.

[0223] While the current method has shown significant promise, further experimental validation is needed to refine certain aspects. For example, empirical testing of linker selection will help confirm computational predictions and guide future improvements. An in-depth analysis and testing of the leaky stop codons would also enable further optimization.Broader Implications

[0224] This study demonstrates the potential to stabilize synthetic genes in both industrial and environmental contexts. In biomanufacturing, the ability to maintain stable GOI expression over extended periods can reduce costs, improve scalability, and simplify regulatory processes. In environmental applications, robust gene fusions could support long-term deployment in dynamic and uncontrolled conditions, enabling breakthroughs in bioremediation and biosensing.Conclusion

[0225] By integrating gene fusion, machine learning-guided gene selection, protein disorder predictions for minimizing misfolding, sequence optimization, and leaky stop codons, the inventors present a robust framework for addressing the evolutionary instability of synthetic genes. This system achieves high GOI expression, ensures host viability, and enhances mutational stability across generations. The reliance on biologically interpretable features, such as tAI and RNA folding energy, underscores the model's capacity to identify high-performing designs.

[0226] Its organism-agnostic nature and focus on sequence-level optimization make this approach a versatile solution for diverse synthetic biology applications. With further validation, refinement, and the incorporation of additional data, this framework holds significant promise for advancing biotechnology and synthetic biology.

[0227] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.

Claims

CLAIMSWhat is claimed is:

1. A chimeric DNA molecule comprising: a. a first nucleic acid sequence encoding a polypeptide of interest, wherein a 3' end of said first nucleic comprises a stop codon; and b. a second nucleic acid sequence encoding an endogenous polypeptide of a target cell, wherein said first nucleic acid sequence is located upstream to said second nucleic acid sequence in said chimeric DNA molecule, and wherein said first and second nucleic acid sequences are operably linked.

2. The chimeric DNA molecule of claim 1, further comprising a third nucleic acid sequence located between said first nucleic acid sequence and said second nucleic acid sequence.

3. The chimeric DNA molecule of claim 1 or 2, being codon optimized for expression in said target cell.

4. The chimeric DNA molecule of any one of claims 1 to 3, wherein said endogenous gene is a gene highly expressed in said target cell, a gene being essential for viability / fitness of said target cell, or both.

5. The chimeric DNA molecule of any one of claims 1 to 4, wherein said first nucleic acid sequence and said second nucleic acid sequence are in-frame or not in-frame.

6. An expression vector or a plasmid comprising the chimeric DNA molecule of any one of claims 1 to 5.

7. A cell comprising: a. the chimeric DNA molecule of any one of claims 1 to 5; b. the expression vector or plasmid of claim 6; or c. both (a) and (b).

8. The cell of claim 7, being a prokaryote cell or a eukaryote cell.

9. The cell of claim 7 or 8, being a recombinant cell, a transgenic cell, a transfected cell, or a transduced cell.

10. The cell of any one of claims 7 to 9, characterized by increased or over expression of said polypeptide of interest compared to a control being devoid of said chimeric DNA molecule.

11. The cell of any one of claims 7 to 10, characterized by expression level of said endogenous polypeptide being essentially similar or increased compared to a control cell being devoid of said chimeric DNA molecule; said expression vector or plasmid; or both.

12. An extract, lysate, homogenate, and any fraction thereof, of the cell of any one of claims 7 to 11.

13. The extract, lysate, homogenate, and any fraction thereof, of claim 12, comprising any one of: said polypeptide of interest, said endogenous polypeptide, an mRNA transcribed from said first nucleic acid sequence, an mRNA transcribed from said second nucleic acid sequence, and any combination thereof.

14. The extract, lysate, homogenate, and any fraction thereof, of claim 12 or 13, consisting essentially of said polypeptide of interest.

15. A composition comprising any one of: a. the chimeric DNA molecule of any one of claims 1 to 5; b. the expression vector or plasmid of claim 6; c. the cell of any one of claims 7 to 11; d. the extract, lysate, homogenate, and any fraction thereof of any one of claims 12 to 14; or e. any combination of (a) to (d), and an acceptable carrier.

16. The composition of claim 15, consisting essentially of said polypeptide of interest and said acceptable carrier.

17. A method for increasing expression of a polypeptide of interest in a target cell, the method comprising culturing the cell of any one of claims 7 to 11 such that said polypeptide of interest encoded by said first nucleic acid sequence is expressed.

18. The method of claim 17, wherein said increased is compared to a control being devoid of said chimeric DNA molecule.

19. The method of claim 17 or 18, wherein said increasing expression, comprises increasing any one of: abundance and / or secretion of mRNA transcribed from said first nucleic acid sequence, abundance and / or secretion of said polypeptide of interest, and both.

20. The method of claim 19, wherein said increased abundance of mRNA comprises increased: mRNA transcription, mRNA stability, or both.

21. The method of any one of claims 17 to 20, wherein an mRNA transcribed from said second nucleic acid sequence, said third nucleic acid sequence, or both, is not translated.

22. A method for increasing genetic stability of a cell producing a polypeptide of interest, the method comprising culturing the cell of any one of claims 7 to 11 such that said polypeptide of interest encoded by said first nucleic acid sequence is expressed.

23. The method of claim 22, wherein said increasing is compared to a control cell.

24. The method of claim 22 or 23, wherein said culturing is for a period of 25 to 75 days.

Citation Information

Patent Citations

  • Chimeric polypeptides and methods of preparing same

    US20230167476A1

  • Methods for producing a polypeptide using a crippled translational initiator sequence

    US6548274B2