Method for optimizing a nucleotide sequence for expression of an amino acid sequence in a target organism - Patent Application 20070122997
By optimizing nucleotide sequences using codon n-sets with higher relative frequencies in the target organism's genome, the method enhances protein yield and solubility, addressing the limitations of existing methods in heterologous expression systems.
Patent Information
- Application Number
- JP2025526880
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-22
- Filing Date
- 2023-07-21
- Publication Date
- 2025-09-17
AI Technical Summary
Existing methods for optimizing nucleotide sequences for heterologous protein expression in target organisms often fail to achieve significant improvements in protein yield and solubility, particularly focusing on quantitative optimization while neglecting the importance of solubility for scientific, medical, and industrial applications.
The method involves replacing base triplets in a nucleotide sequence with synonymous base triplets to optimize codon n-sets based on their relative frequency in the target organism's genome, focusing on modifying positions with low codon n-set relative frequencies to enhance translation efficiency and solubility.
The method significantly increases the yield and solubility of proteins in heterologous expression systems, enabling the biosynthesis of previously unexpressed proteins with improved efficiency and sustainability.
Smart Images

Figure 2025530875000001_ABST
Abstract
Description
Reference to earlier application
[0001] This application claims priority to German Patent Application No. 10 2022 118 459.5, filed July 22, 2022, the disclosure of which is incorporated herein by reference in its entirety.
[0002] [Sequence Listing] This application contains as part of its description an electronic sequence listing in xml format conforming to the WIPO ST.26 standard containing 21 sequences, the contents of which are incorporated herein by reference in their entirety. [Technical Field]
[0003] The present invention relates to the generation of synthetic nucleotide sequences and their use for producing proteins by introducing these nucleotide sequences into an expression system using a suitable host organism that expresses the protein encoded by the nucleotide sequence. The present invention particularly relates to methods for optimizing nucleotide sequences for expression in a given host organism. [Background technology]
[0004] Heterologous expression systems are of great importance in the production of recombinant proteins in biotechnology. An "expression system" is a biological system that is able to carry out protein biosynthesis in a targeted and controlled manner, i.e., to produce or "express" specific proteins according to their nucleotide sequences.
[0005] Here, "heterologous expression" refers to the expression of a gene or a part of a gene in a host organism that does not naturally have this gene or gene fragment. The corresponding nucleotide sequence is introduced into the host organism by genetic engineering or recombinant DNA technology, such as vector or genome editing, and the host organism is then stimulated to grow and overproduce the protein. "Homologous expression" refers to the expression of a gene or part of a gene in the host organism or system from which it originates.
[0006] Heterologous protein expression can be carried out in many types of host organisms. Host organisms can be, for example, bacteria, fungi, yeast, insect cells, mammalian cells, or plant cells. A frequent problem in heterologous protein expression is the low transcription and translation rate of foreign nucleotide sequences in a particular host organism (also referred to herein as the target organism). One cause is the degeneracy of the genetic code, which results in the existence of multiple codons with the same meaning (also referred to herein as "synonymous codons") for the majority of amino acids incorporated during translation. A codon is a sequence of three consecutive nucleic acid bases, or "base triplet," in a nucleic acid sequence that can encode an amino acid. There are a total of 64 potential codons, 61 of which encode the 20 standard proteinogenic amino acids, and the other three encode stop codons.
[0007] Different organisms can utilize codons to express amino acids at different frequencies, a phenomenon also known as "codon usage" or "codon bias." Furthermore, it has been found that codons are not randomly arranged within an organism's genes; the observed frequencies of codon pairs can deviate unexpectedly from the product of their respective single frequencies, resulting in statistical "under-representation" or "over-representation." This relationship, also known as "codon context," can further affect translation efficiency.
[0008] Thus, known possibilities for increasing expression rates include codon optimization by replacing a single codon in the expressed nucleotide sequence with a synonymous codon that is more frequent in the target organism (also known as "codon optimization" or "codon usage optimization"), and codon optimization by transferring the codon frequencies of highly expressed genes in the target organism to the target protein (also known as "codon adaptation index"). Furthermore, there is the possibility of codon context optimization, whereby codons in the nucleic acid sequence to be expressed are replaced with over- or under-represented synonymous codons to adapt them to the codon context of the target organism, without changing the encoded amino acid sequence.
[0009] WO 2020 / 024917 discloses a computer-implemented method for optimizing a nucleic acid sequence for expression of a protein in a host based, inter alia, on a codon adaptation index and codon context, using a computer-assisted NSGA-III algorithm.
[0010] WO 2004 / 059556 discloses a computational method for optimizing a given nucleic acid sequence for expression in a given target organism using a quality function that can take into account, inter alia, codon usage and codon context as quality criteria.
[0011] WO 2008 / 000632 proposes a further method for generating new coding sequences from a given nucleotide sequence encoding a given amino acid sequence by substituting one or more synonymous codons in several iterative steps, and further developing them based on fitness values taking into account, inter alia, the codon context of the host organism.
[0012] A method for determining an optimized nucleotide sequence that encodes a given amino acid sequence and is optimized for expression in a particular target organism is known from WO 2018 / 104385, in which multiple candidate nucleotide sequences are generated and evaluated using statistical machine learning algorithms.
[0013] A further method for optimizing nucleotide sequences for heterologous protein expression, based on the codon context of the target organism, is known from WO 2007 / 130650.
[0014] These conventional optimizations have often proven insufficient to achieve significant improvements in protein yield in heterologous systems. Furthermore, known methods generally focus on quantitative optimization of the yield of transcribed mRNA or protein, while ignoring the solubility of the expressed protein, which is essential for use in scientific, medical, pharmaceutical, and industrial purposes.
[0015] It is therefore an object of the present invention to provide an improved method and apparatus for optimizing nucleotide sequences for expression of a given amino acid sequence in a target organism, thereby at least partially resolving the above-mentioned problems and, in particular, increasing the proportion of soluble protein in heterologous protein expression.
[0016] This object is solved by the subject matter of the independent claims. Preferred embodiments of the invention are the subject matter of the dependent claims and the following description. Summary of the Invention [Problem to be solved by the invention]
[0017] The subject of the present invention is a method for optimizing a nucleotide sequence for the expression of a given amino acid sequence in at least one target organism. Expression can in principle be heterologous or homologous, with heterologous expression being preferred. [Means for solving the problem]
[0018] The nucleotide sequence comprises multiple base triplets, and in at least one modified position of the nucleotide sequence, a base triplet encoding an amino acid of a predetermined amino acid sequence is replaced with a synonymous base triplet encoding the same amino acid of the predetermined amino acid sequence, with the aim of optimizing the nucleotide sequence for expression in at least one target organism.
[0019] wherein the modified position according to the invention comprises a direct succession of n base triplets constituting a first codon n-set and encoding a sequence portion of n amino acids of a predetermined amino acid sequence constituting an amino acid n-set, said amino acid n-set being encoded by a predetermined number of amino acid n-set events in the genome or part thereof of at least one target organism and / or in the genome or part thereof of a virus capable of infecting at least one target organism.
[0020] The method of the present invention comprises replacing at least one of n base triplets in direct succession at at least one modification position with a synonymous base triplet, the synonymous base triplet being selected so as to obtain a second codon n-set having a higher codon n-set relative frequency than the first codon n-set in terms of the number of amino acid n-set events in the genome or part thereof of at least one target organism and / or the genome or part thereof of a virus capable of infecting at least one target organism.
[0021] Here, n is a natural number greater than or equal to 2, and in particular less than or equal to the total number N of amino acids in the given amino acid sequence.
[0022] The present invention is based on the recognition that the effect of a direct series of n base triplets, referred to herein as a codon n-set, on the translation efficiency of a series of n amino acids (referred to herein as an amino acid n-set) encoded by the n base triplets can be represented in terms of the relative frequency in a particular target organism at which the direct series of n base triplets encodes the direct series of n amino acids in the genome of the host organism (referred to herein as the codon n-set relative frequency).
[0023] Without being bound by theoretical considerations, it is assumed herein that the translation process is at least transiently affected by two or more consecutive base triplets simultaneously binding to the ribosome during translation. The inventors have found evidence that codon n-pairs with favorable effects are found more frequently in genomes than synonymous codon n-pairs with less favorable or unfavorable effects.
[0024] In other words, the inventors have recognized that the ratio of the absolute frequency of a codon n-set to the absolute frequency of the corresponding amino acid n-set encoded by the codon n-set is an advantageous measure for quantifying the suitability of a given nucleic acid sequence for expression in a specific target organism. The relative frequency can take a value between 0 and 1 or 0% and 100%, and if a particular codon n-set is not used at all to encode the corresponding amino acid n-set in the genome of the target organism or a portion thereof, the relative frequency will be equal to 0. If a particular amino acid n-set is exclusively encoded by a particular codon n-set in the genome of the target organism or a portion thereof, the relative frequency will be equal to 1 or 100%. In this way, by selectively replacing codon n-sets with low relative frequencies in the genome of a given target organism with codon n-sets with high relative frequencies, nucleic acids can be optimized to be better suited to the target organism for expression of a given amino acid sequence.
[0025] If a particular amino acid n-set of an amino acid sequence does not exist at all in the genome or part thereof of the target organism or the genome or part thereof of the virus, then in principle, any codon n-set encoding this amino acid n-set can be selected. For example, each related synonymous codon n-set can be assigned a uniform relative frequency of 1 / i, where i is the number of synonymous codon n-sets encoding the amino acid n-set. Alternatively, the codon n-set encoding this amino acid n-set can be assigned a relative frequency of 0. Another option is to exclude this modification position or this amino acid n-set from optimization. [Effects of the Invention]
[0026] It has been found that for a given target organism, with the aid of the method of the invention, the expression of heterologous proteins, in particular the yield of soluble proteins, can be increased many-fold compared to previously known methods, see also in this regard the results of comparative experiments in the following embodiments.
[0027] Furthermore, the methods of the present invention allow the biosynthesis of proteins that have not previously been possible or have only rarely been possible to heterologously express, thus leading to improved efficiency and sustainability of biotechnological protein production for scientific, medical, technical or industrial purposes. DETAILED DESCRIPTION OF THE INVENTION
[0028] In preferred embodiments of this method, n is 50 or less, 40 or less, 30 or less, 20 or less, or 10 or less. Particularly preferred is n selected from the group consisting of n=2, n=3, n=4, n=5, n=6, and all combinations thereof. In one embodiment, n=2. In one embodiment, n=3. In one embodiment, n=4. In one embodiment, n=5. In one embodiment, n=6. Particularly preferred is n=3 or greater. These n's achieve advantageous matching of the nucleotide sequence for expression in the target organism.
[0029] As used herein, the term "predetermined" means that the number of events of an n-type amino acid pair is predetermined by the genome or proteome or a portion thereof of at least one target organism, or the genome or a portion thereof of a virus capable of infecting at least one target organism. Therefore, determining the number of events (hereinafter also referred to as absolute frequency) in which a particular n-type amino acid pair is encoded in the genome or a portion thereof of at least one target organism, or the genome or a portion thereof of a virus capable of infecting at least one target organism, is understood to be a step that can be performed during the method, but does not have to be performed. Rather, information regarding the absolute frequency at which an n-type amino acid pair is encoded in the genome or a portion thereof of at least one target organism and / or the genome or a portion thereof of a virus capable of infecting at least one target organism can also be included in the method of the present invention from other means, such as a database. Of course, determining the number of events in which an n-type amino acid pair is encoded in the genome or a portion thereof of at least one target organism, or the genome or a portion thereof of a virus capable of infecting at least one target organism, can also be performed as a step in the method.
[0030] The same principle applies to determining the number of events, i.e., the absolute frequency with which a particular n-pair of codons occurs in the genome or part thereof of at least one target organism or in the genome or part thereof of a virus capable of infecting at least one target organism, and / or the resulting relative frequencies of said n-pairs of codons according to the invention. For example, the absolute and / or relative frequencies of essentially all possible combinations of n-pairs of codons in the genome or part thereof of at least one target organism or in the genome or part thereof of a virus capable of infecting at least one target organism can be stored in a database and included in the method of the invention in the form of database information.
[0031] In the method of the present invention, synonymous base triplets are specifically selected from the perspective of satisfying the condition required by the present invention, that is, the codon n-set relative frequency of the second codon n-set is higher than that of the first codon n-set. Therefore, selecting synonymous base triplets may particularly include determining and / or evaluating the codon n-set relative frequency of the second codon n-set. As mentioned above, the term "determining" can be performed, for example, in the form of calculating the codon n-set relative frequency of the second codon n-set, or in the form of comparing data or incorporating database information. Therefore, a preferred embodiment of this method includes at least the following steps: a) determining at least one modification position; b) replacing at least one base triplet at the at least one modification position with a synonymous base triplet; and c) determining the codon n-set relative frequency of the resulting second codon n-set. Preferably, the method further comprises the step of d) evaluating the codon n-set relative frequency of the second codon n-set determined in step c), wherein said evaluation can, for example, involve a comparison with the codon n-set relative frequency of said first codon n-set and / or can be carried out based on a target criterion such as a minimum value. In this regard, see also the following description and exemplary embodiments. Steps b) and c) and, if applicable, d) can also be repeated until a second codon n-set required by the present invention is obtained, the codon n-set relative frequency of which is higher than that of the first codon n-set.
[0032] As mentioned above, the determined number of amino acid n-set events or the determined codon n-set relative frequency does not necessarily have to be based on the complete genome of at least one target organism or the complete genome of a virus capable of infecting at least one target organism. Rather, in certain embodiments of the method, it may be sufficient and advantageous if the determined amino acid n-set absolute frequency or codon n-set relative frequency includes only a portion of said genome or genomes, especially since most genomes contain to a large extent non-coding regions, respectively, that are less relevant for assessing the suitability of a given nucleic acid sequence for expression in a target organism according to the invention.
[0033] In a preferred embodiment, the number of amino acid n-pair events is obtained from some, preferably all, protein-coding genes and / or proteins of at least one target organism or virus that can infect at least one target organism, or is determined based on some protein-coding genes and / or proteins of at least one target organism or virus that can infect at least one target organism.The protein that is constitutively expressed by target organism or the protein that shows high transient expression or high abundance is particularly suitable for this purpose.Particularly preferably, at least 25%, at least 50%, at least 75%, at least 80%, at least 90% or at least 95% of the coding part of the genome of target organism is included in determining the number of amino acid n-pair events.
[0034] Thus, the codon n-set relative frequency is obtained from the number of events of each of the first or second codon n-set in some, preferably all, protein-coding genes and / or proteins of at least one target organism or a virus capable of infecting at least one target organism, or is determined based on the number of events of at least one amino acid n-set in some protein-coding genes and / or proteins of at least one target organism or a virus capable of infecting at least one target organism, wherein the number of events is based on the number of events of the amino acid n-set. Particularly preferably, at least 25%, at least 50%, at least 75%, at least 80%, at least 90%, or at least 95% of the coding portion of the genome of the target organism is included in the determination of the codon n-set relative frequency.
[0035] In principle, the methods of the present invention do not require that the relative frequencies of the first or second codon n-pairs be values calculated arithmetically from absolute frequencies. For example, the relative frequencies of the first and / or second codon n-pairs can be expressed by or replaced by another probability distribution or another probability measure (e.g., interval estimate). For example, a confidence interval for the success probability of a binomial or multinomial distribution can be determined from random observations in the form of a random sample of the genome or a portion thereof of at least one target organism or a random sample of the genome or a portion thereof of a virus capable of infecting at least one target organism, and the observed frequencies of the first or second codon n-pairs and / or the observed frequencies of all synonymous codon n-pairs encoding the same amino acid n-pairs can be used as the number of successes, and the frequencies of the amino acid n-pairs can be used as the number of trials, thereby determining, for example, a Clopper-Pearson confidence interval or a simultaneous confidence interval for a multinomial ratio at a 95% confidence level. The codon n-pair relative frequencies can then be represented by or replaced by, for example, interval centers, to provide a more accurate estimate. This is particularly advantageous when the frequency of the amino acid n-set is low and when the codon n-set relative frequency is very high or very low. Of course, the codon n-set relative frequency can also be calculated or represented using other values from the interval. In addition to the center of the interval, for example, the minimum value of the interval or the average value of the weighting function in the interval (e.g., -1 / x), especially in the case of very conservative estimates, can represent the codon n-set relative frequency.
[0036] Thus, one embodiment of the method of the present invention provides that at least one modification position comprises a direct succession of n base triplets that constitute a first codon n set and encode a sequence segment of n amino acids of a predetermined amino acid sequence that constitutes the amino acid n set, and at least one of the n base triplets in the direct succession is replaced with a synonymous base triplet, and the synonymous base triplet is selected by a prediction function so as to obtain a second codon n set that encodes the amino acid n set with a higher probability than the first codon n set in the genome or part thereof of at least one target organism and / or the genome or part thereof of a virus capable of infecting at least one target organism. Suitable prediction functions that can be implemented in the method of the present invention are known to those skilled in the art.
[0037] As used herein, the term "at least one modified position" includes the possibility that the method includes several or multiple modified positions, in each of which at least one of n immediately consecutive base triplets is replaced with a synonymous base triplet, the synonymous base triplet being selected so that the codon n-set relative frequency is higher in at least some of the resulting second codon n-sets than in the corresponding first codon n-set. Of course, for this purpose, some or all of the n base triplets can also be replaced with synonymous base triplets, respectively, at one or more modified positions.
[0038] In a preferred embodiment, at least in the modification position where the first codon n-set has the lowest codon n-set relative frequency among all modification positions, at least one of the immediately consecutive base triplets is replaced with a synonymous base triplet selected so that the resulting second codon n-set has a higher codon n-set relative frequency than the first codon n-set. The present inventors considered that the codon n-set with the lowest relative frequency in a nucleotide sequence has a limiting effect on the entire translation process, and therefore, replacing this first codon n-set with a second codon n-set with a higher relative frequency can have a particularly advantageous effect on translation and folding, resulting in a particularly strong improvement in the expression of a soluble protein in a target organism.
[0039] In the case where there are several or multiple modified positions in a nucleotide sequence, the method of the present invention provides the possibility of overlapping direct successions of n base triplets of at least two modified positions, where the base triplet to be replaced with the synonymous base triplet is simultaneously included in at least these two modified positions.
[0040] In such an embodiment, the method of the present invention may further provide that the synonymous base triplets are selected such that, at one of the two modification positions, the resulting second codon n-set has a lower codon n-set relative frequency than the corresponding first codon n-set, and at the other of the two modification positions, the resulting second codon n-set has a higher codon n-set relative frequency than the corresponding first codon n-set.
[0041] Alternatively, or in addition, the methods of the invention can provide overlapping modification positions such that the resulting second codon n-set has a higher codon n-set relative frequency than the corresponding first codon n-set at both modification positions. Furthermore, the codon n-set relative frequency can be reduced at both modification positions.
[0042] When n is selected to be 2 or greater, it is possible that direct succession of n base triplets at more than two modification positions overlap, and that the base triplet substituted with a synonymous base triplet may simultaneously be included at more than two modification positions. The synonymous base triplets can be selected so that the codon n-set relative frequency of the resulting second codon n-set is lower at at least one modification position and higher at other modification positions compared to the corresponding first codon n-set.
[0043] Surprisingly, the methods of the present invention show that when there are many at least partially overlapping modification positions in a nucleotide sequence, it may be necessary, even essential, to accept deterioration of individual codon n-sets, such that the codon n-set relative frequency of a second codon n-set is reduced compared to a first codon n-set, in order to achieve an overall improvement in the translation efficiency of the entire nucleotide sequence. This is particularly true when a modification position by a first codon n-set that is particularly unfavorable for expression in the target organism overlaps with a modification position by a first codon n-set that is favorable for expression in the target organism, making it impossible to select a synonymous base triplet that would increase the relative frequency of both second codon n-sets.
[0044] For example, the codon n-set relative frequency of the second codon n-set can be reduced relative to the corresponding first codon n-set at at least about 1%, 5%, or 10% and / or at most about 40%, 30%, or 20% of the modified positions.
[0045] In summary, the inventors have recognized that, to improve the expression rate of a soluble protein from a nucleotide sequence in a target organism, it may be far more important to increase a particularly low codon n-set relative frequency with the help of the method of the present invention than to achieve a particularly high codon n-set relative frequency in individual cases or even in the majority of cases. Thus, in a preferred embodiment, the method of the present invention provides, as an optimization requirement or goal, that the codon n-set relative frequency of the first codon n-set and the codon n-set relative frequency of the second codon n-set at the modified position each have a minimum value, and that the minimum value of the second codon n-set is greater than the minimum value of the first codon n-set.
[0046] In a further preferred embodiment, synonymous base triplets are selected so that the codon n-set relative frequency of the second codon n-set reaches the greatest possible minimum, i.e., the greatest possible global minimum, or is at least 50%, preferably 40%, or 30%, preferably 20%, and particularly preferably 10% below the greatest possible minimum. Those skilled in the art are familiar with suitable mathematical approximation and / or optimization methods for finding the greatest possible minimum. Again, the method of the present invention deviates from conventional technical teachings on codon optimization or codon context optimization in that it is primarily aimed at globally optimizing unfavorable modification positions in a nucleotide sequence that can be identified by low codon n-set frequencies according to the present invention, rather than achieving a local maximum in terms of codon usage or codon pair preference. This method allows for highly reliable heterologous expression of certain proteins that could not be expressed in heterologous target organisms using conventional optimization methods, either poorly or not at all.
[0047] Preferably, the synonymous base triplets are selected so that the average relative frequency of the second codon n-pair achieves the maximum value or is at least 50%, preferably at most 40% or at most 30%, preferably at most 20%, particularly preferably at most 10% below the maximum achievable value. In this way, the optimization method can be used to provide nucleotide sequences that are particularly well suited for expression in the target organism and can therefore be expressed in the target organism particularly reliably and with high soluble protein expression rates.
[0048] In the above-described embodiments, it is understood that the achievement of the above-described optimization criteria already means that the codon n-set relative frequency of at least some of the second codon n-sets is higher than that of the corresponding first codon n-set, and therefore determining the relative frequency of the first codon n-set is not necessary. In other words, the basic condition of the present invention that at least one of the n base triplets in direct succession at at least one modification position is replaced with a synonymous base triplet specifically selected to obtain at least one second codon n-set with a higher codon n-set relative frequency than the first codon n-set relative to the number of amino acid n-set events in the genome or part thereof of at least one target organism and / or the genome or part thereof of a virus capable of infecting at least one target organism is always met when the above-described criteria are met. However, the relative frequency of the first codon n-set can, in principle, be used as a control parameter for implementing the method of the present invention or, for example, to compare intermediate or final optimization results with the initial state. Furthermore, in such embodiments, the relative frequencies of the first codon n-sets can remain undetermined, since the optimization objectives are not assessed by the initial starting sequence.
[0049] As already explained several times above, the inventors have considered that in optimizing nucleotide sequences, it is often advantageous to focus the method more strongly on optimizing modification positions with low codon n-set relative frequencies rather than increasing the average or already high codon n-set relative frequencies. To this end, in a preferred embodiment of the method of the present invention, the average value includes gradually decreasing the weighting of the codon n-set relative frequencies of the first and second codon n-sets, so that the impact of higher codon n-set relative frequencies on the average value or its calculation is disproportionately greater than that of lower codon n-set relative frequencies. Conversely, increasing a low codon n-set relative frequency, even if only slightly, can have a greater impact on the average value than increasing a medium or high codon n-set relative frequency.
[0050] Of course, the method can be, or alternatively can be, designed so that the codon n-tuple relative frequency achieves a maximum value at each of at least some of the modification positions.
[0051] It is understood that the natural number n of n base triplets, codon n-pairs, and amino acid n-pairs at a modification position must be the same natural number. For example, if a modification position includes a direct sequence of three base triplets, this direct sequence also encodes a sequence portion of three amino acids in a given amino acid sequence, i.e., the first and second codon n-pairs are each a codon triad, and the amino acid n-pairs are corresponding amino acid triads. In any case, the number n may be different at different modification positions, e.g., at least two modification positions, or the number n may be selected differently at different modification positions, e.g., at least two modification positions. Furthermore, it is provided that the n of modification positions can be changed during the method, i.e., the modification positions can be expanded or contracted during the method, i.e., a large modification position can be divided into several smaller modification positions, and vice versa. This can be advantageous, for example, in regions of the nucleotide sequence encoding amino acid n-pairs that are rarely or completely present in the genome or part thereof of a target organism or group of target organisms or viruses, in order to rely on a smaller n with more reliable statistics. For example, if a related amino acid triad occurs rarely, or in extreme cases, not at all, a section of a nucleotide sequence with three codons may be more advantageously optimized with two overlapping modification positions (n=2) than with a single modification position (n=3). A further advantage is that by using a larger n, the methods of the present invention have a high probability of eliminating nucleotide runs that are deleterious to expression in the target organism, such as restriction sites. Because such sequence motifs generally have no basis in the genome of the target organism itself, i.e., because the relative frequency of the codon n-tuples in the genome of the target organism or target organisms is close to zero, the methods of the present invention implicitly lead to their systematic elimination. Such deleterious sequences are automatically excluded from the optimized nucleotide sequence, for example, up to a length of 3n-2 if n is 3 or greater.
[0052] In a preferred embodiment, the replacement of a base triplet with a synonymous base triplet is carried out in multiple iterative steps using a computer-based optimization method. In this way, it is possible to continuously optimize a nucleotide sequence, particularly one with a large number of overlapping modification positions. In the iterative steps, at least one of the n base triplets can be replaced with a synonymous base triplet at all modification positions or only at some modification positions. Preferably, in each iterative step, at least one of the n base triplets is replaced with a synonymous base triplet only at some modification positions. Preferably, in at least some of the iterative steps, particularly preferably in most or all of the iterative steps, only one base triplet is replaced with a synonymous base triplet. The base triplet can be contained in one modification position or in multiple overlapping modification positions. The replacement of base triplets with synonymous base triplets can be repeated, for example, until one of the above-mentioned target criteria is achieved, i.e., until the codon n-pair relative frequency of the second codon n-pair achieves the greatest possible minimum, the greatest possible average, in particular the greatest possible weighted average or maximum, or until these values approach the magnitudes defined above.
[0053] It is understood that the methods of these embodiments may include one or more iterative steps that move away from each target criterion, particularly by at least temporarily reducing the relative frequency of the second codon n-set at one or more modification positions compared to the first codon n-set, so as to overcome a minimum or local maximum of the average value that prevents the target criterion from being reached. Thus, it is also provided that at one or more modification positions, at least one of the n base triplets may be replaced multiple times with different synonymous base triplets until the target criterion is achieved. In particular, it is not provided that the modification positions are fixed to a specific second codon n-set after the iterative steps are performed.
[0054] Preferably, the computer-based optimization method includes an approximation method, particularly a simulated annealing method. These methods have proven particularly suitable and advantageous for finding an approximate solution for the optimal nucleotide sequence for a target organism in terms of the relative frequency of n codon pairs, especially for long nucleotide sequences with many overlapping modification positions, e.g., 30 or more codons, when the complexity of the sequence makes it impossible to thoroughly check all possible synonymous base triplets and mathematical optimization methods. Naturally, other heuristic approximation methods, such as the Flood algorithm or genetic algorithms, are also possible. Further suitable approximation methods are known to those skilled in the art. Other computer-based optimization techniques, such as applications based on artificial intelligence (AI), are also contemplated.
[0055] In a preferred embodiment of the method, the modified positions comprise a total of at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% of the base triplets of the nucleotide sequence encoding the amino acids of the predetermined amino acid sequence. Preferably, the modified positions comprise a total of at least 60%, at least 70%, or at least 80%, particularly preferably at least 90% or at least 95% of the base triplets of the nucleotide sequence encoding the amino acids of the predetermined amino acid sequence. In this way, the method of the present invention ensures particularly reliable optimization of the nucleotide sequence for expression in the target organism. As mentioned above, the modified positions can also comprise base triplets that are not replaced with synonymous base triplets, i.e., each base triplet contained in the modified positions does not need to be replaced with a synonymous base triplet.
[0056] Preferably, the synonymous base triplets are selected so that the second codon n-set has a higher codon n-set relative frequency than the corresponding first codon n-set in at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, or at least 35%, preferably at least 40% or at least 45%, preferably at least 50%, at least 55%, at least 60%, at least 65%, or at least 70% of the modified positions, particularly preferably at least 75%, at least 80%, at least 85%, or at least 90% of the modified positions.
[0057] The term "at least one target organism" as used herein includes the possibility that a nucleotide sequence can be simultaneously optimized for expression of a given amino acid sequence in multiple different target organisms. Thus, the method of the present invention provides, for example, that the codon n-set relative frequency is obtained from the number of codon n-set events relative to the number of corresponding amino acid n-set events in the genomes or portions thereof of multiple different target organisms. In this way, the method of the present invention can be used to provide optimized nucleotide sequences that are more suitable for heterologous expression in different expression systems. This is not possible with conventional methods based on, for example, codon usage or codon adaptation index.
[0058] In principle, at least one target organism can be a specific host cell or any specific organism suitable for expressing a specific amino acid sequence.The host cell can be a prokaryotic or eukaryotic host cell.The host cell can be a host cell suitable for culturing in liquid or solid medium.Alternatively, the host cell can be a cell that is part of a multicellular organism, such as a multicellular tissue or a plant, particularly a transgenic plant, animal, or human.
[0059] Host cells can be microbial or non-microbial. Microbial host cells can be bacterial, yeast, or fungal cells. Suitable bacterial host cells include both Gram-positive and Gram-negative bacteria. Examples of suitable bacterial host cells include bacteria from the genera Bacillus, Actinomyces, Escherichia coli, and Streptomyces, as well as lactic acid bacteria such as Lactobacillus, Streptococcus, Lactococcus, Oenococcus, Leuconostoc, Pediococcus, Carnobacterium, Propionibacterium, Enterococcus, and Bifidobacterium. Particularly preferred are Bacillus subtilis, Bacillus amyloliquefaciens, Bacillus licheniformis, Escherichia coli, Streptomyces coelicolor, Streptomyces clavuligerus, Lactobacillus plantarum, and Lactococcus lactis, especially Escherichia coli.
[0060] Alternatively, the host cell can be a eukaryotic microorganism such as yeast or fungi, particularly filamentous fungi. Preferred yeast host cells belong to the genera Saccharomyces, Kluyveromyces, Candida, Pichia, Schizosaccharomyces, Hansenula, Kloeckera, Schwanniomyces, and Yarrowia. Particularly preferred Debaryomyces host cells are Saccharomyces cerevisiae and Kluyveromyces lactis.
[0061] In a further preferred embodiment, the host cell of the present invention is a filamentous fungal cell. Filamentous fungi include all filamentous fungi of the phylum Eumycota and Oomycota. Filamentous fungi are characterized by a mycelial wall composed of chitin, cellulose, glucan, chitosan, mannan, and other complex polysaccharides. Vegetative growth occurs by hyphal elongation, and carbon decomposition is necessarily aerobic. Filamentous fungal strains that can be used as host cells in the present invention include, but are not limited to, strains of the genera Acremonium, Aspergillus, Aureobasidium, Cryptococcus, Filibasidium, Fusarium humicola, Magnaporthe, Mucor, Myceliophthora, Neocallimastyx, Neurospora, Paecilomyces, Penicillium, Piromyces, Schizophyllum, Chrysosporium, Talaromyces, Thermoascus, Thielavia, Tolypocladium, and Trichoderma. Preferred species of filamentous fungi are selected from the group consisting of Aspergillus niger, Aspergillus oryzae, Aspergillus sojae, Trichoderma reesei, and Penicillium chrysogenum. Examples of suitable host strains are known to those skilled in the art.
[0062] Suitable non-microbial host cells include, for example, mammalian host cells such as hamster cells (e.g., Chinese hamster ovary (CHO) cells; baby hamster kidney (BHK) cells), mouse cells, monkey cells, or human cells or cell lines such as HeLa or HEK293; insect cells such as Drosophila cells or lepidopteran cell lines Hi5, Sf21; and plant cells such as cells derived from tobacco, tomato, potato, rapeseed, cabbage, pea, wheat, corn, rice, Taxus species such as Pacific yew, Arabidopsis species such as Arabidopsis thaliana, and Nicotiana species such as tobacco (Nicotiana tabacum). Furthermore, non-pathogenic Leishmania species are also suitable for protein expression. Such non-microbial cells are particularly suitable for the production of mammalian or human proteins for use in mammalian or human therapy.
[0063] The predetermined amino acid sequence is preferably a protein or part thereof, in particular a naturally occurring eukaryotic protein. The method of the invention has proven particularly advantageous for the expression of eukaryotic proteins, such as insect, plant or mammalian proteins, in bacterial, in particular prokaryotic, expression systems, such as E. coli.
[0064] In embodiments of the method in which nucleotide sequences are simultaneously optimized for expression of a given amino acid sequence in multiple different target organisms, the different target organisms may have significantly different genome sizes, with larger genomes generally encoding proteomes containing more amino acid n-tuples than smaller genomes. This can result in the codon n-tuple relative frequency being disproportionately influenced by the larger genomes, and thus nucleotide sequence optimization will necessarily favor expression in target organisms with larger genomes over those with smaller genomes. In such cases, the methods of the present invention provide that the codon n-tuple relative frequency includes genome-dependent weighting, e.g., configured to at least partially compensate for differences in the size of the genomes or portions thereof, particularly the different sizes of the coding regions, of the different target organisms. In this way, nucleotide sequences optimized according to the present invention are guaranteed to be more suitable for expression in different target organisms.
[0065] Surprisingly, it has been found that the method of the present invention can significantly improve not only the quantitative protein yield in heterologous expression systems compared to conventional optimization methods, but also significantly increase the proportion of soluble protein in the yield. This is a unique advantage of the method of the present invention, since the soluble form of a protein generally constitutes its native, biochemically active state, which is of great importance, especially for high-value proteins for scientific, pharmaceutical, and biotechnological purposes. Thus, a unique feature of the method of the present invention is that, after optimization of the nucleotide sequence, the solubility of the amino acid sequence expressed in at least one target organism is higher and / or a greater proportion of it is present in soluble form than before optimization.
[0066] As can already be seen from the above, the optimization of the nucleotide sequence for expression in at least one target organism can alternatively or additionally be based on the genome or part thereof of a virus capable of infecting at least one target organism, said virus in particular may comprise a bacteriophage.
[0067] For this purpose, it is essential to perform optimization based on multiple or numerous viral genomes or portions thereof, because a single viral genome or transcriptome typically does not have the size and therefore the necessary statistical significance required for effective optimization of nucleotide sequences based on the codon n-set relative frequency according to the present invention. Therefore, the inventors combine the genomes or transcriptomes of different viruses to form a kind of "supergenome" or "supertranscriptome" that serves as the basis for determining the codon n-set relative frequency. Preferably, the methods of these embodiments include genomes or corresponding portions of at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 150, or at least 200 different viral genomes.
[0068] Here, the inventors take advantage of the fact that viral or phage genomes are already naturally optimized for high-throughput protein expression in infected target organisms. In this regard, the inventors recognized a particular advantage: viral or phage genomes often have low base complexity, and as a result, mRNA transcribed from nucleotide sequences optimized according to the present invention forms little or no secondary structure. As a result, the optimization method of the present invention is clearly superior in practical terms to other methods that use artificial algorithms to optimize the secondary structure of mRNA. It has been found that the expression rate of amino acid sequences that can be expressed with very good solubility in heterologous systems without optimization can also be significantly improved by optimization based on viral genomes according to the present invention.
[0069] Finally, the method can include reducing or eliminating randomly generated nucleotide sequences and / or motifs that may adversely affect expression in a target organism by substituting base triplets present within the nucleotide sequence and / or at modified positions. In some embodiments of the method, such unfavorable nucleotide sequences and / or motifs are at least partially removed from the nucleotide sequence. Non-limiting examples of unfavorable nucleotide sequences and / or motifs include cis-acting mRNA destabilizing motifs, RNase splice sites, ribosome binding sites, repetitive elements, and restriction enzyme recognition sequences. Furthermore, for example, the GC content and mRNA secondary structure of the transcribed nucleotide sequence can be considered to further improve the nucleotide sequence for expression in a target organism. However, as described above, a distinctive advantage is that nucleotide sequences that are deleterious to expression in a target organism are already essentially largely eliminated by the method of the present invention, particularly by codon n-tuples where n is 3 or greater. In this way, in certain embodiments, additional steps such as optimization of mRNA secondary structure or GC content, removal of mRNA destabilizing motifs, ribosome binding sites, repetitive elements and / or restriction enzyme recognition sequences can be omitted from the method.
[0070] Another subject of the present invention is the use of the nucleotide sequences optimized according to the method described above for the production of synthetic DNA and / or for protein expression in target organisms.
[0071] Therefore, another subject of the present invention is a nucleic acid molecule, particularly an isolated nucleic acid molecule, comprising an optimized nucleotide sequence obtained by one of the methods described herein.Preferably, the nucleic acid is DNA.In a further embodiment, a vector is provided that comprises a nucleic acid molecule, particularly an isolated nucleic acid molecule.The nucleic acid molecule optimized according to the present invention can be clearly distinguished from the sequence optimized by conventional methods by sequence comparison.In this regard, please also refer to the following comparative example.
[0072] A further subject of the present invention is a recombinant host cell which contains an above-mentioned, in particular isolated, nucleic acid molecule or an above-mentioned vector.
[0073] Thus, the present invention also includes a method for expressing a protein, particularly a recombinant protein, in a target organism, comprising providing a nucleotide sequence encoding the protein and optimized according to the above-described method. Furthermore, the method may comprise one or more of the following steps: synthesizing a nucleic acid molecule comprising the optimized nucleic acid sequence; introducing the nucleic acid molecule into the target organism; and culturing the target organism under conditions that allow expression of the protein from the optimized nucleic acid sequence. Preferably, expression is at least partially carried out at temperatures below 30°C, below 25°C, or below 20°C. It has been found that nucleotide sequences optimized according to the present invention are significantly more advantageous for heterologous protein expression at relatively low temperatures than nucleotide sequences optimized by conventional methods. In this regard, see also the following example embodiments. Thus, the method of the present invention is particularly suitable for expressing sensitive, high-value proteins, and at the same time, the potential for energy savings leads to improved sustainability.
[0074] A further subject of the invention is a computer program comprising program code means adapted to carry out the method according to the above description when the computer program is run on a computer. The computer program may comprise an interface for connecting to a DNA and / or RNA synthesizing machine.
[0075] Another subject of the invention is a computer-readable storage medium on which the above-mentioned computer program is stored in computer-readable form.
[0076] Finally, a further subject of the present invention is an apparatus for optimizing and / or generating nucleotide sequences for the expression of a given amino acid sequence in at least one target organism, said apparatus comprising a computing device configured to carry out one of the methods described above. In particular, said apparatus may be a DNA and / or RNA synthesis device, also referred to as a "DNA / RNA synthesizer".
[0077] Furthermore, it is understood that preferred and advantageous embodiments of the method of the present invention may also relate to other subject matters of the present invention, where applicable. Thus, features disclosed above and below in relation to the method of the present invention may also relate to other subject matters of the present invention, and vice versa. In case of doubt, the use of the word "respectively" denotes an "and / or" relationship.
[0078] [Brief explanation of the sequence] SEQ ID NO: 1: Nucleotide sequence of I. sakaiensis PETase (wild-type sequence) encoding amino acids (AS) 28-290; SEQ ID NO: 2: Nucleotide sequence of synthetically produced I. sakaiensis PETase (encoding AS28-290) with a double Strep tag at the C-terminus after conventional optimization for expression in E. coli according to the state of the art (control); SEQ ID NO: 3: Nucleotide sequence of synthetically produced I. sakaiensis PETase with a double Strep tag at the C-terminus (encoding AS28-290) after sequence optimization for expression in E. coli with n=2 according to the present invention; SEQ ID NO: 4: Nucleotide sequence of synthetically produced I. sakaiensis PETase with a double Strep tag at the C-terminus (encoding AS28-290) after sequence optimization for expression in E. coli with n=3 according to the present invention; SEQ ID NO: 5: Nucleotide sequence of the OTP86-DYW domain of A. thaliana (AS826-960) (wild-type sequence); SEQ ID NO: 6: Nucleotide sequence of the synthetically produced A. thaliana OTP86-DYW domain (AS826-960) with a double Strep tag and a Tobacco Etch Virus (TEV) cleavage site at the N-terminus after conventional optimization for expression in E. coli according to the state of the art (control); SEQ ID NO: 7: Nucleotide sequence of the synthetically produced A. thaliana OTP86-DYW domain with a double Strep tag and a Tobacco Etch Virus (TEV) cleavage site at the N-terminus (AS826-960) after optimization for expression in E. coli with n=3 according to the present invention; SEQ ID NO: 8: Nucleotide sequence of the synthetically generated A. thaliana OTP86-DYW domain (AS826-960) with a double Strep tag and a Tobacco Etch Virus (TEV) cleavage site at the N-terminus after optimization for expression in E. coli based on the 226 phage genome, with n=3 according to the present invention; SEQ ID NO: 9: Protein sequence of citrin (encoding AS1-239), the desired amino acid sequence for expression in H. sapiens; SEQ ID NO: 10: Nucleotide sequence of synthetically produced citrine (encoding AS1-239) with an N-terminal FLAG tag and a double Strep tag and flanked at 5' by a Kozak sequence (GCCACC) after conventional optimization for expression in H. sapiens according to the state of the art (control); SEQ ID NO: 11: Nucleotide sequence of synthetically produced citrine (encoding AS1-239) with an N-terminal FLAG tag and a double Strep tag and flanked at 5' by a Kozak sequence (GCCACC) after sequence optimization for expression in H. sapiens with n=3 according to the present invention; SEQ ID NO: 12: Nucleotide sequence of H. sapiens STING1 ER exit protein 1 ("STEEP1") encoding AS1-222 (wild type); SEQ ID NO: 13: Nucleotide sequence of synthetically produced H. sapiens STEEP1 (encoding AS1-222) with an N-terminal FLAG tag and a double Strep tag and flanked by a Kozak sequence (GCCACC) at 5', after conventional optimization for expression in H. sapiens according to the state of the art (control); SEQ ID NO: 14: Nucleotide sequence of synthetically produced H. sapiens STEEP1 with an N-terminal FLAG tag and a double Strep tag, flanked at 5' by a Kozak sequence (GCCACC), after sequence optimization for expression in H. sapiens with n=3 according to the present invention (encoding AS1-222); SEQ ID NO: 15: Nucleotide sequence of H. sapiens nitric oxide synthase-interacting protein (NOSIP) encoding AS1-304 (wild type); SEQ ID NO: 16: Synthetically generated H. sapiens NOSIP nucleotide sequence (encoding AS1-304) with an N-terminal FLAG tag and a double Strep tag and flanked by a Kozak sequence (GCCACC) at 5', after conventional optimization using the codon adaptation index and optimization of the mRNA secondary structure for expression in H. sapiens according to the state of the art (control);
[0079] SEQ ID NO: 17: Nucleotide sequence of synthetically produced H. sapiens NOSIP (encoding AS1-304) with an N-terminal FLAG tag and a double Strep tag and flanked at 5' by a Kozak sequence (GCCACC) after conventional optimization according to the state of the art and WO 2020 / 024917 for expression in H. sapiens; SEQ ID NO: 18: Nucleotide sequence of synthetically produced H. sapiens NOSIP (encoding AS1-304) with an N-terminal FLAG tag and a double Strep tag and flanked at 5' by a Kozak sequence (GCCACC) after sequence optimization for expression in H. sapiens with n=3 according to the present invention; SEQ ID NO: 19: Protein sequence of EqFP611 (AS1-231), the desired amino acid sequence for expression in S. elongatus; SEQ ID NO: 20: Nucleotide sequence of synthetically produced EqFP611 (encoding AS1-231) with a double Strep tag at the N-terminus, flanked at 5' by the restriction site NdeI and at 3' by a transcription terminator and a KpnI restriction site, after conventional optimization for expression in S. elongatus according to the state of the art (control); SEQ ID NO: 21: Nucleotide sequence of synthetically produced EqFP611 (encoding AS1-231) with a double Strep tag at the N-terminus, flanked at 5' by the restriction site NdeI and at 3' by a transcription terminator and a KpnI restriction site, after optimization for expression in S. elongatus with n=3 according to the invention.
[0080] The invention will now be described in more detail based on exemplary embodiments with reference to the accompanying drawings, in which the invention is not limited to the exemplary embodiments. [Brief explanation of the drawings]
[0081] [Figure 1] 1 is a flow chart showing a general sequence of an embodiment of the method of the present invention. [Figure 2] Figure 1 shows SDS-PAGE (A) and quantitative analysis (B) of heterologous expression of Ideonella sakaiensis PETase in E. coli at 20°C. [Figure 3] SDS-PAGE (A) and quantitative analysis (B) of heterologous expression of Ideonella sakaiensis PETase in E. coli at 30°C. [Figure 4] 1 is a graphical representation of the relative frequencies of the first and second codon triads in sections of the wild-type sequence of Arabidopsis thaliana OTP86-DYW (A) and the nucleotide sequence optimized according to the present invention (B). [Figure 5]SDS-PAGE (A) and quantitative analysis (B) of heterologous expression of Arabidopsis thaliana OTP86-DYW in E. coli at 17°C. [Figure 6] Western blot (A), quantitative analysis of Western blot (B) and fluorescence (C) of heterologous expression of citrin in HeLa cells. [Figure 7] Western blot (A), quantitative analysis of Western blot (B) and fluorescence (C) of heterologous expression of citrin in HEK293 cells. [Figure 8] Western blot (A) and quantitative analysis (B) of the expression of H. sapiens STEEP1 in HeLa cells. [Figure 9] Western blot (A) and quantitative analysis (B) of the expression of H. sapiens STEEP1 in HEK293 cells. [Figure 10] Figure 1 shows Western blot (A) and quantitative analysis (B) of the expression of H. sapiens NOSIP in HeLa cells with SEQ ID NO: 16 (control) and SEQ ID NO: 18. [Figure 11] Figure 1 shows Western blot (A) and quantitative analysis (B) of the expression of H. sapiens NOSIP in HeLa cells with SEQ ID NO: 17 (control) and SEQ ID NO: 18. [Figure 12] Figure 1 shows Western blot (A) and quantitative analysis (B) of the expression of H. sapiens NOSIP in HEK293 cells with SEQ ID NO: 16 (control) and SEQ ID NO: 18. [Example]
[0082] The invention will now be described in more detail on the basis of exemplary embodiments with reference to the accompanying drawings, in which the invention is by no means limited to the exemplary embodiments.
[0083] Comparative Example 1: Heterologous expression of Ideonella sakaiensis PET hydrolase (PETase) in Escherichia coli In a first comparative experiment, the nucleotide sequence of I. sakaiensis PETase (SEQ ID NO: 1), encoding amino acid positions 28 to 290 (molecular weight 27.9 kDa), was optimized for heterologous expression in E. coli according to the method of the present invention and the method of WO 2020 / 024917 as a control.
[0084] For expression of PETase in the target organism, E. coli, three pET28a expression plasmids were obtained from Genscript, Inc., encoding the amino acid sequence of PETase with a double Strep tag at the C-terminus under an inducible T7 promoter. The amino acid sequence contained an N-terminal Met (start codon) and the amino acids Ala and Ser. The plasmids contained the nucleotide sequence encoding PETase with a double Strep tag after optimization by the manufacturer's optimization service according to the method of WO 2020 / 024917 (SEQ ID NO: 2), optimization with n=2 according to the method of the present invention, in which the relative frequency of codon n pairs is determined based on the protein-coding portion of the E. coli genome (SEQ ID NO: 3), and optimization with n=3 according to the method of the present invention (SEQ ID NO: 3).
[0085] In this regard, FIG. 1 shows a schematic sequence of an exemplary embodiment of the method 100 of the present invention, in a computer-implemented embodiment for optimizing I. sakaiensis PETase. Field 102 represents the input of the nucleotide sequence to be optimized into the computer. Of course, it is also possible to input an amino acid sequence and then convert it into a nucleotide sequence for optimization. Here, in this example, the entire coding sequence is divided into consecutive modification positions, i.e., one codon offset is provided between adjacent modification positions. In this way, N-2 codon triplets or N-1 codon dyads are generated, relative to the total number N of amino acids in the given amino acid sequence. In this exemplary embodiment, the modification positions are defined by the input of the nucleotide sequence to be optimized.
[0086] In field 104, a list of E. coli protein-coding genes is generated using a DNA sequence database. A suitable DNA sequence database is, for example, GenBank (Nucleic Acids Research 41, 2013, D36-42). A suitable platform for linking sequence data with genetic and functional information is the Reference Sequence (RefSeq) database (The NCBI Handbook, 2nd edition, Chapter 18: The Reference Sequence (RefSeq) Database, Bethesda (MD), National Center for Biotechnology Information, USA, 2013). Based on this, the absolute frequency of each n-combinable codon pair and each n-combinable amino acid pair in protein-coding genes were determined. Here, n = 2 in one embodiment, and n = 3 in another embodiment.
[0087] Using this information, in field 106, the quotient of the absolute frequency of each combinable codon n-pair and the absolute frequency of the corresponding amino acid n-pair that it encodes was determined and incorporated into the subsequent process.
[0088] As an optional step, undesirable nucleotide sequences known to those skilled in the art that may impair expression in E. coli, such as the TATA box "TATAA" or the ribosome binding site "AGGAGG", were entered in field 108. Additional undesirable sequence motifs were AAAAAA, TTTTT, AGGAGGT, TATAAA, ATCTGTT, GGAGGT, and GGTGGT.
[0089] Next, a computer-assisted optimization procedure using a "simulated annealing" algorithm successively substituted base triplets at modified positions of the wild-type sequence in field 110 in multiple iteration steps 112 until the largest possible weighted average codon n-pair relative frequency of all codon n-pairs in the nucleotide sequence was achieved while minimizing the number of unwanted nucleotide sequences. Only the start codon was excluded from the optimization, although in principle the start codon and / or the stop codon could have been included in the optimization.
[0090] The weighting W of n codon pairs according to their relative frequency P was performed as follows: W = -1 / P when P was greater than 0.0001, and W = -10,000 when P was less than 0.0001. The weighted average F W To obtain W, we summed the Ws for all n codon pairs and then divided by that number, L-(n-1), where L is the number of codons included in the nucleotide sequence optimization (here, L=864). The weighting W=-1 / P is subtractive and can reach very high values for very small P, so we limited it below -10,000.
[0091] In this example, the weighted average was further corrected for occurrence of unwanted sequence motifs. For each unwanted sequence motif, F is calculated, which corresponds to the number of unwanted sequence motifs in the nucleotide sequence to be optimized multiplied by -1. E Furthermore, the value F E can be weighted by a distinct value r, and in this example we chose r=0.035 for the n=2 embodiment and r=0.058 for the n=3 embodiment.
[0092] To optimize the nucleotide sequence, a "simulated annealing" algorithm is used to calculate the function F = F W +r·F E Find the maximum value of the weighted average value using F W was set to the largest possible value, while at the same time minimizing the number of undesired sequences contained in the nucleotide sequence at the end of optimization.
[0093] After optimization, the weighted average value for the embodiment of the invention where n=2 (SEQ ID NO: 3) is F W As a result of a similar comparison calculation, the weighted average of the relative frequencies of the two codon pairs of the wild type (SEQ ID NO: 1) was F W =-13.2, and in the control sequence (SEQ ID NO: 2) according to WO 2020 / 024917, F W =-7.3. The weighted average value after optimization according to the present invention in the embodiment of n=3 (SEQ ID NO: 4) was F W =-10.4. As a result of a similar comparison, F W =-526.8, the control sequence (SEQ ID NO: 2) according to WO 2020 / 024917 W The weighted average relative frequency of the three codon pairs was obtained as =-336.4.
[0094] After the target criteria were achieved, the optimized nucleotide sequences SEQ ID NO:3 and SEQ ID NO:4 were output in field 114 and then synthesized appropriately.
[0095] It is immediately apparent from the sequence protocol that the nucleotide sequences optimized according to the present invention are significantly different from the wild-type sequences even at the nucleotide level. Thus, the wild-type sequence (SEQ ID NO: 1) has only 82.3% nucleotide identity with the sequence optimized according to the method of the present invention (SEQ ID NO: 3) in the n=2 embodiment, and only 82.9% nucleotide identity with the sequence optimized according to the method of the present invention in the n=3 embodiment. Furthermore, there are very clear differences between the nucleotide sequences optimized according to the method of the present invention and those optimized according to the state of the art. In contrast, the sequence optimized according to WO 2020 / 024917 (SEQ ID NO: 2) has only 89.2% nucleotide identity with the sequence optimized according to the method of the present invention (SEQ ID NO: 3) in the n=2 embodiment, and only 84.7% nucleotide identity with the sequence optimized according to the method of the present invention (SEQ ID NO: 4) in the n=3 embodiment. These significant differences are surprising to those skilled in the art given the amino acid sequence, and demonstrate that the optimization method of the present invention produces results that are completely different from those of the state of the art.
[0096] The expression system used was the Escherichia coli strain BL21 (New England Biolabs GmbH, Frankfurt am Main, Germany). Three different aliquots of the expression host were transformed with one of the expression plasmids by electroporation using 0.1 μg of DNA each and streaked onto LB agar plates supplemented with 100 mg / mL kanamycin. The plates were grown overnight at 37°C. Precultures were prepared from several colonies from each plate in 50 mL of LB medium containing 100 mg / mL kanamycin and grown overnight at 37°C in a shaker at 180 rpm.
[0097] Expression cultures were prepared from 1 mL of preculture and 99 mL of TB medium, respectively. Expression of recombinant PETase in the cultures was induced at an OD of 0.6 by adding IPTG to a final concentration of 1 mM. Expression was carried out at 20°C for 14 hours in one embodiment and at 30°C for 5 hours in another embodiment.
[0098] Prior to cell harvest, the OD600 of the expression cultures was measured, and an equal amount of cells was harvested from each culture to normalize protein yield based on cell mass. The cell pellet was resuspended in 10 mL of Buffer A (20 mM Tris-Cl, pH 7.5, 150 mM NaCl, 1 mM DTT) and disrupted using sonication. The cell lysate was then centrifuged at 20,000 g for 1 hour to separate the insoluble cellular components into a pellet.
[0099] The supernatant containing the soluble fraction was mixed with 200 μL of Streptactin beads equilibrated with Buffer A (IBA Lifesciences, Göttingen, Germany) in an Eppendorf tube. The beads were washed twice with 1 mL of Buffer A in the Eppendorf tube and centrifuged to remove the supernatant. Bound proteins were eluted with 200 μL of Buffer A containing 10 mM desthiobiotin.
[0100] Protein identity was analytically verified by SDS-polyacrylamide gel electrophoresis (SDS-PAGE). Protein amounts in SDS-PAGE gel bands were quantified using ImageJ software (National Institutes of Health, USA). Furthermore, protein concentrations in each supernatant were measured using a Bradford assay (Thermo Fisher Scientific, Bremen, Germany) according to the manufacturer's instructions.
[0101] The results are shown in Figures 2 and 3. Figure 2A shows an image of an SDS-PAGE analysis of expression at 20°C. Lane 1 contains size markers, lane 2 contains the expression product of a plasmid containing an inserted SEQ ID NO:2 (control), lane 3 contains the expression product of a plasmid containing an inserted SEQ ID NO:3, and lane 4 contains the expression product of a plasmid containing an inserted SEQ ID NO:4. PETase was successfully expressed from all three plasmids, as indicated by a band at 30.4 kDa, corresponding to the molecular weight of the protein containing the double Strep tag. However, the band intensity already clearly shows that the plasmid containing the nucleotide sequence SEQ ID NO:3, optimized according to the present invention (n=2 selected for optimization), produced a significantly increased protein yield compared to the plasmid containing the nucleotide sequence SEQ ID NO:2, optimized by conventional methods. The plasmid containing the nucleotide sequence SEQ ID NO:4, optimized according to the present invention (n=3 selected for optimization), produced the highest overall yield of recombinant PETase.
[0102] The bar graph in Figure 2B shows the relative protein yield of soluble PETase for each nucleotide sequence used, with the quantitatively measured protein amounts normalized to the protein amount of a control experiment with SEQ ID NO: 2. The shaded bars show the results of quantification using SDS-PAGE, and the open bars show the results of quantification using a Bradford assay. Quantification shows that optimization of the PETase nucleotide sequence according to the present invention (n=2) doubled the soluble protein yield compared to control optimization, and optimization according to the present invention (n=3) increased protein yield by approximately six-fold.
[0103] Figure 3A shows an image of SDS-PAGE analysis of expression at 30°C. Lane 1 contains size markers, lane 2 contains the expression product of a plasmid containing an inserted SEQ ID NO:2 (control), and lane 3 contains the expression product of a plasmid containing an inserted SEQ ID NO:4. PETase was again successfully expressed from both plasmids, as indicated by the 30.4 kDa band. Based on the band intensity, it can be seen that the plasmid containing the nucleotide sequence optimized according to the present invention, SEQ ID NO:4, resulted in a higher protein yield than the plasmid containing the control sequence, SEQ ID NO:2.
[0104] Figure 3B shows the relative protein yield of soluble PETase for each nucleotide sequence used; the quantitatively measured protein amount was normalized to the protein amount in a control experiment using SEQ ID NO: 2. The shaded bars show the results of quantification using SDS-PAGE, and the open bars show the results of quantification using a Bradford assay. Quantitative analysis demonstrated that the use of SEQ ID NO: 4, a nucleotide sequence optimized according to the present invention, more than doubled the expression of PETase compared to the sequence optimized according to WO 2020 / 024917.
[0105] Those skilled in the art were surprised that the method for optimizing the nucleotide sequence encoding a protein of the present invention resulted in such a significant improvement in the heterologous expression rate of the protein compared to the current state of the art. What is surprising here is that even with the lowest parameter n=2, the expression rate increased by almost 200% compared to the state of the art, thus strongly demonstrating that the optimization method of the present invention is superior to established concepts.
[0106] Comparative Example 2: Heterologous expression of Arabidopsis thaliana OTP86-DYW in E. coli In a further comparative experiment, the nucleotide sequence of the OTP86-DYW domain from amino acid positions 826 to 960 of A. thaliana (SEQ ID NO: 5) was optimized for heterologous expression in E. coli according to the method of the present invention and the method of WO 2020 / 024917 as a control. The OTP86-DYW domain is a sensitive plant protein known to be difficult to express in heterologous systems.
[0107] To express the OTP86-DYW domain in the target organism, E. coli, three pET41 expression plasmids containing the amino acid sequence 826-960 of the OTP86-DYW domain were cloned under an inducible T7 promoter, along with a double Strep tag and a TEV protease cleavage site. A Met (start codon) and a Gly were also added to the N-terminus of the amino acid sequence. The inserts were obtained from Genscript. The plasmids contained the OTP86-DYW-encoding nucleotide sequence (SEQ ID NO: 6) after optimization according to the method of WO 2020 / 024917 by the manufacturer's optimization service, the OTP86-DYW-encoding nucleotide sequence (SEQ ID NO: 7) after optimization according to the method of the present invention with the parameter n=3 (the coding portion of the E. coli genome is taken as the basis for determining the relative frequencies of codon triads according to the present invention), and the OTP86-DYW-encoding nucleotide sequence (SEQ ID NO: 8) after optimization according to the method of the present invention with n=3 (the coding portion of the genome of the following viruses or phages capable of infecting E. coli is taken as the basis for determining the relative frequencies of codon triads according to the present invention):
[0108] Enterobacteriaceae phage 13a, Enterobacteriaceae phage 285P, Enterobacteriaceae phage 933W, Enterobacteriaceae phage 9g, Enterobacteriaceae phage BA14, Enterobacteriaceae phage BP-4795, Enterobacteriaceae phage Bp7, Enterobacteriaceae phage EcoDS1, Enterobacteriaceae phage G4, Enterobacteriaceae phage GA, Enterobacteriaceae phage GEC-3S, Enterobacteriaceae phage HK106, Enterobacteriaceae phage HK140, Enterobacteriaceae phage HK225, Enterobacteriaceae phage ID2 Moscow / ID / 2001, Enterobacteriaceae phage IME08, Enterobacteriaceae phage IME10, Enterobacteriaceae phage If1, Enterobacteriaceae phage Ike, Enterobacteriaceae phage J8-65, Enterobacteriaceae phage JS10, Enterobacteriaceae phage JenK1, Enterobacteriaceae phage JenP1, Enterobacteriaceae phage JenP2, Enterobacteriaceae phage K1F, Enterobacteriaceae phage M, Enterobacteriaceae phage MS2, Enterobacteriaceae phage MX1, Enterobacteriaceae phage P4, Enterobacteriaceae phage P88, Enterobacteriaceae phage PRD1, Enterobacteriaceae phage Phi1, Enterobacteriaceae phage RB27, Enterobacteriaceae phage RB49, Enterobacteriaceae phage RB51, Enterobacteriaceae phage RB68, Enterobacteriaceae Enterobacteriaceae phage RB69, Enterobacteriaceae phage SP, Enterobacteriaceae phage ST104, Enterobacteriaceae phage Sf101, Enterobacteriaceae phage SfI, Enterobacteriaceae phage SfV, Enterobacteriaceae phage St-1, Enterobacteriaceae phage T3, Enterobacteriaceae phage T7, Enterobacteriaceae phage UAB_Phi20, Enterobacteriaceae phage UAB_Phi78, Enterobacteriaceae phage VT2-Sakai, Enterobacteriaceae phage VT2phi_272, Enterobacteriaceae phage WA13, Enterobacteriaceae phage YYZ-2008, Enterobacteriaceae phage alpha3, Enterobacteriaceae phage cdtI, Enterobacteriaceae phage fd, Enterobacteriaceae phage fiAA91-ss, Enterobacteriaceae phage mEp043 c-1, Enterobacteriaceae phage mEp235, Enterobacteriaceae phage mEp237, Enterobacteriaceae phage mEp460, Enterobacteriaceae phage phi80, Enterobacteriaceae phage phi92, Enterobacteriaceae phage phiEcoM-GJ1, Enterobacteriaceae phage phiP27, Enterobacteriaceae phage vB_EcoM_VR5, Enterobacteriaceae phage vB_EcoP_ACG-C91, Enterobacteriaceae phage vB_EcoS_NBD2, Enterobacteriaceae phage vB_EcoS_Rogue1, Enterobacteriaceae phage vB_KleM-RaK2, Escherichia phage 121Q,Escherichia phage 172-1, Escherichia phage 4MG, Escherichia phage 64795_ec1, Escherichia phage ADB-2, Escherichia phage APCEc01, Escherichia phage AR1, Escherichia phage Av-05, Escherichia phage Bp4, Escherichia phage CAjan, Escherichia phage CICC 80001, Escherichia phage D108, Escherichia phage EB49, Escherichia phage EC6, Escherichia phage ECBP1, Escherichia phage ECBP2, Escherichia phage ECBP5, Escherichia phage ECML-117, Escherichia phage ECML-134, Escherichia phage ECML-4, Escherichia phage EK99P-1, Escherichia phage Envy, Escherichia phage FFH2, Escherichia phage FV3, Escherichia phage Gluttony, Escherichia phage HK446, Escherichia phage HK542, Escherichia phage HK544, Escherichia phage HK578, Escherichia phage HK629, Escherichia phage HK630, Escherichia phage HK633, Escherichia phage HK639 ...446, Escherichia phage HK542, Escherichia phage HK544, Escherichia Escherichia phage HK75, Escherichia phage HX01, Escherichia phage HY01, Escherichia phage HY02, Escherichia phage HY03, Escherichia phage IME11, Escherichia phage JES2013, Escherichia phage JH2, Escherichia phage JS98, Escherichia phage JSE, Escherichia phage K1-dep(1), Escherichia phage K 1-dep(4), Escherichia phage KBNP1711, Escherichia phage LM33_P1, Escherichia phage Lw1, Escherichia phage MX01, Escherichia phage Min27, Escherichia phage NJ01, Escherichia phage P13374, Escherichia phage P483, Escherichia phage P694, Escherichia phage PA2, Escherichia phage PBECO 4, Escherichia phage PE3-1, Escherichia phage PhaxI, Escherichia phage Pollock, Escherichia phage QL01, Escherichia phage RB3, Escherichia phage SUSP1, Escherichia phage SUSP2,Escherichia phage Seurat, Escherichia phage Stx2 II, Escherichia phage TL-2011b, Escherichia phage TL-2011c, Escherichia phage UFV-AREG1, Escherichia phage V5, Escherichia phage WG01, Escherichia phage YD-2008.s, Escherichia phage e4 / 1c, Escherichia phage ime09, Escherichia phage mEp234, Escherichia phage mEpX1, Escherichia phage mEpX2, Escherichia phage phAPEC8, Escherichia phage phi191, Escherichia phage phiK, Escherichia phage β- ... Escherichia phage phiKT, Escherichia phage phiV10, Escherichia phage pro147, Escherichia phage pro483, Escherichia phage slur01, Escherichia phage slur02, Escherichia phage slur05, Escherichia phage slur14, Escherichia phage slur16, Escherichia phage vB_EcoM-UFV13, Escherichia phage vB_EcoM-VpaE1, Escherichia phage vB_EcoM-ep3, Escherichia phage vB_EcoM_11 2, Escherichia phage vB_EcoM_ACG-C40, Escherichia phage vB_EcoM_AYO145A, Escherichia phage vB_EcoM_Alf5, Escherichia phage vB_EcoM_ECO1230-10, Escherichia phage vB_EcoM_JS09, Escherichia phage vB_EcoM_PhAPEC2, Escherichia phage vB_EcoM_VR20, Escherichia phage vB_EcoM_VR25, Escherichia phage vB_EcoM_VR26, Escherichia phage vB_EcoM _VR7, Escherichia phage vB_EcoP_24B, Escherichia phage vB_EcoP_G7C, Escherichia phage vB_EcoP_GA2A, Escherichia phage vB_EcoP_PhAPEC5, Escherichia phage vB_EcoP_PhAPEC7, Escherichia phage vB_EcoP_SU10, Escherichia phage vB_EcoS_AHP42, Escherichia phage vB_EcoS_AHS24, Escherichia phage vB_EcoS_AKS96, Escherichia phage vB_EcoS_FFH1,Escherichia phage vB_Eco_ACG-M12, Escherichia phage wV7, Escherichia phage wV8, Escherichia virus 186, Escherichia virus AKFV33, Escherichia virus CBA120, Escherichia virus DT57C, Escherichia virus EPS7, Escherichia virus HK022, Escherichia virus HK97, Escherichia virus I22, Escherichia virus JL1, Escherichia virus K1-5, Escherichia virus K1E, Escherichia virus Lambda, Escherichia Escherichiavirus M13, Escherichiavirus Mu, Escherichiavirus N15, Escherichiavirus N4, Escherichiavirus P1, Escherichiavirus P2, Escherichiavirus RB16, Escherichiavirus RB32, Escherichiavirus Rtp, Escherichiavirus SSL2009a, Escherichiavirus T1, Escherichiavirus T4, Escherichiavirus T5, Escherichiavirus TLS, Escherichiavirus Wphi, Escherichiavirus phiEco32, and Escherichiavirus phiX174. The viral genomes were retrieved from the Refseq database on May 31, 2019.
[0109] Otherwise, the sequence optimization experiments according to the present invention were essentially carried out as described in Comparative Example 1. For optimization based on a phage genome, which has a cumulative size significantly smaller than that of the E. coli genome, the relative frequency P of each codon triplet was calculated as the midpoint of the Clopper-Pearson confidence interval at a 95% confidence level. Furthermore, the undesirable sequence motifs in this example were AAAAAA, TTTTT, AGGAGGT, TATAAA, ATCTGTT, GGAGGT, GGTGGT, CCATGG, and AAGCTT, and the number of each of these sequences was weighted by r = 0.06 for optimization based on the E. coli genome and the phage genome. In this example, the start codon and the subsequent glycine codon were not included in the optimization, so the number of codons included in the optimization within the nucleotide sequence, L, is 501. Expression was carried out at 17°C using Buffer A without DTT.
[0110] After optimization, the weighted average value for an embodiment of the invention (SEQ ID NO: 7) based on n=3 E. coli genomes is F W = -7.1. As a result of a similar comparison calculation, the weighted average of the relative frequencies of the three codon pairs of the wild type (SEQ ID NO: 5) based on the E. coli genome was F W = -2533.5, and the weighted average relative frequency of the three codon pairs in the control sequence (SEQ ID NO: 6) according to WO 2020 / 024917 is F W = 451.4. In an embodiment based on the genomes of viruses or phages capable of infecting E. coli with n = 3, the weighted average value after optimization according to the present invention was F W = -9.2 (SEQ ID NO: 8). As a result of a similar comparison calculation, the weighted average of the relative frequencies of the three codon pairs of the wild type (SEQ ID NO: 5) based on the virus or phage genome was F W =-56.8, whereas the control sequence (SEQ ID NO: 6) according to WO 2020 / 024917 W =-37.2.
[0111] The sequence listing also reveals that there are already significant differences at the nucleotide level between the wild-type nucleotide sequence and the nucleotide sequence optimized according to the method of the present invention: for example, the wild-type sequence (SEQ ID NO: 5) has only 72.7% nucleotide identity with the sequence optimized according to the method of the present invention (SEQ ID NO: 7) in an embodiment where n=3, in which the determination of the relative frequencies of codon triads according to the present invention was based on the coding part of the E. coli genome, and only 76.9% nucleotide identity with the sequence optimized according to the method of the present invention (SEQ ID NO: 8) in an embodiment where n=3, in which the coding part of the genome of a virus or phage capable of infecting E. coli was taken as the basis for the determination of the relative frequencies of codon triads according to the present invention.
[0112] Furthermore, there are also very clear differences between the nucleotide sequences optimized according to the method of the invention and sequences optimized according to the state of the art: By way of comparison, the sequence optimized according to the state of the art according to WO 2020 / 024917 (SEQ ID NO: 6) has only 85.2% nucleotide identity with the sequence optimized according to the method of the invention based on the E. coli genome (SEQ ID NO: 7), and only 74.6% nucleotide identity with the sequence optimized according to the method of the invention based on the genome of a virus or phage capable of infecting E. coli (SEQ ID NO: 8).
[0113] 4 shows the relative frequency of the first codon triad (A) in SEQ ID NO: 5, the wild-type sequence of OTP86-DYW, and the relative frequency of the second codon triad (B) in SEQ ID NO: 7, the nucleotide sequence of OTP86-DYW optimized according to the present invention, for the sequence portion corresponding to nucleotide positions 211 to 330 in the Sequence Listing. The sequence portion shown includes codons 71 through 110 of OTP86-DYW. Each of the three adjacent codons constitutes a modification position with a codon triad, thereby dividing the sequence portion shown into a total of 38 modification positions or codon triads, here n1 through n 38 The codons are designated as . There is an offset of one codon between each successive modification position. In the figure, each codon triplet (x-axis) is assigned, by a horizontal line, the relative frequency (y-axis) in percent that each codon triplet encodes the corresponding amino acid triplet in E. coli protein-coding genes. Each vertical line indicates the range of relative frequencies of all codon triplets considered at a particular modification position that encode the corresponding n amino acids at that modification position.
[0114] It is clear that the method of the present invention significantly increased the relative frequency of the codon triads at most of the modification positions shown. 25At each of the modification positions shown, except for , at least one of the base triplets was replaced with a synonymous base triplet; primarily, at the modification positions, two or three base triplets were each replaced with a synonymous base triplet to optimally increase the relative frequency of the second codon triad. It can also be seen that at only some of the modification positions, the second codon triad corresponds to the codon triad with the highest relative frequency in E. coli. Furthermore, for example, at modification position n 35 In the example, a second codon triplet with a lower relative frequency than the original first codon triplet was formed so that the relative frequency of the more important first codon triplet could be increased at other modification positions, thus achieving the highest possible weighted average of the relative frequencies of the codon triplets.
[0115] The corresponding heterologous expression results in E. coli are shown in Figure 5. Figure 5A shows an image of SDS-PAGE analysis of expression at 17°C. Lane 1 contains size markers, lane 2 contains the expression product of the plasmid with SEQ ID NO:6 (control), lane 3 contains the expression product of the plasmid with SEQ ID NO:7, and lane 4 contains the expression product of the plasmid with SEQ ID NO:8. The OTP86-DYW domain was successfully expressed in all three plasmids, as indicated by the 19.5 kDa band. Comparison of band intensities already demonstrated that the plasmids with the nucleotide sequences optimized according to the present invention, SEQ ID NO:7 and SEQ ID NO:8, produced significantly higher protein yields of the recombinant OTP86-DYW domain than the plasmid with the nucleotide sequence optimized by conventional methods, SEQ ID NO:6.
[0116] The histogram in Figure 5B shows the relative protein yield of the soluble OTP86-DYW domain for the nucleotide sequence used; the quantitatively determined protein amount was normalized to the protein amount from a control experiment using SEQ ID NO: 6. The shaded bars represent quantification by SDS-PAGE, and the open bars represent quantification by photometric UV absorbance measurements at 260 nm and 280 nm (Nanodrop). This quantification confirmed that optimization of the nucleotide sequence according to the present invention based on the E. coli transcriptome with parameter n = 3 nearly tripled the soluble protein yield compared to the control optimization, while optimization according to the present invention based on viral or phage genomes with parameter n = 3 nearly doubled the protein yield.
[0117] Such a significant improvement in the heterologous expression rate of proteins compared to the current state of the art would not have been expected by those skilled in the art with the aid of the method for optimizing the nucleotide sequences encoding the proteins of the present invention.
[0118] Comparative Example 3: Heterologous expression of citrin in H. sapiens (HeLa) cell cultures In a further comparative experiment, the nucleotide sequence of the fluorescent protein citrin (SEQ ID NO: 9) was optimized for heterologous expression in human HeLa cells, a target organism, according to the method of the present invention and, as a control, a conventional method based on codon adaptation index and local mRNA secondary structure optimization. Citrin is a variant of the green fluorescent protein (GFP) from the jellyfish Aequorea victoria, and is frequently used in reporter assays and fluorescence microscopy.
[0119] For the expression of citrine in HeLa cells, two pTwist CMV expression plasmids were obtained from Twist Bioscience (San Francisco, CA, USA), encoding the amino acid sequence of citrine with an N-terminal FLAG tag followed by a double Strep tag under the constitutive cytomegalovirus promoter. At the N-terminus, the amino acid sequence was complemented with Met (the initiation codon) and the amino acid Ala. A Kozak sequence was inserted before the initiation codon. One plasmid contained a nucleotide sequence encoding citrine with a FLAG and double Strep tag (SEQ ID NO: 10) optimized according to the manufacturer's protocol as a control. The other plasmid contained a nucleotide sequence encoding citrine with a FLAG and double Strep tag (SEQ ID NO: 11) after optimization with n=3 according to the method of the present invention, in which a protein-encoding portion of the H. sapiens genome was used as the basis for determining the relative frequencies of codon n pairs in accordance with the present invention.
[0120] Otherwise, the sequence optimization experiments according to the present invention were carried out essentially as described in Comparative Example 1. As with the optimization using E. coli as the host, for H. sapiens, undesirable sequence motifs, such as the TATA box TATAAA, known to those skilled in the art to potentially impair expression in H. sapiens, were entered as an optional step in field 108. Further undesirable sequence motifs were ATTTA, GCCACC, GCCGCC, AATAAA, and ATTAAA, with the number of each weighted by r = 0.62. The number of codons L included in the nucleotide sequence optimization in this example was 819, and the 5'-terminal Kozak sequence was not optimized.
[0121] From the sequence protocol, it is immediately clear that the nucleotide sequence optimized according to the invention (SEQ ID NO: 11) differs significantly at the nucleotide level from the sequence optimized according to the state of the art (SEQ ID NO: 10). For example, the sequence optimized according to the state of the art (SEQ ID NO: 10) has only 81.4% nucleotide identity with the sequence optimized according to the method of the invention (SEQ ID NO: 11) in the embodiment where n=3. Considering the given amino acid sequence, these large differences are surprising from the perspective of a person skilled in the art and show that the optimization method of the invention also leads to completely different results than the methods according to the current state of the art for optimization in mammalian cells (in this case, human).
[0122] After optimization, the weighted average value for n=3 of the embodiment of the invention (SEQ ID NO: 11) is F W =-9.6. As a result of a similar comparison calculation, the weighted average value of the relative frequencies of the three codon pairs in the control sequence (SEQ ID NO: 10) was F W =-41.8.
[0123] For citrine expression, HeLa cells were transferred to 6-well plates containing DMEM high-glucose medium (Biowest SAS, Nuailles, France) containing 10% FCS (Biochrom AG, Berlin, Germany) and 1% penicillin / streptomycin (Biowest) 24 h prior to transfection. Transfection was performed using 2 μg of plasmid and Rotifect (Carl Roth GmbH, Karlsruhe, Germany) according to the manufacturer's instructions. 70 h after transfection, the medium was removed, the cells were washed with 1 mL of ice-cold phosphate-buffered saline (PBS), and resuspended in RIPA lysis buffer. The lysates were mixed with 6X SDS loading buffer and separated by size on a 15% SDS polyacrylamide gel. Protein samples from the gel were then transferred to a nitrocellulose membrane by Western blotting. Nonspecific binding sites on the membrane were blocked with 2% BSA, and the membrane was incubated overnight with primary antibodies against the FLAG-tagged expressed target protein or the housekeeping gene GAPDH (loading control). The membranes were washed with TBS-Tween and incubated with horseradish peroxidase (HRP)-conjugated secondary antibodies against rabbit (FLAG) or mouse (GAPDH). Proteins were visualized using an ECL kit (Pierce, Waltham, MA, USA), and bands were quantified using ImageQuantTL (Cytiva, Marlborough, MA, USA). For comparative evaluation of protein expression, the applied protein amount was normalized to the cell mass by setting the band intensity of Citrine to the respective ratio of the band intensity of the loading control, GAPDH.
[0124] The cell lysates were then centrifuged at 13,000 g for 2 minutes, and the citrine fluorescence in the supernatant was measured in triplicate using a Tecan Spark Plate Reader at an excitation wavelength of 516 nm and an emission wavelength of 529 nm. To account for the different cell densities in the cultures, the fluorescence intensity was again set according to the intensity of each band in the loading control (GAPDH).
[0125] The results are shown in Figure 6. Figure 6A shows an image of a Western blot in which FLAG-tagged citrine was stained with an HRP-conjugated secondary antibody and the loading control GAPDH was stained with an HRP-conjugated secondary antibody. Lane 1 contains the expression product of a plasmid containing an inserted SEQ ID NO: 10 (control), and lane 2 contains the expression product of a plasmid containing an inserted SEQ ID NO: 11, which was optimized according to the present invention. Citrine was successfully expressed from both plasmids, as indicated by the bands stained with the specific HRP-conjugated secondary antibody. However, the band intensity already clearly shows that the plasmid containing the nucleotide sequence SEQ ID NO: 11, which was optimized according to the present invention, produced a significantly increased protein yield compared to the plasmid containing the nucleotide sequence SEQ ID NO: 10, which was optimized by conventional methods.
[0126] The bar graph in Figure 6B shows the relative protein yield of citrine versus the nucleotide sequence used; the quantitatively measured protein amounts were normalized to cell mass using the loading control GAPDH as an internal standard and to the protein amount from a control experiment using SEQ ID NO: 10. The shaded bars show the results of quantification based on Western blot band intensity. Quantification revealed that optimizing the citrine nucleotide sequence according to the present invention (n=3) resulted in more than three-fold higher protein yield compared to sequences optimized by conventional methods.
[0127] The bar graph in Figure 6C shows the fluorescence of citrine relative to the nucleotide sequence used. Citrine fluorescence is a measure of the percentage of dissolved functional protein. Fluorescence was normalized to cell mass as described in Figure 6B and normalized to the fluorescence from a control experiment using SEQ ID NO: 10. Fluorescence quantification revealed that optimizing the citrine nucleotide sequence according to the present invention (n=3) increased the yield of dissolved functional protein by approximately 4.7-fold. Combined with the Western blot results, it is clear that optimizing the nucleotide sequence of the fluorescent protein citrine according to the present invention yielded significantly more soluble protein (51%) than conventional optimization.
[0128] It was unexpected for those skilled in the art that the method for optimizing the protein-encoding nucleotide sequence of the present invention resulted in such a significant improvement in the heterologous expression rate of soluble proteins in mammalian cell culture compared to the current state of the art, clearly demonstrating that the optimization method of the present invention is superior to established concepts for eukaryotic expression systems.
[0129] Comparative Example 4: Heterologous expression of citrin in H. sapiens (HEK293) cell cultures In a further comparative experiment, the two optimized nucleotide sequences of citrine used in Comparative Example 3 were expressed in HEK293 cells. Otherwise, the experiment was carried out essentially as described in Comparative Example 3.
[0130] The results are shown in Figure 7. Figure 7A shows an image of a Western blot in which FLAG-tagged citrine was stained with an HRP-conjugated secondary antibody and the loading control GAPDH was stained with an HRP-conjugated secondary antibody. Lane 1 contains the expression product of a plasmid containing an inserted SEQ ID NO: 10 (control), and lane 2 contains the expression product of a plasmid containing an inserted SEQ ID NO: 11, which was optimized according to the present invention. As indicated by the bands stained with the specific HRP-conjugated secondary antibody, citrine was successfully expressed from both plasmids in HEK293 cells. The band intensities clearly show that the plasmid containing the nucleotide sequence SEQ ID NO: 11, which was optimized according to the present invention, produced a significantly increased protein yield compared to the plasmid containing the nucleotide sequence SEQ ID NO: 10, which was optimized by conventional methods.
[0131] The bar graph in Figure 7B shows the relative protein yield of citrine versus the nucleotide sequence used. The quantitatively measured protein amounts were normalized to cell mass using the loading control GAPDH as an internal standard and to the protein amount from a control experiment using SEQ ID NO: 10. The shaded bars show the results of quantification based on Western blot band intensity. Quantification showed that optimizing the citrine nucleotide sequence (n=3) according to the present invention increased protein yield by approximately 75%.
[0132] The bar graph in Figure 7C shows the citrine fluorescence relative to the nucleotide sequence used. Citrine fluorescence is a measure of the proportion of dissolved functional protein. Fluorescence was normalized to cell mass as described in Figure 7B and normalized to the fluorescence from the control experiment of SEQ ID NO: 10. The shaded bars show the results of quantification using citrine fluorescence, which reflects functional, soluble protein. Quantification showed that optimizing the citrine nucleotide sequence in accordance with the present invention (n=3) increased the yield of soluble protein by approximately 2.3-fold. Combined with the Western blot results, it is clear that optimizing the nucleotide sequence of the fluorescent protein citrine in accordance with the present invention yielded significantly more soluble protein (32%) than conventional optimization.
[0133] This significant improvement in heterologous expression rate in HEK293 cells using the protein-encoding nucleotide sequence optimization method of the present invention supports the results of Comparative Experiment 3 in another target organism and suggests that the method of the present invention is universally applicable in mammalian cell cultures. Therefore, it is clear that the optimization method of the present invention is superior to established concepts for eukaryotic expression systems.
[0134] Comparative Example 5: Expression of H. sapiens STING1 ER Exit Protein 1 (STEEP1) in HeLa cell cultures In a further comparative experiment, the nucleotide sequence of the H. sapiens protein STEEP1 was optimized for expression in HeLa cells using the method of the present invention and, as a control, using a conventional method according to Comparative Example 3. STEEP1 is a human protein present in the membrane of the endoplasmic reticulum. Mutations in STEEP1 cause several diseases. In general, proteins derived from H. sapiens are difficult to express in host systems.
[0135] For expression of STEEP1 in target organism HeLa cells, we obtained two pTwist CMV expression plasmids from Twist Biosciences, encoding the amino acid sequence of STEEP1 with an N-terminal FLAG tag followed by a double Strep tag under the control of a constitutive cytomegalovirus promoter. At the N-terminus, the amino acid sequence was complemented with Met (the initiation codon) and the amino acid Ala. A Kozak sequence was inserted before the initiation codon. One plasmid contained the nucleotide sequence encoding STEEP1 with a FLAG and double Strep tag (SEQ ID NO: 13) after conventional optimization by the manufacturer. The other plasmid contained the nucleotide sequence encoding STEEP1 with a FLAG and double Strep tag (SEQ ID NO: 14) after optimization with n=3 according to the method of the present invention, in which a portion of H. sapiens encoding a protein was used as the basis for determining the relative frequencies of codon n pairs in accordance with the present invention.
[0136] Otherwise, the sequence optimization experiment according to the present invention was essentially carried out starting from the wild-type sequence of STEEP1 from H. sapiens (SEQ ID NO: 12) as described in Comparative Example 3, taking into account the undesired sequences mentioned in Comparative Example 3, weighting the number of each by r=0.60. The number L of codons included in the optimization in this example nucleotide sequence was 762, and the Kozak sequence at the 5' end was not optimized.
[0137] The sequence listing shows that there are significant differences between the wild-type nucleotide sequence (SEQ ID NO: 12) and the nucleotide sequence optimized according to the methods of the present invention (SEQ ID NO: 14). For example, the wild-type sequence has only 81.6% nucleotide identity with the sequence optimized according to the methods of the present invention in an embodiment where n=3.
[0138] It can be seen directly from the sequence protocol that the nucleotide sequence optimized according to the invention (SEQ ID NO: 14) also differs significantly at the nucleotide level from the sequence optimized according to the state of the art (SEQ ID NO: 13). For example, the sequence optimized according to the state of the art (SEQ ID NO: 13) has only 77.3% nucleotide identity with the sequence optimized according to the method of the invention (SEQ ID NO: 14) in the embodiment n=3. Considering the given amino acid sequence, these large differences are surprising from the point of view of a person skilled in the art and show that the optimization method of the invention also leads to completely different results than the current state of the art methods for optimization in mammalian (in this case human) cells.
[0139] After optimization, the weighted average value for n=3 of the embodiment of the invention (SEQ ID NO: 14) is F W As a result of a similar comparison calculation, the weighted average of the relative frequencies of the three codon pairs of the wild type (SEQ ID NO: 12) was F W =-36.9, and the control sequence (SEQ ID NO: 13) is F W =-87.1.
[0140] The expression experiment was also performed essentially as described in Comparative Example 3. The results are shown in Figure 8. Figure 8A shows an image of a Western blot in which FLAG-tagged STEEP1 was stained with an HRP-conjugated secondary antibody and the loading control GAPDH was stained with an HRP-conjugated secondary antibody. Lane 1 contains the expression product of a plasmid with inserted SEQ ID NO: 13 (control), and lane 2 contains the expression product of a plasmid with SEQ ID NO: 14, which was optimized according to the present invention. As indicated by the bands stained with the specific HRP-conjugated secondary antibody, STEEP1 was successfully expressed with both plasmids. However, the band intensity already clearly indicates that the protein yield was significantly increased with the plasmid with SEQ ID NO: 14, the nucleotide sequence optimized according to the present invention, compared to the plasmid with SEQ ID NO: 13, the nucleotide sequence optimized by conventional methods.
[0141] The bar graph in Figure 8B shows the relative protein yield of STEEP1 for the nucleotide sequence used. The quantitatively determined protein amount was normalized to cell mass using a loading control, as described above, and related to the protein amount from the control experiment of SEQ ID NO: 13. The shaded bars show the results of quantification based on Western blot band intensity. Quantification showed that optimizing the STEEP1 nucleotide sequence by n=3 according to the present invention increased protein yield by more than 75% compared to conventional optimization.
[0142] It was unexpected for those skilled in the art that the method for optimizing the nucleotide sequence encoding the protein of the present invention resulted in such a significant improvement in the expression rate of a membrane-bound human protein compared to the current state of the art, clearly demonstrating that the optimization method of the present invention is superior to established concepts for eukaryotic expression systems.
[0143] Comparative Example 6: Expression of H. sapiens STING1 ER Exit Protein 1 (STEEP1) in HEK293 Cell Cultures In a further comparative experiment, the two optimized nucleotide sequences of STEEP1 used in Comparative Experiment 5 were expressed in HEK293 cells. In all other respects, the experiment was carried out essentially as described in Comparative Example 5.
[0144] The results are shown in Figure 9. Figure 9A shows an image of a Western blot in which FLAG-tagged STEEP1 was stained with an HRP-conjugated secondary antibody and the loading control GAPDH was stained with an HRP-conjugated secondary antibody. Lane 1 contains the expression product of a plasmid containing an inserted SEQ ID NO: 13 (control), and lane 2 contains the expression product of a plasmid containing an inserted SEQ ID NO: 14, which was optimized according to the present invention. STEEP1 was also successfully expressed in HEK293 cells with both plasmids, as indicated by the bands stained with the specific HRP-conjugated secondary antibody. The band intensity already indicates that the plasmid containing the nucleotide sequence SEQ ID NO: 14, which is optimized according to the present invention, produced a significantly increased protein yield compared to the plasmid containing the nucleotide sequence SEQ ID NO: 13, which was optimized by conventional methods.
[0145] The bar graph in Figure 9B shows the relative protein yield of STEEP1 for the nucleotide sequence used. The quantitatively determined protein amount was normalized to the cell amount using the loading control, as described above, and related to the protein amount from the control experiment of SEQ ID NO: 13. The shaded bars show the results of quantification based on the band intensity in Western blots. Quantification showed that optimizing the STEEP1 nucleotide sequence according to the present invention (n=3) increased protein yield by 22% compared to expression of the conventionally optimized sequence.
[0146] These results confirm that the protein-encoding nucleotide sequence optimization method of the present invention significantly improves the expression rate of STEEP1 in HEK293 cells and supports the universal applicability of the method of the present invention to protein expression in mammalian cells, thereby again demonstrating that the optimization method of the present invention is superior to established concepts for eukaryotic expression systems.
[0147] Comparative Example 7: Expression of H. sapiens nitric oxide synthase-interacting protein (NOSIP) in HeLa cell cultures In a further comparative experiment involving eukaryotic expression systems, the nucleotide sequence of the H. sapiens protein NOSIP was optimized for expression in HeLa cells according to the method of the present invention and the state of the art as a control. Control optimization was performed in the embodiment according to Comparative Example 3 and in the second embodiment according to WO 2020 / 024917. NOSIP regulates the activity and localization of nitric oxide synthase, controlling the production of nitric oxide, which is crucial for the development of the human brain, eyes, and face.
[0148] For expression of NOSIP in target HeLa cells, we obtained three pTwist CMV expression plasmids from Twist Bioscience, encoding the amino acid sequence of NOSIP with an N-terminal FLAG tag followed by a double Strep tag under the constitutive cytomegalovirus promoter. At the N-terminus, the amino acid sequence was complemented with Met (start codon) and the amino acid Ala. A Kozak sequence was inserted before the start codon. As a control, one of the plasmids contained the nucleotide sequence encoding NOSIP with a FLAG and double Strep tag (SEQ ID NO: 16) after optimization according to the state of the art by the manufacturer's optimization service. As an additional control, the second plasmid contained the nucleotide sequence encoding NOSIP with a FLAG and double Strep tag (SEQ ID NO: 17) after optimization according to WO 2020 / 024917. The third plasmid contained the nucleotide sequence (SEQ ID NO: 18) encoding FLAG- and double-Strep-tagged NOSIP after optimization with n=3 according to the method of the present invention, in which a protein-encoding portion of the H. sapiens genome was used as the basis for determining the relative frequencies of codon n pairs according to the present invention.
[0149] Otherwise, the sequence optimization experiment according to the present invention was essentially carried out starting from the NOSIP wild-type sequence from H. sapiens (SEQ ID NO: 15) as described in Comparative Example 3, taking into account the undesired sequences described in Comparative Example 3, weighting the number of each by r=0.64. The number of codons L included in the optimization in this example nucleotide sequence was 1008, and the 5'-terminal Kozak sequence was not optimized.
[0150] The sequence listing also reveals that in this case there are already considerable differences at the nucleotide level between the wild-type nucleotide sequence and the nucleotide sequence optimized according to the method of the invention: for example, the wild-type sequence (SEQ ID NO: 15) has only 86.0% nucleotide identity with the sequence optimized according to the method of the invention (SEQ ID NO: 18) in the embodiment where n=3.
[0151] The sequence protocol shows that the nucleotide sequence optimized according to the present invention (SEQ ID NO: 18) also differs significantly at the nucleotide level from the sequences optimized according to the state of the art (SEQ ID NOs: 16, 17). For example, the sequence optimized according to the state of the art (SEQ ID NO: 16) shares only 76.3% nucleotide identity with the sequence optimized according to the method of the present invention (SEQ ID NO: 18) in the embodiment where n=3. The sequence optimized according to the state of the art according to WO 2020 / 024917 (SEQ ID NO: 17) shares only 85.2% nucleotide identity with the sequence optimized according to the method of the present invention (SEQ ID NO: 18) in the embodiment where n=3. These significant differences once again demonstrate that the optimization method of the present invention also leads to completely different results than the current state of the art when it comes to optimization in mammalian cells (in this case Homo sapiens).
[0152] After optimization, the weighted average value for n=3 embodiments of the invention (SEQ ID NO: 18) is F W =-13.1. As a result of a similar comparison calculation, the weighted average value of the relative frequencies of the three codon pairs of the wild type (SEQ ID NO: 15) was F W =-39.1, and F in the control sequence SEQ ID NO: 16 W=-148.0, and F in the control sequence SEQ ID NO: 17 W =-36.3.
[0153] Otherwise, the expression experiment was carried out essentially as described in Comparative Example 3. The results are shown in Figures 10 and 11.
[0154] Figure 10 shows a comparison of SEQ ID NO: 16 and SEQ ID NO: 18. Figure 10A shows an image of a Western blot in which FLAG-tagged NOSIP was stained with an HRP-conjugated secondary antibody and the loading control GAPDH was stained with an HRP-conjugated secondary antibody. Lane 1 contains the expression product of a plasmid containing an inserted SEQ ID NO: 16 (control), and lane 2 contains the expression product of a plasmid containing an inserted SEQ ID NO: 18 optimized according to the present invention. As evidenced by the bands stained with the specific HRP-conjugated secondary antibody, NOSIP was successfully expressed with both plasmids, but the intensity of the bands indicates significantly better protein yields by expressing the nucleotide sequence of SEQ ID NO: 18 optimized according to the present invention compared to the plasmid with the nucleotide sequence of SEQ ID NO: 16 optimized by conventional methods.
[0155] In the bar graph shown in Figure 10B, the relative protein yield of NOSIP is shown for each nucleotide sequence used. Quantitatively measured protein amounts were first normalized to cell mass using a GAPDH loading control, as described above, and then related to the protein amount from a control experiment using SEQ ID NO: 16. Shaded bars show the results of quantification based on Western blot band intensity. Quantification demonstrated that optimizing the NOSIP nucleotide sequence (n=3) according to the present invention increased protein yield by 31-fold compared to conventional optimization.
[0156] Those skilled in the art would never have expected that the method for optimizing the nucleotide sequence encoding the protein of the present invention would result in such a significant improvement in the expression rate of a membrane-bound human protein compared to the current state of the art, thus clearly demonstrating once again that the optimization method of the present invention is superior to established concepts for eukaryotic expression systems.
[0157] Figure 11 shows a comparison of SEQ ID NO: 17 and SEQ ID NO: 18. Figure 11A shows a Western blot image of FLAG-tagged NOSIP stained with an HRP-conjugated secondary antibody and the loading control GAPDH stained with an HRP-conjugated secondary antibody. Lane 1 contains the expression product of a plasmid with SEQ ID NO: 17, the control sequence optimized in accordance with WO 2020 / 024917, while lane 2 contains the expression product of a plasmid with SEQ ID NO: 18, the sequence optimized in accordance with the present invention. As indicated by the bands stained with the specific HRP-conjugated secondary antibody, NOSIP was successfully expressed from both plasmids. The band intensities already indicate that the plasmid with SEQ ID NO: 18, the nucleotide sequence optimized in accordance with the present invention, resulted in increased protein yield compared to the plasmid with SEQ ID NO: 17, the nucleotide sequence optimized by conventional methods.
[0158] In the bar graph shown in Figure 11B, the relative protein yield of NOSIP is shown relative to the nucleotide sequence used; quantitatively measured protein amounts were first normalized to cell mass using a GAPDH loading control, as described above, and then related to the protein amount from a control experiment using SEQ ID NO: 17. The shaded bars show the results of quantification based on Western blot band intensity. Quantification showed that optimizing the NOSIP nucleotide sequence (n=3) according to the present invention increased protein yield by more than 60%.
[0159] This shows that the optimization method of the present invention is superior to the technical teachings of the current state of the art (in this case WO 2020 / 024917) also for eukaryotic expression systems.
[0160] Comparative Example 8: Expression of NOSIP in HEK293 cell culture In a further comparative experiment, the nucleotide sequences of two optimized NOSIPs, SEQ ID NO: 16 (control) and SEQ ID NO: 18 used in Comparative Example 7, were expressed in HEK293 cells. Otherwise, the experiment was carried out essentially as described in Comparative Example 7.
[0161] The results are shown in Figure 12. Figure 12A shows an image of a Western blot in which FLAG-tagged NOSIP was stained with an HRP-conjugated secondary antibody and the loading control GAPDH was stained with an HRP-conjugated secondary antibody. Lane 1 contains the expression product of a plasmid having SEQ ID NO: 16 (control) optimized by conventional methods, and lane 2 contains the expression product of a plasmid having SEQ ID NO: 18 optimized according to the present invention. As indicated by the bands stained with the specific HRP-conjugated secondary antibody, NOSIP was successfully expressed in HEK293 cells from both plasmids, and the band intensity indicates significantly better expression from SEQ ID NO: 18, the nucleotide sequence optimized according to the present invention.
[0162] In the bar graph shown in Figure 12B, the relative protein yield of NOSIP is shown for each nucleotide sequence used. Quantitatively measured protein amounts were first normalized to cell mass using a GAPDH loading control, as described above, and then related to protein amounts from a control experiment using SEQ ID NO: 16. Shaded bars show the results of quantification based on Western blot band intensity. Quantification also showed that optimizing the NOSIP nucleotide sequence (n=3) according to the present invention increased protein yield in the HEK293 expression system by approximately four-fold compared to conventional sequence optimization.
[0163] As in Comparative Experiments 4 and 6, it was confirmed that the optimization method of the nucleotide sequence encoding the protein of the present invention significantly improved the expression rate in HEK293 cells. Therefore, the method of the present invention is universally applicable to mammalian cell cultures. Thus, it was again confirmed that the optimization method of the present invention is superior to the optimization concepts established for eukaryotic expression systems.
[0164] Comparative Example 9: Heterologous expression of eqFP611 in the cyanobacterium Cyanococcus elongatus In a further comparative experiment, the nucleotide sequence of the fluorescent protein eqFP611 was optimized for heterologous expression in S. elongatus using the method of the present invention and, as a control, the method described in WO 2020 / 024917. eqFP611 is a red fluorescent protein (RFP) derived from the coral sea anemone Entacmaea quadricolor, and is frequently used in reporter assays and fluorescence microscopy.
[0165] For expression of eqFP611 in the target organism, S. elongatus, two pSyn-6 expression plasmids (Thermo Fisher, Waltham, MA, USA) encoding the amino acid sequence of eqFP611 with a double N-terminal strep tag under the constitutive psbA1 promoter were obtained from Genscript. One plasmid contained the nucleotide sequence encoding eqFP611 with a double strep tag (SEQ ID NO: 20) after optimization according to WO 2020 / 024917. The other plasmid contained the nucleotide sequence encoding eqFP611 with a double strep tag (SEQ ID NO: 21) after optimization with n=3 according to the method of the present invention, in which a protein-encoding portion of the S. elongatus genome was used as the basis for determining the relative frequencies of codon n-tuples according to the present invention.
[0166] The optimization procedure was carried out as described in Comparative Example 1. Similar to the optimization for the host E. coli, undesirable sequence motifs, such as the ribosome binding site AGGAGG, known to those skilled in the art to potentially impair expression in S. elongatus, were entered as an optional step in field 108. Further undesirable sequence motifs were GGGGGGG, CGCGCG, ATTTA, CATATG, GGTACC, and GTCGAC. The number of undesirable sequence motifs was weighted by r=0.62. The number of codons L included in the optimization in this example nucleotide sequence was 768, and terminator sequences and terminal restriction sites were not optimized.
[0167] The sequence listing shows that the nucleotide sequence (SEQ ID NO: 21) optimized for heterologous expression of eqFP611 in S. elongatus according to the present invention is significantly different at the nucleotide level from the sequence optimized according to the state of the art (SEQ ID NO: 20). For example, the sequence optimized according to the state of the art (SEQ ID NO: 20) has only 89.8% nucleotide identity with the sequence optimized according to the method of the present invention (SEQ ID NO: 21) in the embodiment where n=3. Considering the given amino acid sequence, these large differences are surprising from the perspective of a person skilled in the art and demonstrate that the optimization method of the present invention leads to completely different results for optimization in cyanobacteria (in this case, S. elongatus) than optimization methods according to the current state of the art.
[0168] After optimization, the weighted average value for n=3 of the embodiment of the invention (SEQ ID NO: 21) is F W =-8.7. As a result of a similar comparison calculation, the weighted average value of the relative frequencies of the three sets of codons in the control sequence (SEQ ID NO: 20) according to WO 2020 / 024917 was F W =-496.2.
[0169] Expression of eqFP611 in S. elongatus was performed according to the manufacturer's protocol for the "GeneArt Algal Protein Expression System" (Thermo Fisher Scientific). Further processing of the cyanobacterial biomass was performed as described for E. coli in Comparative Example 1. The expressed eqFP611 protein was purified with streptactin and quantified by SDS-PAGE. Furthermore, fractions of eqFP611 cell lysates were centrifuged at 13,000 g for 2 minutes, and the fluorescence of eqFP611 in the supernatant was measured in triplicate using a Tecan Spark plate reader at an excitation wavelength of 559 nm and an emission wavelength of 611 nm.
[0170] As a result of this experiment, as already described in Comparative Examples 1 to 8 for bacterial and eukaryotic expression systems, it is expected that the method of the present invention will also significantly increase soluble protein yields in cyanobacteria.
[0171] The present invention is not limited to the description based on the exemplary embodiments, but rather includes each and every novel feature and each and every combination of features, and in particular each and every combination of features in the claims and the specification, even if this feature or combination of features itself is not explicitly recited in the claims, the specification or the exemplary embodiments.
Claims
1. 1. A method for optimizing a nucleotide sequence for expression of a predetermined amino acid sequence in at least one target organism, comprising: the nucleotide sequence comprises a plurality of base triplets; 1. A method for optimizing the nucleotide sequence for expression in at least one target organism, wherein at least one modified position of the nucleotide sequence, a base triplet encoding an amino acid of the predetermined amino acid sequence is replaced with a synonymous base triplet encoding the same amino acid of the predetermined amino acid sequence, the at least one modification position comprises a direct succession of n base triplets constituting a first codon n-set and encoding a sequence portion of n amino acids of the predetermined amino acid sequence constituting an n-set of amino acids, the n-set of amino acids being encoded by a predetermined number of n-set of amino acid events in the genome or part thereof of the at least one target organism and / or in the genome or part thereof of a virus capable of infecting the at least one target organism; At least one of the n immediately consecutive base triplets is replaced with a synonymous base triplet, and the synonymous base triplet is selected so as to obtain a second codon n-set that has a higher codon n-set relative frequency than the first codon n-set in the genome or part thereof of the at least one target organism and / or the genome or part thereof of a virus capable of infecting the at least one target organism, in light of the number of events of the amino acid n-set; A method characterized in that n is a natural number greater than or equal to 2 and less than or equal to the total number N of amino acids in the predetermined amino acid sequence.
2. 2. The method of claim 1, wherein n is 50 or less, 40 or less, 30 or less, 20 or less, or 10 or less.
3. 3. The method of claim 1, wherein n is selected from the group consisting of n=2, n=3, n=4, n=5, and any combination thereof.
4. 4. The method according to any one of claims 1 to 3, wherein the number of amino acid n-tuple events is determined and / or obtained from a plurality of protein-coding genes and / or proteins of the at least one target organism and / or a virus capable of infecting the at least one target organism.
5. The method according to any one of claims 1 to 4, The method, characterized in that the codon n-set relative frequency is determined and / or obtained from the number of events of each of the first codon n-set and / or the second codon n-set in multiple protein-coding genes of the at least one target organism and / or a virus capable of infecting the at least one target organism, based on the number of events of the amino acid n-set.
6. The method according to any one of claims 1 to 5, below: a) determining the at least one modification position; b) replacing at least one base triplet at said at least one modification position with said synonymous base triplet; and c) determining the codon n-set relative frequency of the resulting second codon n-set.
7. The method according to any one of claims 1 to 6, the at least one modified position is a plurality of modified positions; at each of said modified positions, at least one of said n immediately successive base triplets is replaced with a synonymous base triplet; wherein the synonymous base triplets are selected such that at least some of the resulting second codon n-sets have a higher codon n-set relative frequency than the corresponding first codon n-sets.
8. 8. The method of claim 7, At least one of the n immediately successive base triplets at a modification position having a first codon n-set with the lowest codon n-set relative frequency is replaced with a synonymous base triplet selected such that the resulting second codon n-set has a higher codon n-set relative frequency than the first codon n-set.
9. 9. The method of claim 7 or 8, A method characterized in that the direct succession of n base triplets at at least two modification positions overlap, and the base triplet replaced with the synonymous base triplet is simultaneously contained in both of the at least two modification positions.
10. 10. The method of claim 9, The method is characterized in that the synonymous base triplets are selected such that at one of the at least two modification positions, the resulting second codon n-set has a lower codon n-set relative frequency than the corresponding first codon n-set, and at the other of the at least two modification positions, the resulting second codon n-set has a higher codon n-set relative frequency than the corresponding first codon n-set.
11. The method according to any one of claims 7 to 10, The method, characterized in that the codon n-set relative frequency of the first codon n-set and the codon n-set relative frequency of the second codon n-set at the modification position each have a minimum value, and the minimum value of the second codon n-set is greater than the minimum value of the first codon n-set.
12. The method according to any one of claims 7 to 11, The method, characterized in that the synonymous base triplets are selected so that the codon n-set relative frequency of the second codon n-set achieves the greatest possible minimum value or is at most 50% below the greatest possible minimum value.
13. The method according to any one of claims 7 to 12, The method, characterized in that the synonymous base triplets are selected so that the average value of the codon n-pair relative frequency of the second codon n-pair achieves a maximum value or is at most 50% below the maximum achievable value.
14. 14. The method of claim 13, The method is characterized in that the average value includes gradually decreasing the weighting of codon n-set relative frequencies, and is configured so that the impact of high codon n-set relative frequencies on the average value is disproportionate in amount compared to lower codon n-set relative frequencies.
15. The method according to any one of claims 7 to 14, The method according to claim 1, wherein the number n is different at at least two of said modification positions or n is selected to be different for at least two of said modification positions.
16. The method according to any one of claims 7 to 15, 10. A method according to claim 9, wherein the replacement of base triplets with synonymous base triplets is carried out in multiple iterative steps using a computer-assisted optimization method.
17. 17. The method of claim 16, A method, wherein the computer-aided optimization method comprises approximation methods and / or simulated annealing.
18. The method according to any one of claims 7 to 17, The method, wherein the modified positions comprise, in total, at least 50% of the base triplets of the nucleotide sequence encoding the amino acids of the predetermined amino acid sequence.
19. The method according to any one of claims 1 to 18, the at least one target organism is a plurality of different target organisms; the codon n-set relative frequencies include genome-dependent weightings configured to at least partially compensate for differences in size of the genomes or portions thereof of the different target organisms.
20. The method according to any one of claims 1 to 19, A method characterized in that after optimization of the nucleotide sequence, the solubility of the amino acid sequence expressed in the at least one target organism is higher and / or the proportion of the amino acid sequence present in soluble form is increased compared to before optimization.
21. 1. A method for optimizing a nucleotide sequence for expression of a predetermined amino acid sequence in at least one target organism, comprising: the nucleotide sequence comprises a plurality of base triplets; 1. A method for optimizing the nucleotide sequence for expression in at least one target organism, wherein at least one modified position of the nucleotide sequence, a base triplet encoding an amino acid of the predetermined amino acid sequence is replaced with a synonymous base triplet encoding the same amino acid of the predetermined amino acid sequence, the at least one modification position comprises a direct succession of n base triplets constituting a first set of n codons and encoding a sequence portion of n amino acids of the predetermined amino acid sequence constituting an n set of amino acids; at least one of the n immediately consecutive base triplets is replaced with a synonymous base triplet, and the synonymous base triplet is selected using a prediction function so as to obtain a second set of codons n that encodes the set of amino acids with a higher probability in the genome of the at least one target organism or a part thereof and / or in the genome of a virus capable of infecting the at least one target organism than the first set of codons n; A method, wherein n is a natural number greater than or equal to 2 and less than or equal to the total number N of amino acids in the predetermined amino acid sequence.
22. Use of a nucleotide sequence optimized according to the method of any one of claims 1 to 21 for the production of synthetic DNA and / or protein expression in at least one target organism.
23. 1. A computer program having program code means, comprising: Computer program, characterized in that the program code means are arranged to perform the method according to any one of claims 1 to 21 when the computer program is run on a computer.
24. A computer-readable storage medium storing the computer program according to claim 23 in a computer-readable form.
25. 1. An apparatus for optimizing and / or generating a nucleotide sequence for expression of a predetermined amino acid sequence in at least one target organism, comprising: An apparatus comprising a computing device configured to perform the method of any one of claims 1 to 21.
26. providing a nucleotide sequence optimized for the expression of a protein in at least one target organism according to the method of any one of claims 1 to 21; and expressing said protein in said at least one target organism.
27. A nucleic acid molecule comprising a nucleotide sequence obtained by the method of any one of claims 1 to 21.
28. A vector comprising the nucleic acid molecule of claim 27.
29. 29. A recombinant host cell comprising the nucleic acid molecule of claim 27 or the vector of claim 28.
Citation Information
Patent Citations
Systems and methods for increasing synthetic protein stability
JP2022531295A
Codon optimization of a synthetic gene(s) for protein expression
US20140244228A1
Codon optimization
US20190325989A1
A method for achieving improved polypeptide expression
WO2008000632A1
Codon optimization
WO2020024917A1