Codon-Optimized Diverged DNA Sequences for Stable Protein Expression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for optimizing gene expression in host organisms face challenges when dealing with proteins containing repeated amino acid domains, as these sequences can lead to instability and expression problems due to secondary DNA structures and recombination issues, and current gene design processes struggle to accommodate large, diverged codon-biased DNA sequences for multiple repeats.
Innovation Solution
A method for designing diverged and codon-optimized nucleic acid sequences that encode amino acid repeat regions, involving the use of computer-implemented software programs like OPTGENE and CLUSTALW to align and optimize codon usage, while avoiding undesirable secondary structures and recombination sites, allowing for the creation of synthetic sequences that enhance expression levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If codon optimization is applied to large repeated DNA sequences, then translation efficiency is improved, but sequence stability deteriorates due to secondary structures and recombination
Solution Approach 1:
The patent applies different optimization strategies to different regions of the repeated sequences. Each repeat unit is codon-optimized individually to maximize translation efficiency, while maintaining sufficient sequence divergence between repeats to prevent secondary structure formation and recombination. This local differentiation allows simultaneous achievement of high translation efficiency and sequence stability.
Solution Approach 2:
The patent systematically varies codon usage parameters across repeat units to achieve optimal translation while preventing instability. By controlling the degree of codon optimization and sequence identity parameters, the method creates a balance where each repeat is optimized for translation but maintains enough divergence to avoid harmful secondary structures and recombination events.
2Reliability
If highly similar repeated sequences are used, then protein function is maintained, but recombination and secondary structure formation increase
Solution Approach 1:
The patent maintains high sequence similarity within individual repeat units to preserve protein function, while introducing controlled divergence between different repeat units. This local quality differentiation ensures that each repeat maintains the necessary sequence features for proper protein folding and function, while the inter-repeat divergence prevents harmful interactions.
Solution Approach 2:
The patent introduces asymmetric variations between repeated sequences through differential codon optimization. While the amino acid sequences remain functionally equivalent, the nucleotide sequences are made asymmetric through codon usage variations, which prevents the formation of symmetric secondary structures and reduces recombination susceptibility while maintaining protein function.
3Speed
If codon usage is maximized for frequent codons, then translation speed increases, but sequence homogeneity increases leading to instability
Solution Approach 1:
The patent applies codon optimization locally to each repeat unit rather than uniformly across all repeats. Each repeat unit uses frequent codons optimized for translation speed, but the specific codon choices vary between repeats. This local optimization approach maintains high translation speed while creating sufficient sequence heterogeneity to prevent instability.
Solution Approach 2:
The patent segments the repeated DNA sequence into individual repeat units, each independently codon-optimized. This segmentation allows each segment to achieve optimal translation speed through frequent codon usage while the collective diversity of segmented repeats prevents overall sequence homogeneity, thereby maintaining stability.
Data Source
AI summary
This disclosure concerns methods for the design of synthetic nucleic acid sequences that encode polypeptide amino acid repeat regions. This disclosure also concerns the use of such sequences to express a polypeptide of interest that comprises amino acid repeat regions, and organisms comprising such sequences.


