Codon Optimization via Sequence Space Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods lack a reliable strategy for selecting codons in synthetic genes to achieve high protein expression levels, and there is no systematic algorithm to assess the likely expression of proteins from synthetic genes, as existing codon optimization algorithms do not account for expression control elements like promoters and ribosome binding sites.
Innovation Solution
The use of computational biology and data mining techniques to map codon sequence space for polynucleotide sequences, allowing for the design and synthesis of codon variants that optimize expression characteristics by considering factors such as codon bias, mRNA secondary structures, and GC content, and iteratively refining these designs based on measured expression properties.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If codon optimization algorithms are used to select codons based on genomic sequence characteristics, then protein expression levels may be improved, but the reliability of expression prediction is insufficient because existing algorithms do not systematically account for expression control elements
Solution Approach 1:
The patent segments the gene sequence into distinct functional regions (5' UTR, coding sequence, 3' UTR) and analyzes codon usage patterns separately in each region. This segmentation allows for region-specific optimization strategies that account for different functional requirements, improving both expression levels and prediction reliability.
Solution Approach 2:
The patent systematically varies multiple parameters including codon adaptation index (CAI), GC content, codon pair bias, and mRNA secondary structure characteristics to identify optimal combinations. By changing these parameters in a controlled manner and measuring their effects on expression, the patent establishes reliable predictive relationships between sequence features and expression outcomes.
2Ease of manufacture
If natural gene sequences are used as a basis for codon optimization, then codon selection can be performed, but the manufacturing precision is insufficient because natural genes contain variable elements not relevant to expression optimization
Solution Approach 1:
The patent extracts only the relevant expression-determining features from natural gene sequences, such as codon usage patterns in the coding region and regulatory elements in UTRs. By taking out and analyzing only these functional elements separately from non-expressive portions of natural genes, the patent achieves more precise control over expression levels while maintaining ease of codon selection.
Solution Approach 2:
The patent develops a universal codon optimization framework that can be applied to any gene sequence regardless of its natural origin. The methodology universally evaluates codon choices based on multiple independent parameters (CAI, GC content, secondary structure) rather than relying on gene-specific natural patterns, enabling precise expression control across diverse genetic sequences.
Data Source
AI summary
A method of designing a polynucleotide sequence encoding a polypeptide sequence of a predetermined polypeptide is provided. A frequency lookup table corresponding to an expression system is obtained. The table comprises a plurality of sequence elements and a plurality of frequency ranges, each frequency range for a corresponding sequence element. Each frequency range is a range of frequencies with which a corresponding sequence element can occur in a polynucleotide. The polynucleotide sequence is defined using the frequency lookup table by determining, for each respective sequence element in the frequency lookup table, whether the respective sequence element encodes a portion of the polypeptide sequence. When the respective sequence element encodes a portion of the polypeptide sequence, the sequence element is incorporated into the polynucleotide at a frequency of occurrence that is within the frequency range specified for the respective sequence element in the lookup table.


