Codon Optimization Algorithm for Protein Expression Reliability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods lack a systematic strategy for selecting codons in synthetic genes to achieve high protein expression levels, as they do not account for the specific control elements like promoter strength and ribosome binding sites, leading to unreliable protein expression outcomes.

Innovation Solution

The use of computational biology and data mining techniques to map codon sequence space for polynucleotide sequences, allowing for the design and synthesis of codon variants that optimize expression characteristics by considering factors such as codon bias, mRNA secondary structures, and GC content, and iteratively refining these designs based on expression data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If codon optimization algorithms use common codon frequency analysis, then protein expression levels may be increased, but the reliability of expression prediction remains poor due to not accounting for specific control elements

Engineering Contradiction:
Improveprotein expression levelVSAvoidexpression prediction reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent transforms the codon optimization approach from using simple codon frequency parameters to using comprehensive sequence parameters including promoter strength, ribosome binding site strength, mRNA secondary structure, and codon context. This multi-parameter model (Equation 1) integrates multiple factors that actually control expression, making predictions more reliable while maintaining high expression levels.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements an iterative feedback process where initial codon optimization is performed, then expression is measured experimentally, and the results are used to refine the model parameters. This feedback loop (described in the systematic study section) allows the algorithm to learn from actual expression data and improve prediction reliability for future optimizations.

Inventive Principle:
Principle #23Feedback

2Reliability

If comprehensive sequence analysis is performed to improve expression prediction, then the complexity of the design process increases

Engineering Contradiction:
Improveexpression prediction reliabilityVSAvoiddesign process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates a universal computational platform that handles multiple functions: sequence analysis, parameter calculation, optimization, and prediction. This integrated system (referred to as the algorithm or software platform) can process different gene sequences and expression systems using the same core methodology, reducing the apparent complexity by providing a standardized approach.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces manual, trial-and-error codon optimization with an automated computational algorithm. The software performs comprehensive sequence analysis, calculates multiple parameters simultaneously, and generates optimized sequences without human intervention, transforming a complex manual process into a streamlined computational procedure.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS7561972B1Synthetic nucleic acids for expression of encoded proteins
Publication Date: 2009.07.14 DNA TWOPOINTO INC
  • US7561972B1 patent drawing
  • US7561972B1 patent drawing
  • US7561972B1 patent drawing

AI summary

A method of designing a polynucleotide sequence encoding a polypeptide sequence of a predetermined polypeptide is provided. A frequency lookup table corresponding to an expression system is obtained. The table comprises a plurality of sequence elements and a plurality of frequency ranges, each frequency range for a corresponding sequence element. Each frequency range is a range of frequencies with which a corresponding sequence element can occur in a polynucleotide. The polynucleotide sequence is defined using the frequency lookup table by determining, for each respective sequence element in the frequency lookup table, whether the respective sequence element encodes a portion of the polypeptide sequence. When the respective sequence element encodes a portion of the polypeptide sequence, the sequence element is incorporated into the polynucleotide at a frequency of occurrence that is within the frequency range specified for the respective sequence element in the lookup table. The polynucleotide sequence is then outputted.