Codon Optimization via Sequence Space Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods lack a reliable strategy for selecting codons in synthetic genes to achieve high protein expression levels, and there is no systematic algorithm to assess the likely expression of proteins from synthetic genes, as existing codon optimization algorithms do not account for expression control elements like promoters and ribosome binding sites.

Innovation Solution

The use of computational biology and data mining techniques to map codon sequence space for polynucleotide sequences, allowing for the design and synthesis of codon variants that optimize expression characteristics by considering factors such as codon bias, mRNA secondary structures, and GC content, and iteratively refining these designs based on measured expression properties.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If codon optimization algorithms are used to select codons based on genomic sequence characteristics, then protein expression levels may be improved, but the reliability of expression prediction is insufficient because existing algorithms do not systematically account for expression control elements

Engineering Contradiction:
Improveprotein expression levelVSAvoidexpression prediction accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the gene sequence into distinct functional regions (5' UTR, coding sequence, 3' UTR) and analyzes codon usage patterns separately in each region. This segmentation allows for region-specific optimization strategies that account for different functional requirements, improving both expression levels and prediction reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent systematically varies multiple parameters including codon adaptation index (CAI), GC content, codon pair bias, and mRNA secondary structure characteristics to identify optimal combinations. By changing these parameters in a controlled manner and measuring their effects on expression, the patent establishes reliable predictive relationships between sequence features and expression outcomes.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If natural gene sequences are used as a basis for codon optimization, then codon selection can be performed, but the manufacturing precision is insufficient because natural genes contain variable elements not relevant to expression optimization

Engineering Contradiction:
Improvecodon selection capabilityVSAvoidexpression level control
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent extracts only the relevant expression-determining features from natural gene sequences, such as codon usage patterns in the coding region and regulatory elements in UTRs. By taking out and analyzing only these functional elements separately from non-expressive portions of natural genes, the patent achieves more precise control over expression levels while maintaining ease of codon selection.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent develops a universal codon optimization framework that can be applied to any gene sequence regardless of its natural origin. The methodology universally evaluates codon choices based on multiple independent parameters (CAI, GC content, secondary structure) rather than relying on gene-specific natural patterns, enabling precise expression control across diverse genetic sequences.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP2294407B1Systems and methods for determining properties that affect an expression property value of polynucleotides in an expression system
Publication Date: 2017.03.15 DNA TWOPOINTO INC
  • EP2294407B1 patent drawing
  • EP2294407B1 patent drawing
  • EP2294407B1 patent drawing

AI summary

A method of designing a polynucleotide sequence encoding a polypeptide sequence of a predetermined polypeptide is provided. A frequency lookup table corresponding to an expression system is obtained. The table comprises a plurality of sequence elements and a plurality of frequency ranges, each frequency range for a corresponding sequence element. Each frequency range is a range of frequencies with which a corresponding sequence element can occur in a polynucleotide. The polynucleotide sequence is defined using the frequency lookup table by determining, for each respective sequence element in the frequency lookup table, whether the respective sequence element encodes a portion of the polypeptide sequence. When the respective sequence element encodes a portion of the polypeptide sequence, the sequence element is incorporated into the polynucleotide at a frequency of occurrence that is within the frequency range specified for the respective sequence element in the lookup table.