Genomic Design Language for Scalable Nucleotide Sequence Workflows

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for high-throughput nucleotide sequence design and manufacturing face challenges such as generating unmanageable amounts of data, leading to complex memory management issues and slow processing times, particularly in large-scale genomic design and manufacturing systems.

Innovation Solution

The development of recursive data structures and a genomic design language that allows for the efficient generation and management of large sets of nucleotide sequences, enabling the specification of workflows and inputs like primers, enzymes, and environmental factors, thereby optimizing the manufacturing process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional methods are used for high-throughput nucleotide sequence design and manufacturing, then large sets of sequences can be generated, but unmanageable amounts of data are produced leading to complex memory management issues and slow processing times

Engineering Contradiction:
Improvesequence generation throughputVSAvoidmemory management complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the large set of nucleotide sequences into multiple smaller batches or groups that can be managed independently. This segmentation allows the system to process and store sequences in manageable units rather than attempting to handle all sequences simultaneously, thereby reducing memory management complexity while maintaining high throughput production capabilities

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical or multi-level data structure dimension to organize sequences. Instead of flat storage, sequences are organized in nested structures (e.g., collections containing DNA components containing sequence annotations), enabling efficient memory management through selective loading and caching strategies that reduce the active memory footprint while preserving access to large sequence datasets

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If conventional methods are used for high-throughput nucleotide sequence design and manufacturing, then large sets of sequences can be generated, but processing times become slow

Engineering Contradiction:
Improvesequence generation throughputVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing and organizing sequence data into optimized data structures before actual manufacturing operations. This includes pre-validating sequence annotations, pre-organizing DNA components into reusable libraries, and pre-caching frequently accessed sequence data, thereby reducing processing time during high-throughput production while maintaining the ability to generate large sequence sets

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying strategies by creating simplified representations or proxies of sequence data that can be manipulated more efficiently. Instead of directly processing all original sequence data, the system works with compressed or summarized copies for certain operations, then references the full data only when necessary, thereby reducing processing time while maintaining access to complete sequence information

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20220076177A1Microbial strain design system and methods for improved large-scale production of engineered nucleotide sequences
Publication Date: 2022.03.10 ZYMERGEN INC
  • US20220076177A1 patent drawing
  • US20220076177A1 patent drawing
  • US20220076177A1 patent drawing

AI summary

The generation of a factory order includes receiving an expression indicating an operation on a first sequence operand and a second sequence operand. The first sequence operand represents multiple biological sequence parts, and the second sequence operand represents at least one biological sequence part. The expression is evaluated to a sequence specification, which represents modifications to at least one biological sequence, and comprises a data structure representing (a) the first and second sequence operands, (b) one or more first-level operations to be performed on one or more first-level sequence operands, and (c) one or more second-level operations, the execution of at least one of which resolves values of at least one of the first-level sequence operands. A factory order is generated based upon execution of at least one first-level operation and at least one second-level operation.