Genomic Information Storage via Stable Coding Regions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
DNA storage in living organisms faces challenges due to complex mechanisms of data corruption and evolutionary processes, where slight changes in nucleotide sequences can affect organism behavior, growth rate, and even lead to death, necessitating a new model for reliable information storage.
Innovation Solution
A method involving the selection of specific genomic sequences with low indel probabilities and intermediate nucleotide sequence identity across orthologs, followed by inducing genetic mutations that encode information without being detrimental to the organism's health, using error-correcting codes like Reed-Solomon codes to ensure data fidelity across generations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If DNA is used as a storage medium, then vast amounts of data can be stored for very long periods with high reliability, but complex evolutionary processes create data corruption mechanisms that can profoundly affect the organism
Solution Approach 1:
The patent applies local quality by selecting specific coding regions with low indel probability and intermediate nucleotide identity across orthologs. Instead of treating the entire genome uniformly, the method identifies and stores information in specific localized regions that have been evolutionarily selected for stability, while avoiding regions prone to harmful mutations.
Solution Approach 2:
The patent employs preliminary action by pre-selecting and characterizing suitable coding regions before storing information. The method involves estimating indel probability and analyzing nucleotide sequence identity across orthologs in advance to identify safe storage locations, preventing future corruption before it occurs.
2Loss of information
If genetic mutations are induced to store information, then information can be encoded in the genome, but the mutations may be detrimental to the health of the living organism
Solution Approach 1:
The patent applies parameter changes by carefully controlling the type and location of mutations. The method changes nucleotide parameters (selecting specific bases) and positional parameters (selecting specific coding regions) to encode information while maintaining the mutations within safe boundaries that prevent harm to the organism.
Solution Approach 2:
The patent uses intermediary measures by introducing error-correcting codes as a protective layer between the stored information and the biological system. This intermediary structure allows the genome to tolerate certain mutations while maintaining information integrity, acting as a buffer against harmful effects.
3Reliability
If coding regions with high nucleotide sequence identity across orthologs are selected, then storage reliability is improved, but the flexibility to induce meaningful mutations is reduced
Solution Approach 1:
The patent applies local quality by differentiating between regions of high conservation (for reliability) and regions with acceptable variability (for mutation flexibility). The method selects coding regions with intermediate nucleotide identity - conserved enough to ensure reliability across species, but variable enough to allow meaningful mutations for information storage.
Data Source
AI summary
Methods of storing information in a living organism are provided. Systems and computer program products for performing the methods are also provided.


