DNA Key-Value Store Using Fluorescent Sorting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing in-vivo DNA storage methods require sequencing the entire population of live microorganisms to retrieve a single dataset, lacking selective or random-access capabilities, which is inefficient and time-consuming.
Innovation Solution
A method that uses a key-value store system to map digital data to unique genes expressing fluorescent proteins, allowing for the selective retrieval of subsets by isolating microorganisms with specific plasmids containing the desired data using fluorescence-activated cell sorting and DNA sequencing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If whole DNA sequencing is performed to retrieve data from live microorganisms, then complete data retrieval is achieved, but retrieval time is excessively long and efficiency is low
Solution Approach 1:
The patent extracts only the specific plasmid containing the desired data from the population of microorganisms using fluorescence-activated cell sorting, rather than sequencing the entire DNA population. This selective extraction of the relevant component (specific plasmid) eliminates the need to process irrelevant DNA sequences, thereby dramatically reducing retrieval time while maintaining complete data retrieval for the requested dataset.
Solution Approach 2:
The patent segments the DNA storage system into distinct functional components: chromosomal DNA for essential functions and separate plasmids for data storage. This segmentation allows independent manipulation and selective retrieval of data-containing plasmids without affecting or requiring analysis of the entire genomic DNA, enabling efficient selective access to specific datasets.
2Adaptability or versatility
If artificial DNA sequences are introduced into live microorganisms for data storage, then data can be stored in living systems, but the artificial sequences may interfere with normal genetic and biological mechanisms
Solution Approach 1:
The patent separates data storage functions from essential biological functions by using extrachromosomal plasmids rather than integrating data sequences into the chromosome. This segmentation isolates artificial data-containing sequences from critical genomic regions, minimizing potential interference with normal genetic mechanisms while maintaining the ability to store and retrieve data in living microorganisms.
Solution Approach 2:
The patent uses plasmids as intermediary vehicles to carry artificial data sequences away from the chromosomal DNA. These plasmids act as mediators that hold the artificial sequences in a separate, controllable compartment within the cell, reducing direct interaction between artificial data sequences and essential genomic elements, thereby minimizing genetic interference.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables rapid and selective retrieval of specific data subsets, significantly reducing retrieval time and making it feasible for big data analytics, while maintaining data integrity and efficiency through biological storage.
Implementation Method 1
a first sequence is a fluorescent protein sequence that expresses a fluorescent protein
Implementation Method 2
a second sequence is a guide RNA sequence that directs cleavage of a chromosomal DNA sequence
Implementation Method 3
the Cas9 gene and the guide RNA sequence direct cleavage of a chromosomal DNA sequence
Data Source
AI summary
A digital store comprising of a method to store digital data in live micro-organisms, and a method to selectively retrieve subsets of stored data, is disclosed. Digital data is represented as a plurality of key-value pairs. The proposed system stores copies of key-value pairs in a plurality of live micro-organisms. Upon presentation of a retrieval key, the proposed digital store retrieves the value associated with the retrieval key. Storage method for a key-value pair comprises of (a) mapping the key to a gene that expresses a unique fluorescent protein so that no two keys map to the same gene, (b) encoding the key-value pair as base-pair sequences, (c) synthesizing oligonucleotide chains from base-pairs for the key-value pair and the gene, (d) synthesizing recombinant DNA plasmids that have oligonucleotide chains for the key-value pair, the gene, and two primers, as foreign DNA inserts, (e) incorporation of recombinant DNA plasmids into live micro-organisms, (f) isolation of live micro-organisms that have absorbed the recombinant DNA plasmids, and (g) safe storage of population of live micro-organisms with embedded key-value pairs in a common pool. Retrieval of the value paired with a key comprises of (a) taking as input the retrieval key, and mapping the key to the specific gene for fluorescent protein, (b) taking a sample from the safe storage pool that contains live micro-organisms embedded with key-value pairs, (c) isolating the live micro-organisms that have expressed the gene by using high-speed fluorescence activated cell sorting or flow cytometry, (d) extracting DNA from the recombinant DNA plasmid in the isolated live micro-organisms, (d) selectively amplifying and sequencing only those DNA strands that contain the value for the key, and (e) decoding the base-pair sequence obtained after DNA sequencing to yield the value associated with the retrieval key.We also disclose two important variations. The first variation relates to the storage step. The recombinant DNA plasmid is constructed to include additional non-fluorescent oligonucleotides and genes so that during the data retrieval step, the live micro-organisms that have absorbed the said plasmid can be sorted by cell sorters based on parameters of individual cells such as cell size, cell complexity, cell phenotype, cell structure, cell function, and magnetic or electrical properties. The second variation relates to both storage and retrieval of key-value pairs with large values. To store such a key-value pair, the large value is split into smaller blocks so that a block can fit into a recombinant DNA plasmid, and a distinct pair of primers is used for each block. A block's primer pair is used to selectively amplify and sequence only the DNA that encodes the data in the block, thereby enabling the retrieval of a specific block of the value, as opposed to retrieving the entire value associated with a key.


