DNA Data Storage Multiplexing for Parallel Similarity Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional DNA storage technologies lack full database functionality, particularly in performing similarity searches, which are essential for content-based retrieval and are often limited to single-query operations, leading to inefficiencies in time and resource usage.
Innovation Solution
Implementing multiplex similarity search techniques in DNA data storage, where multiple queries can be performed in a single reaction, with linked DNA strands allowing for longer result strands to enhance accuracy and differentiate between captured and non-captured items, utilizing hybridization and ligation reactions to connect query and data strands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If single-query operations are performed in DNA data storage, then simplicity of operation is maintained, but productivity is reduced due to inability to perform multiple searches simultaneously
Solution Approach 1:
The patent combines multiple query operations into a single reaction by introducing a multiplexing mechanism where multiple query strands with different feature sequences are pooled together and processed simultaneously against the DNA database, enabling parallel search operations without requiring separate reactions for each query
Solution Approach 2:
The patent creates a universal search reaction system that can handle multiple different queries through a common reaction framework. The DNA database elements are designed with universal feature sequences that can interact with various query types, allowing the same reaction mixture to perform multiple search functions simultaneously
2Measurement precision
If DNA strands are connected via linking strands, then measurement precision is improved through length differentiation, but device complexity increases due to additional molecular components
Solution Approach 1:
The patent introduces DNA linking strands as intermediary molecules that connect captured DNA database elements to query strands. These linking strands serve as mediators that enable the formation of extended result strands, allowing for length-based differentiation between captured and non-captured items without directly modifying the core data or query sequences
Solution Approach 2:
The patent adds a length dimension to the search results by connecting strands. Instead of relying solely on sequence matching, the system creates result strands of different lengths based on whether capture occurred, enabling an additional layer of discrimination and verification through size-based separation or detection methods
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables efficient and accurate multiple-query operations, improving data retrieval and storage efficiency by allowing simultaneous searches across numerous queries, enhancing accuracy through length differentiation and resource optimization.
Implementation Method 1
matching comprises hybridizing a feature nucleotide sequence of the DNA database element to a complementary feature nucleotide sequence of the query strand
Implementation Method 2
matching comprises arranging a DNA linking strand into a connection between a data strand of the DNA database element and the query strand
Data Source
AI summary
Multiplex similarity search can be performed in a DNA data storage context. The described technologies can support a plurality of different DNA data storage queries in a single query run. A linking strand can be used to connect a query to its matching data element. After the query finds a matching data element, a result strand can be sequenced to the reveal the matching data element as well as which of the queries resulted in the match. Thus, in a multiplex similarity search scenario, a plurality of result strands from a single query run can be correlated to a plurality of different queries. Also, the result strand can be of significantly longer length than both the unmatched data strands and the unmatched query strands. Therefore, filtering based on length can provide more accurate results.


