Blockchain Genetic Data Compression via Reference Genome
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems lack an efficient and secure method for storing and managing large amounts of genetic data, particularly in a decentralized manner, which is essential for genetic research and data sharing while ensuring privacy and security.
Innovation Solution
A blockchain-based system that compresses genetic data relative to a reference genome, allowing for secure and decentralized storage and exchange by using pointers to transactions storing genetic data, enabling efficient access and decompression of target genome data while ensuring data privacy through encryption and secure protocols.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If genetic data is stored in raw form on a blockchain, then data accessibility and completeness are improved, but storage costs and data management complexity increase significantly
Solution Approach 1:
The patent extracts only the essential genetic data (variations from reference genome) and stores it on the blockchain, while leaving the complete reference genome elsewhere. This extraction approach reduces storage requirements dramatically while maintaining data accessibility for research purposes.
Solution Approach 2:
The patent introduces a reference genome as an intermediary element that enables compression of genetic data. By storing variations relative to this reference rather than raw sequences, the system reduces storage needs while maintaining information completeness through the reference as a mediator.
2Quantity of substance
If genetic data is compressed relative to reference genome, then storage costs are reduced, but data decompression and access complexity increase
Solution Approach 1:
The patent performs preliminary compression of genetic data relative to a reference genome before storage on the blockchain. This preliminary action reduces the data size that needs to be stored and transmitted, while the decompression process is simplified by having the reference genome readily available for comparison.
3Loss of information
If complete raw genetic data is stored and transferred, then data completeness is maintained, but transmission time and network bandwidth requirements increase
Solution Approach 1:
The patent extracts only the variable portions of genetic data (mutations and variations) from the complete genome sequence and stores only these differences on the blockchain. This extraction maintains data completeness for research purposes while dramatically reducing transmission requirements and time.
Solution Approach 2:
The patent changes the representation parameter of genetic data from raw complete sequences to compressed variation data relative to a reference. This parameter change maintains information completeness while reducing data volume, thereby decreasing transmission time and network bandwidth requirements.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method performed by computer equipment of a consuming party, comprising: accessing an electronic document comprising a plurality of pointers, each pointer comprising a respective transaction identifier of a respective destination transaction stored on a blockchain, wherein the destination transactions comprise one or more first transactions storing respective genetic data of at least part of a reference genome, and one or more second transactions storing respective genetic data of at least a corresponding part of a target genome in compressed form compressed relative to the reference genome; accessing the genetic data from at least one of the first destination transactions and at least a corresponding one of the second destination transactions based on the respective identifiers accessed from the electronic document; and decompressing the accessed genetic data of the target genome based on the accessed genetic data of the reference genome.