Genome Data De-duplication via Conformed Format Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Genome sequencing generates massive amounts of redundant data, leading to significant storage challenges due to the redundancy in file formats like FASTQ, SAM, BAM, and CRAM, which require large amounts of storage space.
Innovation Solution
A computer-implemented method and system for data deduplication that creates conformed representations of genome data, identifies and retains unique copies, and releases duplicates across different file formats, utilizing hardware processors to manage and process genome data files, thereby reducing storage needs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If genome data is stored in multiple file formats (FASTQ, SAM, BAM, CRAM), then data compatibility and accessibility are improved, but storage space requirements increase significantly
Solution Approach 1:
The patent creates conformed representations (copies) of genome data in a standardized format that can represent multiple original formats. Instead of storing multiple format-specific files, the system creates a single conforming copy that preserves the essential information from FASTQ, SAM, BAM, and CRAM formats, thereby reducing storage space while maintaining data accessibility across different format requirements
Solution Approach 2:
The patent develops a universal conformed representation that serves multiple functions: it can represent data from different input formats, enable cross-format comparison, and support various downstream analyses. This single universal structure replaces the need to maintain separate format-specific storage, achieving multi-functionality while eliminating redundant storage
2Reliability
If all genome data files are retained in original formats, then data integrity and format-specific features are preserved, but storage costs and management complexity increase
Solution Approach 1:
The patent extracts the essential genomic information from multiple format-specific files and consolidates it into a single conformed representation. By taking out only the critical data elements needed for analysis and comparison while discarding format-specific redundancies, the system preserves data integrity for analytical purposes while significantly reducing storage requirements
Solution Approach 2:
The patent transforms data from multiple format-specific parameter structures into a unified parameter set in the conformed representation. This parameter transformation maintains the essential genetic information and analytical capabilities while eliminating the storage overhead of maintaining multiple parameter representations for the same underlying data
3Loss of information
If genome data from multiple samples is stored with full sequence information, then analytical completeness is improved, but redundancy and storage requirements increase
Solution Approach 1:
The patent merges sequence information from multiple samples into conformed representations that share common structural elements. By combining the data representation into a unified format with shared parameters and structures, the system eliminates redundancy across samples while preserving all necessary sequence information for complete analytical coverage
Data Source
AI summary
A computer-implemented method for data-deduplication of genome data that is in different file formats is described. Representative data from different genome file formats is conformed to a selected file format and compared. Duplicate files are identified and duplicate files are released, with at least one file copy being retained.


