Genomic Annotation File Compression With Field-Specific Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current compression algorithms for genomic data lack selectivity, interoperability, and generality, failing to efficiently compress and decompress genomic annotation data due to incompatible file formats and the inability to selectively compress or encrypt specific data fields, leading to suboptimal compression and inefficient data management.

Innovation Solution

A method and system that access genomic annotation data in various file formats, extract attributes, divide them into chunks, process these attributes and chunks for selective compression using different compressors, and generate a unified file format allowing selective decompression and encryption, enabling efficient compression and decompression of specific data fields while integrating metadata and linkages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing compression algorithms compress all fields together, then compression is applied to the entire dataset, but selectivity is lost and the ability to extract specific fields without decompressing all attributes is eliminated

Engineering Contradiction:
ImproveselectivityVSAvoidcompression structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent divides genomic annotation data into separate fields (attributes) and applies different compression algorithms to each field based on its statistical characteristics. This segmentation enables selective compression and extraction of specific fields without decompressing the entire dataset, directly resolving the contradiction between selectivity and compression structure complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different compression algorithms to different fields based on their local statistical characteristics. Each field is analyzed individually and assigned the most appropriate compression method, achieving optimal compression ratios for each attribute while maintaining the ability to selectively access specific fields.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If multiple incompatible file formats are used for genomic annotation data, then data representation flexibility is maintained, but interoperability issues arise and frequent format conversions are required

Engineering Contradiction:
Improvedata representation flexibilityVSAvoidinteroperability
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent creates a universal compressed file format that can store multiple types of genomic annotation data (VCF, BED, WIG, and other formats) in a single standardized structure. This universal format eliminates the need for frequent format conversions while maintaining the ability to represent diverse data types, resolving the contradiction between data representation flexibility and interoperability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces a standardized compressed file format as an intermediary that translates multiple incompatible genomic data formats into a common structure. This intermediary format enables interoperability between different tools and platforms while preserving the original data representation characteristics, avoiding the need for repeated conversions.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If specialized compression methods based on disk-based array management tools are used, then compression is optimized for specific data types, but high-level features like metadata, linkages, and attribute-specific indexing are lacking

Engineering Contradiction:
Improvecompression efficiencyVSAvoidfeature set
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent merges the advantages of specialized compression methods with high-level data management features into a single integrated system. It combines field-specific compression algorithms with metadata storage, attribute-specific indexing, and data linkage capabilities, achieving both high compression efficiency and comprehensive feature support simultaneously.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11916576B2System and method for effective compression, representation and decompression of diverse tabulated data
Publication Date: 2024.02.27 KONINKLIJKE PHILIPS NV
  • US11916576B2 patent drawing
  • US11916576B2 patent drawing
  • US11916576B2 patent drawing

AI summary

A method for controlling compression of data includes accessing genomic annotation data in one of a plurality of first file formats, extracting attributes from the genomic annotation data, dividing the genomic annotation data into multiple chunks, and processing the extracted attributes and chunks into correlated information. The method also includes selecting different compressors for the attributes and chunks identified in the correlated information and generating a file in a second file format that includes the correlated information and information indicative of the different compressors for the chunks and attributes indicated in the correlated information. The information indicative of the different compressors is processed into the second file format to allow selective decompression of the attributes and chunks indicated in correlated information.