Hybrid Genomic Data Storage Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems struggle to efficiently store genomic data while maintaining I/O performance, reliability, and data redundancy, particularly due to the large and varied storage requirements across different stages of the genomics workflow.

Innovation Solution

A method that designates multiple data storage techniques for genomic data based on file type, pipeline stage, access frequency, and most accessed blocks, allowing each block of a file to be stored using the most optimal technique, thereby utilizing a hybrid approach at both the file and block levels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple data storage techniques are used for different blocks of genomic data files, then storage efficiency is optimized, but system complexity increases

Engineering Contradiction:
Improvestorage efficiencyVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent divides genomic data files into blocks and applies different storage techniques to different blocks based on their characteristics. This segmentation allows the system to optimize storage efficiency for each block type while managing complexity through modular organization of storage strategies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality by tailoring storage techniques to specific blocks of data based on their access patterns, size, and importance. Different blocks within the same file can use different storage techniques, allowing optimization at the local level while maintaining overall system functionality.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If data is stored using multiple different data storage techniques, then storage efficiency improves, but I/O performance may deteriorate

Engineering Contradiction:
Improvestorage efficiencyVSAvoidI/O performance
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The patent applies different storage techniques to different blocks based on their specific characteristics such as access frequency and data importance. This ensures that frequently accessed blocks use techniques optimized for speed, while less critical blocks use techniques optimized for space efficiency.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamic block migration between different storage techniques based on changing access patterns. Blocks can be moved between storage tiers as their access characteristics evolve, ensuring optimal performance is maintained over time while maximizing storage efficiency.

Inventive Principle:
Principle #15Dynamics

3Quantity of substance

If storage is optimized for genomic workloads, then storage efficiency increases, but adaptability to other workloads decreases

Engineering Contradiction:
Improvestorage efficiencyVSAvoidworkload adaptability
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent creates a storage system with multiple storage techniques that can handle different data characteristics. This multi-functional architecture allows the same storage system to efficiently handle genomic workloads while also being adaptable to other types of data workloads with different access patterns and requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses configurable parameters and metadata to describe data characteristics, allowing the storage system to adapt its behavior based on the specific workload. By changing parameters such as access frequency thresholds and block size configurations, the system can optimize for different workload types including genomic and non-genomic data.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12210904B2Hybridized storage optimization for genomic workloads
Publication Date: 2025.01.28 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12210904B2 patent drawing
  • US12210904B2 patent drawing
  • US12210904B2 patent drawing

AI summary

A method for more efficiently storing genomic includes designating multiple different data storage techniques for storing genomic data generated by a genomic pipeline. The method further identifies a file, made up of multiple blocks, generated by the genomic pipeline. The method determines which data storage technique is most optimal to store each block of the file. In doing so, the method may consider the type of the file, the stage of the genomic pipeline that generated the file, the access frequency for blocks of the file, the most accessed blocks of the file, and the like. The method stores each block using the data storage technique determined to be most optimal after completion of a designated stage of the genomic pipeline, such that blocks of the file are stored using several different data storage techniques. A corresponding system and computer program product are also disclosed.