Object-Based Storage Header Segmentation for Direct Data Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current networked storage systems face inefficiencies in storing and retrieving data, particularly in object-based data stores, where data is not optimally managed across performance and capacity tiers, leading to suboptimal access times and storage utilization.

Innovation Solution

The system generates objects with header and data segments, where the header segment provides offset addresses and compression group sizes, enabling efficient decompression and retrieval of data chunks, and uses a unified format for storing both compressed and uncompressed data chunks, allowing direct access without re-reading headers for uncompressed data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If data is stored in object-based data stores with traditional formats, then data can be stored at capacity tiers, but data access requires re-reading headers which reduces read performance

Engineering Contradiction:
Improvedata access speedVSAvoidtime spent re-reading headers
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The data object is segmented into distinct components: a header segment containing metadata and offset information, and a data segment containing the actual data chunks. This segmentation allows the system to store offset addresses in the header that enable direct access to data chunks without requiring the header to be re-read during data retrieval operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The header segment is prepared in advance with pre-calculated offset addresses and compression group size information. This preliminary action of storing access metadata upfront eliminates the need for repeated header reading during data access, as the offset information is already available for direct data chunk retrieval.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If compression is applied to data chunks, then storage utilization improves, but decompression processing is required which adds complexity

Engineering Contradiction:
Improvestorage capacity utilizationVSAvoidcompression management complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system applies compression selectively to specific groups of data chunks (compression groups) rather than uniformly to all data. The header segment contains information about compression group sizes and offsets, allowing the system to identify which data chunks are compressed and manage them appropriately, reducing overall complexity while maintaining storage efficiency.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the state of data chunks by compressing them and storing compression metadata in the header segment. The header contains parameters such as compression group size and offset addresses that track the compressed data's location and characteristics, enabling efficient management of compressed data without requiring complex real-time compression/decompression operations.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If a unified format is used for both compressed and uncompressed data, then storage management is simplified, but the format must accommodate variable data states

Engineering Contradiction:
Improveformat flexibilityVSAvoidformat structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The object-based storage format is designed to be universal, accommodating both compressed and uncompressed data chunks within the same data segment. The header segment contains flexible fields that can represent different data states, allowing the same storage structure to handle variable data formats without requiring separate storage paths or complex format conversion mechanisms.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9798496B2Methods and systems for efficiently storing data
Publication Date: 2017.10.24 NETAPP INC
  • US9798496B2 patent drawing
  • US9798496B2 patent drawing
  • US9798496B2 patent drawing

AI summary

Methods and systems for a networked system are provided. One method includes generating an object by a processor for storing a plurality of data chunks at a storage device, where the object includes a header segment and a data segment, the header segment providing a first offset address where an uncompressed data chunk is stored within the object and a second offset address of the object indicating a beginning of a compressed group having compressed data chunks and providing an indicator of a compression group size; reading the header segment by the processor to retrieve the second offset and the compressed group size in response to a first request for a data chunk within the compressed group; and decompressing the data chunk of the compressed group by the processor and providing the uncompressed data chunk for completing the first read request.