ML-Based Compression Selection for Storage and Compute Tradeoffs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The inefficiency in managing large data storage due to the requirement of substantial computing resources, where existing methods fail to optimize compression algorithms based on production host performance objectives, leading to suboptimal storage solutions.

Innovation Solution

A method and system that utilize a compression selection model, generated through machine learning algorithms, to select an optimal compression algorithm based on data attributes and production host performance objectives, enabling efficient compression and storage of data objects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional compression methods are used without optimization, then storage capacity is maintained, but computing resources are wasted due to inefficient compression algorithms

Engineering Contradiction:
Improvestorage efficiencyVSAvoidcomputing resources
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system changes the parameter of compression algorithm selection by using machine learning models to dynamically select the most appropriate compression algorithm based on data characteristics and performance objectives, rather than using fixed or manual selection methods. This optimization reduces computing resources while improving storage efficiency

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces manual or rule-based compression algorithm selection with an automated machine learning-based selection system. The ML model automatically analyzes data characteristics and selects optimal compression algorithms, substituting human decision-making or simple heuristics with intelligent automation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If compression algorithms are selected manually or using simple rules, then implementation complexity is low, but storage optimization is insufficient

Engineering Contradiction:
Improvestorage optimizationVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a machine learning-based compression selection service as an intermediary between raw data and compression algorithms. This service acts as a mediator that analyzes data characteristics and recommends optimal compression algorithms, adding intelligence without requiring complex integration throughout the entire system

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system segments the compression optimization task into distinct components: data characteristic analysis, ML-based algorithm selection, and compression execution. This modular approach allows each component to be optimized independently while maintaining overall system manageability

Inventive Principle:
Principle #1Segmentation

3Productivity

If a single compression algorithm is used for all data types, then system simplicity is maintained, but compression performance is suboptimal

Engineering Contradiction:
Improvecompression performanceVSAvoidalgorithm adaptability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by selecting different compression algorithms tailored to specific data characteristics. Instead of using a uniform compression approach for all data, the ML model identifies local patterns in the data and selects compression algorithms optimized for those specific characteristics, achieving superior compression performance

Inventive Principle:
Principle #3Local quality

4Productivity

If compression algorithm selection is automated without ML, then automation level is moderate, but selection accuracy is low

Engineering Contradiction:
Improveselection accuracyVSAvoidautomation level
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The patent replaces rule-based or heuristic automation with machine learning-based automation for compression algorithm selection. The ML models learn from data characteristics and automatically select optimal algorithms with high accuracy, surpassing the capabilities of traditional automated methods while maintaining operational efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11394397B2System and method for selecting a lossless compression algorithm for a data object based on performance objectives and performance metrics of a set of compression algorithms
Publication Date: 2022.07.19 EMC IP HLDG CO LLC
  • US11394397B2 patent drawing
  • US11394397B2 patent drawing
  • US11394397B2 patent drawing

AI summary

A method for managing data includes obtaining a compression algorithm selection request for a data object, wherein the data object is generated by a production host, identifying, in response to the compression algorithm selection request, a set of production host performance objectives of the production host, performing a compression algorithm selection analysis using the set of production host performance objectives and a compression selection model to obtain a compression algorithm selection for a compression algorithm, specifying the compression algorithm to the production host using a data agent, wherein the data agent is operatively connected to the production host, initiating a compression on the data object using the data agent by applying the compression algorithm to obtain a compressed data object, and initiating a storage of the compressed data object.