Storage Device ML Facet Filtering for Dataset Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing storage devices struggle to efficiently filter out noisy, redundant, and superfluous data from datasets, leading to performance bottlenecks and reduced quality of machine learning models due to the transfer of large, unfiltered datasets.

Innovation Solution

Storage devices utilize ML facets to automatically filter datasets by applying dataset preparation operations based on ML facet mappings, providing high-quality, filtered datasets to machine learning applications, and allowing user customization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If large unfiltered datasets are transferred from storage to ML models, then complete data is provided for training, but retrieval performance is severely bottlenecked and training time increases

Engineering Contradiction:
Improvedata completenessVSAvoidretrieval performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system extracts and removes noisy, redundant, and superfluous data from datasets before transfer, keeping only the useful portions needed for ML training. This extraction process resolves the contradiction by providing complete relevant data while eliminating unnecessary data that causes retrieval bottlenecks.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The storage device performs dataset filtering and preparation operations before the ML application requests the data. By preliminarily processing the dataset to remove irrelevant portions, the system ensures that only useful data is transferred, improving retrieval performance without compromising data completeness for training.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If large unfiltered datasets are transferred, then all available data is available for training, but network and computational resources are wasted on irrelevant data

Engineering Contradiction:
Improvedata availabilityVSAvoidresource efficiency
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The system extracts and removes noisy, redundant, and superfluous data from datasets before transfer, keeping only the useful portions needed for ML training. This extraction process resolves the contradiction by providing complete relevant data while eliminating unnecessary data that causes retrieval bottlenecks.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system discards irrelevant, noisy, and redundant data portions from datasets, recovering and retaining only the valuable information needed for ML training. This discarding process improves resource efficiency by preventing waste of network and computational resources on useless data while maintaining data availability for effective training.

Inventive Principle:
Principle #34Discarding and recovering

3Manufacturing precision

If storage devices filter datasets automatically, then dataset quality improves and training efficiency increases, but system complexity increases

Engineering Contradiction:
Improvedataset qualityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The storage device autonomously performs dataset filtering and preparation operations using predefined ML facet mappings without requiring external intervention. This self-service capability improves dataset quality and training efficiency while managing system complexity by automating the filtering process through pre-established rules and mappings.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The storage device performs dataset filtering and preparation operations before the ML application requests the data. By preliminarily processing the dataset to remove irrelevant portions, the system ensures that only useful data is transferred, improving retrieval performance without compromising data completeness for training.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If users can customize dataset preparation, then flexibility and adaptability improve, but ease of operation decreases

Engineering Contradiction:
Improvecustomization flexibilityVSAvoiduser effort
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The storage device autonomously performs dataset filtering and preparation operations using predefined ML facet mappings without requiring external intervention. This self-service capability improves dataset quality and training efficiency while managing system complexity by automating the filtering process through pre-established rules and mappings.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12450003B2Machine learning facets for dataset preparation in storage devices
Publication Date: 2025.10.21 HEWLETT PACKARD ENTERPRISE DEV LP
  • US12450003B2 patent drawing
  • US12450003B2 patent drawing
  • US12450003B2 patent drawing

AI summary

Examples described herein relate to preparing datasets in a storage device for machine learning (ML) applications. Examples include maintaining ML facet mappings between ML facets and dataset preparation tags, deriving ML facets of a dataset stored in the storage device, and generating filtered datasets from the datasets using the ML facets and ML facet mappings. The filtered dataset is associated with improved dataset quality compared to unfiltered dataset. The storage device transmits the filtered dataset to ML applications requesting the dataset. Some examples include recommending, by the storage device, ML facets to the ML application based on performance metrics.