Custom Metadata Extraction Across Heterogeneous Storage Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computing storage systems face challenges in accessing and understanding large amounts of data spread across multiple heterogeneous storage systems, as they often lack the ability to extract metadata from files and objects, which hinders data retrieval and insights.

Innovation Solution

A processor-based solution that extracts custom metadata from data across heterogeneous storage systems, indexing it into a centralized search index, using mechanisms like policy engines, natural language processing, and AI to correlate and enhance metadata, enabling efficient data access and insights.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If metadata extraction is implemented across heterogeneous storage systems, then data accessibility and insights are improved, but system complexity and processing overhead increase

Engineering Contradiction:
Improvemetadata extraction capabilityVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary metadata extraction and indexing system that sits between heterogeneous storage systems and users/applications. This mediator automatically extracts metadata from diverse storage formats, standardizes it, and indexes it in a centralized repository, thereby improving data accessibility without requiring users to directly handle the complexity of heterogeneous storage systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system segments the metadata extraction process into independent modular components that can handle different storage formats separately. Each component is specialized for specific file types or storage systems, allowing the overall system to manage complexity through division while providing unified metadata access across all heterogeneous sources.

Inventive Principle:
Principle #1Segmentation

2Productivity

If real-time metadata extraction is performed across large volumes of data, then data retrieval efficiency is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvedata retrieval efficiencyVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary metadata extraction and indexing operations in advance, proactively building and maintaining a centralized metadata repository as data is written to or moved between storage systems. This preliminary action ensures that when data retrieval is needed, the metadata is already prepared and indexed, eliminating the need for time-consuming extraction operations at the moment of data access.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If custom metadata is extracted and indexed from heterogeneous data sources, then search capability and data understanding are improved, but the complexity of managing diverse data formats increases

Engineering Contradiction:
Improvesearch capabilityVSAvoiddata format compatibility
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal metadata schema and indexing framework that can accommodate multiple data formats and storage systems through a common interface. The system is designed to be multi-functional, handling various file types, storage protocols, and data structures while maintaining a consistent metadata extraction and indexing approach, thereby improving search capability without sacrificing adaptability to diverse data sources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10885007B2Custom metadata extraction across a heterogeneous storage system environment
Publication Date: 2021.01.05 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10885007B2 patent drawing
  • US10885007B2 patent drawing
  • US10885007B2 patent drawing

AI summary

Embodiments for triggering custom metadata extraction by a processor. Information may be extracted from an event so as to access data across a plurality of heterogeneous storage systems. Metadata may be extracted from the data that is accessed such that the metadata is assigned as custom metadata and indexed into a centralized search index, wherein the custom metadata is correlated to existing metadata associated with the data in the centralized search index.