Unstructured Data Virtual Mart via Metadata Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for creating a data mart from unstructured data sources, such as those used in big-data analysis, face challenges due to the requirement of designating source data addresses, making it difficult to apply snapshot creation methods effectively.

Innovation Solution

Associating first type metadata with unstructured data sources, creating second type metadata that includes content information, and using this metadata to efficiently extract and manage data for a virtual data set, allowing for rapid data mart creation without duplicating the original data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If snapshot creation technique is used to present data to host computer, then data presentation speed is improved, but applicability to unstructured data sources deteriorates due to address designation requirement

Engineering Contradiction:
Improvedata presentation speedVSAvoidapplicability to unstructured data sources
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent segments the data management process by separating structured data (in data marts) from unstructured data (in object storage). Metadata is segmented into different types (first type for object storage, second type for data mart) to enable independent management and efficient access without requiring address designation of source data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces metadata as an intermediary layer between unstructured data in object storage and the host computer. This metadata acts as a mediator that enables data access and management without requiring direct address designation of the source data, thus making snapshot techniques applicable to unstructured data sources.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If data duplication process is performed to create data mart, then data extraction is achieved, but time consumption increases significantly

Engineering Contradiction:
Improvedata extraction completenessVSAvoidDM creation time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent creates virtual copies of data by storing metadata that references the original unstructured data in object storage, rather than physically duplicating the data. This virtual copying mechanism enables data mart creation without time-consuming data duplication while maintaining data extraction completeness.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary actions by creating and maintaining metadata that describes unstructured data before data mart creation is needed. This preliminary metadata preparation enables rapid data mart creation without requiring time-consuming data extraction and duplication at the time of need.

Inventive Principle:
Principle #10Preliminary action

3Loss of time

If virtual data set creation is implemented, then data mart creation time is reduced, but system complexity increases due to metadata management

Engineering Contradiction:
Improvedata mart creation timeVSAvoidmetadata management complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent segments metadata into distinct types (first type metadata for object storage, second type metadata for data mart) with clearly defined roles and relationships. This segmentation simplifies metadata management by establishing a structured hierarchy and reducing the complexity of managing virtual data sets.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10685046B2Data processing system and data processing method
Publication Date: 2020.06.16 HITACHI VANTARA LTD
  • US10685046B2 patent drawing
  • US10685046B2 patent drawing
  • US10685046B2 patent drawing

AI summary

First type metadata is associated with unstructured data included in an unstructured data source. A data processing system performs an extraction process. This extraction process includes: (a) creating, for each of a plurality of selected pieces of unstructured data in the unstructured data source, second type metadata, which is metadata including content information representing one or more content attributes of the piece of unstructured data; and (b) associating the created second type metadata with the first type metadata of the piece of unstructured data.