Unstructured Data Virtual Mart via Metadata Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for creating a data mart from unstructured data sources, such as those used in big-data analysis, face challenges due to the requirement of designating source data addresses, making it difficult to apply snapshot creation methods effectively.
Innovation Solution
Associating first type metadata with unstructured data sources, creating second type metadata that includes content information, and using this metadata to efficiently extract and manage data for a virtual data set, allowing for rapid data mart creation without duplicating the original data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If snapshot creation technique is used to present data to host computer, then data presentation speed is improved, but applicability to unstructured data sources deteriorates due to address designation requirement
Solution Approach 1:
The patent segments the data management process by separating structured data (in data marts) from unstructured data (in object storage). Metadata is segmented into different types (first type for object storage, second type for data mart) to enable independent management and efficient access without requiring address designation of source data.
Solution Approach 2:
The patent introduces metadata as an intermediary layer between unstructured data in object storage and the host computer. This metadata acts as a mediator that enables data access and management without requiring direct address designation of the source data, thus making snapshot techniques applicable to unstructured data sources.
2Loss of information
If data duplication process is performed to create data mart, then data extraction is achieved, but time consumption increases significantly
Solution Approach 1:
The patent creates virtual copies of data by storing metadata that references the original unstructured data in object storage, rather than physically duplicating the data. This virtual copying mechanism enables data mart creation without time-consuming data duplication while maintaining data extraction completeness.
Solution Approach 2:
The patent performs preliminary actions by creating and maintaining metadata that describes unstructured data before data mart creation is needed. This preliminary metadata preparation enables rapid data mart creation without requiring time-consuming data extraction and duplication at the time of need.
3Loss of time
If virtual data set creation is implemented, then data mart creation time is reduced, but system complexity increases due to metadata management
Solution Approach 1:
The patent segments metadata into distinct types (first type metadata for object storage, second type metadata for data mart) with clearly defined roles and relationships. This segmentation simplifies metadata management by establishing a structured hierarchy and reducing the complexity of managing virtual data sets.
Data Source
AI summary
First type metadata is associated with unstructured data included in an unstructured data source. A data processing system performs an extraction process. This extraction process includes: (a) creating, for each of a plurality of selected pieces of unstructured data in the unstructured data source, second type metadata, which is metadata including content information representing one or more content attributes of the piece of unstructured data; and (b) associating the created second type metadata with the first type metadata of the piece of unstructured data.


