Semantic Syntactic Metadata Linking for Heterogeneous Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data management systems face challenges in accessing and organizing highly dimensional, heterogeneous data due to the lack of standards for describing semantic and syntactic properties, which hinders collaboration across institutional boundaries and is exacerbated by rapid data generation.

Innovation Solution

A method for managing semantic and syntactic metadata involves receiving heterogeneous data, capturing and linking semantic and syntactic metadata, and storing it in a repository, using a system that determines data source, captures metadata attributes, and generates parsers for standardized syntax, enabling efficient data organization and access.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional data management systems are used to store heterogeneous data, then data storage is simple, but data accessibility and organization efficiency deteriorate due to lack of semantic and syntactic standards

Engineering Contradiction:
Improvedata accessibilityVSAvoiddata management system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments data into multiple hierarchical levels including raw data, annotated data with syntactic metadata, and structured data with semantic metadata. This segmentation allows different levels of data processing and access, improving accessibility without requiring the entire system to handle all complexity simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces metadata as an intermediary layer between raw heterogeneous data and data access operations. Syntactic metadata describes data structure and format, while semantic metadata describes data meaning and context. This intermediary layer enables standardized access without requiring direct handling of raw data complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If data standards for semantic and syntactic properties are implemented, then data organization and collaboration improve, but system complexity and implementation difficulty increase

Engineering Contradiction:
Improvecollaboration capabilityVSAvoiddata standard implementation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates universal metadata schemas that can describe multiple types of heterogeneous data (images, audio, video, text) using the same syntactic and semantic frameworks. This universality enables collaboration across different data types and institutional boundaries without requiring separate standards for each data type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies syntactic annotation and semantic tagging to data during the data ingestion and processing stages, before data access operations occur. This preliminary action ensures that metadata is already available when needed for collaboration, rather than requiring complex real-time analysis during data access.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If comprehensive metadata capture is performed on all heterogeneous data, then data description accuracy improves, but processing time and computational resources increase

Engineering Contradiction:
Improvedata property description accuracyVSAvoidmetadata processing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent implements a tiered metadata capture approach where syntactic metadata (basic structure and format) is captured for all data, while semantic metadata (detailed meaning and context) is captured selectively based on data type, importance, and available resources. This partial action approach ensures essential description accuracy without requiring exhaustive metadata for every data element.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent performs metadata capture during data ingestion and initial processing stages, rather than on-demand during data access. This preliminary action allows batch processing of metadata extraction, reducing the time impact on individual data access operations while maintaining comprehensive description accuracy.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If rapid data generation is accommodated without standardized metadata, then data ingestion speed is maintained, but data usability and access efficiency deteriorate

Engineering Contradiction:
Improvedata ingestion speedVSAvoiddata usability
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent implements automated syntactic annotation and semantic tagging that occurs during or immediately after data ingestion, rather than as a separate post-processing step. This preliminary action ensures that even rapidly ingested data carries standardized metadata, maintaining both ingestion speed and data usability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs automated metadata extraction systems that can independently analyze and annotate data without requiring manual intervention. This self-service approach to metadata generation maintains data usability standards even as data ingestion rates increase, without proportionally increasing processing overhead.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9483464B2Method and system for managing semantic and syntactic metadata
Publication Date: 2016.11.01 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US9483464B2 patent drawing
  • US9483464B2 patent drawing
  • US9483464B2 patent drawing

AI summary

A method and system for managing semantic and syntactic metadata. Heterogeneous data is received. After the heterogeneous data is received, the semantic metadata associated with the received heterogeneous data is captured and syntactic metadata associated with the received heterogeneous data is captured. The semantic metadata describes contextually relevant or domain-specific information about data based on an industry-specific or enterprise-specific metadata model or ontology. The syntactic metadata included grammatical rules and structural patterns governing an ordered use of formats and arrangement pertaining to specified data. The received heterogeneous data and said captured semantic metadata and said syntactic metadata are logically linked. The heterogeneous data is stored in a repository.