Data Join System with Metadata Capture for Semi-Structured Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for data joining and metadata configuration struggle with integrating semi-structured and unstructured data, particularly in a parallelized environment, as they were not designed for massive parallelization and user participation, leading to inefficiencies in understanding and processing diverse data forms.

Innovation Solution

A method and system that capture metadata information from semi-structured and unstructured data, define a flattened structure, extract entities, and join them with relational data, utilizing a computer network and processors to enable efficient data processing and integration across various data formats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If semi-structured and unstructured data are flattened for integration with relational data, then data joining capability is improved, but metadata and schema information become invalid and require redefinition

Engineering Contradiction:
Improvedata joining capabilityVSAvoidmetadata redefinition complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by capturing and storing metadata information before the flattening process occurs. This allows the original semi-structured metadata to be preserved and referenced during integration, avoiding the need to completely redefine schema information after flattening.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary mechanism that handles the transition between semi-structured and relational data formats. This intermediary layer manages the metadata transformation and validation processes, separating the complexity of metadata redefinition from the core data joining operation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If existing data joining methods are used, then single-program execution is simplified, but massive parallelization and user participation are not supported

Engineering Contradiction:
Improvesingle-program execution simplicityVSAvoidparallel processing capability
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The data joining process is segmented into independent, parallelizable operations. The system divides the data integration workflow into discrete tasks that can be executed concurrently across multiple processors and users, while maintaining coordination through the centralized metadata capture mechanism.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal data joining framework that serves multiple functions: it supports both single-program execution and massive parallelization, accommodates multiple user participations, and handles various data formats. This multi-functional design resolves the contradiction between operational simplicity and processing capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If data is flattened for integration, then diverse data forms can be joined, but the original data structure and metadata information are lost

Engineering Contradiction:
Improvedata integration capabilityVSAvoidmetadata information loss
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system creates copies of metadata information during the flattening process. The captured metadata is stored separately and can be referenced to reconstruct or validate the original data structure, preventing permanent loss of information while enabling data integration.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

Metadata capture is performed as a preliminary action before data flattening occurs. This ensures that all necessary metadata information is preserved in advance, allowing the flattened data to be integrated without losing the structural context needed for interpretation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10776357B2System and method of data join and metadata configuration
Publication Date: 2020.09.15 INFOSYS LTD
  • US10776357B2 patent drawing
  • US10776357B2 patent drawing
  • US10776357B2 patent drawing

AI summary

A method and system of a data join includes capture of metadata information associated with one of semi-structured data and unstructured data. A flattened structure for one of the semi-structured data and the unstructured data is defined, and an entity is extracted from the unstructured data. Further, one of the semi-structured data and an entity extracted unstructured data are flattened based on the flattened structure, and flattened semi-structured data and flattened entity extracted unstructured data with relational data are joined.