Unstructured Data Association via Schema Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data extraction systems face challenges in effectively associating and structuring unstructured data sets with structured data sets, as unstructured data lacks organization according to predefined schemas, making it difficult to identify relevant information and maintain accurate associations.

Innovation Solution

A data extraction system that analyzes unstructured data sets based on groups of structured data sets and schemas, using natural language processing to identify matches and establish associations, allowing for concurrent display and updating of structured data sets from unstructured data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If unstructured data sets are processed without schema-based organization, then data capture flexibility is improved, but data association accuracy deteriorates

Engineering Contradiction:
Improvedata capture flexibilityVSAvoiddata association accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary analysis of unstructured data sets by comparing them against predefined schemas before final association. This advance processing allows the system to identify potential matches and prepare association candidates, resolving the contradiction by maintaining flexibility in data capture while ensuring accuracy through pre-scheduled schema validation and matching operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces schema-based structured data sets as an intermediary layer between unstructured data capture and final data association. This intermediary structure allows flexible unstructured data to be temporarily organized according to schemas, enabling accurate association while preserving the original flexibility of unstructured data input methods.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If unstructured data sets are analyzed against multiple schemas, then data association completeness is improved, but processing time increases

Engineering Contradiction:
Improvedata association completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system segments the schema matching process by dividing multiple schemas into separate comparison operations. Instead of analyzing all schemas simultaneously, the system processes schemas in segments or batches, allowing comprehensive data association while managing processing time through divided computational tasks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by performing schema analysis on subsets of schemas initially, then progressively analyzing additional schemas based on preliminary match results. This approach ensures data association completeness by eventually analyzing all relevant schemas while reducing processing time through staged analysis that focuses computational resources on the most promising matches first.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If schema modifications are made to structured data sets, then data organization adaptability is improved, but association maintenance complexity increases

Engineering Contradiction:
Improvedata organization adaptabilityVSAvoidassociation maintenance complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements feedback mechanisms that automatically detect schema modifications and trigger re-analysis of associated unstructured data sets. When schemas are modified, the system receives feedback about the change, automatically updates associations based on the new schema structure, and maintains data integrity without requiring manual intervention, thus resolving the complexity issue.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent enables the data association system to self-service by automatically detecting schema changes and re-performing association operations without external intervention. The system monitors its own structured data sets for modifications and autonomously maintains accurate associations between unstructured and structured data, reducing maintenance complexity through self-managing capabilities.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10303704B2Processing a data set that is not organized according to a schema being used for organizing data
Publication Date: 2019.05.28 ORACLE INT CORP
  • US10303704B2 patent drawing
  • US10303704B2 patent drawing
  • US10303704B2 patent drawing

AI summary

Techniques are disclosed for processing a data set that is not organized according to a schema being used for organizing data (referred to herein as an “unstructured data set”). An unstructured data set is analyzed based on a group of structured data sets that are organized according to the schema. A particular structured data set is determined to be associated with the unstructured data set. The unstructured data set is stored in association with the particular structured data set. Periodically, the unstructured data set is re-analyzed based on a current version of the group of structured data sets. Additionally or alternatively, an unstructured data set is analyzed based on a particular schema of a set of schemas. A subset of information is extracted from the unstructured data set, and stored in accordance with the particular schema. Periodically, the unstructured data set is re-analyzed based on a current version of the set of schemas.