Unstructured Data Association via Schema Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data extraction systems face challenges in effectively associating and structuring unstructured data sets with structured data sets, as unstructured data lacks organization according to predefined schemas, making it difficult to identify relevant information and maintain accurate associations.
Innovation Solution
A data extraction system that analyzes unstructured data sets based on groups of structured data sets and schemas, using natural language processing to identify matches and establish associations, allowing for concurrent display and updating of structured data sets from unstructured data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If unstructured data sets are processed without schema-based organization, then data capture flexibility is improved, but data association accuracy deteriorates
Solution Approach 1:
The system performs preliminary analysis of unstructured data sets by comparing them against predefined schemas before final association. This advance processing allows the system to identify potential matches and prepare association candidates, resolving the contradiction by maintaining flexibility in data capture while ensuring accuracy through pre-scheduled schema validation and matching operations.
Solution Approach 2:
The patent introduces schema-based structured data sets as an intermediary layer between unstructured data capture and final data association. This intermediary structure allows flexible unstructured data to be temporarily organized according to schemas, enabling accurate association while preserving the original flexibility of unstructured data input methods.
2Loss of information
If unstructured data sets are analyzed against multiple schemas, then data association completeness is improved, but processing time increases
Solution Approach 1:
The system segments the schema matching process by dividing multiple schemas into separate comparison operations. Instead of analyzing all schemas simultaneously, the system processes schemas in segments or batches, allowing comprehensive data association while managing processing time through divided computational tasks.
Solution Approach 2:
The patent applies partial action by performing schema analysis on subsets of schemas initially, then progressively analyzing additional schemas based on preliminary match results. This approach ensures data association completeness by eventually analyzing all relevant schemas while reducing processing time through staged analysis that focuses computational resources on the most promising matches first.
3Adaptability or versatility
If schema modifications are made to structured data sets, then data organization adaptability is improved, but association maintenance complexity increases
Solution Approach 1:
The system implements feedback mechanisms that automatically detect schema modifications and trigger re-analysis of associated unstructured data sets. When schemas are modified, the system receives feedback about the change, automatically updates associations based on the new schema structure, and maintains data integrity without requiring manual intervention, thus resolving the complexity issue.
Solution Approach 2:
The patent enables the data association system to self-service by automatically detecting schema changes and re-performing association operations without external intervention. The system monitors its own structured data sets for modifications and autonomously maintains accurate associations between unstructured and structured data, reducing maintenance complexity through self-managing capabilities.
Data Source
AI summary
Techniques are disclosed for processing a data set that is not organized according to a schema being used for organizing data (referred to herein as an “unstructured data set”). An unstructured data set is analyzed based on a group of structured data sets that are organized according to the schema. A particular structured data set is determined to be associated with the unstructured data set. The unstructured data set is stored in association with the particular structured data set. Periodically, the unstructured data set is re-analyzed based on a current version of the group of structured data sets. Additionally or alternatively, an unstructured data set is analyzed based on a particular schema of a set of schemas. A subset of information is extracted from the unstructured data set, and stored in accordance with the particular schema. Periodically, the unstructured data set is re-analyzed based on a current version of the set of schemas.


