Unified Query Interface for Structured and Unstructured Data Integration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for accessing structured and unstructured data separately limit the integration of information, failing to establish connections between related entities across different data sources, resulting in incomplete answers when evidence is spread across both types of data.
Innovation Solution
A computer-implemented method and system that uses open domain information extraction to recognize patterns in unstructured data, create schemas, and associate elements with structured data entities based on similarity, establishing links between structured and unstructured data sources for seamless querying and integration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If separate independent queries are performed for structured and unstructured data sources, then each source type can be queried using its native interface, but the integration of information from both sources is limited and connections between related entities cannot be established
Solution Approach 1:
The patent introduces an intermediary layer that translates structured queries into keyword queries for unstructured data sources. This mediator enables seamless querying across both structured and unstructured data without requiring separate query interfaces, while maintaining the ability to establish connections between entities through schema mapping and entity resolution techniques.
Solution Approach 2:
The system implements a universal query interface that can handle both structured queries (e.g., SPARQL) and unstructured data retrieval through a single paradigm. The query processor is designed to evaluate structured queries against structured sources while automatically translating them for evaluation against unstructured text sources, providing multi-functional capability in one system.
2Ease of operation
If information extraction techniques are used to extract structured data from unstructured data, then the problem of accessing both data types is reduced to accessing only structured data, but structured data remains disconnected from other available structured data if extraction is not restricted to a fixed set of relationship types
Solution Approach 1:
The system performs preliminary schema mapping and entity resolution before the actual query execution. By pre-establishing mappings between extracted entities and existing structured schemas, and resolving entity identities across sources in advance, the system simplifies the query access process while maintaining comprehensive connection capabilities without requiring fixed relationship type restrictions.
3Ease of operation
If a common query interface is used for both structured and unstructured data, then a single querying paradigm is involved providing convenient integration at the user interface layer, but only shallow integration at the data layer occurs with no connections established between related entities
Solution Approach 1:
The patent introduces an intermediary layer that translates structured queries into keyword queries for unstructured data sources. This mediator enables seamless querying across both structured and unstructured data without requiring separate query interfaces, while maintaining the ability to establish connections between entities through schema mapping and entity resolution techniques.
Solution Approach 2:
The system implements feedback mechanisms where query results from unstructured data are enriched with connections to structured data entities through entity resolution. The system continuously refines entity mappings and relationships based on query patterns and results, ensuring deep integration and reliable connections between entities across data sources.
Data Source
AI summary
A computer-implemented method, system, and article of manufacture for querying and integrating structured and unstructured data. The method includes: receiving entity information that is extracted from a first set of unstructured data using an open domain information extraction system, wherein the entity in-formation comprises relationship information between a first entity and a second entity of the first set of unstructured data; recognizing a pattern based on the relationship information and creating a schema for the first set of unstructured data based on the pattern; and associating an element of the created schema with (i) an entity of a second set of unstructured data or (ii) a schema element of an existing set of structured data if there is sufficient overall similarity between the created schema element and either the second unstructured data entity or the schema element of the existing structured data.


