Unified Query Engine for Structured and Unstructured Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data management solutions struggle to handle both structured and unstructured data efficiently, often requiring multiple systems and formats, making it difficult to manage and access diverse data types effectively.
Innovation Solution
A query engine that supports both structured and unstructured data processing through a common interface, using query languages like SQL and SPL, allowing for translation between formats and executing queries across different data types without the need for data transformation, and utilizing a provider network with structured and unstructured data processing services, data storage, and ingestion services to manage and process data in a unified manner.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple different data management and storage solutions are used to handle different data types, then each data type can be processed appropriately, but system complexity increases and data becomes distributed across different locations in different formats
Solution Approach 1:
The patent implements a universal query engine that can process both structured and unstructured data through a single interface. The system uses a unified data model that represents both data types consistently, allowing the same query processing logic to handle different data formats without requiring separate management systems for each data type.
Solution Approach 2:
The patent introduces an intermediary layer consisting of adapters and a unified data model that mediates between diverse data sources and the query engine. These adapters translate different data formats into a common representation, enabling the query engine to interact with various data types through a standardized interface without direct complexity exposure.
2Reliability
If data is stored in different formats across different locations, then each format can be optimized for its specific type, but data access requires multiple different systems
Solution Approach 1:
The query engine is designed to be multi-functional, capable of executing queries against both structured and unstructured data through a single unified interface. This eliminates the need for users to switch between different systems or learn different query languages for different data types, while still maintaining format-specific optimizations internally.
Solution Approach 2:
The system segments the complexity into distinct layers: data storage maintains format-specific optimizations, adapters handle format translation, and the query engine provides unified access. This segmentation allows each layer to be optimized independently while presenting a simplified interface to users.
3Adaptability or versatility
If separate systems are used for structured and unstructured data, then each system can be specialized, but it becomes difficult to select a data management solution that satisfies most storage and processing needs
Solution Approach 1:
The patent creates a unified data management solution that combines the capabilities of separate structured and unstructured data systems into a single platform. The query engine can process various data types using a common query language, eliminating the need for organizations to evaluate and integrate multiple specialized systems while maintaining specialized processing capabilities internally.
4Productivity
If data is distributed across different locations and formats, then data can be stored efficiently, but query execution requires coordination across multiple systems
Solution Approach 1:
The patent uses adapters as intermediaries that translate queries into format-specific operations and translate results into a unified format. This mediation layer handles the coordination complexity internally, allowing efficient distributed storage while presenting a unified query interface that executes queries across multiple data sources without requiring users to manage the coordination manually.
Data Source
AI summary
Queries received at a query engine may be executed for structured data and not-structured data. A query execution plan may be generated for the query that includes stateless operations to apply the query to the not-structured data at remote query processing engines. The remote query processing engines may perform the stateless operations and return results to the query engine. The query engine may generate a result for the query based on the results received from the remote query engine as well as results determined as part of applying the query to structured data. The result to the query may be returned.


