Query-Time Sampling for Approximate Analytics in Column-Oriented Databases
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data analysis systems face challenges in providing fast and accurate exploratory data analysis due to the time-consuming nature of analytical queries in petabyte-scale data sets, which is not suitable for real-time response requirements in business intelligence and public service applications.
Innovation Solution
Implementing query-time sampling with stratified and hash-based samplers as query plan operators within a disk-based column-oriented RDBMS, using Polymorphic Table Function technology to facilitate approximate query processing, allowing for fast and accurate exploratory data analytics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If analytical queries are executed on petabyte-scale data sets using conventional OLAP systems, then comprehensive data analysis can be performed, but the query execution time becomes excessively long (several seconds)
Solution Approach 1:
The patent applies partial action by executing analytical queries on a sampled subset of data rather than the complete petabyte-scale dataset. Multiple sampling techniques (uniform random sampling, stratified sampling, and histogram-based sampling) select representative portions of data that capture the essential characteristics needed for exploratory data analysis, thereby reducing query execution time from several seconds to near real-time while preserving sufficient analytical accuracy for EDA purposes
2Productivity
If query-time sampling is implemented to reduce execution time, then near real-time response can be achieved, but query result accuracy may be compromised
Solution Approach 1:
The patent changes the sampling parameters dynamically based on the specific query characteristics and data distribution. Different sampling techniques are selected and configured with appropriate parameters (sample size, stratification levels, histogram bins) to optimize the balance between query response speed and result accuracy for each specific EDA task, allowing the system to adapt to different analytical requirements
3Ease of operation
If conventional sampling methods are used, then query execution is simplified, but early-row bias and outlier detection issues occur
Solution Approach 1:
The patent segments the sampling process into multiple distinct techniques, each addressing specific sampling challenges. Uniform random sampling handles general cases, stratified sampling addresses early-row bias by dividing data into strata and sampling from each, and histogram-based sampling captures outlier characteristics through data distribution binning. This segmentation allows the system to select the appropriate sampling method based on the specific analytical requirements
4Productivity
If approximate query processing is used for fast visualization, then near real-time response is achieved, but existing OLAP systems cannot easily support these requirements
Solution Approach 1:
The patent makes the OLAP system universally capable of supporting both exact analytical queries and approximate exploratory queries through a unified architecture. The query optimizer automatically determines whether to execute exact queries or approximate sampled queries based on the query type and requirements, allowing the same system infrastructure to serve multiple functions without requiring separate specialized systems
Data Source
AI summary
A system and method are disclosed to facilitate exploratory data analytics for an enterprise. A storage area network, for a column-oriented relational database management system, may contain electronic records that store enterprise information. A query engine may receive, from a user via an interactive user interface, query parameters associated with the enterprise information. The query engine may then automatically generate an approximate query for exploratory data analytics using query-time sampling, the approximate query being associated with at least one of: (1) a stratified sampler with randomized row access, and (2) a hash-based, outlier aware join sampler. The approximate query may then be executed in connection with the enterprise information in the storage area network, and results of the executed approximate query may be provided to the user via the user interface.


