Query-Time Sampling for Approximate Analytics in Column-Oriented Databases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data analysis systems face challenges in providing fast and accurate exploratory data analysis due to the time-consuming nature of analytical queries in petabyte-scale data sets, which is not suitable for real-time response requirements in business intelligence and public service applications.

Innovation Solution

Implementing query-time sampling with stratified and hash-based samplers as query plan operators within a disk-based column-oriented RDBMS, using Polymorphic Table Function technology to facilitate approximate query processing, allowing for fast and accurate exploratory data analytics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If analytical queries are executed on petabyte-scale data sets using conventional OLAP systems, then comprehensive data analysis can be performed, but the query execution time becomes excessively long (several seconds)

Engineering Contradiction:
Improvedata analysis accuracyVSAvoidquery execution time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by executing analytical queries on a sampled subset of data rather than the complete petabyte-scale dataset. Multiple sampling techniques (uniform random sampling, stratified sampling, and histogram-based sampling) select representative portions of data that capture the essential characteristics needed for exploratory data analysis, thereby reducing query execution time from several seconds to near real-time while preserving sufficient analytical accuracy for EDA purposes

Inventive Principle:
Principle #16Partial or excessive action

2Productivity

If query-time sampling is implemented to reduce execution time, then near real-time response can be achieved, but query result accuracy may be compromised

Engineering Contradiction:
Improvequery response speedVSAvoidquery result accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the sampling parameters dynamically based on the specific query characteristics and data distribution. Different sampling techniques are selected and configured with appropriate parameters (sample size, stratification levels, histogram bins) to optimize the balance between query response speed and result accuracy for each specific EDA task, allowing the system to adapt to different analytical requirements

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If conventional sampling methods are used, then query execution is simplified, but early-row bias and outlier detection issues occur

Engineering Contradiction:
Improvesampling implementation simplicityVSAvoidsampling representativeness
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the sampling process into multiple distinct techniques, each addressing specific sampling challenges. Uniform random sampling handles general cases, stratified sampling addresses early-row bias by dividing data into strata and sampling from each, and histogram-based sampling captures outlier characteristics through data distribution binning. This segmentation allows the system to select the appropriate sampling method based on the specific analytical requirements

Inventive Principle:
Principle #1Segmentation

4Productivity

If approximate query processing is used for fast visualization, then near real-time response is achieved, but existing OLAP systems cannot easily support these requirements

Engineering Contradiction:
Improvevisualization response timeVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent makes the OLAP system universally capable of supporting both exact analytical queries and approximate exploratory queries through a unified architecture. The query optimizer automatically determines whether to execute exact queries or approximate sampled queries based on the query type and requirements, allowing the same system infrastructure to serve multiple functions without requiring separate specialized systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11269886B2Approximate analytics with query-time sampling for exploratory data analysis
Publication Date: 2022.03.08 SAP SE
  • US11269886B2 patent drawing
  • US11269886B2 patent drawing
  • US11269886B2 patent drawing

AI summary

A system and method are disclosed to facilitate exploratory data analytics for an enterprise. A storage area network, for a column-oriented relational database management system, may contain electronic records that store enterprise information. A query engine may receive, from a user via an interactive user interface, query parameters associated with the enterprise information. The query engine may then automatically generate an approximate query for exploratory data analytics using query-time sampling, the approximate query being associated with at least one of: (1) a stratified sampler with randomized row access, and (2) a hash-based, outlier aware join sampler. The approximate query may then be executed in connection with the enterprise information in the storage area network, and results of the executed approximate query may be provided to the user via the user interface.