Distributed Pseudo-Random Subset Generation for Low-Latency Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing database analytic tools are inefficient, costly, and require substantial configuration and training, struggling to handle the exponential growth of data effectively in low-latency data analysis systems.
Innovation Solution
Implementing distributed pseudo-random subset generation methods in low-latency data analysis systems using distributed in-memory databases, which involve pseudo-random filtering and bitmask techniques to efficiently process data queries by reducing the cardinality of rows and improving resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If distributed pseudo-random subset generation is implemented, then resource utilization and performance are improved, but the complexity of data processing increases
Solution Approach 1:
The patent segments the data processing task by generating pseudo-random subsets of rows from the distributed in-memory database. Instead of processing the entire dataset, the system divides it into manageable random subsets that can be analyzed independently, reducing the overall complexity while maintaining analytical effectiveness.
Solution Approach 2:
The patent introduces pseudo-random subset generation as an intermediary step between data retrieval and full data analysis. This intermediary mechanism filters and reduces the data volume before it reaches the analysis stage, improving resource utilization without requiring complex changes to the underlying database architecture.
2Loss of energy
If the amount of data transported and stored is reduced, then resource utilization improves, but measurement precision of analysis results may deteriorate
Solution Approach 1:
The patent employs dynamic pseudo-random subset generation where the subset composition changes based on query requirements. This dynamic approach allows the system to adapt the data sampling strategy to specific analysis needs, maintaining measurement precision while reducing data transport costs through intelligent, context-aware subset selection.
Solution Approach 2:
The patent changes the parameter of data representation by transforming full dataset queries into pseudo-random subset queries. By modifying the query execution parameters to work with reduced data subsets rather than complete datasets, the system achieves lower data transport costs while preserving analytical accuracy through proper subset generation methods.
Data Source
AI summary
Distributed pseudo-random subset generation includes obtaining a data-query indicating a first table having a first column including unique values, a second table having a second column including unique values, a join clause joining the first table and the second table on the first column and the second column, and a limit value, pseudo-random filtering the first table to obtain left intermediate data and left filtering criteria, pseudo-random filtering the second table to obtain right intermediate data and right filtering criteria, obtaining intermediate results data by full outer joining the left intermediate data and the right intermediate data, obtaining results data by filtering the intermediate results data using most-restrictive filtering criteria among the left filtering criteria and the right filtering criteria, and outputting the results data, wherein outputting the results data includes limiting the cardinality of rows of the results data to be at most the limit value.


