Distribution Function Pattern Recognition for Data Granularity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Business Intelligence (BI) tools face limitations in predicting and simulating object values due to the unavailability of raw data at desired granularity, caused by memory constraints in databases, leading to influenced distribution functions that do not accurately represent the object's distribution without the influence of other attributes.
Innovation Solution
A method is developed to compare influenced distribution functions with other distribution functions, determine correlations, and extract the raw distribution function by subtracting the influencing distribution function, allowing for the isolation of the object's distribution without external influences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored in database with memory constraints, then storage capacity is improved, but data granularity and availability are worsened
Solution Approach 1:
The patent extracts the distribution function pattern from the aggregated data stored in the database, separating the essential statistical characteristics from the raw data that cannot be fully stored. This allows the system to retrieve and use distribution function information without requiring access to the complete granular raw data, thus resolving the contradiction between storage capacity and data granularity availability.
2Productivity
If data is aggregated to save database resources, then storage efficiency is improved, but analytical accuracy is worsened
Solution Approach 1:
The patent transforms the raw data into a different parameter representation - the distribution function - which captures the essential statistical characteristics (mean, variance, shape) of the data. This parameter transformation allows aggregated data to retain sufficient information for accurate analytical predictions and simulations, resolving the contradiction between storage efficiency and analytical accuracy.
3Device complexity
If influenced distribution function is used for prediction, then computational simplicity is improved, but prediction accuracy is worsened
Solution Approach 1:
The patent introduces the raw distribution function as an intermediary between the influenced distribution function (from aggregated data) and the prediction process. By first extracting the raw distribution function that represents the object's inherent distribution without external influences, and then using this as the basis for predictions, the system achieves both computational feasibility and high prediction accuracy, resolving the contradiction between computational simplicity and prediction accuracy.
Data Source
AI summary
Various embodiments of systems and methods for pattern recognition of a distribution function are described herein. An influenced distribution function corresponding to an influenced attribute is compared with other distribution functions corresponding to other attributes. Based on the comparison, a correlation is determined between the influenced distribution function and an influencing distribution function from the other distribution functions. Based on the determination, a raw distribution function corresponding to an influenced attribute is extracted using the influenced distribution function and the influencing distribution function. The extracted raw distribution function and the influencing distribution function may be classified.


