Uniform and Outlier Row Sampling for File Display
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems fail to effectively display uniform and outlier rows from a file, leading to inefficient processing and storage issues, as they do not adequately differentiate and prioritize rows based on their characteristics.
Innovation Solution
A method and system that identify uniform and outlier rows by comparing characteristics to expected values, using reservoir and congressional sampling techniques to select and display subsets of rows, prioritizing storage and display based on likelihood of being displayed, and dynamically adjusting criteria to optimize processing speed and storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all rows from a file are processed and stored for display, then complete data representation is achieved, but storage requirements and processing time increase significantly
Solution Approach 1:
The patent extracts only the most representative rows (uniform and outlier rows) from the file rather than processing all rows. By identifying and extracting specific subsets of rows that capture the essential characteristics of the data, the system achieves effective data representation with significantly reduced storage requirements.
Solution Approach 2:
The patent applies partial action by processing only a portion of the rows (uniform and outlier rows) rather than all rows. This selective approach maintains the essential information needed for data representation while avoiding the excessive processing and storage of redundant uniform rows.
2Loss of information
If more rows are kept in memory for display, then display completeness improves, but memory usage and processing overhead increase
Solution Approach 1:
The patent extracts only the necessary rows (uniform and outlier rows) that contribute meaningfully to display completeness. By removing redundant rows from consideration, the system maintains informative displays while reducing memory usage and processing overhead.
Solution Approach 2:
The patent applies different quality standards to different rows by identifying uniform rows (representing typical data) and outlier rows (representing exceptional data). This local differentiation allows the system to prioritize rows based on their informational value, improving processing efficiency while maintaining display completeness.
3Quantity of substance
If uniform rows are heavily sampled to reduce data volume, then storage efficiency improves, but risk of losing representative uniform patterns increases
Solution Approach 1:
The patent uses reservoir sampling to create a representative copy of uniform rows that preserves the statistical properties of the original dataset. This sampling method ensures that the reduced set of uniform rows accurately reflects the distribution and characteristics of all uniform rows in the file.
Solution Approach 2:
The patent changes the sampling parameters dynamically by adjusting the sample size and selection criteria based on the characteristics of uniform rows. This allows the system to optimize between data volume reduction and representation accuracy by tuning the sampling process to the specific dataset.
4Quantity of substance
If outlier rows are selectively sampled based on deviation from median, then storage efficiency improves, but risk of missing significant outliers increases
Solution Approach 1:
The patent applies different sampling strategies to different types of rows by identifying outlier rows based on their deviation from the median. This local quality approach ensures that rows with significant deviations are captured with higher priority, maintaining outlier detection accuracy while improving storage efficiency through selective sampling.
Solution Approach 2:
The patent uses congressional sampling to create a representative copy of outlier rows that preserves the distribution of deviations from the median. This sampling method ensures that the reduced set of outlier rows accurately reflects the variety and significance of outliers in the original dataset.
5Measurement precision
If processing continues until all rows are analyzed, then complete classification is achieved, but processing time increases
Solution Approach 1:
The patent performs preliminary classification by identifying uniform and outlier rows as they are read from the file, without waiting to process all rows. This preliminary action allows the system to start sampling and preparing data for display before the complete file is processed, significantly reducing total processing time while maintaining classification accuracy.
Solution Approach 2:
The patent maintains continuous useful action by processing and classifying rows in a single pass through the file, without requiring multiple passes or waiting for complete data collection. This continuous processing approach achieves complete classification accuracy while minimizing processing time.
Data Source
AI summary
A system and method identifies rows in a file as uniform rows or outlier rows based on statistics from the file, and displays a sampling of uniform and outlier rows.


