Skew Detection in Massively Parallel Processing Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Massively parallel processing systems face inefficiencies due to unacceptably skewed query processing, where conventional manual analysis by experienced analysts is time-consuming and delays the identification of queries needing tuning.
Innovation Solution
An electronic method to identify unacceptably skewed queries by extracting processing data from computer logs, filtering unsuitable queries, and calculating skew using equations based on CPU and I/O operations, flagging queries for tuning when actual skew exceeds acceptable levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual analysis by experienced analysts is used to determine acceptable skew, then accuracy in identifying skewed queries is improved, but analysis time and operational complexity increase significantly
Solution Approach 1:
The system performs self-diagnosis by automatically extracting processing data from logs, calculating actual skew values using the formula (max processing time - min processing time) / average processing time, and comparing against acceptable thresholds without requiring external analyst intervention. This automated self-assessment resolves the contradiction by eliminating manual analysis time while maintaining identification accuracy through systematic computational methods.
Solution Approach 2:
The patent replaces the mechanical manual analysis process with an automated electronic system that extracts data from logs, computes skew metrics using standardized formulas, and flags queries for tuning. This substitution eliminates the time-consuming human review process while maintaining measurement precision through consistent algorithmic evaluation of processing time variations across processors.
2Measurement precision
If manual analysis by experienced analysts is used to determine acceptable skew, then expertise-based judgment is improved, but device complexity and operational difficulty increase
Solution Approach 1:
The system encapsulates expert knowledge within automated algorithms that extract processing data from logs, calculate skew values using standardized formulas, and compare against pre-established acceptable thresholds. This self-service approach replaces complex human expertise with systematic computational procedures, reducing operational complexity while maintaining judgment quality through consistent application of tuning criteria.
Solution Approach 2:
The patent transforms the complex qualitative judgment process into quantitative parameter evaluation by calculating specific skew metrics (actual skew value and acceptable skew threshold) from processing time data. This parameter-based approach simplifies the operational process by reducing expert judgment to straightforward numerical comparisons, thereby reducing device complexity while preserving measurement precision.
3Productivity
If queries are not identified and tuned promptly, then processing efficiency deteriorates due to backed up queries, but automated identification systems add complexity to the system
Solution Approach 1:
The patent implements an automated electronic system that substitutes manual monitoring with continuous log analysis, automatic skew calculation using processing time data, and real-time flagging of queries exceeding acceptable thresholds. This automation maintains high processing efficiency by promptly identifying skewed queries while managing system complexity through standardized algorithms and integrated log extraction capabilities.
Solution Approach 2:
The system establishes a feedback loop by continuously extracting processing data from logs, calculating actual skew values, comparing against acceptable thresholds, and flagging queries for tuning when skew exceeds limits. This automated feedback mechanism maintains processing efficiency by providing real-time identification of problematic queries while managing complexity through systematic data flow and decision logic embedded in the electronic identification process.
Data Source
AI summary
Apparatus and methods for determination of unacceptable skew for query in a massively parallel processing system. The apparatus and methods may use data associated with processing the query that has been stored in computer logs. The processing data may be used to determine the actual level of skew for the query. The apparatus and methods may calculate an acceptable level of skew. If the actual skew exceeds the acceptable skew, the query may be considered unacceptably skewed and may be flagged for tuning.


