Database Cluster Time Series Analysis via Semantic Tagging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Legacy solutions fail to effectively process and analyze the vast amounts of raw data from modern database clusters, leading to inefficiencies in identifying system health and performance bottlenecks, with high false alarm rates and missed critical information due to reliance on naive algorithms and inability to handle dynamic changes.

Innovation Solution

Transforming raw sensory and measurement data into meaningful time series signals using semantic tagging and statistical analysis to isolate dominant signals, enabling the application of robust learning models for predicting system health and availability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If legacy solutions use simple threshold techniques to process raw data, then the processing is easy to implement, but the false alarm rate increases and critical information is missed

Engineering Contradiction:
Improveease of implementationVSAvoidfalse alarm rate
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent transforms raw sensory data into time series signals and applies statistical processing techniques to change the parameters of data representation. This allows the system to move beyond simple threshold techniques to more sophisticated analysis methods that reduce false alarms while maintaining implementation feasibility through automated processing pipelines.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces time series signals as an intermediary representation between raw sensory data and system health assessment. This intermediary layer enables sophisticated statistical analysis and machine learning models to process the data, filtering out false alarms while preserving critical information about system health and performance.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If legacy solutions process raw data with simple algorithms, then the system complexity is low, but the ability to discern meaningful information is insufficient

Engineering Contradiction:
Improvesystem complexityVSAvoidinformation discernment capability
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent segments the data processing into distinct stages: raw data collection, time series transformation, statistical analysis, and health assessment. This segmentation allows each stage to be optimized independently, maintaining overall system manageability while enabling sophisticated information extraction through specialized processing modules.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces manual or simple algorithmic processing with automated time series analysis and machine learning models. This substitution enables the system to automatically discern meaningful patterns and information from raw data that would be impossible to identify using simple algorithms, while the automation keeps the system complexity manageable.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Quantity of substance

If the system collects detailed process logs and service measurements, then the data volume increases, but the ability to identify system health state improves

Engineering Contradiction:
Improvedata volumeVSAvoidsystem health detection accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent extracts and transforms the most critical information from the vast volume of raw data into time series signals that represent system health state. This extraction process filters out redundant information while preserving the essential patterns and trends needed for accurate health detection, reducing data volume without sacrificing detection accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameters of the data by transforming raw measurements into time series signals with specific statistical properties. This transformation consolidates detailed process logs into summarized time series representations that maintain the essential information for health detection while significantly reducing the overall data volume.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If the system uses robust learning models to analyze time series data, then the prediction accuracy improves, but the processing time increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing of raw data into time series signals before applying machine learning models. This preliminary action pre-processes and structures the data in a way that reduces the computational burden on the learning models, allowing accurate predictions to be made more efficiently by working with pre-transformed data rather than raw data.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9424288B2Analyzing database cluster behavior by transforming discrete time series measurements
Publication Date: 2016.08.23 ORACLE INT CORP
  • US9424288B2 patent drawing
  • US9424288B2 patent drawing
  • US9424288B2 patent drawing

AI summary

A method, system, and computer program product for analyzing performance of a database cluster. Disclosed are techniques for analyzing performance of components of a database cluster by transforming many discrete event measurements into a time series to identify dominant signals. The method embodiment commences by sampling the database cluster to produce a set of timestamped events, then pre-processing the timestamped events by tagging at least some of the timestamped events with a semantic tag drawn from a semantic dictionary and formatting the set of timestamped events into a time series where a time series entry comprises a time indication and a plurality of values corresponding to signal state values. Further techniques are disclosed for identifying certain signals from the time series to which is applied various statistical measurement criteria in order to isolate a set of candidate signals which are then used to identify indicative causes of database cluster behavior.