Big Data Activity Analysis Platform for Proactive Issue Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Big data systems are complex, making it difficult for developers and operations staff to identify, diagnose, and fix issues related to application behavior, resource allocation, data layout, and job scheduling, leading to reduced productivity and inefficiencies.

Innovation Solution

A system and method for analyzing big data activities, comprising a distributed file system and a data processing platform with an application manager that identifies slow-running, failed, killed, or malfunctioning applications, and provides a holistic view to proactively address issues, including a suite of managers for different types of applications and a user-friendly interface for deep analysis and solution recommendations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a data processing platform with multiple application managers is implemented to monitor and analyze big data activities, then the ability to identify and diagnose application issues is improved, but the system complexity increases

Engineering Contradiction:
Improveissue identification capabilityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system divides the monitoring function into specialized application managers for different big data technologies (MapReduce manager, Hive manager, Spark manager, HBase manager, Storm manager, Solr manager). Each manager focuses on specific application types, enabling precise issue identification without requiring a single complex monolithic system. This segmentation allows the platform to gather and analyze data specific to each application type while maintaining manageable complexity through modular design.

Inventive Principle:
Principle #1Segmentation

2Productivity

If comprehensive data gathering and analysis is performed on application behavior, resource allocation, and job scheduling, then productivity is improved through proactive issue identification, but the time and resources required for data processing increase

Engineering Contradiction:
Improvedeveloper productivityVSAvoiddata processing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system continuously gathers and pre-processes data about application behavior, resource allocation, and job scheduling in the background before issues occur. The application managers maintain ongoing monitoring of key metrics, so when problems arise, the analysis is already prepared or can be quickly performed. This preliminary data collection enables proactive issue identification and reduces the time needed for diagnostic analysis when productivity-critical events occur.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If the system provides detailed analysis and solution recommendations through a user interface, then ease of operation is improved for developers, but the amount of information processing and storage required increases

Engineering Contradiction:
Improvedeveloper ease of useVSAvoiddata storage requirements
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The system extracts and presents only the most relevant information to developers through the user interface, rather than displaying all collected data. The application managers filter and synthesize monitoring data to highlight critical issues, performance bottlenecks, and actionable recommendations. This extraction approach provides developers with ease of operation by presenting curated, actionable insights while avoiding the need to store and display every raw data point, thus managing storage requirements efficiently.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10521326B2System and method for analyzing big data activities
Publication Date: 2019.12.31 UNRAVEL DATA SYSTEMS INC
  • US10521326B2 patent drawing
  • US10521326B2 patent drawing
  • US10521326B2 patent drawing

AI summary

A system and method for analyzing big data activities are disclosed. According to one embodiment, a system comprises a distributed file system for the entities and applications, wherein the applications include one or more of script applications, structured query language (SQL) applications, Not Only (NO) SQL applications, stream applications, search applications, and in-memory applications. The system further comprises a data processing platform that gathers, analyzes, and stores data relating to entities and applications. The data processing platform includes an application manager having one or more of a MapReduce Manage, a script applications manager, a structured query language (SQL) applications manager, a Not Only (NO) SQL applications manager, a stream applications manager, a search applications manager, and an in-memory applications manager. The application manager identifies if the applications are one or more of slow-running, failed, killed, unpredictable, and malfunctioning.