Network Security Anomaly Alerting via Data Lake Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Analyzing and searching massive quantities of machine data generated from diverse sources poses challenges due to its vastness and complexity, as existing tools typically pre-process and discard data, limiting analysis to pre-specified sets rather than allowing investigation of all data.

Innovation Solution

An event-based data intake and query system with a late-binding schema that collects, indexes, and stores machine data as events, enabling flexible schema development and extraction rules application at search time, allowing for field-searchability and correlation across disparate data sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data is pre-processed and only specified data items are extracted and stored, then data retrieval and analysis efficiency is improved, but the ability to analyze all generated data is lost

Engineering Contradiction:
Improvedata retrieval and analysis efficiencyVSAvoiddiscarded data
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent segments data into two storage paths: extracted data items stored in traditional data systems for efficient querying, and raw data stored separately in data lakes for comprehensive analysis. This segmentation allows both efficiency and completeness to coexist by directing different data types to appropriate storage locations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a nested storage architecture where extracted data items are stored within the structured data system, while raw data is stored in the broader data lake environment. The nested doll principle is applied by placing the extracted data structure within the larger raw data context, enabling both detailed analysis and comprehensive review.

Inventive Principle:
Principle #7Nested doll (Nesting)

2Adaptability or versatility

If massive quantities of raw data are stored for later retrieval and analysis, then flexibility and ability to analyze all data are improved, but storage costs and data management complexity increase

Engineering Contradiction:
Improveflexibility in data analysisVSAvoiddata management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces data lakes as an intermediary storage layer between data generation and analysis. Data lakes serve as a mediator that accepts raw data in various formats, preserves it without pre-processing, and makes it accessible for later analysis while managing the complexity of storing diverse data types.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the storage parameter from structured, pre-processed data to unstructured or minimally processed raw data. By altering the data format and organization parameters, the system enables flexible analysis of all data while using data lakes to manage the complexity through appropriate storage mechanisms.

Inventive Principle:
Principle #35Parameter changes

3Speed

If traditional data systems pre-process data based on anticipated analysis needs, then data retrieval speed is improved, but the ability to derive new insights from unexpected data patterns is reduced

Engineering Contradiction:
Improvedata retrieval speedVSAvoidunexpected data patterns
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

The patent applies preliminary action by extracting and storing key data items in advance in traditional data systems for fast retrieval, while simultaneously preserving raw data in data lakes. This preliminary processing enables both quick access to known important data and the ability to later discover unexpected patterns in the preserved raw data.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20220141188A1Network Security Selective Anomaly Alerting
Publication Date: 2022.05.05 CISCO TECHNOLOGY INC
  • US20220141188A1 patent drawing
  • US20220141188A1 patent drawing
  • US20220141188A1 patent drawing

AI summary

Described herein, is a technique of data reduction and focusing for system and network security. Anomaly alerts pertain to specific risk objects that are network devices or users that triggered the associated anomaly. Threat objects are entities used by the risk object that include the specific activity of the risk object that triggered the anomaly. Once identified, threat objects are linked to the risk objects that they respectively pertain to. The link between a risk object and a threat object is generated via searchable metadata. Through linking, relationships are built between threat objects and risk objects. Links are between a number (N) risk objects and a number (M) of threat objects. The relationships are surfaced to a user based on satisfaction of predetermined thresholds. Examples of display to the user may include generation of a threat report, anomaly alerts, or graphical presentations depicting the links in the relationship(s). Where alerts are limited (via searches or reports) to relationships between threat objects and risk objects that are of a predetermined character, the excessive amount of data is reduced to a manageable number of notices.