Network Security Event Data Clustering and Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing network security systems generate large and complex data sets of event messages, overwhelming analysts and consuming significant resources, while prior methods for reducing these data sets are inadequate for real-time analysis and may elide crucial information for attack identification and mitigation.
Innovation Solution
A clustering and compressing algorithm that reduces data sets by grouping and summarizing messages using attribute-oriented-induction techniques, with customizable per-attribute thresholds to preserve essential information, allowing for intelligent reduction of redundant messages and identification of relevant signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If network security systems monitor and record all security events, then complete security coverage and detection capability are improved, but data volume and system resource consumption increase significantly
Solution Approach 1:
The patent extracts and removes redundant information from security event data through clustering analysis. By identifying and eliminating duplicate or highly similar events, the system retains only essential security information, reducing data volume while preserving detection capability.
Solution Approach 2:
The patent transforms security event data by changing its representation parameters through clustering. Events are grouped based on similarity metrics, and representative samples are selected to represent entire clusters, fundamentally altering how data is stored and processed while maintaining security integrity.
2Loss of information
If all security messages are displayed to analysts, then complete information availability is improved, but operator attention and decision-making quality deteriorate due to information overload
Solution Approach 1:
The patent merges similar security events into clustered groups, presenting consolidated views to operators. Instead of displaying individual redundant messages, the system combines them into meaningful clusters with representative characteristics, reducing visual clutter while preserving essential security information.
Solution Approach 2:
The patent applies partial action by selectively displaying only the most relevant or representative events from each cluster to operators. Rather than showing all events equally, the system presents a curated subset that provides sufficient information for effective security analysis without overwhelming operators.
3Quantity of substance
If prior art clustering algorithms are used to reduce messages, then data volume reduction is achieved, but critical security information is lost due to insufficient flexibility
Solution Approach 1:
The patent implements dynamic clustering where cluster formation and representation selection adapt based on event characteristics and security context. The system can adjust clustering parameters and representative selection in real-time, ensuring that critical security information is preserved while achieving effective data reduction.
Solution Approach 2:
The patent applies different clustering strategies and information preservation rules to different types of security events based on their local characteristics. Critical events receive special handling to ensure their information is preserved, while less critical events can be more aggressively clustered and reduced.
Data Source
AI summary
This document describes techniques for reducing a size of data sets related to network security alarms or logs, or other messages. Preferably, the reduction is performed via a clustering and compressing algorithm that, among other things, enables an operator to provide customized control in the form of ordered, per-attribute thresholds, or “stop” points. These thresholds function to preserve important information while still achieving excellent clustering and compression results. In some embodiments, the technique described herein can be used to reliably produce reduced-size data sets composed entirely of unique entries. The unique entries can thus be used as keys into a database, e.g., for storage and later analysis or other purposes.


