Unstructured Data Clustering via Automatic Attribute Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems lack an efficient method for processing and analyzing unstructured data, particularly in managed infrastructures, as they require different processes for various formats and lack a single consumer that can understand multiple formats, making it difficult to decompose events and extract meaningful insights.
Innovation Solution
A system that executes automatic attribute inference, comprising a processor, memory, an extraction engine, and a signaliser engine with NMF, k-means clustering, and topology proximity engines, which receives data from managed infrastructures, determines common characteristics, and produces clusters of events, enabling better understanding and analysis of unstructured data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional structured data processing methods are used, then data organization is simple and database integration is easy, but the system cannot effectively process unstructured data from managed infrastructures
Solution Approach 1:
The system employs a universal extraction engine that can process multiple data formats (structured, semi-structured, and unstructured) through a single interface. This engine uses configurable parsers and normalization routines to handle diverse event formats from different managed infrastructure sources, eliminating the need for separate processing pipelines for each data type while maintaining system integrity
Solution Approach 2:
The patent introduces an intermediary layer consisting of the extraction engine and signalizer engine that sits between raw data sources and the analysis system. This intermediary automatically infers attributes, normalizes data structures, and converts unstructured events into a standardized format, bridging the gap between diverse data formats and the structured processing requirements without exposing complexity to end users
2Measurement precision
If multiple separate processes are used for different data formats, then each process can be optimized for its specific format, but the system lacks a single consumer that can understand multiple formats
Solution Approach 1:
The system merges attribute inference, data normalization, and event clustering functions into an integrated processing pipeline. The extraction engine combines multiple inference techniques (statistical analysis, pattern matching, and contextual reasoning) to accurately determine event attributes, while the signalizer engine simultaneously performs clustering and anomaly detection, reducing processing steps while maintaining or improving accuracy
Solution Approach 2:
The system performs preliminary attribute inference and data normalization automatically as data enters the system, before analysis or storage operations. This pre-processing stage infers missing attributes, validates data integrity, and standardizes formats in advance, eliminating the need for manual intervention or repeated processing and thereby improving both accuracy and overall productivity
3Loss of information
If unstructured data is processed without automatic attribute inference, then processing speed is maintained, but meaningful insights and event clustering cannot be extracted
Solution Approach 1:
The system replaces manual attribute definition and data tagging processes with automated machine learning-based attribute inference. The signalizer engine uses unsupervised learning algorithms to automatically identify patterns, infer event attributes, and cluster similar events without human intervention, preserving information that would otherwise be lost while reducing the time investment required for data preparation and analysis
Data Source
AI summary
A system texecutes automatic attribute inference and includes: a processor; a memory coupled to the memory; a first engine that executes automatic attribute inference; an extraction engine in communication with a managed infrastructure and the first engine, the extraction engine configured to receive managed infrastructure data; and a signaliser engine that includes one or more of an NMF engine, a k-means clustering engine and a topology proximity engine, the signaliser engine inputting a list of devices and a list a connections between components or nodes in the managed infrastructure, the signaliser engine determining one or more common characteristics and produces one or more dusters of events.


