Unstructured Data Clustering via Automatic Attribute Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems lack an efficient method for processing and analyzing unstructured data, particularly in managed infrastructures, as they require different processes for various formats and lack a single consumer that can understand multiple formats, making it difficult to decompose events and extract meaningful insights.

Innovation Solution

A system that executes automatic attribute inference, comprising a processor, memory, an extraction engine, and a signaliser engine with NMF, k-means clustering, and topology proximity engines, which receives data from managed infrastructures, determines common characteristics, and produces clusters of events, enabling better understanding and analysis of unstructured data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional structured data processing methods are used, then data organization is simple and database integration is easy, but the system cannot effectively process unstructured data from managed infrastructures

Engineering Contradiction:
Improvedata format compatibilityVSAvoidprocessing system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs a universal extraction engine that can process multiple data formats (structured, semi-structured, and unstructured) through a single interface. This engine uses configurable parsers and normalization routines to handle diverse event formats from different managed infrastructure sources, eliminating the need for separate processing pipelines for each data type while maintaining system integrity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary layer consisting of the extraction engine and signalizer engine that sits between raw data sources and the analysis system. This intermediary automatically infers attributes, normalizes data structures, and converts unstructured events into a standardized format, bridging the gap between diverse data formats and the structured processing requirements without exposing complexity to end users

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple separate processes are used for different data formats, then each process can be optimized for its specific format, but the system lacks a single consumer that can understand multiple formats

Engineering Contradiction:
Improveattribute inference accuracyVSAvoiddata processing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system merges attribute inference, data normalization, and event clustering functions into an integrated processing pipeline. The extraction engine combines multiple inference techniques (statistical analysis, pattern matching, and contextual reasoning) to accurately determine event attributes, while the signalizer engine simultaneously performs clustering and anomaly detection, reducing processing steps while maintaining or improving accuracy

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system performs preliminary attribute inference and data normalization automatically as data enters the system, before analysis or storage operations. This pre-processing stage infers missing attributes, validates data integrity, and standardizes formats in advance, eliminating the need for manual intervention or repeated processing and thereby improving both accuracy and overall productivity

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If unstructured data is processed without automatic attribute inference, then processing speed is maintained, but meaningful insights and event clustering cannot be extracted

Engineering Contradiction:
Improveinformation retentionVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system replaces manual attribute definition and data tagging processes with automated machine learning-based attribute inference. The signalizer engine uses unsupervised learning algorithms to automatically identify patterns, infer event attributes, and cluster similar events without human intervention, preserving information that would otherwise be lost while reducing the time investment required for data preparation and analysis

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11924018B2System for decomposing events and unstructured data
Publication Date: 2024.03.05 DELL PROD LP
  • US11924018B2 patent drawing
  • US11924018B2 patent drawing
  • US11924018B2 patent drawing

AI summary

A system texecutes automatic attribute inference and includes: a processor; a memory coupled to the memory; a first engine that executes automatic attribute inference; an extraction engine in communication with a managed infrastructure and the first engine, the extraction engine configured to receive managed infrastructure data; and a signaliser engine that includes one or more of an NMF engine, a k-means clustering engine and a topology proximity engine, the signaliser engine inputting a list of devices and a list a connections between components or nodes in the managed infrastructure, the signaliser engine determining one or more common characteristics and produces one or more dusters of events.