Network Data Anomaly Detection via Unified Document Grouping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current network analysis and troubleshooting systems face challenges in managing and analyzing vast, complex network data from diverse sources, including issues with data inconsistency, compartmentalization, and the inability to handle unstructured data, leading to difficulties in extracting useful insights for network performance improvement.
Innovation Solution
A method that processes network data records into searchable documents, performs anomaly detection using statistical and quantitative algorithms, and ranks anomalies based on abnormality scores, enabling the identification of anomalous terms, document groups, and documents across different data sources, including structured and unstructured data formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If network data is stored in multiple separate data sources with different formats, then data coverage and comprehensiveness are improved, but data consistency and ease of analysis deteriorate
Solution Approach 1:
The patent merges data from multiple separate network data sources into a unified data structure. The system collects data records from diverse sources including structured data (databases, data streams) and unstructured data (logs, text files), then processes and stores them in a unified format that maintains consistency across all sources while preserving the comprehensive coverage of the original diverse data.
2Speed
If traditional search engines are used to extract data records, then individual data retrieval speed is improved, but effectiveness for network analysis deteriorates due to relevance-based ranking
Solution Approach 1:
The patent changes the ranking parameter from relevance-based (traditional search engines) to anomaly-based ranking. The system calculates anomaly scores for data records based on statistical deviations from normal network behavior patterns, then ranks records by these anomaly scores. This allows the system to maintain fast data retrieval while significantly improving analysis effectiveness for network troubleshooting, as anomalous records are prioritized regardless of keyword relevance.
3Measurement precision
If manual review of data documents is performed, then analysis accuracy is improved, but time consumption increases significantly
Solution Approach 1:
The patent implements automated anomaly detection that performs the analysis work itself without requiring manual review. The system automatically calculates anomaly scores, identifies anomalous data records, and ranks them for operator review. This self-service approach maintains high analysis accuracy through automated statistical analysis while dramatically reducing the time operators need to spend manually reviewing data documents.
4Measurement precision
If existing anomaly detection systems are used, then structured data analysis capability is improved, but ability to handle unstructured data deteriorates
Solution Approach 1:
The patent creates a universal anomaly detection system that can handle multiple data types including structured data (databases, data streams) and unstructured data (logs, text files). The system uses a unified processing approach that applies statistical anomaly detection across all data types, making the system adaptable to diverse data formats while maintaining detection precision through consistent analytical methods.
Data Source
AI summary
A method for analysing performance of a network by managing network data relating to operation of the network is disclosed. The method comprises receiving a plurality of network data records from at least one network data source, processing the received plurality of network data records into a plurality of network data documents, each network data document corresponding to a received network data record, assembling the plurality of network data documents into document groups, generating statistical data for terms appearing in the document groups, and for at least one term, performing anomaly detection upon the statistical data for the term in the document groups, and detecting an anomaly in the term. The method further comprises performing at least one of identifying the term as an anomalous term, identifying the document group containing the anomaly as an anomalous group, and/or identifying a document containing the anomalous term as an anomalous document.


