Distributed Reverse Indexing for Network Flow Logs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for managing flow logs in data centers rely on collecting and processing these logs at a central location, which can be resource-intensive and inefficient, especially as network traffic increases.
Innovation Solution
The solution involves generating metadata at network appliances that indexes the flow logs, allowing these logs and metadata to be transmitted to a central analyzer for merging, thereby creating a searchable flow log database.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If flow logs are collected and processed at a central location, then centralized analysis is achieved, but computational resources and processing time increase significantly
Solution Approach 1:
The patent segments the metadata generation process by creating distributed indexers at each network appliance that independently generate metadata locally. This divides the centralized processing task into multiple distributed units, reducing the computational burden on any single system while maintaining centralized analysis capability through subsequent merging of metadata at the central location.
2Loss of information
If flow logs are collected and processed at a central location, then centralized analysis is achieved, but processing time increases
Solution Approach 1:
The patent applies preliminary action by having network appliances generate metadata and indexes for flow logs before centralized collection. This preprocessing step creates organized, searchable metadata structures in advance, so that when logs are aggregated at the central location, the analysis can proceed much faster without needing to process raw logs from scratch.
3Productivity
If specialized data stores are used for flow logs, then search efficiency is improved, but system complexity increases
Solution Approach 1:
The patent implements self-service by enabling each network appliance to autonomously generate its own metadata and indexes for its flow logs. This distributed self-service approach improves search efficiency locally without requiring complex centralized management, as each appliance independently maintains its own searchable structure that can be merged with others.
Data Source
AI summary
Rather than collecting flow logs at a central location, and then processing these flow logs to create general purpose or specialized data stores, the embodiments herein rely on the network appliances to create the flow logs and metadata that indexes these flow logs. The flow logs and the metadata can then be collected at the central location (e.g., a central analyzer) and merged with flow logs and metadata generated by other network appliances to yield a data store that can be used to analyze the flow logs in computing environment (e.g., a data center).


