Distributed Key-Value Repository for Low-Latency Data Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data warehousing systems are inadequate for handling large-scale data sets, experiencing high latency in searches, suffering from data siloing, and losing original data context, which hampers efficient data analysis and cyber security investigations.
Innovation Solution
A distributed key-value data repository system that ingests data from disparate sources, provides efficient indexing, and supports low-latency searches by using a horizontally-scalable architecture, allowing data to be stored in its original form and enabling flexible, adaptive querying across large volumes of constantly updated data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional data warehousing systems are used to aggregate and analyze large amounts of data, then data consolidation is achieved, but search latency increases to hours or days
Solution Approach 1:
The patent segments the monolithic data warehouse into multiple distributed database systems that can operate independently and in parallel. Each database system handles a portion of the data, enabling simultaneous processing of multiple search queries across different data segments, thereby reducing overall search latency while maintaining the ability to consolidate large volumes of data.
2Productivity
If data is distributed across multiple disparate database systems, then data consolidation is reduced, but system complexity and integration requirements increase
Solution Approach 1:
The patent introduces a data exchange mechanism that acts as an intermediary between disparate database systems. This intermediary enables standardized data sharing and communication protocols, allowing multiple database systems to operate independently while still achieving coordinated data analysis, thus reducing integration complexity while maintaining search efficiency.
3Adaptability or versatility
If data is transformed during analysis, then data compatibility across systems is improved, but original data context is lost
Solution Approach 1:
The patent creates and maintains copies of original data across the distributed database systems in their native formats. Instead of transforming original data, the system replicates data copies that can be queried independently, preserving the original data context and format while enabling compatibility through standardized access mechanisms.
4Ease of operation
If custom information technology components are developed to integrate disparate database systems, then data exchange between systems is enabled, but development time and cost increase
Solution Approach 1:
The patent develops a universal data exchange mechanism that can interface with multiple types of disparate database systems through standardized protocols. This multi-functional approach eliminates the need for custom integration components for each database pair, reducing development time and cost while maintaining ease of data exchange across different system types.
Data Source
AI summary
A data analysis system is proposed for providing fine-grained low latency access to high volume input data from possibly multiple heterogeneous input data sources. The input data is parsed, optionally transformed, indexed, and stored in a horizontally-scalable key-value data repository where it may be accessed using low latency searches. The input data may be compressed into blocks before being stored to minimize storage requirements. The results of searches present input data in its original form. The input data may include access logs, call data records (CDRs), e-mail messages, etc. The system allows a data analyst to efficiently identify information of interest in a very large dynamic data set up to multiple petabytes in size. Once information of interest has been identified, that subset of the large data set can be imported into a dedicated or specialized data analysis system for an additional in-depth investigation and contextual analysis.


