Distributed Key-Value Repository for Low-Latency Data Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data warehousing systems are inadequate for handling large-scale data sets, experiencing high latency in searches, suffering from data siloing, and losing original data context, which hampers efficient data analysis and cyber security investigations.

Innovation Solution

A distributed key-value data repository system that ingests data from disparate sources, provides efficient indexing, and supports low-latency searches by using a horizontally-scalable architecture, allowing data to be stored in its original form and enabling flexible, adaptive querying across large volumes of constantly updated data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional data warehousing systems are used to aggregate and analyze large amounts of data, then data consolidation is achieved, but search latency increases to hours or days

Engineering Contradiction:
Improvedata volumeVSAvoidsearch latency
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent segments the monolithic data warehouse into multiple distributed database systems that can operate independently and in parallel. Each database system handles a portion of the data, enabling simultaneous processing of multiple search queries across different data segments, thereby reducing overall search latency while maintaining the ability to consolidate large volumes of data.

Inventive Principle:
Principle #1Segmentation

2Productivity

If data is distributed across multiple disparate database systems, then data consolidation is reduced, but system complexity and integration requirements increase

Engineering Contradiction:
Improvesearch efficiencyVSAvoidsystem integration complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a data exchange mechanism that acts as an intermediary between disparate database systems. This intermediary enables standardized data sharing and communication protocols, allowing multiple database systems to operate independently while still achieving coordinated data analysis, thus reducing integration complexity while maintaining search efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If data is transformed during analysis, then data compatibility across systems is improved, but original data context is lost

Engineering Contradiction:
Improvedata compatibilityVSAvoidoriginal data context
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent creates and maintains copies of original data across the distributed database systems in their native formats. Instead of transforming original data, the system replicates data copies that can be queried independently, preserving the original data context and format while enabling compatibility through standardized access mechanisms.

Inventive Principle:
Principle #26Copying

4Ease of operation

If custom information technology components are developed to integrate disparate database systems, then data exchange between systems is enabled, but development time and cost increase

Engineering Contradiction:
Improvedata exchange capabilityVSAvoiddevelopment time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent develops a universal data exchange mechanism that can interface with multiple types of disparate database systems through standardized protocols. This multi-functional approach eliminates the need for custom integration components for each database pair, reducing development time and cost while maintaining ease of data exchange across different system types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11392550B2System and method for investigating large amounts of data
Publication Date: 2022.07.19 PALANTIR TECHNOLOGIES INC
  • US11392550B2 patent drawing
  • US11392550B2 patent drawing
  • US11392550B2 patent drawing

AI summary

A data analysis system is proposed for providing fine-grained low latency access to high volume input data from possibly multiple heterogeneous input data sources. The input data is parsed, optionally transformed, indexed, and stored in a horizontally-scalable key-value data repository where it may be accessed using low latency searches. The input data may be compressed into blocks before being stored to minimize storage requirements. The results of searches present input data in its original form. The input data may include access logs, call data records (CDRs), e-mail messages, etc. The system allows a data analyst to efficiently identify information of interest in a very large dynamic data set up to multiple petabytes in size. Once information of interest has been identified, that subset of the large data set can be imported into a dedicated or specialized data analysis system for an additional in-depth investigation and contextual analysis.