Relationship Graph Failover in Holistic Data Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Enterprises face challenges in managing and correlating data from diverse sources due to format differences and security limitations, which are exacerbated by cluster failures in distributed databases, leading to difficulties in maintaining data uptime and relationships.
Innovation Solution
A holistic data management platform processes data from various sources, generates relationship graphs, and employs a detach-and-promote failover process to maintain graph availability during cluster failures, using encryption and tokenization to handle sensitive data and leveraging machine learning for relationship identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data from different third-party providers is stored in separate databases, then data security and format compatibility are maintained, but data relationship correlation and system reliability deteriorate
Solution Approach 1:
The patent merges data from multiple third-party providers into a unified distributed database system with a centralized relationship graph. This allows data to be stored in its original formats while enabling cross-source relationship correlation through the unified graph structure, resolving the contradiction between maintaining data security/separation and enabling data correlation.
Solution Approach 2:
The relationship graph acts as an intermediary layer between the distributed database clusters and the data relationship analysis. It mediates the connection between disparate data sources, enabling correlation without requiring direct integration of the underlying databases, thus reducing system complexity while maintaining reliability.
2Reliability
If a primary cluster is used to store relationship graphs, then data access efficiency is improved, but system reliability deteriorates due to single point of failure
Solution Approach 1:
The system segments the relationship graph storage across multiple database clusters, with each cluster holding a portion of the graph data. This segmentation eliminates the single point of failure while maintaining efficient access through distributed query routing, resolving the contradiction between reliability and operational simplicity.
Solution Approach 2:
The system dynamically changes the operational parameters of cluster operations based on failure detection. When a primary cluster fails, the system automatically adjusts by promoting a secondary cluster, changing the access parameters to route queries to the new primary, thus maintaining reliability without permanently complicating data access.
3Productivity
If manual processes are used to standardize and collect data, then data format control is improved, but time consumption and productivity deteriorate
Solution Approach 1:
The system implements self-service data standardization through automated format detection and adaptation layers. When data is ingested from third-party providers, the system automatically identifies the format and applies appropriate transformation rules without manual intervention, dramatically improving productivity while minimizing time loss.
Solution Approach 2:
The system performs preliminary data standardization and format normalization during the data ingestion phase rather than during analysis. By pre-processing and standardizing data formats upfront, the system eliminates the need for time-consuming manual standardization later, thus improving overall productivity.
Data Source
AI summary
Systems, methods, and apparatuses are described for providing a holistic data management platform. A computing device may receive different sets of data from different third-party data providers. The computing device may then process that data and store it in a distributed database. That data may be encrypted and/or tokenized. A relationship graph may be generated by processing entities indicated in the stored data. For instance, a relationship type may be identified based on properties of entities indicated by different sets of data stored in the distributed database. That relationship graph may be stored and updated based on new data. A failover of a primary cluster of the distributed database my cause performance of an automatic detach-and-promote process, whereby the relationship graph may be provided by a promoted secondary cluster.


