Cloud Data Replication via User Geolocation Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cloud providers face network congestion and resource overload due to uneven data replication across recovery data centers, leading to potential failures during geographic events, as they often replicate data from a concentrated geographic area without considering user device geolocation.
Innovation Solution
A leader node in the cloud infrastructure system determines the optimal recovery data center for replicating customer applications and data by analyzing the geolocation of user devices, generating distributions of geolocations, and selecting data centers with lower frequency scores for those geolocations to distribute the load more evenly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If data is replicated from a concentrated geographic area without considering user device geolocation, then data replication speed is improved, but network congestion and resource overload occur
Solution Approach 1:
The patent applies local quality by selecting recovery data centers based on the specific geolocation of user devices. Instead of uniform data replication to all data centers, the system identifies data centers in or near the user's geographic region and prioritizes replication to those locations, making the replication strategy location-specific and adaptive to local conditions.
Solution Approach 2:
The system changes the parameter of data center selection from a static, uniform approach to a dynamic, geolocation-based approach. By incorporating user device geolocation data and calculating frequency scores for different data centers, the system adapts the replication target selection based on geographic distribution parameters, avoiding concentration in single locations.
2Reliability
If data is replicated to multiple data centers, then reliability is improved, but load distribution becomes uneven causing hotspots
Solution Approach 1:
The patent implements feedback by calculating frequency scores for each data center based on the geolocations of user devices and the distribution of existing recovery data. This feedback mechanism allows the system to identify which data centers are experiencing high concentrations of recovery data and adjust replication decisions accordingly, distributing load more evenly across multiple data centers while maintaining reliability.
Solution Approach 2:
The system dynamically adjusts the selection of recovery data centers based on real-time or near-real-time analysis of geolocation distributions and frequency scores. Rather than using a static replication strategy, the system adapts its behavior based on the current state of data distribution across data centers, preventing hotspots from forming by balancing the load dynamically.
3Productivity
If geolocation analysis is performed for each replication decision, then load distribution is improved, but system complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-calculating and storing geolocation information for user devices and organizing data center capabilities before replication decisions are needed. The system maintains geolocation databases and frequency score calculations in advance, so that when a replication decision is required, the leader node can quickly query pre-processed data rather than performing complex analysis from scratch for each decision.
Data Source
AI summary
A method of detecting hotspots in a cloud infrastructure via aggregate geolocation information of user devices is described. The method includes receiving a request to launch a virtual machine executing on behalf of a first user device and retrieving a first set of identifiers of recovery data from a first data center and a second set of identifiers of recovery data from a second data center. The recovery data may be associated with a plurality of virtual machines previously executed on behalf of a plurality of user devices. The method further includes generating a first distribution of geolocations based on the first set of identifiers and a second distribution of geolocations based on the second set of identifiers. The method includes selecting the first data center and replicating, at the first data center, recovery data associated with the virtual machine executing on behalf of the first user device.


