Geographically Distributed Hybrid Cloud Clusters

By dividing data into partitions and migrating them to locations with lower access loads within the same geographic area, the system addresses the challenge of maintaining seamless data access across distributed locations and public cloud storage, ensuring high availability and minimizing network latency issues.

JP7678892B2Active Publication Date: 2025-05-16HITACHI VANTARA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023559138
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-15
Publication Date
2025-05-16
Estimated Expiration
2041-04-15

AI Technical Summary

Technical Problem

Existing data storage systems face challenges in maintaining seamless access to data across geographically distributed locations and public cloud storage, particularly in wide area networks where heartbeat communication for resource monitoring is unreliable due to latency and network issues.

Method used

The system divides data into multiple partitions, with at least two computing devices in each geographic location holding a copy of the data and exchanging periodic heartbeat communications. It migrates data to a location with a lower access load frequency, ensuring that all copies of a partition are moved to the same geographic location to maintain local system containment and avoid WAN communication for heartbeat signals.

Benefits of technology

This approach ensures high availability and redundancy of data and metadata across geographically distributed systems, while minimizing the impact of network latency and ensuring seamless data access by restricting heartbeat communications to the same geographic location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007678892000001
    Figure 0007678892000001
  • Figure 0007678892000002
    Figure 0007678892000002
  • Figure 0007678892000003
    Figure 0007678892000003
Patent Text Reader

Abstract

In some examples, a computing device of the plurality of computing devices at a first geographic location may partition data into a plurality of partitions. For example, at least two computing devices at the first geographic location may maintain copies of data of a first partition of the plurality of partitions and exchange periodic heartbeat communications associated with the first partition. The computing devices may determine that a computing device at a second geographic location has a lower access load frequency than a computing device at the first location that maintains the first partition. The computing devices may migrate the data of the first partition to a computing device at the second geographic location to cause the at least two computing devices at the second geographic location to maintain the data of the first partition and exchange periodic heartbeat communications associated with the first partition.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to the field of data storage. [Background technology]

[0002] Data may be spread across multiple geographic locations to provide efficient access to users who access or store data. In addition, with the advent of public cloud storage, it is possible to reduce the cost of maintaining multiple data centers by storing some data in public cloud storage and keeping other data in private cloud or local storage. In addition, storage users want seamless access to data that may be spread across their own geographically distributed storage locations as well as public cloud storage locations. For example, clusters of computing devices are often used to provide efficient and seamless storage services to users. Some clustered systems may rely on periodic heartbeat communication signals exchanged between members of a cluster, such as to monitor other members of the cluster and other cluster resources. However, the basic technique of using periodic heartbeat communication to monitor resources in a clustered system and to address loss of heartbeats does not typically scale to a wide area network (WAN) such as the Internet. Summary of the Invention [Means for solving the problem]

[0003] Some implementations include a computing device of a plurality of computing devices at a first geographic location partitioning data into a plurality of partitions. For example, at least two computing devices at the first geographic location may maintain copies of data of a first partition of the plurality of partitions and may exchange periodic heartbeat communications associated with the first partition. The computing devices at the first location may determine that a computing device at a second geographic location has a lower access load frequency than a computing device at the first location that maintains the first partition. The computing devices may migrate the data of the first partition to a computing device at the second geographic location to cause at least two computing devices at the second geographic location to maintain the data of the first partition and exchange periodic heartbeat communications associated with the first partition. [Brief description of the drawings]

[0004] The detailed description will now be described with reference to the accompanying drawings, in which the left-most digit of a reference number identifies the drawing in which the reference number first appears. Use of the same reference number in different drawings indicates similar or identical items or features.

[0005] [Figure 1] FIG. 1 illustrates an exemplary logical layout of a system capable of storing and managing metadata according to some implementations.

[0006] [Diagram 2] FIG. 2 illustrates an example architecture of a geographically distributed system capable of storing data and metadata in heterogeneous storage systems according to some implementations.

[0007] [Diagram 3] FIG. 3 is a block diagram illustrating an example logical configuration of a partition group according to some implementations.

[0008] [Figure 4]FIG. 4 is a block diagram illustrating an example of partition adjustment in a system according to some implementations.

[0009] [Diagram 5] FIG. 5 is a diagram illustrating an example of partitioning according to some implementations.

[0010] [Figure 6] FIG. 6 illustrates an example data structure of a composite partition map according to some implementations.

[0011] [Figure 7] FIG. 7 illustrates an example of partitioning implementation according to some implementations.

[0012] [Figure 8] FIG. 8 illustrates an example of a post-partition migration according to some implementations.

[0013] [Figure 9] FIG. 9 illustrates an example of the state after partitioning and migration is complete according to some implementations.

[0014] [Figure 10] FIG. 10 illustrates example pseudocode for load balancing according to some implementations.

[0015] [Figure 11] FIG. 11 is a flow diagram illustrating an example process for partitioning and load balancing according to some implementations.

[0016] [Figure 12] FIG. 12 illustrates a selection of components of a service computing device that may be used to implement at least a portion of the functionality of the systems described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0017] Some examples herein relate to techniques and arrangements for distributed computer systems, including hybrid cloud structures that can deploy clusters with multiple nodes across multiple cloud systems, multiple sites, and multiple geographic locations. Additionally, some implementations may include efficient solutions for maintaining availability in large-scale clusters across multiple geographic locations over a WAN. For example, the systems herein may synchronize data across hybrid cloud systems to provide storage users with cost management flexibility, such as by enabling backup and data movement across a variety of different storage services in one or more network locations. For example, in a hybrid cloud system disclosed herein, the system may employ network storage available from a variety of different service providing entities to store and synchronize data.

[0018] Some examples separate system resources and management, such as by separating the compute, network, and storage used for cluster management services and data management services. Similarly, some examples herein include separation of data and metadata management services, such that data may be stored in an entirely different storage location than the storage location in which the metadata associated with that data is stored. This allows metadata to be stored all in one location, or alternatively distributed across many locations, regardless of where the corresponding data is stored.

[0019] Implementations herein include a metadata management layer that operates by separating metadata associated with data into partitions distributed across geographic locations. For example, each partition may operate as a partition group, where, for example, there may be at least three or more copies of the same partition of metadata stored on different computing devices (also referred to as "nodes") across the system to provide data protection and high availability. In some cases, members of the same partition group periodically send heartbeat communications to each other to monitor each other's availability and to take action based on the heartbeat response or lack thereof. However, sending heartbeat communications over a WAN, such as the Internet, for example, can be problematic due to intermittent latency, delayed response times, or other network issues that may arise to delay or otherwise interfere with the heartbeat communications. Thus, to avoid these issues, some examples herein may include configuration criteria for metadata partition groups that may limit the use of heartbeat communications to nodes within the same geographic location, even though the cluster itself (e.g., which may include the entire metadata database) may span multiple distributed geographic locations.

[0020] Some examples herein may include separating metadata storage and data storage, such as by providing separate monitoring and management for each of the metadata storage and data storage. For example, a system herein may include a cluster monitor to monitor the data nodes separately from monitoring and managing the metadata nodes. Additionally, in some examples, the metadata may be divided into smaller, easily manageable, independent segments, referred to herein as partitions. For example, partitions may be made highly available and redundant by creating at least three copies of the metadata contained in each partition and by ensuring that each copy of the same partition is distributed to different nodes. As mentioned above, a node may be a computing device, such as a physical or virtual machine computing device, that manages one or more partitions. Additionally, nodes herein may monitor each other using heartbeat communication between respective nodes in the same group that manage the same partition, which may be referred to as a "partition group", for example.

[0021] When a partition herein grows beyond a certain size, e.g., so large that the performance of each node managing the partition may degrade, or is greater than a threshold size, the partition may be split into two or more partitions. When a partition is split, in some examples, one or more of the resulting partitions may be moved to a different location, such as to avoid hot spots. For example, the partition resulting from the split may be marked as a candidate for migration by a local coordination service. The partition identifier (ID) may be stored in a priority queue, such as based on the criteria used to select the original parent partition for the split. The local coordination service may determine the location to move the partition to based at least in part on the current load of each node in the system participating in the cluster.

[0022] As an example, the coordination service may determine the most lightly loaded location by referring to a composite partition map assembled from multiple partition maps received from all systems distributed across a geographic location. If all locations are equally loaded, a particular location may be selected randomly. For example, each coordination service at each geographic location may maintain load information specific to that location (e.g., for the number of gateways at each end). When a threshold load is reached at a particular location, the coordination service at that location may send a request to all coordination services to find candidates for offloading one or more of the partitions from that location.

[0023] Requests to migrate partitions received from other coordination services are queued and serviced based on the order of arrival. When a particular coordination service accepts to be a migration destination for receiving an offload of a particular partition, the corresponding location accepting the migration may be locked and configured to not receive additional load balancing requests from other coordination services. The locked state may include the identity of the remote coordination service and / or node that is currently forwarding the partition to the particular coordination service. Furthermore, when a partition is moved to a different geographic location, implementations herein may ensure that all copies of that partition are moved to the same geographic location such that heartbeats do not need to be sent over a WAN and remain contained within the same local system (e.g., a local area network (LAN)). The locked state of a particular coordination service at a particular geographic location accepting a partition migration may be removed from the respective coordination service after all copies for the migrated partition have been fully migrated.

[0024] In some examples herein, each system at each geographic location may provide a metadata gateway service and communicate with one or more networked storage systems over a network. For example, a first networked storage system may be provided by a first service provider employing a first storage protocol and / or software stack, and a second networked storage system may be provided by a second service provider employing a second storage protocol and / or software stack that is different from the first storage protocol / stack, etc. The system may receive an object and determine a storage location for the object in the first networked storage system, the second networked storage system, etc. The system may update a metadata database based on the receipt and storage of the object.

[0025] Storing and synchronizing data and metadata across multiple heterogeneous systems can be difficult. For example, the software stacks used in these heterogeneous systems may be configured and controlled by different entities. Furthermore, replication-specific changes may not always work as desired across all of the various different systems. Nonetheless, some examples herein are scalable to store trillions of objects on a hybrid cloud topology that may include multiple public and private cloud computing devices distributed across multiple different geographic locations.

[0026] For purposes of discussion, some example implementations are described in the context of multiple geographically distributed systems for managing metadata storage. However, implementations herein are not limited to the specific examples provided and may be extended to other types of computing system architectures, other types of storage environments, other storage providers, other types of client configurations, other types of data, etc., as will be apparent to those of skill in the art upon review of the disclosure herein.

[0027] 1 illustrates an exemplary logical arrangement of a system 100 capable of storing and managing metadata according to some implementations. As mentioned above, in some examples herein, object data may be stored in an entirely different storage location than the metadata corresponding to the object data. Furthermore, metadata may in some cases be stored in metadata databases spread across multiple distributed geographic locations that may collectively function as a cluster.

[0028] In the illustrated example, a first system 102(1) is located at geographic location A and includes multiple metadata sources 104(1), 104(2), and 104(3), such as metadata gateways. Additionally, the first system 102(1) may be in communication with a second system 102(2) located at geographic location B, which may be a different geographic location than geographic location A, via one or more networks 106. The second system 102 may include multiple metadata sources 104(4), 104(5), and 104(6). Additionally, while two systems 102 are shown in this example, implementations herein are not limited to any particular number of systems 102 or geographic locations, as implementations herein are scalable to very large systems capable of storing trillions of data objects. Additionally, while three metadata sources 104 are shown in each system 102, in other examples there may be multiple metadata sources 104 associated with each system.

[0029] Additionally, systems 102(1), 102(2), ... may be capable of communicating with user devices 108 over one or more networks. For example, each user device 108 may be any suitable type of computing device, such as a desktop, laptop, tablet computing device, mobile device, smartphone, wearable device, terminal, and / or any other type of computing device capable of transmitting data over a network. Users 112 may be associated with user devices 108 through respective user accounts, user login credentials, or the like. Furthermore, user devices 108 may be capable of communicating with system 102 over one or more networks 106, over a separate network, or over any other suitable type of communication connection.

[0030] In some examples, there may be a first group 114 of users 112 that access a first system 102(1) based on, for example, their geographic proximity to the first system 102(1), a second group 116 of users 112 that access a second system 102(2) based on, for example, their geographic proximity to the second system 102(2), etc. As mentioned above, in some cases, allowing users 112 to access a metadata gateway that is physically closer to each user 112 may reduce latency compared to accessing a system 102 that is geographically farther from each user's geographic location.

[0031] Each metadata source 104 may include one or more physical or virtual computing devices (also referred to as "nodes") that may be configured to store and provide metadata corresponding to the stored data. For example, users 112 may store data, such as object data, in the system 102 and, in some examples, may access, search, modify, update, delete, or migrate the stored data. In some cases, data stored by users 112 may be stored in one or more data storage locations 118. For example, data storage locations 118 may include one or more of public network storage (e.g., public cloud storage), private network storage (e.g., private cloud storage), local storage (e.g., storage provided locally to the respective system 102), and / or other private or public storage options.

[0032] The one or more networks 106 may include any suitable network, including a wide area network such as the Internet, a local area network (LAN) such as an intranet, a wireless network such as a cellular network, a local wireless network such as Wi-Fi, and / or a wired network including short-range wireless communication such as BLUETOOTH, Fibre Channel, optical fibre, Ethernet, or any other such network, a direct wired connection, or any combination thereof. Thus, the one or more networks 106 may include both wired and / or wireless communication technologies. The components used for such communication may depend at least in part on the type of network, the selected environment, or both. Protocols for communicating over such networks are well known and will not be discussed in detail herein. Thus, the system 102, the metadata source 104, the user device 108, and the data storage location 118 can communicate over the one or more networks 106 using wired or wireless connections, and combinations thereof.

[0033] In the illustrated example, the metadata stored by the metadata source 104 may be divided into multiple partitions (not shown in FIG. 1). Each partition may operate as a group, e.g., there may be at least three or more copies of the same partition distributed across different nodes in the same system to provide redundancy protection and high availability. In the current design, members of the same partition group may periodically send heartbeat communications 120 to each other to monitor each other's availability and to take one or more actions based on the heartbeat response. For example, the heartbeat communication 120 may be an empty message, a message with a heartbeat indicator, or the like, that may be sent periodically from a leader node to its follower nodes, such as every 200 ms, every 500 ms, every second, or every two seconds. In some cases, if one or more follower nodes do not receive a heartbeat message 120 within a threshold time, the follower nodes may initiate a process of selecting a new leader node, such as by establishing a consensus based on a RAFT algorithm or the like. As mentioned above, because maintaining heartbeat communications over a WAN may be prohibitive, examples herein may include metadata partition groups formed using a number of formation criteria, including size and location, such that heartbeat communications are restricted to the same geographic location (e.g., the same LAN), even though multiple systems 102 may otherwise form a cluster configuration to distribute data across multiple different geographic locations.

[0034] 2 illustrates an example architecture of a geographically distributed system 200 capable of storing data and metadata in heterogeneous storage systems according to some implementations. In some examples, system 200 may correspond to system 100 discussed above with respect to FIG. 1. System 200 in this example includes a first system 102(1), a second system 102(2), a third system 102(3), ..., etc. that are geographically distributed with respect to one another. Additionally, in this example, details of the configuration of first system 102(1) are described, although the configurations of second system 102(2), third system 102(3), etc. may be similar.

[0035] The first system 102(1) includes multiple service computing devices 202(1), 202(2), 202(3), ..., at least some of which may communicate with multiple networked storage systems, such as a first provider networked storage system 204(1), a second provider networked storage system 204(2), ..., over one or more networks 106. As mentioned above, in some cases, each provider of the networked storage systems 204 may be a distinct entity unrelated to the other providers. Some examples of commercial networked storage providers include AMAZON WEB SERVICES, MICROSOFT AZURE, GOOGLE CLOUD, IBM CLOUD, and ORACLE CLOUD. The networked storage systems 204 may be referred to as "cloud storage" or "cloud-based storage" in some examples, and may enable a lower cost per gigabyte storage solution than local storage that may be available at the service computing device 202 in some cases. Additionally or alternatively, in some examples, the storage provider may be a private or other proprietary storage provider, such as to provide access only to particular users or entities associated with the service computing device 202, system 102, etc. An example of a proprietary system may include configurations of the HITACHI CONTENT PLATFORM.

[0036] The service computing device 202 can communicate with one or more user computing devices 108 and one or more administrator computing devices 210 via the network 106. In some examples, the service computing device 202 may include one or more servers, which may be embodied in any number of ways. For example, at least a portion of the programs, other functional components, and data storage of the service computing device 202 may be implemented in at least one server, such as a cluster of servers, a server farm, a data center, a cloud-hosted computing service, etc., although other computer architectures may additionally or alternatively be used. Further details of the service computing device 202 are discussed below with reference to FIG. 12. In addition, the service computing devices 202 may be able to communicate with each other via one or more networks 207. In some cases, the one or more networks 207 may be a LAN, a local intranet, a direct connection, a local wireless network, etc.

[0037] Service computing device 202 may be configured to provide storage management services and data management services, respectively, to users 112 via user devices 108. As some non-limiting examples, users 112 may include users performing functions for a company, business, organization, government agency, academic institution, or the like, which may in some examples include the storage of very large amounts of data. Nevertheless, implementations herein are not limited to any particular use or application of system 200, and other systems and arrangements described herein.

[0038] Each user device 108 may include a respective instance of a user application 214 that may execute on the user device 108, such as to transmit user data for storage on the networked storage system 204 and / or to receive stored data from the networked storage system 204, such as via data instructions 218, and to communicate with a user web application 216 executable on one or more of the service computing devices 202. In some cases, the application 214 may include or operate via a browser, while in other cases the application 214 may include any other type of application having communication capabilities that enable communication with the user web application 216 over one or more networks 106.

[0039] In some examples of the system 200, the users 112 may store data on and receive data from the service computing device 202 with which their respective user devices 108 are in communication. Thus, the service computing device 202 may provide storage for the users 112 and their respective user devices 108. In steady state operation, there may be users 108 that communicate with the service computing device 202 periodically.

[0040] Additionally, the administrator device 210 may be any suitable type of computing device, such as a desktop, laptop, tablet computing device, mobile device, smartphone, wearable device, terminal, and / or any other type of computing device capable of transmitting data over a network. An administrator 220 may be associated with the administrator device 210 through a respective administrator account, administrator login credentials, or the like. Further, the administrator device 210 may be capable of communicating with the service computing device 202 over one or more networks 106, 207, over a separate network, or over any other suitable type of communications connection.

[0041] Each administrator device 210 may include a respective instance of an administrator application 222 that may execute on the administrator device 210, such as to communicate with an administrative web application 224 executable on one or more of the service computing devices 202, such as to send management instructions for managing the system 200 and to send management data for storage on the networked storage system 204 and / or to receive stored management data from the networked storage system 204, such as via management instructions 226. In some cases, the administrator application 222 may include or operate via a browser, while in other cases the administrator application 222 may include any other type of application having communication capabilities that enable communication with the administrative web application 224 over one or more networks 106.

[0042] The service computing device 202 may execute storage programs 230 that may provide access to the networked storage system 204, such as to transmit data stored in the networked storage system 204 using data synchronization operations 232, and to retrieve requested data from the networked storage system 204 or local storage. Additionally, the storage programs 230 may manage the data stored by the system 200, such as to manage data retention periods, data protection levels, data replication, etc.

[0043] The service computing device 202 may further include a metadata database (DB) 234, which may be divided into multiple metadata DB portions called partitions 240 and distributed across multiple service computing devices 202. For example, the metadata included in the metadata DB 234 may be used to manage object data 236 stored in the network storage system 204 and local object data 237 stored locally in the system 102. The metadata DB 234 may include a number of metadata regarding the object data 236, such as information about individual objects, how to access individual objects, storage protection level of objects, storage retention period, object owner information, object size, object type, etc. Furthermore, a metadata gateway program 238 may manage and maintain the metadata DB 234 to respond to requests to access data, create new partitions 240, provide coordination services between systems 102 in different geographic locations, etc., as well as update the metadata DB 234 according to the storage of new objects, deletion of old objects, migration of objects, etc.

[0044] FIG. 3 is a block diagram illustrating an example logical configuration of a partition group 300 according to some implementations. The partition group 300 may include a first service computing device 202(1) (first node), a second service computing device 202(2) (second node), and a third service computing device 202(3) (third node). In this example, the partition group 300 relates to a first partition 302 that includes a metadata database portion 304. As mentioned above, there may be multiple partitions in the system, each of which may include a different portion of the metadata database 134 discussed above. When the metadata database 134 becomes large, the overgrown partition may be split into multiple new partitions. Additionally, heartbeat communications 306 may be exchanged between members of the partition group 300 relating to the first partition 302.

[0045] Partition group 300 may be included in system 200 discussed above with respect to FIG. 2. The entire metadata database 134 may resemble a large logical space that contains metadata items that can be located using respective keys. In some examples, the keys may be cryptographic hashes of data paths to the respective metadata items, or any other relevant input provided by a user. The metadata key space range may depend in part on the width of the key employed. A suitable example may include a SHA-256 key that contains 256 bits, thus providing a range of 0 to 2256.

[0046] As an example, metadata may be managed by dividing the (0-2256) metadata space into smaller manageable ranges, e.g., 0-228, 228-256, etc. The size of the range is configurable and may be dynamically changed to a smaller or larger size. As mentioned above, each partition may be replicated to a certain number of redundant copies spread across the set of available nodes that make up the partition group for that partition. In general, three copies are sufficient to form a highly available active group with one of the three nodes in the partition group selected or otherwise designated as the leader, and the remaining two nodes in the partition group designated as followers. For example, all reads and writes of data in a given partition may be performed through the leader node of the respective partition group. The leader node may be configured to keep each of the follower nodes updated and ensure that replication occurs to each of the follower nodes in order to maintain consistency and redundancy of data within the partition group. In some examples, the systems herein may employ a consensus protocol for leader selection. Examples of suitable consensus algorithms may include the RAFT algorithm. The leader and follower nodes of each partition may operate independently and do not need to be aware of other partitions in the system, so the failure of an entire partition does not affect other partitions in the system.

[0047] As mentioned above, leaders and followers in a partition group may communicate and monitor each other via periodic heartbeats 306. Furthermore, nodes in the system may include multiple partitions, with a different set of nodes for each partition forming a partition group.

[0048] Additionally, the metadata in metadata database 134 may be organized as a set of user-defined schemas (not shown in FIG. 3). Each schema may be partitioned with its own keyspace to accommodate different growth patterns based on the respective schema. For example, the number of objects in the system may grow very differently than the number of users or the number of folders (also called buckets in some examples).

[0049] 4 is a block diagram illustrating an example of partition coordination in system 200 according to some implementations. For example, a partition coordination service 402 may be executed within each system 102 of the geographically distributed systems 102. For example, a first system 102(1) may include a first coordination service 402(1), a second system 102(2) may include a second coordination service 402(2), a third system 102(3) may include a third coordination service 402(3), etc. As an example, each respective coordination service 402 may be provided by execution of a metadata gateway program 238 on one of the service computing devices 202 included in the respective system 102.

[0050] The coordination service 402 may manage the creation of partitions, splitting partitions, deleting partitions, and balancing of partitions on each system 102 on which it executes, as shown at 404. The coordination service 402 may build a global view of all partitions on each system 102 on which it executes as well as by receiving partition information from other systems in the geographically distributed systems 102.

[0051] For example, in the illustrated example, the first coordination service 402(1) may build a first partition map 406 by collecting information from all leader nodes throughout the first system 102(1) at geographic location A. In this example, there are four service computing devices 202(1)-202(4), each of which manages three partitions. In particular, the first service computing device 202(1) includes a first partition 408, a second partition 410, and a fourth partition 412. The second service computing device 202(2) includes a first partition 408, a third partition 414, and a fourth partition 412. The third service computing device 202 includes a first partition 408, a second partition 410, and a third partition 414. The fourth service computing device 202 includes a second partition 410, a third partition 414, and a fourth partition 412.

[0052] Service computing devices 202(1)-202(4) may provide metadata gateway services 416, such as by each executing an instance of a metadata gateway program 238 to store metadata in each partition and to retrieve metadata from each partition managed by each service computing device 202. Additionally, leaders and followers of each partition group in each partition may communicate with and monitor each other via periodic heartbeats with respect to each partition. Several examples of heartbeat communications are shown, including heartbeat communications 420 between the first service computing device 202(1) and the second service computing device 202(2) with respect to the first partition 408 and the fourth partition 412, heartbeat communications 422 between the second service computing device 202(2) and the fourth service computing device 202(4) with respect to the third partition 414 and the fourth partition 412, heartbeat communications 424 between the fourth service computing device 202(4) and the third service computing device 202(3) with respect to the second partition 410 and the third partition 414, and heartbeat communications 426 between the third service computing device 202(3) and the first service computing device 202(1) with respect to the first partition 408 and the second partition 410.

[0053] Moreover, other heartbeat communications between service computing devices 202(1)-202(4) are not shown for clarity of illustration. Thus, in this example, each node in first system 102(1) may include multiple partitions and form partition groups with different sets of nodes for each partition 408-414.

[0054] As described above, the first coordination service 402(1) may obtain partition information from the service computing devices 202(1)-202(4) to construct a first partition map 406. Exemplary partition maps are discussed below with reference to FIG. 6. In addition, the first coordination service 402(1) may interact with remote coordination services, such as a second coordination service 402(2) in the second system 102(2) and a third coordination service 402(3) in the third system 102(3), to construct a composite partition map 430 for the entire geographically distributed system 200 and receive respective partition maps from all systems 102(2), 102(3), ... in other geographic locations. For example, the first coordination service 402 may receive a second partition map 432 from the second coordination service 402(2) and a third partition map 434 from the third coordination service 402(3). In exchange, the first coordination service 402(1) may send the first partition map 406 to the second coordination service 402 and the third coordination service 402. Similarly, the second coordination service 402(2) and the third coordination service 402(3) may exchange partition maps 432 and 434, respectively. Thus, each geographically distributed system 102 may build its own composite partition map 430.

[0055] Each metadata coordination service 402(1), 402(2), 402(3), ... may be executed to determine when to split partitions at each local system 102(1), 102(2), 102(3), .... For example, the coordination service 402 may constantly monitor the size of the partitions at each system 102 and use a size threshold to determine when to split partitions. In addition, after partitions are split, the coordination service 402 may employ the composite partition map 430 to re-balance the partitions across systems 102 at various different geographic locations to optimize performance. To enable efficient information exchange over the WAN between different geographic locations, each system 102 may run an instance of the coordination service 402 locally. The coordination service 402 executing at each geographic location communicates with a local metadata gateway service provided by running a separate instance of the metadata gateway program 238 on the service computing device 202 to send commands at the appropriate time to initiate splits and moves for load balancing. Each coordination service 402 may employ a lightweight messaging protocol to periodically exchange partition maps for its location with those of other locations within the geography.

[0056] In addition, because partitions may span multiple geographic locations, in some instances it may be ensured that connectivity to all geographic locations is sufficiently available to fulfill a request arriving at one of the geographic locations. As an example, assume that a first system 102(1) at geographic location A receives a request for a partition stored in, for example, a second system 102(2) at geographic location B. For example, the request may be received by the metadata gateway service 416. The metadata gateway service 416 may first attempt to determine whether the request maps to a key space range (i.e., a partition) that exists at the current location. If not, the metadata gateway service 416 may consult the first coordination service 402(1) to determine the geographic location of the partition associated with the request. Once identified, the metadata gateway service 416 may temporarily cache the partition information of the target partition to enable efficient lookup for subsequent requests that may be received. For example, the metadata gateway service 416 may then look up the entry by directly querying a remote metadata gateway service 440 or 442 for that partition keyspace range and respond to the end user when the lookup is successfully completed.

[0057] 5 illustrates an example partitioning 500 according to some implementations. Some examples may employ dynamic partitioning for elastic and scalable growth. For example, not all data may be ingested at deployment time. As a result, the system may start with a small set of partitions, e.g., at least one partition, such as a first partition 502(1). As the amount of data ingested increases, partitions may be split and the number of partitions may increase as the amount of data expands.

[0058] In the illustrated example, assume that initially there is metadata (MD) for N objects in a first partition 502(1), as shown at 504. As the number of objects (and resulting metadata) continues to grow, the coordination service may decide to split the first partition 502(1), as shown at 506. Thus, the metadata in the first partition 502(1) may be split into a second partition 502(2) and a third partition 502(3). For example, assume then there is metadata (MD) for 2N objects, as shown at 508. Additionally, because the number of objects and corresponding metadata continues to grow over time, at least the second partition 502(2) may be split to generate a fourth partition 502(4) and a fifth partition 502(5), as shown at 510. Assume then there is metadata (MD) for approximately 4N objects, as shown at 512. Additionally, as the number of objects and metadata continues to grow over time, at least the third partition 502(3) may be split, as shown at 510, to create a sixth partition 502(6) and a seventh partition 502(7), such that the amount of metadata may then be equal to the metadata for approximately 6N objects, as shown at 516. In this manner, partitions may continue to be split and expanded as the amount of data and resources of the system increase.

[0059] FIG. 6 illustrates an exemplary data structure of a composite partition map 430 according to some implementations. In this example, the composite partition map 430 includes a partition ID 602, a key space name 604, a key space range 606, a node map 608, a partition size 610, a location ID 612, and an access load frequency 614. The partition ID 602 ​​may identify an individual partition. The key space name 604 may identify an individual key space. The key space range 606 may identify the start hash value and the end hash value of each key space range of each partition. The node map 608 may identify each of the service computing devices 202 included in each partition group of each partition, and may further identify which of the service computing devices 202 is the leader of the partition group. In this example, the IP address of the computing device is used to identify each node, but in other examples, any other system unique identifier may be used. The partition size 610 may indicate the current size of each partition, such as in megabytes. The location ID 612 may identify the geographic location of each of the identified partitions.

[0060] Access load frequency 614, sometimes referred to as the load of a partition, may indicate the number of data access requests received within the last unit of time, e.g., every hour, every 6 hours, every 12 hours, every day, every few days, every week, etc. As an example, the number of internal calls made to each partition in each client call may be collected as an indicator to determine the access load frequency in each of the partitions. This access frequency indicator may be used as an additional attribute to consider partition balance across the system, along with location and partition size. For example, as mentioned above, a partition is generally a range within a key space. If a particular key space range forms a hot path, a balancing algorithm, discussed further below, may cause the system to split the partitions of this key space range and move at least one of the new partitions to a different physical node (which may be determined based on other balancing criteria, discussed further herein) to balance the load of the physical nodes.

[0061] For example, the composite partition map 430 may include all of the partitions in the system 200 discussed above from each of the geographically distributed systems 102(1), 102(2), 102(3), ... that contain the partitions, identifying each partition according to a partition ID 602 ​​and a location ID. Additionally, the individual system partition maps 406, 432, 434, etc. discussed above with respect to FIG. 4 may have a similar data structure as the composite partition map 430, but may only include the partition IDs 602 and corresponding information 604-612 for the partitions contained in their respective local systems 102.

[0062] As discussed above, with respect to FIG. 4, each instance of the coordination service 402 may rely on the partition map 430, or the local partition maps 406, 432, 434 of each of its systems 102(1), 102(2), 102(3), respectively, to determine when to split a partition maintained on each local system 102. As an example, the coordination service 402 may determine the size of the partition as the primary criterion. Thus, if the size of a partition exceeds a threshold size, the coordination service may determine that the partition should be split into two partitions. In addition, the coordination service 402 may also split a partition if the access rate of the partition keyspace exceeds a threshold for a specified period of time. For example, if the access load frequency 614 of a particular partition, or the number of partitions hosted by a particular one of the service computing devices 202, exceeds a threshold for a specified period of time, the service computing device 202 may become overloaded, and therefore one or more partitions may be split and migrated to other computing devices. The decision to split and migrate partitions may also be based on the amount of remaining storage space available on each of the service computing devices 202.

[0063] The split operation to split the selected partition will initially only update the respective local partition map. Metadata corresponding to the split partition may then be subsequently migrated to a different service computing device 202, either at the same geographic location or a different geographic location, depending on available resources in the system 200. The updated local partition map may be sent to other coordination services in other geographic locations of the system 200 along with a request for data migration for load balancing of the geographically distributed system. A partition balancing algorithm, discussed further below with respect to Figures 10 and 11, may be executed to determine the system 102 and geographic location to which the split partition should be migrated.

[0064] FIG. 7 illustrates an example partition split 700 according to some implementations. FIG. 7 illustrates the initial split of partitions, while FIG. 8 and FIG. 9 illustrate additional subsequent operations performed to complete the migration and rebalancing of data across multiple locations. The example of FIG. 7 illustrates two geographically separated systems: a first system 102(1) at geographic location A and a second system 102(2) at geographic location B. The example of FIG. 7 also illustrates a first local partition map portion 702 for the first system 102(1) before the split is performed and a first local partition map portion 702(1) after the split is performed. FIG. 7 also includes a local partition map portion 704 for the second system 102(2) before the split is performed and before data is migrated to the second system 102(2).

[0065] In this example, the local partition map portions include a column 706 that indicates whether a split threshold has been reached by any of the partitions listed in the respective local partition map portions 702, 702(1), 704. As shown at 708, the split threshold for partition ID "1" has been reached. Examples of criteria for determining whether a split threshold has been reached are discussed above with respect to FIG. 6 and further below with respect to FIGS. 10 and 11. For example, local coordination service 402(1) may determine that the split threshold has been reached with respect to partition 1, and based on this determination, coordination service 402 may split partition 1 into two new partitions, partition 3 and partition 4. Additionally, coordination service 402(1) may send a split notification to the leader node, i.e., "node 1" in this example, to notify the leader node of the split. After the split is confirmed by the leader node, coordination service 402(1) may update partition map 702 to implement the split, such that new partitions 3 and 4 are added to partition map 702(1) and partition 1 is removed from partition map 702(1). In addition, coordination service 402(1) may decide to migrate partition 4 to second system 102(2) at location B and may send migration request 710 to coordination service 402(2) at second system 102(2) at geographic location B. For example, the migration request may include the updated map portion 702(1).

[0066] 8 illustrates an example post-partition migration 800 according to some implementations. For example, FIG. 8 illustrates the next balancing step after partitioning, which in this example is the migration of partition 4 to the second system 102(2). For example, the coordination service 402 may be configured to ensure that the partitions in system 200 are balanced across all geographic locations in system 200 while taking into account available resources in all geographic locations.

[0067] As discussed above with respect to FIGURE 7, the first coordination service 402(1) may send a migration request to the second coordination service 402(2). In response, the second migration service 402(2) may send a migration response 802, which may typically include an acceptance of the migration request. For example, the acceptance of the migration request may include an identification of a destination node in the second system 102(2) to which partition 4 is to be migrated. For example, the identified destination node will be the leader node of the new partition group in the second system 102(2) at geographic location B (node ​​10 in this example).

[0068] When the first coordination service 402(1) at geographic location A receives an acceptance of the migration request and an identification of the destination node from the second coordination service 402(2) at geographic location B, the first coordination service 402(1) at geographic location A may request that the leader node for partition 4 at geographic location A (i.e., node 1 in this example) initiate a migration to the destination node (node ​​10) specified by the second coordination service 402(2). The leader node at geographic location A may then migrate a snapshot 804 of partition 4 to the destination node. Partition 4 may be locked from changes while the snapshot is taken and migrated. Migration of snapshot 804 may, in some examples, only occur between leader nodes. Thereafter, upon receiving snapshot 804, the destination leader node (node ​​10) may create a copy in the second system 102(2) to provide copies of partition 4 to other nodes in the partition group, e.g., nodes 9 and 11 in this example. Partition map portion 704(1) shows the partition map information for second system 102(2) after migration.

[0069] 9 illustrates an example state 900 after completion of a partition split and migration according to some implementations. In this example, the current state is shown after the partitions have been balanced at geographic locations A and B. For example, after the destination leader node (node ​​10) determines that a copy of partition 4 has been created at geographic location B, the first coordination service 402(1) at geographic location A releases the lock and updates partition map portion 702(2) to reflect the changes and show the current state of the partitions at location A after the split and migration. Additionally, partition map portion 704(3) shows the current state of partition 4 at geographic location B after the split and migration from geographic location A is complete.

[0070] The coordination service 402 in each geographic location is responsible for ensuring that the number of copies of each partition in that geographic location always meets a certain number of copies. A reduction in this number places the system in a state where data availability may be compromised in the event of a failure. To this end, the coordination service 402 may periodically monitor and collect from the leader node the last received heartbeat time of all members of each partition group. If a particular node storing a copy of a partition appears to be inactive for more than a threshold period of time, the coordination service may trigger a partition repair sequence. For example, a partition repair sequence may include removing an inactive node from the partition group it belongs to and triggering a copy of a snapshot of the partition from its respective leader to an available active node in the same geographic location. If no active nodes are available in the same geographic location, the coordination service 402 may attempt to migrate all copies of that partition to a different geographic location that contains a sufficient number of healthy nodes with sufficient capacity to receive the partition.

[0071] 10 illustrates example pseudocode 1000 for load balancing according to some implementations. For example, as described above, the system 102 may execute a load balancing algorithm to identify locations to migrate partitions to, such as after a partition split. For example, the load balancing algorithm 1000 may be included in the metadata gateway program 238 and executed as part of the coordination service 402 provided during execution of the metadata gateway program 238.

[0072] Once a partition is split, the system may attempt to move at least one of the new partitions to a different computing device in a different geographic location to avoid hotspotting. The partition may be marked as a possible candidate for migration by the local coordination service 402. The partition ID may be stored in a priority queue based on the splitting criteria. The local coordination service 402 may determine the geographic location to move the partition to, such as based on determining the recent access load frequency at each geographic location. For example, as shown at 1002, the coordination service accesses a composite partition map 430 assembled from partition maps received from all geographic locations in the system 200 to determine the geographic location that is the lightest loaded (e.g., has the lowest access load frequency indicating the fewest data access requests over a period of time). Alternatively, if all geographic locations are equally loaded, the particular geographic location may be selected randomly, by round robin, or using other suitable techniques.

[0073] In some cases, each local coordination service 402 may maintain load information specific to its location, for example, for the number of compute nodes processing partitions of the metadata. When a threshold access load frequency is reached at a particular geographic location, the local coordination service 402 may send requests to all other coordination services 402 in other geographic locations to find candidates to offload at least one partition. The requests may be queued at each receiving location and serviced based on the order of arrival.

[0074] When the local coordination service 402 of the targeted destination geographic location accepts to be the destination for offloading a particular partition, the destination location may be locked by the local coordination service 402 and will not accept any additional load balancing requests from other coordination services 402. The locked state may refer to the identity of the remote coordination service where the migration is received. A copy of the p partition is migrated to the selected destination node, as shown at 1004. The destination node may then replicate the partition to two more local nodes such that the heartbeat continues to be contained within the local system 102. After the migration and replication is complete, the location locked state may be removed from the local coordination service and the source location may remove the copy of the partition from the source node.

[0075] FIG. 11 is a flow diagram illustrating an example process 1100 for partitioning and load balancing according to some implementations. The process is illustrated as a collection of blocks in a logical flow diagram that represents a sequence of operations, some or all of which may be implemented in hardware, software, or a combination thereof. In the context of software, the blocks may represent computer-executable instructions stored on one or more computer-readable media that, when executed by one or more processors, programs the processors to perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the blocks are described should not be construed as a limitation. Any number of the described blocks may be combined in any order and / or in parallel to perform a process or alternative processes, and not all of the blocks need be executed. For purposes of discussion, the process is described with reference to the environments, frameworks, and systems described in the examples herein, although the process may be implemented in a wide variety of other environments, frameworks, and systems. In the example of FIG. 11, the process 1100 may be performed, at least in part, by one or more service computing devices 202 executing a metadata gateway program 238, which includes executing the load balancing algorithm discussed above with respect to FIG.

[0076] At 1102, the systems may run a local coordination service. For example, each system at each geographic location may run an instance of a coordination service locally that can communicate with other instances of the coordination service running at other geographic locations.

[0077] At 1104, the system may determine whether to split the partition. For example, as discussed above, the system may determine whether to split the partition based on one or more considerations, such as whether the partition exceeds a threshold partition size, whether a node in the node group that maintains the partition is overloaded based on a threshold access load frequency, and / or whether the remaining storage capacity available to a node in the node group that maintains the partition has fallen below a threshold minimum. If the partition is ready to be split, the process proceeds to 1106. If the partition is not ready to be split, the process proceeds to 1108.

[0078] At 1106, the system may split the partition identified at 1104. For example, the coordination service may update the local partition map by splitting the key values ​​assigned to the identified partition between the two new partitions in the local partition map.

[0079] At 1108, the system may determine the most overloaded and most underloaded nodes in the geographically distributed system. As an example, the system may refer to a composite partition map to determine which nodes are the most underloaded and which nodes are the most overloaded nodes in the system.

[0080] At 1110, the system may select a partition to migrate if there is another node that is less loaded than the node currently holding the partition.

[0081] At 1112, the system may send a migration request to a coordination service in the geographic location with the least loaded node. For example, the migration request may request to migrate the selected partition to a system in the geographic location with the least loaded node.

[0082] At 1114, the system may determine whether the migration request is accepted. For example, if the migration request is accepted, the acceptance message may include an identifier of the destination node that will act as the leader node undergoing the partition migration. If the migration request is accepted, the process proceeds to 1118. If the migration request is not accepted, the process proceeds to 1116.

[0083] If the migration request is rejected, the system may send another migration request to another geographic location with a less loaded node, at 1116. For example, the system may select the next geographic location with the least loaded node to receive the next migration request.

[0084] At 1118, if the migration request is accepted, the system may send instructions to the source leader node to lock the selected partition and migrate the selected partition to the destination leader node. For example, the partition may remain locked from data writes while the migration is occurring, and received writes, deletes, etc. may be forwarded to the destination node to update the partition after the migration is complete.

[0085] At 1120, the system may receive confirmation of the migration and replication from the destination coordination service. For example, the destination node may replicate the received partition to at least two additional nodes that form a partition group with the destination node. After this replication is complete, the destination coordination service may send a migration completion notification to the source system.

[0086] At 1122, the system may send instructions to the source leader node and other nodes in the partition group to delete the selected partition, so that the selected partition may be deleted from the nodes in the source system.

[0087] The exemplary processes described herein are merely examples of processes provided to facilitate discussion. Numerous other variations will become apparent to those of skill in the art upon reviewing the disclosure herein. Additionally, the disclosure herein describes several examples of suitable frameworks, architectures, and environments for carrying out the processes, but the implementations herein are not limited to the specific examples shown and discussed. Additionally, the disclosure provides various exemplary implementations, as described and illustrated in the drawings. However, the disclosure is not limited to the implementations described and illustrated herein, but may extend to other implementations known or to become known to those of skill in the art.

[0088] FIG. 12 illustrates selected example components of a service computing device 202 that may be used to implement at least some of the functionality of the system described herein. The service computing device 202 may include one or more servers or other types of computing devices that may be embodied in any number of ways. For example, in the case of a server, the programs, other functional components, and data may be implemented on a single server, a cluster of servers, a server farm or data center, a cloud-hosted computing service, etc., although other computer architectures may additionally or alternatively be used. Multiple service computing devices 202 may be collocated or separately deployed and may be organized, for example, as virtual servers, server banks, and / or server farms. The described functionality may be provided by a server of a single entity or company, or by servers and / or services of multiple different entities or companies.

[0089] In the illustrated example, the service computing device 202 may include or be associated with one or more processors 1202, one or more computer-readable media 1204, and one or more communication interfaces 1206. Each processor 1202 may be a single processing unit or may be multiple processing units, and may include single or multiple arithmetic units, or multiple processing cores. The processor 1202 may be implemented as one or more central processing units, microprocessors, microcomputers, microcontrollers, digital signal processors, state machines, logic circuits, and / or any device that manipulates signals based on operational instructions. By way of example, the processor 1202 may include one or more hardware processors and / or logic circuits of any suitable type specifically programmed or configured to execute the algorithms and processes described herein. The processor 1202 may be configured to fetch and execute computer-readable instructions stored on the computer-readable media 1204, which may program the processor 1202 to perform the functions described herein.

[0090] The computer readable medium 1204 may include volatile and non-volatile memory, and / or removable and non-removable media implemented in any type of technology for storing information, such as computer readable instructions, data structures, program modules, or other data. For example, the computer readable medium 1204 may include, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, optical storage, solid state storage, magnetic tape, magnetic disk storage, storage arrays, network attached storage, storage area networks, cloud storage, or any other medium that may be used to store the desired information and that may be accessed by a computing device. Depending on the configuration of the service computing device 202, the computer readable medium 1204 may be a tangible non-transitory computer readable medium, insofar as non-transitory computer readable, when referred to, excludes media such as energy, carrier signals, electromagnetic waves, and / or signals themselves. In some cases, the computer readable medium 1204 may be co-located with the service computing device 202, while in other examples, the computer readable medium 1204 may be partially remote from the service computing device 202. For example, in some cases, the computer-readable medium 1204 may comprise a portion of the storage in the networked storage system 204 discussed above with respect to FIG.

[0091] The computer readable medium 1204 may be used to store any number of functional components executable by the processor 1202. In many implementations, these functional components include instructions or programs executable by the processor 1202 that, when executed, specifically program the processor 1202 to perform operations attributed to the service computing device 202 herein. The functional components stored on the computer readable medium 1204 may include a user web application 216, an administrative web application 224, a storage program 230, and a metadata gateway program 238, each of which may include one or more computer programs, applications, executable code, or portions thereof. Furthermore, although these programs are illustrated together in this example, during use, some or all of these programs may be executed on separate service computing devices 202.

[0092] In addition, the computer-readable medium 1204 may store data, data structures, and other information used to perform the functions and services described herein. For example, the computer-readable medium 1204 may store metadata database 234, which may include partition 240. In addition, the computer-readable medium may be local object data 237. Moreover, although these data structures are illustrated together in this example, some or all of these data structures may be stored by separate service computing devices 202 during use. The service computing device 202 may also include or hold other functional components and data, which may include programs, drivers, etc., and data used or generated by the functional components. Moreover, the service computing device 202 may include many other logical, programmatic, and physical components, of which the above-mentioned are merely examples relevant to the discussion herein.

[0093] The one or more communications interfaces 1206 may include one or more software and hardware components for enabling communication with various other devices, such as via one or more networks 106 and 207. For example, the communications interface 1206 may enable communication via one or more of a LAN, the Internet, a cable network, a cellular network, wireless networks (e.g., Wi-Fi) and wired networks (e.g., Fibre Channel, Fiber Optic, Ethernet), a direct connection, and short-range communications such as BLUETOOTH, etc., as additionally enumerated elsewhere herein.

[0094] Various instructions, methods, and techniques described herein may be considered in the general context of computer-executable instructions, such as computer programs and applications, stored on a computer-readable medium and executed by a processor herein. In general, the terms programs and applications may be used interchangeably and may include instructions, routines, modules, objects, components, data structures, executable code, and the like, for performing particular tasks or implementing particular data types. These programs, applications, and the like may be executed as native code or may be downloaded and executed, such as in a virtual machine or other just-in-time compilation execution environment. In general, the functionality of the programs and applications may be combined or distributed as desired in various implementations. Implementations of these programs, applications, and techniques may be stored on computer storage media or transmitted over some form of communication medium.

[0095] Although the subject matter has been described in terms of specific structural features and / or methodological acts, it is understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

Claims

1. 1. A system comprising: a plurality of computing devices at a first geographic location capable of communicating over a network with a plurality of computing devices at a second geographic location remote from the first geographic location; At least one of the computing devices at the first geographic location: partitioning metadata into a plurality of partitions, the metadata corresponding to object data stored at a storage location, a leader computing device and at least one follower computing device of the plurality of computing devices at the first geographic location maintaining a copy of the metadata included in a first partition of the plurality of partitions, the leader computing device and the at least one follower computing device exchanging periodic heartbeat communications associated with the first partition; determining that at least one computing device at the second geographic location has a lower access load frequency than at least one of the leader computing device or the at least one follower computing device at the first geographic location; migrating the metadata of the first partition to at least one of the computing devices at the second geographic location, wherein a leader computing device and at least one follower computing device at the second geographic location maintain the copy of the metadata of the first partition and exchange periodic heartbeat communications related to the first partition; A system configured to perform operations including:

2. The operation includes: receiving a data structure from one of the computing devices at the second geographic location, the data structure indicating an access load frequency at one or more of the computing devices at the second geographic location; determining that the at least one computing device at the second geographic location has a lower access load frequency based at least in part on the received data structure; and The system of claim 1 further comprising:

3. The operation includes: receiving one or more additional data structures from one or more additional geographic locations, respectively; generating a composite data structure including partition information for partitions maintained at the first geographic location, the second geographic location, and the one or more additional geographic locations; The system of claim 2 further comprising:

4. The operation includes: splitting a second partition to obtain the first partition and a third partition prior to migrating the metadata of the first partition; migrating the metadata of the first partition after the division of the second partition; The system of claim 1 further comprising:

5. 5. The system of claim 4, wherein the splitting of the second partition to obtain the first partition and the third partition is based at least in part on splitting the second partition based at least in part on a key space range of the second partition.

6. The system of claim 1 , wherein the network is a wide area network.

7. locking the first partition to prevent writing of metadata prior to migrating the metadata of the first partition to the computing device at the second geographic location; receiving an indication that the leader computing device and the at least one follower computing device maintain the copy of the metadata of the first partition; sending an instruction to the leader computing device at the first geographic location to delete the copy of the metadata for the first partition; The system of claim 1 further comprising:

8. The operation includes: sending a migration request to at least one computing device in the second geographic location prior to migrating the metadata of the first partition; receiving an identifier associated with the leader computing device at the second geographic location in response to the migration request; The system of claim 1 further comprising:

9. The operation includes: migrating the metadata of the first partition by causing the leader computing device at the first geographic location to transmit a snapshot of the metadata of the first partition to the leader computing device at the second geographic location based at least in part on the identifier received in response to the migration request. The system of claim 8 further comprising:

10. The system of claim 1 , wherein the object data is stored in one or more network storage locations accessible over a wide area network.

11. The system of claim 1 , wherein there are at least two follower computing devices in the first geographic location, and the leader computing device and the at least two follower computing devices exchange periodic heartbeat communications associated with the first partition.

12. 1. A method comprising: partitioning data into a plurality of partitions by a computing device of a plurality of computing devices at a first geographic location, where at least two computing devices at the first geographic location maintain copies of the data in a first partition of the plurality of partitions and exchange periodic heartbeat communications associated with the first partition; determining that at least one computing device at a second geographic location has a lower access load frequency than at least one of the computing devices at the first geographic location that maintains the copy of the data; migrating the copy of the data of the first partition to at least one computing device in the second geographic location to cause at least two computing devices in the second geographic location to maintain the copies of the data of the first partition and to exchange periodic heartbeat communications related to the first partition; A method comprising:

13. receiving a data structure from one of the computing devices at the second geographic location, the data structure indicating an access load frequency at one or more of the computing devices at the second geographic location; determining that the at least one computing device at the second geographic location has a lower access load frequency based at least in part on the received data structure; and The method of claim 12 further comprising:

14. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a computing device at a first geographic location, configure the computing device to perform operations; The operation includes: partitioning data into a plurality of partitions, wherein at least two computing devices at the first geographic location maintain copies of the data in a first partition of the plurality of partitions and exchange periodic heartbeat communications associated with the first partition; determining that at least one computing device at a second geographic location has a lower access load frequency than at least one of the computing devices at the first geographic location that maintains the copy of the data of the first partition; migrating the copy of the data of the first partition to at least one computing device in the second geographic location to cause at least two computing devices in the second geographic location to maintain the copies of the data of the first partition and to exchange periodic heartbeat communications related to the first partition; [0023] In one embodiment, the present invention relates to a computer-readable medium comprising:

15. The operations include: receiving a data structure from one of the computing devices at the second geographic location, the data structure indicating an access load frequency at one or more of the computing devices at the second geographic location; determining that the at least one computing device at the second geographic location has a lower access load frequency based at least in part on the received data structure; and The one or more non-transitory computer-readable media of claim 14, further comprising:

Citation Information

Patent Citations

  • System and method for splitting a replicated data partition

    US8930312B1