Dynamic minimum replication performance

The system dynamically adjusts replication performance based on node loads to optimize resource utilization and ensure timely data protection, addressing inefficiencies in manual scheduling and resource utilization.

WO2025239891A1PCT designated stage Publication Date: 2025-11-20HITACHI VANTARA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/029385
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Existing data replication systems require manual scheduling and understanding of system topologies, leading to inefficient resource utilization and potential delays in data protection, especially when systems are idle or overloaded.

Method used

A computing node dynamically determines the highest node load across systems and adjusts the number of threads for replication based on this load, allowing for maximum resource utilization without disrupting other services.

Benefits of technology

Ensures timely replication with minimal disruption by maximizing resource use when systems are idle and increasing performance when needed to meet replication goals, eliminating the need for manual scheduling and ensuring data protection within specified timeframes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024029385_20112025_PF_FP_ABST
    Figure US2024029385_20112025_PF_FP_ABST
Patent Text Reader

Abstract

In some examples, a first computing node associated with a first computing system is able to communicate over one or more networks with a second computing system having at least one second computing node. The first computing node determines a first node load on the first computing node and at least one second node load on the at least one second computing node. The first computing node determines a highest node load from amongst the first node load and the at least one second node load. Further, based at least on the highest node load, the first computing node determines a number of threads to use for replicating data to the second computing system. The first computing node sends the replication data to the second computing system using the number of threads determined based at least on the highest node load.
Need to check novelty before this filing date? Find Prior Art

Description

DYNAMIC MINIMUM REPLICATION PERFORMANCETECHNICAL FIELD

[0001] This disclosure relates to the technical field of data replication.BACKGROUND

[0002] Data objects and other types of data may be stored in a first storage location and replicated to one or more other storage locations such as for redundancy, disaster recovery, and the like. Conventionally, an administrator or other user may be required to schedule the performance level of replication for a computing system that manages data storage and replication.For example, to accomplish this task and to avoid interfering with other system operations, the user should have an understanding of the timing and topologies of the source computing system that originally receives the data and the destination computing system(s) to which the data will be replicated. Additionally, the user may need to understand and specify a performance level for the replication to ensure that the replication keeps up with the demands on the source computing system. Furthermore, in some cases, when at least some of the computing systems in the topology are idle, replication processing may be unduly limited, such as based on a calendar setting or the like, so that the idle resources are not fully utilized. In addition, in a situation in which replication is not keeping pace with newly added data, the newly added data may not be protected in a timely matter; however, replication may still be limited by the user- specified performance level and is not able to keep up with the replication processing demands for the newly added data.SUMMARY

[0003] Some implementations include a first computing node associated with a first computing system that is able to communicate over one or more networks with a second computing system having at least one second computing node. The first computing node determines a first node load on the first computing node and at least one second node load on the at least one second computing node. The first computing node determines a highest node load from amongst the first node load and the at least one second node load. Further, based at least on the highest node load, the first computing node determines a number of threads to use for replicating data to the second computing system. The first computing node sends the replication data to the second computing system using the number of threads determined based at least on the highest node load.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] The detailed description is set forth with reference to the accompanying figures. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The use of the same reference numbers in different figures indicates similar or identical items or features.

[0005] FIG. 1 illustrates an example architecture of a system able to dynamically adjust replication processing performance according to some implementations.

[0006] FIG. 2 is a block diagram illustrating an example of replication in which a computing system has a fully loaded computing node according to some implementations.

[0007] FIG. 3 is a block diagram illustrating an example of replication in which a computing system has a fully loaded computing node according to some implementations.

[0008] FIG. 4 is a block diagram illustrating an example of replication in which a computing system has a fully loaded computing node according to some implementations.

[0009] FIG. 5 is a block diagram illustrating an example of replication in which the computing system topology is considered idle according to some implementations.

[0010] FIG. 6 is a flow diagram illustrating an example process for performing replication processing according to some implementations.

[0011] FIG. 7 illustrates select example components of the computing nodes that may be used to implement at least some of the functionality of the systems described herein.

[0012] FIG. 8 illustrates select components of an example configuration of a storage subsystem according to some implementations.DESCRIPTION OF THE EMBODIMENTS

[0013] Some implementations herein are directed to techniques and arrangements for utilizing available resources in a computing system when the resources become available to ensure timely replication with minimal or no disruption to other services being provided by the computing system. Some examples herein eliminate the need for the user to understand other activities when scheduling replication and can remove the necessity of scheduling replication. For instance, the examples herein may maximize replication performance without disrupting other services, and can allow replication to use all currently available resources. Furthermore, the examples herein may maintain sufficient replication performance to achieve time to safety goals and can ensure that replication completes within a specified time based on set user requirements.

[0014] Some examples herein include the ability to allow replication to use all or a large amount of available resources. For instance, a process according to some implementations hereinmay use maximum performance and throughput when other services and programs in the system are idle, and may use minimum performance when the system is fully loaded (e.g., as determined based on one or more measured metrics, or the like, indicative of system component utilization). Further, the process may vary performance between the minimum replication performance level and the maximum replication performance level based on a highest- loaded system (e.g., source system or destination system(s)) involved in the replication process. For instance, the process may determine the percentage of the utilization of the resources of each system involved in the replication process and may set the replication performance level based on the percentage of utilization of resources of the most-loaded system.

[0015] In addition, the examples herein allow replication to compete for resources when needed in order to achieve goals. For example, the replication process may use a minimum performance level (e.g., within the resources available from all the involved systems) when replication is not in danger of missing goals. For instance, this may even mean performing no replication when a system is fully loaded, and replication is caught up within a permitted threshold. However, when a replication goal is at risk of not being met, the replication process may increase the minimum performance level and may vary the minimum performance level between an absolute minimum and a scheduled or maximum minimum even though resources may be taken from other programs or services executing on the system. For instance, the replication process may use the maximum minimum performance level when any data object has not been replicated by the established goal. In some cases, a set goal can be more aggressive than an actual goal so this state might be achieved before the actual goal which is preferably never missed. Additionally, in some cases, batch sizes may be adjusted to respond quickly when the load on one or more of the involved systems increases.

[0016] Some examples herein enable elimination of the need for a user to schedule replication. Furthermore, the scheduled performance levels herein are able to default to higher values without having to be concerned about other activity in the replication topology. For example, the replication may complete work as fast as possible with minimal disruption to other services. Additionally, the replication processing is able to meet any desired goal such as ensuring that all items awaiting replication are replicated within a set time period. Further, the replication processing herein is able to use more resources that are available when the systems are less busy, and replication does not compete for resources when the computing systems are busy, and the replication processing has not fallen behind. In addition, the replication processing is able to increase the performance level as necessary to avoid falling behind and without completely blocking the activity of other system services and programs. Furthermore, some examples hereinmay completely eliminate a replication calendar and, instead, may include replication limitations based on link activity details all the way to a destination system so that a user never has to specify an off period for replication. For instance, a maximum minimum replication setting may allow the system to perform replication processing on an as needed basis.

[0017] For discussion purposes, some example implementations are described in the environment of multiple computing systems able to communicate for performing replication of data, such as over a network or the like. However, implementations herein are not limited to the particular examples provided, and may be extended to other types of computing system architectures, other types of storage environments, other types of client configurations, other types of data, and so forth, as will be apparent to those of skill in the art in light of the disclosure herein.

[0018] FIG. 1 illustrates an example architecture of a system 100 able to dynamically adjust replication processing performance according to some implementations. The system 100 includes a first computing system 102 able to communicate with at least one second computing system 104 over one or more networks 106. Furthermore, at least the first computing system 102 or the second computing system 104 is able to communicate with a plurality of client devices 108 over the one or more networks 106.

[0019] In the illustrated example, the first computing system 102 may be physically located, at least in part, at a first site 109, and may include a plurality of computing nodes 110, such as a first computing node 110- 1 , a second computing node 110-2,..., and so forth. Similarly, the second computing system 104 may be physically located, at least in part, at a second site 111, and may include a plurality of computing nodes 112 such as a first computing node 112-1, a second computing node 112-2, ..., and so forth. Each computing node 110, 112 may include one or more physical computing devices configured to execute various services, programs, algorithms, operations, or the like, as discussed additionally below. In some examples, the first site 109 may be geographically located sufficiently distant from the second site(s) 111 to provide for redundancy and disaster recovery of data, such as by being located in different cities, different states, different countries, and so forth.

[0020] The computing nodes 110 at the first computing system 102 may be able to communicate with one or more storage subsystems 114, and the computing nodes 112 at the second computing system 104 may be able to communicate with one or more storage subsystems 116. In some examples, the computing nodes 110, 112 may include access nodes, server nodes, management nodes, and / or other types of service nodes that provide the client devices 108 with access to storage provided by the storage subsystems 114, 116, respectively, for enabling the clientdevices 108 to store and access data, as well as performing other management and control functions, as discussed additionally below.

[0021] In some examples, the computing nodes 110, 112 may include one or more servers that may be embodied in any number of ways. For instance, the programs, other functional components, and at least a portion of data storage of the computing nodes 110, 112 may be implemented on at least one server, such as in a cluster of servers, a server farm, a data center, a cloud-hosted computing service, and so forth, although other computer architectures may additionally or alternatively be used. As another example, the computing nodes 110, 112 may be abstracted as or otherwise treated as a single node that is actually a cluster containing multiple computing nodes 110 or 112. Additional hardware and software configuration details of the computing nodes 110, 112 are discussed below with respect to FIG. 7.

[0022] The one or more networks 106 may include any suitable network, including a wide area network, such as the Internet; a local area network (LAN), such as an intranet; a wireless network, such as a cellular network, a local wireless network, such as Wi-Fi, and / or short-range wireless communications, such as BLUETOOTH®; a wired network including Fibre Channel, fiber optics, Ethernet, or any other such network, a direct wired connection, or any combination thereof. Accordingly, the one or more networks 106 may include both wired and / or wireless communication technologies. Components used for such communications can depend at least in part upon the type of network, the environment selected, or both. Protocols for communicating over such networks are well known and will not be discussed herein in detail. Further, the network(s) 106 may be a public network that may include the Internet in some cases, or a combination of public and private networks. Implementations herein are not limited to any particular type of network as the network(s) 106.

[0023] The computing nodes 110, 112 may be configured to provide storage and data management services to client users 118 via the client devices 108, respectively. As several nonlimiting examples, the users 118 may include users performing functions for businesses, enterprises, organizations, governmental entities, academic entities, or the like, and which may include storage of very large quantities of data in some examples. Nevertheless, implementations herein are not limited to any particular use or application for the system 100 and the other systems and arrangements described herein. Further, the users 118 may include administrative users who use one or more of the client devices to 108 to configure settings of the first computing system 102 and / or the second computing system 104.

[0024] Each client device 108 may be any suitable type of computing device such as a desktop, laptop, tablet computing device, mobile device, smart phone, wearable device, terminal, and / orany other type of computing device able to send data over the one or more networks 106, or otherwise communicate with the first computing system 102 and / or the second computing system 104. Users 118 may be associated with respective client devices 108 such as through a respective user account, user login credentials, or the like. Furthermore, the client devices 108 may be configured to communicate with the computing nodes 110, 112 through the one or more networks 106, through separate networks, or through any other suitable type of communication connection. Numerous other variations will be apparent to those of skill in the art having the benefit of the disclosure herein.

[0025] Each client device 108 may include a respective instance of a user application 120 that may execute on the client device 108, such as for communicating with a client application 122 executable on the computing nodes 110, 112, such as for sending user data for storage on the storage subsystems 114, 116 and / or for receiving stored data from the storage subsystems 114, 116 through a data instruction 124, such as a write operation, read operation, delete operation, or the like. In some cases, the user application 120 may include a browser or may operate through a browser, while in other cases, the user application 120 may include any other type of application having communication functionality enabling communication over the one or more networks 106 with the client application 122 or other application on the computing nodes 110, 112. Accordingly, the computing nodes 110, 112 may provide storage capacity for the users 118 and respective client devices 108. During steady state operation there may be a plurality of users 118 periodically communicating with the computing nodes 110, 112.

[0026] In addition, when a user 118 is acting as an administrator, the user 118 may be associated with one of the client devices 108 such as through a respective administrator account, administrator login credentials, or the like. For example, the user 118 acting as an administrator may be able to communicate through the client device 108 with the computing nodes 110, 112 through the one or more networks 106 (or through any other suitable type of communication connection), such as for sending management instructions for managing the system 100, such as instructions for one or more instances of a replication management program 128 that may execute on at least one computing node 110, 112, in each computing system 102, 104, respectively. In addition, the user 118 acting as an administrator may send management data for storage on the storage subsystems 114, 116 and / or may send one or more instructions for receiving stored management data from the storage subsystems 114, 116.

[0027] At least some of the computing nodes 110, 112 may execute the replication management program 128, which may dynamically manage replication processing of data 130 to be replicated to another computing system using the techniques and processes described herein.In addition, at least some of the computing nodes 110, 112 may execute a storage management program (not shown in FIG. 1) for managing storage of data received from the client devices 108 to the storage subsystems 114, 116 and / or for retrieving data from the storage subsystems 114, 116.

[0028] Each storage subsystem 114, 116 may include a storage 134 for storing received client data 136 and replicated data 132 received from other systems. For instance, the received client data 136 may include data 138 already replicated to a destination system, and the data 130 to be replicated. Further, each storage subsystem 114, 116, may execute a storage program (not shown in FIG. 1) for managing data client data 136 received from the computing nodes 110, 112, and may store the client data 136 on one or more storage devices (not shown in FIG. 1) in the storage 134. Additionally, in some examples, the computing nodes 110, 112 may use on-board storage devices in addition to or in place of the storage subsystems 114, 116, respectively for storage of at least some of the above-discussed data, such as recently received user data.

[0029] The storage subsystems 114, 116 may include any of various types of storage systems such as local storage systems, or network storage systems, such as publicly provided commercial storage, cloud storage, private network storage, or the like. Accordingly, some of the storage subsystems 114, 116 may be located remotely from the sites 109, 111 respectively such as at separate data centers, or the like. Typically however the storage subsystems 114 may be located separately geographically from the storage subsystems 116, such as to provide the above discussed redundancy and disaster recovery ability.

[0030] In the examples herein, at least one computing node 110, 112, at each computing system 102, 104, respectively may execute the replication management program 128 for ensuring that data is replicated from a source computing system 102, 104, to at least one destination computing system 102, 104. For example, in the case of replication of replication data 140 from the first computing system 102 to the second computing system 104, the replication management program 128 may determine a node load on a source computing node 110 and all possible destination computing nodes 112 at the destination computing system. As one example, the node load may be determined as the maximum of several resources used for replication processing, which may include memory usage, database usage, current data ingest traffic, disk usage, network usage, and the replication across all links. Furthermore, the node load may be determined for both the source computing node and all of the possible destination computing nodes at the destination system for determining the maximum node load on each of the computing nodes 110, 112 that might be involved in the replication. Then, the highest maximum load on any of these computingnodes 110 or 112 may be used as the current maximum load for dynamically determining replication performance settings, such as a number of replication threads, or the like.

[0031] Similarly, when replication data 142 is to be replicated from the second computing system 104 to the first computing system 102, the computing node 112 on the second computing system may determine the node load on itself and also on all possible computing nodes at the first computing system 102. The computing node on the second computing system may then determine the maximum node load among all of the computing nodes that may participate in the data replication processing, and may set the replication processing performance level according to the maximum node load from among any of the computing nodes 110 or itself.

[0032] Additionally, in some examples, a calendar setting may be used to set the maximum minimum performance level for replication. For example, a user can use the calendar to adjust this value or to turn replication off entirely if desired for some reason.

[0033] The arrangements herein take advantage of available resources and help ensure timely replication with minimal or no disruption to other services. In some cases, threads are used to establish and dynamically adjust replication performance. For example, a “thread” may include a sequence of programmed instructions for performing an operation (e.g., a replication operation in the examples herein) and that can be managed independently by a scheduler, such as a scheduling algorithm, or the like, that may be included in the replication management program 128.

[0034] In some examples herein, a maximum thread count may be used to achieve the highest possible throughput when the source and destination systems are otherwise idle. Additionally, an absolute minimum thread count may be used when a destination or source system is fully loaded (e.g., in some cases, the absolute minimum may be zero; however, in other cases, minimal replication progress may be required to be performed at all times by default, in which case a different absolute minimum such as a one thread or other desired absolute minimum may be set by the user.

[0035] Typically, replication may be expected to be completed within a specified time threshold to ensure adequate data redundancy or data recovery capability. Furthermore, in some examples, there may be built-in buffer times to ensure that the data replication is completed within the specified time and that the data recoverability specifications are met. For example, if replication is expected to be completed within one day, a completion goal may be set for two thirds of a day. Additionally, replication performance may be adjusted based on the completion goal and further based on system resource availability at any given time in the source and destination systems. As one example, suppose that until one half of the replication goal is met or one third of the day has elapsed, the replication management program 128 may perform replicationusing at least the absolute minimum, and may adjust the replication performance evenly to the maximum minimum.

[0036] For example, suppose that until one half of the replication goal is met or one third of a day has elapsed, the system may use the absolute minimum performance level for replication processing. Subsequently, between elapse of one third of the day and two thirds of the day, the system may adjust the replication performance level to the maximum minimum replication performance level until two thirds of the day has elapsed. If two thirds of the day has elapsed, the replication management program 128 may use at least the maximum minimum processing rate, or, in some cases, may draw additional resources from other parts of the system, to further increase the replication performance. For instance, the replication goal may be based on the age of the oldest data object not yet replicated. The number of threads devoted to replication may be increased as the goal becomes closer to not being met. Additionally, in some examples, a maximum thread count may also be set, and which is not exceeded despite the replication goal being behind schedule.

[0037] Progress toward achieving the replication goal may be determined based at least in part on database checkpoints that show the oldest data object modification that needs to be replicated. In some examples, the number of usable threads may be calculated as a number between the adjusted minimum and maximum threads and the highest load on the source computing node or the computing nodes in the destination system. Consequently, the examples herein eliminate the need for the user to understand other activities that the system may be performing when scheduling replication, and removes the necessity of scheduling replication. Further, the examples herein may maximize replication performance without disrupting other services, and may allow replication to use all available resources. In addition, the examples herein are able to maintain sufficient replication performance to achieve time-to-safety goals to ensure that replication completes within a timeframe based on user-set requirements.

[0038] FIG. 2 is a block diagram illustrating an example 200 of replication in which a computing system has a fully loaded computing node according to some implementations. This example assumes replication is not in any danger of missing a goal. In this example, suppose that a first computing system 102 and a second computing system 104 are configured for replication with each other and have a configuration such as system 100 discussed above with respect to FIG. 1. Furthermore, for simplifying the explanation, suppose that each computing system 102, 104 has two computing nodes; however, in other examples, a smaller or larger number of computing nodes may be included in each computing system 102, 104. Furthermore, while only two computing systems 102, 104 are illustrated in this example (i.e., one replication link for eachsystem 102, 104), in other examples, the replication topology is not limited to any particular number of computing systems to which replication may be performed. Accordingly, the second computing system 104 is the destination replication system for the first computing system 102, and the first computing system 102 is the destination replication system for the second computing system 104.

[0039] In this example, each computing node 110-1, 110-2, 112-1, 112-2 may be configured to replicate to an unspecified computing node in the destination computing system. Further, the thread counts may be determined based on applicable load. In this example, suppose that an absolute minimum thread count of 1 may be used, a scheduled thread count of 5 may be used, and a maximum thread count of 75 may be used. For instance, since this examples supposes that there is no danger of missing the time-to-safety replication goal, the absolute minimum of 1 thread may be used in some examples herein, e.g., for a time period corresponding to a first half of the replication goal. Additionally, the replication goal may be based on the age of the oldest object not yet replicated and the thread counts may be increased progressively as the replication goal becomes progressively more at risk of not being achieved.

[0040] When a computing node 110-1, 110-2, 112-1, 112-2 has replication to perform, the computing node may determine the current node load on itself and on all of the computing nodes in the destination computing system. For example, an instance of the replication management program 128 may be executed on each of the computing nodes 110-1, 110-2, 112-1, 112-2 to enable communication between the computing nodes 110-1, 110-2, 112-1, 112-2 for determining a current load on each computing node 110-1, 110-2, 112-1, 112-2 that may be provided to each other computing node 110-1, 110-2, 112-1, 112-2. Accordingly, actions performed by the computing nodes 110-1, 110-2, 112-1, 112-2 in these examples may be attributed to the respective instances of the replication management program 128. Alternatively of course, in other examples, alternative techniques for determining the load on each computing node in a computing system may be employed. For instance, as one alternative (not shown), each computing system 102, 104 may include a management node that may determine the load on each node in the computing system and provide this information to various nodes of the computing system and to other computing systems. Numerous other variations will be apparent to those of skill in the art having the benefit of the disclosure herein.

[0041] In this example suppose that the second computing node 112-2 at the second computing system 104 currently has 200 ingest connections 202 by which the second computing node 112-2 is receiving data instructions from a plurality of client devices 108. Additionally, suppose that the first computing node 110-1 and the second computing node 110-2 of the firstcomputing system 102 want to determine how many replication threads to use to replicate data to the second computing system 104. Accordingly, the first computing node 110-1 may determine that it is not currently under a load. Further, the first computing node 110-1 may communicate with the first computing node 112-1 and the second computing node 112-2 at the second 5 computing system 104 to determine the current node loads on those computing nodes 112-1 and 112-2. In response to communications from the first computing node 112-1 and the second computing node 112-2, suppose that the first computing node 110-1 determines that the second computing node 112-2 at the second computing system 104 is fully loaded due to the 200 current ingest connections 202 and that the first computing node 112-1 at the second computing system 10 is not currently loaded. Accordingly, as indicated at 204, the first community node 110-1 may perform replication using only one thread based on the second computing node 112-2 being fully loaded, even though the first computing node 112- 1 is not currently loaded. The second computing node 110-2 at the first computing system 102 may make the same determination and similarly may perform replication to the second computing system 104 using only a single thread as 15 indicated at 206.

[0042] On the other hand, suppose that the first computing node 112-1 is currently not loaded beyond a load threshold, e.g., otherwise idle, and desires to perform replication to the first computing system 102. The first computing node 112-1 may communicate with the first computing node 110-1 and the second computing node 110-2 at the first computing system 102 10 to determine the current loads on the computing nodes 110-1, 110-2. In this example, suppose that these computing nodes 110-1 and 110-2 are currently not loaded beyond a load threshold, e.g., idle. Accordingly, the first computing node 112-1 at the second computing system 104 may determine that it can perform replication up to the maximum allowed replication performance level, which in this example corresponds to 75 threads of replication processing as indicated at >5 208.

[0043] Additionally, the second computing node may determine that it is fully loaded and that the first computing node 110-1 and the second computing node 110-2 are not currently loaded within a threshold load indicating idleness. Consequently, based on the second computing node 112-2 currently being fully loaded, the second computing node 112-2 only performs replication 10 using one thread, as indicated at 210, even though the computing nodes 110-1 and 110-2 at the first computing system 102 crude receive significantly more replication processing.

[0044] FIG. 3 is a block diagram illustrating an example 300 of replication in which a computing system has a fully loaded computing node according to some implementations. This example assumes replication is in some danger of missing a goal. In this example, with conditionssimilar to the example 200 of FIG. 2, discussed above, suppose that the second computing node 112-2 has 200 ingest connections 202 similar to the example 200 discussed above with respect to FIG. 2. Additionally, in this example, suppose that the replication goal has been partially missed, e.g., one or more of the computing nodes 110-1, 110-2, 112-1, 112-2 has fallen behind the timing goal for replication processing. For example, suppose that the replication goal has not yet been missed but is getting closer to the potential of missing the goal. As one non-limiting case, this example may correspond to the middle of the second third of a day, e.g., the oldest unreplicated object is now at one-half day. Accordingly, this may correspond to a time during the second half of the goal where the minimum number of threads used progresses evenly between the absolute minimum and the maximum minimum once the goal is actually missed. As discussed above, the replication time goal may be based on the maximum time permitted for replication of a data object change to be completed, and the replication goal may be a percentage of the maximum permitted time, or may use some other technique for providing a buffer for completing replication before the maximum permitted time for replication has elapsed.

[0045] In this example, suppose that the replication time goal is two thirds of the maximum permitted time for replication. Furthermore, suppose that in this example one third of the maximum permitted time has elapsed and that the replication processing has fallen behind the replication goal. Accordingly, the number of replication threads may be increased on the fully loaded computing nodes 110-1, 110-2, and 112-2 to attempt to catch up with the replication goal. For example, as discussed above, based on the 200 ingest connections 202 to the second computing node 112-2, the second computing node 112-2 is fully loaded.

[0046] Additionally, because the second computing node 112-2 at the second computing system 104 is fully loaded, the replication management program 128 considers the first computing node 110-1 and the second computing node 110-2 at the first computing system 102 to also be fully loaded. Nevertheless, because the replication processing has fallen behind the replication timing goal for these computing nodes, the number of threads devoted to replication processing is increased from the previous absolute minimum of one thread to three threads in this example. Therefore, as illustrated, the first computing node 110-1 allocates three threads 302 to replication processing, the second computing node 110-2 allocates three threads 304 to replication processing, and the second computing node 112-2 at the second computing system 100 for also allocates three threads 308 to replication processing even though this may impact some resources being used for the 200 ingest connections 202. Furthermore, the first computing node 112-1 is not affected by the fully loaded condition of the second computing node 112-2 and may continue to allocate as many as 75 threads 306 to replication processing.

[0047] FIG. 4 is a block diagram illustrating an example 400 of replication in which a computing system has a fully loaded computing node according to some implementations. This example assumes replication is in maximum danger of missing a goal. In this example, with conditions similar to the examples 200 of FIG. 2 and 300 of FIG. 3, discussed above, suppose that the second computing node 112-2 has 200 ingest connections 202 similar to the examples 200 and 300 discussed above. Additionally, in this example, suppose that the replication goal has been completely missed, e.g., one or more of the computing nodes 110-1, 110-2, 112-1, 112-2 has failed to complete replication processing of one or more data objects within the replication time goal. As discussed above, the replication time goal may be based on the maximum time permitted for replication of a data object change to be completed, and the replication goal may be a percentage of the maximum permitted time, or may use some other technique for providing a buffer for completing replication before the maximum permitted time for replication has elapsed. For example, using the two thirds goal discussed above, suppose that two thirds of the maximum time permitted for replication has elapsed without the replication processing for one or more of the objects corresponding to the maximum time having been completed.

[0048] To address this situation, the computing nodes that are behind may increase the number threads used for replication processing. In this example, suppose that the user has specified that a scheduled thread count of five threads may be used as a maximum minimum performance level when the node is considered fully loaded. Accordingly, as illustrated, at the first computing system 102, the first computing node 110-1 may dedicate five threads 402 for replication processing and the second computing node 110-2 may designate five threads 404 for replication processing. Similarly, at the second computing system 104, the second computing node 112-2 may dedicate five threads 406 for replication processing, while the first computing node 112-1 may continue to use up to 75 threads 408. Furthermore, while five threads has been designated as the maximum minimum performance level thread count in this example, in other examples, the thread count may be a different number, such as based on a higher performance level that may take resources from other services and programs also being performed by the loaded computing node(s).

[0049] FIG. 5 is a block diagram illustrating an example 500 of replication in which the computing system topology is considered idle according to some implementations. In this example, with conditions similar to the example 200 of FIG. 2, discussed above, suppose that all of the computing nodes 110-1, 110-2, 112-1, 112-2 in the first computing system 102 and the second computing system 104, respectively, have sufficiently low loads (e.g., below a threshold load) to qualify as being designated to be currently idle. For example, suppose that the second computing node 112-2 is no longer receiving ingest from the client devices 108, and that the loadson the second computing node 112-2 as well as the other nodes 110-1, 110-2, and 112-1 are below a threshold load level.

[0050] In this case, in which all of the computing nodes in both the first computing system 102 and the second computing system 104 are designated to be idle, the computing nodes 110-1, 110-2, 112-1, 112-2 may perform replication at the maximum permitted replication performance level, i.e., 75 threads in this example. Accordingly, at the first computing system 102, the first computing node 110-1 may use up to 75 threads 502 for replication processing, and the second computing node 110-2 may also use up to 75 threads 504 for replication processing. Similarly, at the second computing system 104, the first computing node 112-1 may use up to 75 threads 506 for replication processing, and the second computing node 112-2 may also use up to 75 threads 508 for replication processing. Furthermore, while 75 threads is provided as an example of a maximum replication processing performance level in this example, in other examples any other suitable number of threads may be designated to be the maximum performance level for replication processing, as will be apparent to those of skill in the art having the benefit of the disclosure herein. For example, the number threads available may be dependent, at least in part, on the type and number of processors included in each of the computing nodes 110, 112.

[0051] FIG. 6 is a flow diagram illustrating an example process 600 for performing replication processing according to some implementations. The process 600 is illustrated as a collection of blocks in a logical flow diagram, which represents a sequence of operations, some or all of which may be implemented in hardware, software or a combination thereof. In the context of software, the blocks may represent computer-executable instructions stored on one or more computer- readable media that, when executed by one or more processors, program the one or more processors to perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures and the like that perform particular functions or implement particular data types. The order in which the blocks are described should not be construed as a limitation. Any number of the described blocks can be combined in any order and / or in parallel to implement the process, or alternative processes, and not all of the blocks need to be executed. For discussion purposes, the process 600 is described with reference to the environments, frameworks, and systems described in the examples herein, although the process 600 may be implemented in a wide variety of other environments, frameworks, and systems.

[0052] In the example of FIG. 6, the process 600 may be executed at least in part by at least one of the computing nodes 110, 112 executing the replication management program 128 or the like. For instance in some examples any of the computing nodes 110, 112 may execute the process 600 independently of the other computing nodes 110, 112.

[0053] At 602, the computing node may determine whether there is data waiting to be replicated to another system. If so, the process goes to 606. If not, the process goes to 604.

[0054] At 604, when the computing node does not have any data waiting to be replicated, the computing node may wait until replication processing is indicated. For example, when the 5 computing node receives data such as a change to an existing data object, a new data object, or the like, replication may be indicated for providing redundancy and disaster recovery for the corresponding data object.

[0055] At 606, when there is data waiting to be replicated, the computing node may determine current node loads on the source computing node and any possible destination system computing L0 nodes. Determine the current node load on a computing load may include determining one or more of the following: determining a number of incoming connections (e.g., data ingest for the computing node); determining a portion of processor usage above a processor usage threshold; determining a portion of database semaphores currently in use for a metadata database for the stored data (for instance, the use of semaphores enables multiple computing nodes to access the L5 same database simultaneously while maintaining continuity of the database); determining a portion of memory usage above a memory usage threshold; determining a portion of replication activity using a link between the source computing system and the destination computing system; and / or determining an amount of current storage read / write activity in comparison to the total read / write bandwidth for the storage. The node load may be determined from some or all of these 10 parameters as the highest portion of the above discussed parameters and / or any other applicable loads that may be placed on the respective computing nodes.

[0056] At 608, the computing node may determine the highest node load from among the source node and the destination system nodes. For instance, as discussed above, the highest node load may be selected from among all the node loads determined for the source node and all the 15 computing nodes at the destination system.

[0057] At 610, the computing node may determine a number of threads to use for replication based on the selected highest node load determined at 608. For example, the number threads may be calculated based on the dynamic minimum number of threads determined based on a first threshold time goal for completing replication. Further, the process may use the absolute minimum 10 number of threads if the age of the oldest object to be replicated is under a second time threshold.As one example, the dynamic minimum number of threads may be calculated as:(portion of time of oldest object not replicated between second time threshold and the first time threshold) * (configured minimum - absolute minimum) + absolute minimum

[0058] Further, the number of threads may be calculated as:Maximum threads - applicable load * (maximum threads - dynamic minimum threads) wherein the maximum threads is the number of threads that may be allocated to replication processing when all the computing nodes are considered to be idle.5

[0059] The configured or scheduled minimum number of threads may be used for replication after an object has not been replicated by the replication goal. Additionally, if the highest node load indicates that at least one of the computing nodes is currently fully loaded (i.e., has a node load that exceeds a threshold of node load for being designated fully loaded), and the replication goal is not in danger of being missed, then the replication may be performed using the absolute L0 minimum number of threads. On the other hand, if none of the nodes of interest is fully loaded, then the maximum minimum number threads may be used. As still another alternative, if all of the nodes are considered to be idle based on the respective node loads being below a node load minimum threshold, then the maximum number of threads may be used.

[0060] At 612, the computing node may calculate the batch size for replication. For example, L5 a large batch size may be used to optimize replication throughput, but may include a maximum count. On the other hand, a smaller batch size may be used to react more quickly to a load increase on one or more of the computing nodes, and may include a minimum count. As one example, the batch size may be calculated as: applicable load * (maximum count - minimum count) + minimum count 0

[0061] During batch collection, the batch size may also be limited by a total data size for data transfers. However, this limit does not apply to metadata transfers.

[0062] At 614, the computing node may begin transferring the calculated batch of objects using the determined number of threads as replication data. For instance, the replication data may be sent to the destination computing system, which may select one or more computing nodes to 5 receive and process the replication data.

[0063] At 616 the computing node may check the progress of the replication. For example, progress toward achieving the replication goal may be determined based at least in part on database checkpoints that indicate the oldest data that needs to be replicated.

[0064] At 618, based at least on the oldest data that needs to be replicated, the computing node 10 may determine whether there is more work in the current process, e.g., whether the data replication is complete or not. If there is no more work, the process may proceed to 624. On the other hand, if there is still replication processing to be performed the process may proceed to 620.

[0065] At 620, the computing node may recalculate the thread limit. For example, if the replication processing has fallen behind, the computing node may increase the number threads dedicated to replication processing such as discussed above with respect to FIGS. 3 and 4.

[0066] At 622, the computing node may determine whether there are still enough threads available for performing the replication processing according to the recalculated thread limit. If so, the process goes back to 612 to determine a next batch size. If not, the process goes to 624.

[0067] At 624, the computing node may determine whether a current threat is the last thread in the current batch of replication processing. If so, the process goes to 604. If not, the process goes to 626.

[0068] At 626, the computing node may terminate the current thread and return to 604.

[0069] The example processes described herein are only examples of processes provided for discussion purposes. Numerous other variations will be apparent to those of skill in the art in light of the disclosure herein. Additionally, while the disclosure herein sets forth several examples of suitable frameworks, architectures and environments for executing the processes, implementations herein are not limited to the particular examples shown and discussed. Furthermore, this disclosure provides various example implementations, as described and as illustrated in the drawings. However, this disclosure is not limited to the implementations described and illustrated herein, but can extend to other implementations, as would be known or as would become known to those skilled in the art.

[0070] FIG. 7 illustrates select example components of the computing nodes 110, 112 that may be used to implement at least some of the functionality of the systems described herein. The computing nodes 110, 112 may include one or more servers or other types of computing devices that may be embodied in any number of ways. For instance, in the case of a server, the programs, other functional components, and data may be implemented on a single server, a cluster of servers, a server farm or data center, a cloud-hosted computing service, and so forth, although other computer architectures may additionally or alternatively be used. Multiple computing nodes may be located together or separately, and organized, for example, as virtual servers, server banks, and / or server farms. The described functionality may be provided by the servers of a single entity or enterprise, or may be provided by the servers and / or services of multiple different entities or enterprises.

[0071] In the illustrated example, the computing nodes 110, 112 each include, or may have associated therewith, one or more processors 702, one or more computer-readable media 704, and one or more communication interfaces 706. Each processor 702 may be a single processing unit or a number of processing units, and may include single or multiple computing units, or multipleprocessing cores. The processor(s) 702 can be implemented as one or more central processing units, microprocessors, microcomputers, microcontrollers, graphics processors, digital signal processors, system-on-chip processors, Al processors, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. As one example, the processor(s) 702 may include one or more hardware processors and / or logic circuits of any suitable type specifically programmed or configured to execute the algorithms and processes described herein. The processor(s) 702 may be configured to fetch and execute computer-readable instructions stored in the computer-readable media 704, which may program the processor(s) 702 to perform the functions described herein.

[0072] The computer-readable media 704 may include volatile and nonvolatile memory and / or removable and non-removable media implemented in any type of technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. For example, the computer-readable media 704 may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, optical storage, solid state storage, magnetic tape, magnetic disk storage, RAID storage systems, storage arrays, network attached storage, storage area networks, cloud storage, or any other medium that can be used to store the desired information and that can be accessed by a computing device. Depending on the configuration of the computing nodes 110, 112, the computer-readable media 704 may be a tangible non-transitory medium to the extent that, when mentioned, non-transitory computer- readable media exclude media such as energy, carrier signals, electromagnetic waves, and / or signals per se. In some cases, the computer-readable media 704 may be at the same location as the service computing device 102, while in other examples, the computer-readable media 704 may be partially remote from the service computing device 102. For instance, in some cases, the computer-readable media 704 may include a portion of storage in the subsystem(s) 114, 116 discussed above with respect to FIG. 1, and as discussed additionally below with respect to FIG. 8.

[0073] The computer-readable media 704 may be used to store any number of functional components that are executable by the processor(s) 702. In many implementations, these functional components comprise instructions or programs that are executable by the processor(s) 702 and that, when executed, specifically program the processor(s) 702 to perform the actions attributed herein to the service computing device 102. Functional components stored in the computer-readable media 704 may include the client application 122 and the replication management program 128 discussed above, as well as a storage management program 708, each of which may include one or more computer programs, applications, executable code, or portions thereof. For example, the storage management program 708 may configure the computing nodesto store and manage client data, such as for storage in the storage subsystems. Further, while these programs are illustrated together in this example, during use, some or all of these programs may be executed on separate computing nodes 110, 112.

[0074] In addition, the computer-readable media 704 may store data, data structures, and other 5 information used for performing the functions and services described herein. For example, the computer-readable media 704 may store node load information 710 and at least a portion of a metadata database 712. For example, the node load information 710 may include node load information for the computing node as well as node load information received from other computing nodes. Further, while these data and data structures are illustrated together in this L0 example, during use, some or all of these data structures may be stored on separate computing nodes 110, 112. The service computing device 102 may also include or maintain other functional components and data, which may include programs, drivers, etc., and the data used or generated by the functional components. Further, the service computing device 102 may include many other logical, programmatic, and physical components, of which those described above are merely L5 examples that are related to the discussion herein.

[0075] The one or more communication interfaces 706 may include one or more software and hardware components for enabling communication with various other devices, such as over the one or more networks 106. For example, the communication interface(s) 706 may enable communication through one or more of a LAN, the Internet, cable networks, cellular networks, 10 wireless networks (e.g., Wi-Fi) and wired networks (e.g., Fibre Channel, fiber optic, Ethernet), direct connections, as well as close-range communications such as BLUETOOTH®, and the like, as additionally enumerated elsewhere herein.

[0076] FIG. 8 illustrates select components of an example configuration of a storage subsystem 114, 116 according to some implementations. The storage subsystem 114, 116 may 15 include one or more storage computing devices 802, which may include one or more servers or any other suitable computing device, such as any of the examples discussed above with respect to the computing nodes 110, 112. The storage computing device(s) 802 may each include one or more processors 804, one or more computer-readable media 806, and one or more communication interfaces 808. For example, the processor(s) 804 may correspond to any of the examples 10 discussed above with respect to the processors 702, the computer-readable media 806 may correspond to any of the examples discussed above with respect to the computer-readable media 704, and the communication interface(s) 808 may correspond to any of the examples discussed above with respect to the communication interfaces 706.

[0077] In addition, the computer-readable media 806 may include a storage program 810 as a functional component executed by the one or more processors 804 for managing the storage of the data 818 on the storage 134 associated with the storage subsystem 114, 116. The storage 134 may include one or more controllers 812 for storing the data 818 on one or more trays, racks, extent groups, or other types of arrays 814 of storage devices 816. For instance, the controller 812 may control the arrays 814, such as for configuring the arrays 814, such as in an erasure coded protection configuration, or any of various other configurations, such as a RAID configuration, JBOD configuration, or the like, and / or for presenting storage extents, logical units, or the like, based on the storage devices 816 to the storage program 810, and for managing the data 818 stored on the underlying physical storage devices 816. The storage devices 816 may be any type of storage device, such as hard disk drives, solid state drives, optical drives, magnetic tape, combinations thereof, and so forth. Further, as discussed above, the data 818 may include received client data 136, including data 130 to be replicated and data 138 already replicated, and also may include replicated data 132 received from other systems.

[0078] Various instructions, methods, and techniques described herein may be considered in the general context of computer-executable instructions, such as computer programs and applications stored on computer-readable media, and executed by the processor(s) herein. Generally, the terms program and application may be used interchangeably, and may include instructions, routines, modules, objects, components, data structures, executable code, etc., for performing particular tasks or implementing particular data types. These programs, applications, and the like, may be executed as native code or may be downloaded and executed, such as in a virtual machine or other just-in-time compilation execution environment. Typically, the functionality of the programs and applications may be combined or distributed as desired in various implementations. An implementation of these programs, applications, and techniques may be stored on computer storage media or transmitted across some form of communication media.

[0079] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

Claims

CLAIMS1. A system comprising: a first computing node associated with a first computing system and able to communicate over one or more networks with a second computing system having at least one second computing node associated therewith, the first computing node configured by executable instructions to perform operations comprising: determining, by the first computing node, a first node load on the first computing node and at least one second node load on the at least one second computing node; determining, by the first computing node, a highest node load from amongst the first node load and the at least one second node load; determining, by the first computing node, a number of threads to use for replicating data to the second computing system, wherein determining the number of threads is based at least on the highest node load; and sending, by the first computing node, replication data to the second computing system using the number of threads determined based at least on the highest node load.

2. The system as recited in claim 1, the operations further comprising: based at least on determining that the replication of the data to the second computing system has fallen behind a timing goal for completing the replication, increasing the number of threads used by the first computing node for sending the replication data to the second computing system.

3. The system as recited in claim 2, the operation of determining that the replication has fallen behind the timing goal further comprising: determining an oldest data object change in a database; and determining that the replication of the data object change has fallen behind the timing goal based at least on the oldest data object in the database.

4. The system as recited in claim 1, wherein the operation of determining the first node load on the first computing node and the at least one second node load on the at least one second computing node is based on determining at least one of: a number of incoming connections from client devices; a portion of processor usage above a processor usage threshold;a portion of database usage of a metadata database; a portion of memory usage above a memory usage threshold; a portion of replication activity using a link between the first computing system and the second computing system; or an amount of current storage access activity.

5. The system as recited in claim 1, the operations further comprising: based at least on determining that the highest node load exceeds a first node load threshold, determining that the number of threads to use for replicating the data to the second computing system corresponds to a minimum thread number.

6. The system as recited in claim 1, the operations further comprising: based at least on determining that the highest node load exceeds a second node load threshold, but does not exceed a first node load threshold indicating a higher load than the second node load threshold, determining that the number of threads to use for replicating the data to the second computing system corresponds to a thread number between a minimum thread number and a maximum thread number.

7. The system as recited in claim 1, the operations further comprising: based at least on determining that the highest node load fails to satisfy a node load threshold, indicating that a current load on the first computing node and the at least one second computing node is low enough to classify the first computing node and the at least one second computing node as idle, determining that the number of threads to use for replicating the data to the second computing system corresponds to a maximum number of threads permitted to be allocated for replication.

8. The system as recited in claim 1, the operations further comprising determining a batch size for replicating the data to the second computing device based at least in part on an amount of the data for replication and the highest node load, wherein sending the replication data to the second computing system using the number of threads determined based at least on the highest node load further comprises sending a batch of data based on the determined batch size.

9. The system as recited in claim 1, the operations further comprising: receiving, by the first computing node, from a second computing node of the at least one second computing nodes associated with the second computing system, an inquiry regarding a current node load on the first computing node; sending, by the first computing node, to the second computing node, an indication of the first node load on the first computing node; and based at least on sending the first node load, receiving, by the first computing node, second computing node replication data from the second computing node.

10. A method comprising: determining, by a first computing node associated with a first computing system and able to communicate over one or more networks with a second computing system having at least one second computing node associated therewith, a first node load on the first computing node. determining, by the first computing node, at least one second node load on the at least one second computing node; determining, by the first computing node, a highest node load from amongst the first node load and the at least one second node load; determining, by the first computing node, a number of threads to use for replicating data to the second computing system, wherein determining the number of threads is based at least on the highest node load; and sending, by the first computing node, replication data to the second computing system using the number of threads determined based at least on the highest node load.

11. The method as recited in claim 10, further comprising: based at least on determining that the replication of the data to the second computing system has fallen behind a timing goal for completing the replication, increasing the number of threads used by the first computing node for sending the replication data to the second computing system.

12. The method as recited in claim 11, wherein determining that the replication has fallen behind further comprises: determining an oldest data object change in a database; and determining that the replication of the data object change has fallen behind the timing goal based at least on the oldest data object in the database.

13. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a first computing node associated with a first computing system, configure the first computing node to perform operations comprising: determining, by the first computing node, a first node load on the first computing node; determining, by the first computing node, at least one second node load on at least one second computing node, the at least one second computing node associated with a second computing system with which the first computing node is able to communicate over one or more networks; determining, by the first computing node, a highest node load from amongst the first node load and the at least one second node load; determining, by the first computing node, a number of threads to use for replicating data to the second computing system, wherein determining the number of threads is based at least on the highest node load; and sending, by the first computing node, replication data to the second computing system using the number of threads determined based at least on the highest node load.

14. The one or more non-transitory computer-readable media as recited in claim 13, the operations further comprising: based at least on determining that the replication of the data to the second computing system has fallen behind a timing goal for completing the replication, increasing the number of threads used by the first computing node for sending the replication data to the second computing system.

15. The one or more non-transitory computer-readable media as recited in claim 14, the operation of determining that the replication has fallen behind further comprising: determining an oldest data object change in a database; and determining that the replication of the data object change has fallen behind the timing goal based at least on the oldest data object in the database.

Citation Information

Patent Citations

  • Buffered virtual machine replication

    US20210117294A1

  • Network node simulation method based on linux container

    US20230216806A1

  • Load balancing and fault tolerant service in a distributed data system

    US20230325259A1