Information processing device, adjustment method, and adjustment program
The information processing device addresses the challenge of determining the optimal bulk transfer amount in data replication by calculating throughput and adjusting transfer amounts accordingly, resulting in improved efficiency and throughput.
Patent Information
- Application Number
- JP2023182068
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-23
- Publication Date
- 2025-05-08
AI Technical Summary
Existing data replication technologies, such as Raft, struggle to specify the optimal bulk transfer amount of log entries based on hardware and external environment conditions, leading to inefficient network bandwidth usage and low throughput performance.
An information processing device that calculates the throughput of data transferred from a reader to a follower, specifies the amount of data to be transferred based on this throughput, and performs bulk transfers of the specified amount to improve efficiency.
The solution allows for the optimal bulk transfer amount to be determined based on the hardware and external environment, thereby enhancing network bandwidth utilization and improving throughput performance.
Smart Images

Figure 2025071684000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device, an adjustment method, and an adjustment program. [Background technology]
[0002] In data systems that make up services that utilize data, a method is used to replicate data to multiple nodes to improve availability. In this case, a method called state machine replication is used so that all replicas execute the same commands in the same order (see, for example, Non-Patent Document 1).
[0003] One of the replicated state machine algorithms is Raft (see, for example, Non-Patent Document 2), which is used in many data systems. Raft realizes a replicated state machine by having a node assigned the role of leader transfer and replicate each entry of the command log to a group of nodes assigned the role of follower.
[0004] Here, since a certain delay occurs in the transfer of entries from the leader to the followers, a simple implementation in which the leader replicates one log entry at a time results in low network bandwidth utilization efficiency and low throughput performance. For this reason, Raft aims to improve throughput performance by transferring multiple log entries together in bulk. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Fred B. Schneider, "The State Machine Approach: A Tutorial," Fault-Tolerant Distributed Computing, 1986, 18-41 [Non-Patent Document 2] Diego Ongaro, John K. Ousterhout, "In Search of an Understandable Consensus Algorithm," USENIX Annual Technical Conference 2014, 305-319 Summary of the Invention [Problem to be solved by the invention]
[0006] However, the above conventional technology cannot specify the optimal bulk transfer amount depending on the hardware and external environment. For example, the conventional technology cannot specify the optimal bulk transfer amount depending on the hardware and external environment constraints such as the upper limit of the network bandwidth to which the node is connected and the current network usage status, and cannot transfer data of the optimal transfer size from the leader to the follower. [Means for solving the problem]
[0007] In order to solve the above-mentioned problems and achieve the objective, the information processing device of the present invention is characterized by having a calculation unit that calculates the throughput of data transferred from a leader to a follower, a determination unit that determines the amount of data transfer based on the throughput calculated by the calculation unit, and a transfer unit that bulk transfers data of the amount of transfer determined by the determination unit to the follower. Effect of the Invention
[0008] According to the present invention, it is possible to determine an optimal bulk transfer amount depending on the hardware and the external environment. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 shows pseudocode for Raft's processing when no traffic adjustment is performed. [Diagram 2] FIG. 2 is a diagram showing a specific example of Raft processing when the transfer rate is not adjusted. [Diagram 3]FIG. 3 is a diagram showing a system including an information processing device according to the embodiment. [Figure 4] FIG. 4 is a block diagram illustrating an example of the configuration of the information processing device according to the embodiment. [Diagram 5] FIG. 5 is a diagram illustrating an example of data stored in the information processing device according to the embodiment. [Figure 6] FIG. 6 is a diagram illustrating pseudocode relating to a specific example of the process in the identification unit according to the embodiment. [Figure 7] FIG. 7 is a diagram showing a specific example of the process in the identification unit according to the embodiment. [Figure 8] FIG. 8 is a flowchart showing an example of the overall processing flow of the information processing device according to the embodiment. [Figure 9] FIG. 9 is a flowchart showing an example of the overall processing flow of the information processing device according to the embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of a computer that executes an adjustment program. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Hereinafter, an embodiment of an information processing device, an adjustment method, and an adjustment program according to the present application will be described in detail with reference to the drawings. Note that the information processing device, the adjustment method, and the adjustment program according to the present application are not limited to the embodiment.
[0011] 1. Introduction First, the process of Raft (see Non-Patent Document 2), which is an algorithm of a replicated state machine related to this embodiment, will be described. A system to which Raft is applied is composed of multiple information processing devices (hereinafter also referred to as nodes), and when no network or node failure occurs, one node becomes a leader and the remaining nodes become a group of followers.
[0012] The specific processing flow of Raft is explained below. First, a command is sent to the leader from any number of clients such as computers. In response, the leader receives the client commands and appends them to the end of its own log as log entries. After that, the leader calls an RPC (Remote Procedure Call) called AppendEntriesRPC for each follower and transfers its own log entry to each follower.
[0013] The process of adding a log entry may be performed in a different thread from the process of sending the log entry. Also, AppendEntriesRPC can transfer multiple log entries by passing multiple log entries as arguments and transferring them in bulk.
[0014] Next, we will explain an example of the procedure for replicating multiple log entries |E| by bulk transfer in Raft. In this paper, when the leader has an entry with log index i, where i is equal to or less than the value commitIndex managed by the leader, the replication process for that entry is said to be complete. For the definitions of log index and commitIndex, see Non-Patent Document 2.
[0015] In a cluster with N nodes, the leader determines whether there are |E| log entries in its log whose replication process has not been completed. If there are not |E| such log entries, the leader waits until enough commands arrive from the user and such a log entry occurs.
[0016] On the other hand, if there are |E| log entries for which the replication process has not been completed, the leader selects the |E| consecutive log entries with the smallest log number (index) from among them, and bulk transfers them by calling AppendEntriesRPC to each follower. Here, when the commitIndex before the leader calls AppendEntriesRPC for each follower is set to ci_old, log entries with log numbers between ci_old+1 and ci_old+|E| are selected and bulk transferred. As described in Non-Patent Document 2, if a node failure, response delay, or packet loss occurs in a follower while AppendEntriesRPC is being executed, the RPC is re-executed.
[0017] In a replication state machine to which Raft is applied, data can be replicated to followers by selecting |E| consecutive log entries with the smallest log numbers for which replication processing has not been completed, and repeating the process until the replication processing of the log entries is completed. Hereinafter, the process from the selection of the log entries for which replication processing has not been completed (including the waiting process until a log entry is generated when there are no |E| selectable log entries) to the process in which the relevant entries are replicated to the majority of the N servers, and |E| is added to the leader's commitIndex according to the Raft rules (i.e., the leader's commitIndex becomes equal to ci_old+|E|) is referred to as ReplicateEntries(|E|). Note that the AppendEntriesRPC for the remaining nodes may be processed asynchronously with ReplicateEntries(|E|).
[0018] Next, an example of processing when Raft always transfers a fixed number of log entries in bulk will be described. As described above, Raft can transfer multiple log entries in bulk at once by passing multiple log entries as arguments to AppendEntriesRPC. Here, the processing content when always transferring a fixed number of log entries in bulk will be described with reference to Figs. 1 and 2. Fig. 1 is a diagram showing pseudocode related to Raft processing when the transfer amount is not adjusted. Fig. 2 is a diagram showing a specific example of Raft processing when the transfer amount is not adjusted.
[0019] First, referring to Figure 1, we will explain the pseudocode executed by the reader in the process of always bulk transferring a fixed number of log entries. The first line of the pseudocode shown in Figure 1 shows the code that sets the initial value of the number of log entries to be selected (transferred). Here, "e_init" is an integer equal to or greater than 1, and is a constant. The second and third lines of the pseudocode show the code that performs the previously mentioned ReplicateEntries(|E|) process in an infinite loop.
[0020] Next, referring to Figure 2, we will explain the Raft processing when "e_init=2" is set in the pseudocode shown in Figure 1 in a cluster with "N=3" nodes. Figure 2 shows an example in which two log entries are always transferred from the leader "L" to each of the two followers "F1" and "F2" and replication is performed.
[0021] "L" bulk transfers two log entries for each of "F1" and "F2" using AppendEntriesRPC, and receives information from "F1" and "F2" that the log entries have been received. When AppendEntriesRPC for either "F1" or "F2" is successful, replication to the majority is completed and commitIndex is incremented. In Figure 2, AppendEntriesRPC for "F1" and "F2" always completes almost simultaneously. This series of processes is ReplicateEntries(2), and this process is repeated. This allows Raft to always transfer two log entries from leader to follower, replicating data.
[0022] However, in the past, although it was possible to always bulk transfer a fixed number of log entries, it was not possible to bulk transfer the optimal transfer amount for each transfer depending on the hardware and external environment, such as the network bandwidth status to which the node is connected and the current network usage status, resulting in low throughput. Therefore, the information processing device according to this embodiment aims to specify the optimal bulk transfer amount depending on the hardware and external environment.
[0023] 2. Overall structure Next, an overall configuration of a system including an information processing device 100 according to this embodiment will be described. Fig. 3 is a diagram showing a system including an information processing device according to an embodiment. The system shown in Fig. 3 is composed of one information processing device 100, which is a leader, and four followers 200-1 to 200-4.
[0024] Here, the information processing device of the leader and the information processing device of the follower are realized by a PC or the like, and are connected to each other so as to be able to communicate with each other by wire or wirelessly. In addition, both the information processing device 100 of the leader and the information processing device of the follower 200 have the same components, and the leader and the follower can be changed by setting.
[0025] The information processing device 100 calculates the throughput of data transferred from the leader to the follower, specifies the amount of data transfer based on the calculated throughput, and then bulk-transfers the specified amount of data to the follower.
[0026] For example, the information processing device 100 calculates the throughput by dividing the number |E| of log entries transferred by the process of ReplicateEntries(|E|) to the follower by the time required to execute the process. Then, the information processing device 100, for example, compares the calculated throughput with the previous throughput to determine whether the throughput has increased and specifies the number of log entries to be transferred next time. After that, the information processing device 100, for example, bulk transfers the specified number of log entries to each follower.
[0027] The information processing device 100 repeatedly performs a series of processes from the calculation of throughput to the specification of the number of log entries to be transferred, and the transfer of the specified number of log entries. This allows the information processing device 100 to specify the optimal transfer amount taking into consideration the throughput each time it processes ReplicateEntries(|E|) (transfer of log entries), and therefore to specify the optimal bulk transfer amount depending on the hardware and external environment.
[0028] 3. Configuration of Information Processing Device 100 Next, the configuration of the information processing device 100 shown in Fig. 3 will be described with reference to Fig. 4. Fig. 4 is a block diagram showing an example of the configuration of the information processing device according to the embodiment. The information processing device 100 has a communication unit 110, a control unit 120, and a storage unit 130, and is connected to the followers 200 via a network N so that they can communicate with each other.
[0029] The communication unit 110 is realized by, for example, a network interface card (NIC) or the like. The communication unit 110 is connected to a network N and transmits and receives information to and from the followers 200. The communication unit 110, for example, receives commands from an external client and transmits multiple consecutive log entries to each of the multiple followers 200.
[0030] The storage unit 130 is realized by, for example, a storage device such as a RAM (Random Access Memory) or a hard disk. The storage unit 130 stores data and programs necessary for various processes by the control unit 120. The storage unit 130 stores, for example, a throughput value calculated by the calculation unit 121 described below, the amount of data transferred by the transfer unit 123, the values of variables used in the processing of the identification unit 122, and the like.
[0031] Here, the data stored in the storage unit 130 will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of data stored in the information processing device according to the embodiment. The storage unit 130 shown in Fig. 5 is configured with the items "number of transfers", "transfer amount", "throughput", and "search stage".
[0032] "Number of transfers" stores the number of times bulk transfer was performed by the information processing device 100 for replicating a series of log entries. "Transfer volume" stores the volume of data, such as the number of log entries transferred to a follower by bulk transfer. "Throughput" stores the throughput value of the ReplicateEntries (|E|) process. "Search stage" stores the search stage of the transfer volume in the process of identifying the bulk transfer volume, which will be described later.
[0033] Returning to the explanation of Fig. 4, the control unit 120 is realized by a CPU (Central Processing Unit), MPU (Micro Processing Unit), etc., executing various programs stored in a storage device inside the device using a RAM as a working area. The control unit 120 is also realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 120 has a calculation unit 121, an identification unit 122, and a transfer unit 123.
[0034] The calculation unit 121 calculates the throughput of data transferred from the leader to the follower. For example, the calculation unit 121 calculates the throughput value of the immediately preceding process of ReplicateEntries(|E|) and stores it in the storage unit 130.
[0035] Here, the calculation unit 121 calculates the throughput by, for example, dividing the number of transferred log entries |E| by the time required to execute the process of ReplicateEntries(|E|). In addition, the calculation unit 121 can calculate the throughput by dividing the total number of bytes of log entries sent to followers in the process of ReplicateEntries(|E|) (transmissions to different followers are not counted separately, and retransmissions of AppendEntriesRPC are not counted) by the time required to process ReplicateEntries(|E|), or by dividing the number of bytes sent to followers in the process of ReplicateEntries(|E|) by the time required to process ReplicateEntries(|E|).
[0036] In addition, in the above-mentioned method of calculating throughput, the calculation unit 121 can calculate the throughput by using the time required to process ReplicateEntries (|E|) when there are no log entries |E| for which replication processing has not been completed in its own log, and does not include the time required for waiting processing.
[0037] The identifying unit 122 identifies the data transfer amount based on the throughput calculated by the calculating unit 121. Then, the identifying unit 122 notifies the transferring unit 123 (described later) of the identified transfer amount. The identifying unit 122, for example, refers to the throughput value stored in the storage unit 130 and the current search stage, determines whether the throughput has increased or decreased, and identifies the transfer amount for the next bulk transfer.
[0038] Here, a description will be given of an example of a specific process performed by the identifying unit 122. The identifying unit 122 identifies, for example, a value obtained by processing the immediately preceding bulk transfer amount according to three search stages as the next bulk transfer amount.
[0039] In the first stage (exponential), if the current throughput calculated by the calculation unit 121 is higher than the previous throughput, the determination unit 122 determines the next transfer volume as a value obtained by multiplying the current transfer volume determined by the determination unit 122 by a predetermined multiplier, and if the current throughput is lower than the previous throughput, the determination unit 122 determines the next transfer volume as a value obtained by dividing the current transfer volume by a predetermined multiplier.
[0040] For example, the determination unit 122 compares the current throughput and the previous throughput stored in the storage unit 130, and if the current throughput is higher, it determines that the transfer performance has improved and determines the current transfer amount multiplied by k as the next transfer amount. On the other hand, if the current throughput is lower, it determines the next transfer amount by multiplying the transfer amount by 1 / k, where k is an arbitrary constant set in advance.
[0041] In the second stage (linear stage), if the current throughput calculated by the calculation unit 121 is higher than the previous throughput, the determination unit 122 determines a value obtained by increasing the current transfer amount determined by the determination unit 122 by a predetermined amount as the next transfer amount, and if the current throughput is lower than the previous throughput, the determination unit 122 determines a value obtained by decreasing the current transfer amount by a predetermined amount as the next transfer amount.
[0042] For example, the determination unit 122 compares the current throughput and the previous throughput stored in the storage unit 130, and if the current throughput is higher, it determines that the transfer performance has improved and determines the next transfer amount to be the current transfer amount plus C. On the other hand, if the current throughput is lower, it determines the next transfer amount by subtracting C from the transfer amount. Note that C is a preset arbitrary constant.
[0043] In the third stage (saturated stage), if the current throughput calculated by the calculation unit 121 is higher than a threshold value set by multiplying the past throughput by a predetermined value, the determination unit 122 determines the next transfer volume to be the same value as the current transfer volume determined by the determination unit 122.
[0044] For example, the determination unit 122 compares the current throughput stored in the storage unit 130 with a threshold calculated by multiplying the past throughput and "reset_thres", and if the current throughput is higher, determines that the transfer performance is maintained at or above the threshold, and determines the same value as the current transfer amount as the next transfer amount. Note that "reset_thres" is a preset arbitrary constant, and is set to a value greater than 0 and equal to or less than 1.
[0045] Here, the determination unit 122 may be set to perform the processes of the first to third stages in order, or may perform the process of any one of the stages. In addition, the k and C in the above-mentioned process may be values set to find the optimal transfer amount with a small number of searches, taking into account the performance of the information processing device 100, etc.
[0046] In addition, "reset_thres" in the above process is a value used to set the allowable range of throughput, and if the calculated throughput is within the allowable range (above the threshold), the transfer rate is fixed. For example, if "reset_thres:0.5" is set, and the throughput is 50% or more of the throughput "max_tp" recorded just before the transition to the third stage, the throughput is determined to be within the allowable range, and the transfer rate is fixed.
[0047] The transfer unit 123 bulk-transfers data of the transfer amount identified by the identification unit 122 to the follower. For example, the transfer unit 123 bulk-transfers the number of log entries identified by the identification unit 122 described above to each follower 200 via the communication unit 110 by executing the process of ReplicateEntries(|E|). Then, the transfer unit 123 stores the number of transferred log entries in the storage unit 130.
[0048] 4. Specific Examples of Processing by Information Processing Device 100 Here, a process in which the information processing device 100 calculates the throughput while changing the bulk transfer amount and searches for an optimal bulk transfer amount will be described with reference to Fig. 6 and Fig. 7. Fig. 6 is a diagram showing pseudo code relating to a specific example of the process in the information processing device according to the embodiment. Fig. 7 is a diagram showing a specific example of the process in the information processing device according to the embodiment.
[0049] First, the overall processing flow of the information processing device 100 and the specific processing content of each stage will be described with reference to the pseudo code shown in FIG. 6. The first to fourth lines of FIG. 6 describe variables and the setting of initial values. Specifically, |E|, which indicates the number of log entries to be bulk-transferred, is set by the initial value "e_init" as in FIG. 1 described above. Next, "prev_tp" indicates a variable for saving the throughput of the previous ReplicateEntries, and "cur_tp" indicates a variable for saving the throughput of the current ReplicateEntries. And "phase" indicates a variable for saving information on the current search stage. In this specific example, since the processing starts from the initial measurement, "initial" is saved.
[0050] Then, from the fifth line to the last line, the process performed according to the current search stage is described. First, an initial measurement "initial" is executed in preparation for searching the transfer volume. In the initial measurement, the transfer volume is set to the initial value "e_init" and a replication process is executed only once to measure the throughput. First, the transfer unit 123 executes the ReplicateEntries(|E|) process with the transfer volume set to the initial value "e_init". Then, the calculation unit 121 calculates the throughput. After that, the information processing device 100 transitions to the "exponential" stage, which is the first stage of the search.
[0051] In the "exponential" stage, first, the identifying unit 122 multiplies the transfer amount |E| by k. Next, the transferring unit 123 executes the process of ReplicateEntries(|E|) with the k-fold transfer amount. Then, the calculating unit 121 calculates the throughput with the k-fold transfer amount. After that, if the current throughput is lower than the previous throughput, the identifying unit 122 multiplies the transfer amount by 1 / k, and the information processing device 100 transitions to the "linear" stage.
[0052] Next, in the "linear" stage, first, the identification unit 122 adds C to the transfer amount. Next, the transfer unit 123 executes the process of ReplicateEntries(|E|) with the transfer amount including C. Then, the calculation unit 121 calculates the throughput with the transfer amount including C. Thereafter, if the current throughput is lower than the previous throughput, the identification unit 122 subtracts C from the transfer amount. At the same time, the information processing device 100 records the previous throughput as "max_tp" and transitions to the "saturated" stage.
[0053] Next, in the "saturated" stage, first, the transfer unit 123 executes the process of ReplicateEntries(|E|). Next, the calculation unit 121 calculates the throughput. Then, the information processing device 100 calculates a threshold value by multiplying "max_tp" recorded in the above-mentioned "linear" stage by a preset "reset_thres". After that, if the current throughput is lower than the threshold value, the identification unit 122 sets the transfer amount to the initial value "e_init" and returns to the process of initial measurement. Then, the information processing device 100 transitions to the "exponential" stage.
[0054] This allows the information processing device 100 to search for the optimal bulk transfer volume by sequentially transitioning between the aforementioned "exponential" (first stage), "linear" (second stage), and "saturated" (third stage) processing.
[0055] Next, the flow of the transfer amount and the transition of stages by the above-mentioned series of processes will be explained with reference to Fig. 7. In the example shown in Fig. 7, the process when "N (number of nodes) = 3", "e_init = 2", "k = 2", "C = 1", and "reset_thres = 0.8" are set is shown, and an example is shown in which log entries are transferred from the leader "L" to each of the two followers "F1" and "F2" and replicated, as in Fig. 1. Note that in Fig. 7, it is assumed that the AppendEntriesRPC to "F1" and "F2" is always completed almost simultaneously.
[0056] First, the information processing device 100 performs ReplicateEntries(2) for the first time as an initial measurement with an initial value "e_init=2", and then moves to the search stage "exponential". Next, since the current stage is "exponential", the information processing device 100 multiplies "transfer amount 2" by "2(k)" and performs ReplicateEntries(4) for the second time with a "transfer amount 4". Here, since the throughput for the second time has increased compared to the first time, the information processing device 100 remains in "exponential", multiplies "transfer amount 4" by "2(k)", and performs ReplicateEntries(8) for the third time with a "transfer amount 8".
[0057] Here, since the throughput for the third time has decreased compared to the second time, the information processing device 100 multiplies the "transfer amount 8" by "1 / 2 (1 / k)" and transitions to "linear" at a "transfer amount 4". Next, since the information processing device 100 is currently in "linear" mode, it adds "1(C)" to the "transfer amount 4" and performs the fourth ReplicateEntries(5) at a "transfer amount 5".
[0058] Here, since the throughput for the fourth time has increased compared to the third time, the information processing device 100 remains in “linear”, adds “1(C)” to “transfer volume 5”, and performs the fifth ReplicateEntries(6) with “transfer volume 6”.
[0059] Here, since the throughput for the fifth time has decreased compared to the fourth time, the information processing device 100 subtracts "1(C)" from "Transfer rate 6" to obtain "Transfer rate 5", records the throughput for the fourth time as "max_tp", and transitions to "saturated".
[0060] Next, since the current state is "saturated", the information processing device 100 calculates the threshold by multiplying "max_tp (fourth throughput)" and "0.8 (reset_thres)". After that, the information processing device 100 fixes the transfer rate at "transfer rate 5" and performs the sixth ReplicateEntries(5). Here, the sixth throughput is equal to or greater than the threshold, and the information processing device 100 remains in the "saturated" state. Therefore, the information processing device 100 performs the seventh ReplicateEntries(6) at the fixed "transfer rate 5".
[0061] The information processing device 100 can determine an increase or decrease in throughput while changing the transfer amount by performing a series of processes as shown in Fig. 7, and can identify an optimal bulk transfer amount. Furthermore, when the throughput falls below a threshold in the "saturated" state due to a change in the network environment or the like, the information processing device 100 returns to the initial measurement and performs a series of search processes, and can therefore identify an optimal bulk transfer amount corresponding to the change in the network environment.
[0062] 5. Example of Processing of Information Processing Device 100 Next, a flow of processing by the information processing device 100 will be described with reference to Fig. 8 and Fig. 9. Fig. 8 and Fig. 9 are flowcharts showing an example of an overall flow of processing by the information processing device according to the embodiment.
[0063] First, the information processing device 100 judges whether or not there are |E| consecutive log entries in the log for which the replication process has not been completed (S101). If there are |E| consecutive log entries in the log for which the replication process has not been completed (S101; Yes), the transfer unit 123 bulk transfers the |E| consecutive log entries to each follower (S102). On the other hand, if there are not |E| consecutive log entries in the log for which the replication process has not been completed (S101; No), the information processing device 100 waits until there are |E| consecutive log entries in the log for which the replication process has not been completed.
[0064] Subsequently, after the process of S102, the calculation unit 121 calculates the throughput of the transferred data (S103). Next, the identification unit 122 sets the number of log entries to be transferred to k times (S104). After that, the information processing device 100 judges whether or not the log contains a set number of log entries whose replication process has not been completed (S105). If the log contains a set number of log entries whose replication process has not been completed (S105; Yes), the transfer unit 123 transfers the set number of log entries in bulk (S106). On the other hand, if the log does not contain the set number of log entries whose replication process has not been completed (S105; No), the information processing device 100 waits until the log contains a set number of log entries whose replication process has not been completed.
[0065] After the process of S106, the calculation unit 121 calculates the throughput, and the information processing device 100 judges whether the throughput has increased (S107). If the throughput has increased (S107; Yes), the information processing device 100 returns to S104 and continues the process. On the other hand, if the throughput has not increased (S107; No), the identification unit 122 sets the number of log entries to be transferred to 1 / k times (S108).
[0066] Thereafter, the identifying unit 122 sets the number of log entries to be transferred to a value increased by C (S109). Then, the information processing device 100 judges whether or not the log contains the set number of log entries for which the replication process has not been completed (S110). If the log contains the set number of log entries for which the replication process has not been completed (S110; Yes), the transferring unit 123 bulk transfers the set number of log entries (S111). On the other hand, if the log does not contain the set number of log entries for which the replication process has not been completed (S110; No), the information processing device 100 waits until the log contains the set number of log entries for which the replication process has not been completed.
[0067] Then, after the process of S111, the calculation unit 121 calculates the throughput, and the information processing device 100 judges whether the throughput has increased (S112). If the throughput has increased (S112; Yes), the information processing device 100 returns to S109 and continues the process. On the other hand, if the throughput has not increased (S112; No), the identification unit 122 sets a value that reduces the number of log entries to be transferred by C (S113). Then, the information processing device 100 records the previous throughput as "max_tp" (S114).
[0068] Thereafter, the information processing device 100 judges whether or not the log contains a set number of log entries for which the replication process has not been completed (S115). If the log contains a set number of log entries for which the replication process has not been completed (S115; Yes), the transfer unit 123 bulk transfers the set number of log entries (S116). On the other hand, if the log does not contain the set number of log entries for which the replication process has not been completed (S115; No), the information processing device 100 waits until the log contains the set number of log entries for which the replication process has not been completed.
[0069] Then, after the process of S116, the calculation unit 121 calculates the throughput, and the information processing device 100 judges whether the throughput is higher than the threshold calculated based on "max_tp" (S117). If the throughput is higher than the threshold (S117; Yes), the information processing device 100 returns to S115 and continues the process. On the other hand, if the throughput is lower than the threshold (S117; No), the information processing device 100 returns to S101 and continues the process.
[0070] 6. Effects of the embodiment As described above, the information processing device 100 according to this embodiment includes the calculation unit 121, the determination unit 122, and the transfer unit 123. The calculation unit 121 calculates the throughput of data transferred from the leader to the follower. The determination unit 122 determines the amount of data transfer based on the throughput calculated by the calculation unit 121. The transfer unit 123 bulk transfers data of the amount of data transfer determined by the determination unit 122 to the follower.
[0071] As a result, the information processing device 100 can calculate the throughput each time it performs a bulk transfer and adjust the transfer volume of the next bulk transfer according to the throughput, thereby identifying the optimal bulk transfer volume depending on the hardware and external environment.
[0072] In addition, when the current throughput calculated by the calculation unit 121 is higher than the previous throughput, the determination unit 122 of the information processing device 100 determines the next transfer volume as a value obtained by multiplying the current transfer volume determined by the determination unit 122 by a predetermined multiplier, and when the current throughput is lower than the previous throughput, the determination unit 122 determines the next transfer volume as a value obtained by dividing the current transfer volume by a predetermined multiplier.
[0073] This allows the information processing device 100 to adjust the bulk transfer amount to an optimum amount corresponding to changes in the external environment, such as the network environment, by increasing or decreasing the bulk transfer amount in response to increases or decreases in throughput.
[0074] Furthermore, when the current throughput calculated by the calculation unit 121 is higher than the previous throughput, the determination unit 122 of the information processing device 100 determines a value obtained by increasing the current transfer amount determined by the determination unit 122 by a predetermined amount as the next transfer amount, and when the current throughput is lower than the previous throughput, the determination unit 122 determines a value obtained by decreasing the current transfer amount by a predetermined amount as the next transfer amount.
[0075] This allows the information processing device 100 to increase or decrease the bulk transfer amount in relatively small increments in response to increases or decreases in throughput, thereby enabling optimal adjustment of the bulk transfer amount in response to changes in the external environment, such as the network environment.
[0076] In addition, if the current throughput calculated by the calculation unit 121 is higher than a threshold value set by multiplying the past throughput by a predetermined value, the determination unit 122 of the information processing device 100 determines the next transfer volume to be the same value as the current transfer volume determined by the determination unit 122.
[0077] As a result, the information processing device 100 can repeat bulk transfer at a fixed transfer volume when the throughput is higher than a threshold calculated by multiplying the past throughput set as the upper limit by a predetermined value, thereby preventing repeated search processing due to a decrease in throughput within the allowable range for bulk transfer at a fixed transfer volume.
[0078] [7. System configuration, etc.] Of the processes described in the above embodiments, some of the processes described as being performed automatically can be performed manually. Alternatively, all or some of the processes described as being performed manually can be performed automatically by a known method. In addition, the information including the processing procedures, specific names, various data and parameters shown in the above documents and drawings can be changed arbitrarily unless otherwise specified. For example, the various information shown in each drawing is not limited to the illustrated information.
[0079] In addition, each component of each device shown in the figure is a functional concept, and does not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or a part of it can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions. Furthermore, each processing function performed by each device can be realized in whole or in any part by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic.
[0080] 4 may be held in a storage server or the like, rather than being held by the information processing device 100. In this case, the information processing device 100 acquires various pieces of information by accessing the storage server.
[0081] [8. Hardware Configuration] 10 is a diagram showing an example of a hardware configuration The information processing device 100 according to the embodiment described above is realized by a computer 1000 having a configuration as shown in FIG.
[0082] 10 is a diagram showing an example of a computer that executes an adjustment program. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0083] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1041. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1041. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.
[0084] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the information processing device 100 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing the same process as the functional configuration of the information processing device 100 is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced with an SSD (Solid State Drive).
[0085] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads out program module 1093 or program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary, and executes the program.
[0086] Note that the program module 1093 and the program data 1094 are not limited to being stored in the hard disk drive 1090, but may be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1041 or the like. Alternatively, the program module 1093 and the program data 1094 may be stored in another computer connected via a network (LAN, WAN, etc.). Then, the program module 1093 and the program data 1094 may be read by the CPU 1020 from the other computer via the network interface 1070. [Explanation of symbols]
[0087] 100 Information processing device 110 Communications Department 120 Control section 121 Calculation section 122 Specific part 123 Transfer Department 130 Storage section 200, 200-1~200-4 Followers
Claims
1. A calculation unit that calculates a throughput of data transferred from a leader to a follower; a determination unit that determines the data transfer amount based on the throughput calculated by the calculation unit; a transfer unit that bulk-transfers the data of the transfer amount specified by the specification unit to the follower; 13. An information processing device comprising:
2. When the current throughput calculated by the calculation unit is higher than the previous throughput, the determination unit determines a value obtained by multiplying the current transfer amount determined by the determination unit by a predetermined factor as the next transfer amount, and when the current throughput is lower than the previous throughput, the determination unit determines a value obtained by dividing the current transfer amount by the predetermined factor as the next transfer amount.
2. The information processing apparatus according to claim 1,
3. When the current throughput calculated by the calculation unit is higher than the previous throughput, the determination unit determines a value obtained by increasing the current transfer amount determined by the determination unit by a predetermined amount as the next transfer amount, and when the current throughput is lower than the previous throughput, the determination unit determines a value obtained by decreasing the current transfer amount by a predetermined amount as the next transfer amount.
2. The information processing apparatus according to claim 1,
4. When the current throughput calculated by the calculation unit is higher than a threshold value set by multiplying the past throughput by a predetermined value, the determination unit determines the same value as the current transfer amount determined by the determination unit as the next transfer amount.
2. The information processing apparatus according to claim 1,
5. An adjustment method executed by an information processing device, comprising: A calculation unit that calculates a throughput of data transferred from a leader to a follower; a determination unit that determines the data transfer amount based on the throughput calculated by the calculation unit; a transfer unit that bulk-transfers the data of the transfer amount specified by the specification unit to the follower; The adjustment method comprising the steps of:
6. A calculation unit that calculates a throughput of data transferred from a leader to a follower; a determination unit that determines the data transfer amount based on the throughput calculated by the calculation unit; a transfer unit that bulk-transfers the data of the transfer amount specified by the specification unit to the follower; A coordination program for causing a computer to execute the above.