A Method for Locally Linearizable Reads on Replicas Based on the Raft Algorithm
By adopting the local linear and consistent reading method on replicas based on Raft algorithm in distributed systems, the best-performing replicas are dynamically selected, and combined with cuckoo filters and multi-threaded parallel search technology, the problem of low linear and consistent reading performance in the existing technology is solved, and the reading effect of low latency and high throughput is achieved.
Patent Information
- Application Number
- CN202211579736.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-12-09
AI Technical Summary
When the prior art realizes linear and consistent reading in distributed systems, performance is seriously affected, and the requirements for synchronous replication are high, resulting in insufficient scalability of the system.
The local linear consistent reading method on replicas is adopted based on Raft algorithm, and the best-performing replica is dynamically selected by periodically detecting and dynamically, cuckoo filters and multi-threaded parallel search technology are introduced, consensus protocols are optimized, leader load is reduced, local read operations are provided, and linear consistency requirements are met.
It reduces the system's write latency, improves the system's throughput, reduces the request load on leaders, and realizes local linear and consistent reading on replicas with low latency and high throughput, which has good fault tolerance and scalability.
Smart Images

Figure CN115794854B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed systems, and in particular to a local linear consistency reading method on a replica based on a Raft algorithm. Background Art
[0002] Herlihy et al. gave a formal definition and proof of linearizability. From the perspective of distributed systems, linearizability means that for read and write operations between replicas, even if a network partition or machine node failure occurs, the data of the entire cluster is like having only one copy, and each operation will occur atomically at a certain point in time between its call and completion. However, this strong consistency guarantee seriously affects the performance of the system. Saito et al. found through investigation that the weak consistency replication method allows for temporary inconsistent data in replicas for higher availability and performance. The latest write will be updated in the end, and it does not require high synchronous replication and has better scalability, but this is not suitable for systems with high consistency requirements.
[0003] There are three main methods for replica reading: Read Index, Lease Read, and Quorum Read. In the Read Index method, the replica obtains the index of the log entry commit position from the leader within the validity period through two rounds of messages. (1) Obtain the leader's commit index ( Figure 2 (a) R1). Because the leader needs to handle all write operations, the last commit index position indicates that all previous commands have been completed, which can be seen in the linear consistency read. (2) Confirm the current leader ( Figure 2 (R2 in (a)). Only the response from the current leader is reliable. Therefore, after receiving the request in the first phase, the replica that is considered to be the leader sends a message to other replicas to confirm its leadership status. When the replica receives a read request and obtains a valid commit index, it will respond to the client after waiting for all log items before the current latest commit index to be applied in the replica. Therefore, this method requires three RTTs (Round Trip Time), which greatly increases the overhead of reading data from the replica.
[0004] In the Lease Read method, the second round of communication overhead is removed from the Read Index method by introducing a lease mechanism. The lease mechanism ensures that no new leader will appear in the cluster for a period of time, so the leader can be directly considered valid. This method reduces the overhead of leadership confirmation, but still cannot avoid the communication overhead with the leader ( Figure 2(R1) in (b), which increases the burden on the leader.
[0005] The Quorum Read method achieves linearizable reads through majority guarantees. The latest write operation requires confirmation from a majority. If a read operation is also verified by a majority, there will be an intersection between the two majorities, and the latest written log entry can be read. By collecting the latest log entry indices from a majority other than the leader ( Figure 2 (Q1) in (c), and waiting for this log entry to be committed before returning the result. Since majority confirmation requires sending more messages, the current method brings more request counts and waiting delays. Summary of the Invention
[0006] The object of the present invention is to provide a method for local linearizable reads on replicas based on the Raft algorithm in view of the deficiencies of the prior art. By optimizing the consensus protocol in a distributed database system, the method dynamically selects the replica with the best performance through periodic detection to provide services externally, reduces the write latency of the system as much as possible, and introduces the Cuckoo Filter and multi-threaded parallel search technology to accelerate the log entry retrieval during the read process. Using different replica sets to solve the single-point bottleneck problem in the consensus protocol, reducing the request load on the leader, enabling local read operations on replicas, meeting the requirements of linearizability, and providing fault tolerance guarantees by dynamically adjusting the replica set strategy, avoiding affecting the overall performance of the system due to the failure of a certain server, and being able to reduce the latency of write operations, providing a technical solution for local linearizable reads on replicas in a distributed system with low latency and high throughput, having certain prospects and value for popularization and application.
[0007] The specific technical solution to achieve the object of the present invention is: a method for local linearizable reads on replicas based on the Raft algorithm. This method utilizes some service replicas existing in the distributed system. This type of replica has the Log Recency property, thereby providing local linearizable read services on replicas that bypass the leader, expanding the throughput of the system. By periodically detecting to dynamically select the replica with the best performance to provide services externally, reducing the write latency of the system as much as possible, and introducing the Cuckoo Filter and multi-threaded parallel search technology to accelerate the log entry retrieval during the read process. The present invention provides a technical solution for local linearizable reads on replicas in a distributed system with low latency and high throughput, which specifically includes the following steps:
[0008] Step S01: Selection of Nodes
[0009] After the cluster is started, all servers are initialized to the state of Tolerance Follower, and are divided into service replicas and fault-tolerant replicas according to the Follower role in the Raft algorithm, and a node is selected as the Leader for the cluster.
[0010] Step S02: Division of replicas
[0011] The Leader obtains information such as performance and status on each node through a round of Separate RPC messages, and scores and sorts the nodes through a weighted evaluation method. Subsequently, the top K nodes with the highest scores complete the conversion of the role status from Tolerance Follower to Service Follower through the Change Status RPC message, so as to divide the replicas.
[0012] Step S03: Fault tolerance processing
[0013] A certain replica in the cluster may fail or a network partition may occur. The fault tolerance solution of the present invention is as follows:
[0014] 1) When the Leader fails, a new Leader will be reselected through the leader election strategy.
[0015] 2) When a service replica fails, it is first removed from the service replica set and converted into a fault-tolerant replica. Then, the replica with the highest score is selected from other fault-tolerant replicas. After the log entries are completed, it is converted into a service replica and added to the service replica set.
[0016] 3) When a fault-tolerant replica fails, the present invention does not handle it.
[0017] Step S04: Selection of the best replica
[0018] After the Leader receives a command from the client, according to the size relationship between K and this command is written to the replica in the form of a log entry. In addition, L shadow replicas are randomly selected from the service replicas to be responsible for asynchronously sending the log entries to the fault-tolerant replicas in batches.
[0019] Step S05: Replica consistency reading
[0020] The read operations of M clients can be sent to the Leader and the service replica set. The latest written data is directly obtained from the data replica, meeting the requirements of linear consistency replica reading. A cuckoo filter is established for the log entries to accelerate the search process of the log entries. At the same time, the technology of multi-threaded parallel reading is used to shorten the latency of the read operation.
[0021] The present invention provides a method for partitioning different replica sets and fault tolerance. In a distributed system, it includes N servers, M clients, a replica partitioning configuration parameter K, and a shadow replica configuration parameter L. Among them, N, M, K, and L are all positive integers, N≥2 and L≤K<N.
[0022] In the present invention, the Follower role is subdivided into a Service Follower and a Tolerance Follower. The Service Follower has the attribute of log currency, synchronously receives logs from the leader, and can directly provide local linear consistency read results to the outside; the Tolerance Follower is responsible for providing the fault tolerance guarantee of the system. All Service Followers form an Adaptive Service Set, and all Tolerance Followers form an Adaptive Tolerance Set.
[0023] The Leader periodically sends Separate RPC messages to other replicas to collect the remaining storage space size (p space ) of the server, the length of the log entry (p entry ) and the disk write speed (p velocity ), and performs weighted evaluation scoring according to the following formula (a):
[0024]
[0025] Among them, S space , S entry , S velocity are the corresponding measurement units respectively, and W space , W entry , W velocity are the corresponding weights respectively.
[0026] The nodes are scored by quantifying and weighted summing the attributes. Subsequently, the Leader selects the K nodes with the highest scores, and simultaneously carries the missing log entries and the new role status through the Change Status RPC message to complete the conversion of the role status from Tolerance Follower to Service Follower, so as to partition the replicas.
[0027] A certain replica in the cluster may fail or a network partition may occur. The fault tolerance solution of the present invention is as follows:
[0028] Fault tolerance solution:
[0029] 1) When the Leader fails, a new Leader will be reselected through the leader election strategy.
[0030] 2) After a service replica fails, the Leader will remove it from the service replica set after waiting for 5 heartbeat timeouts. During the waiting process, the Leader will select a replica with the highest score from the dynamic fault tolerance set to complete the log entry filling work. After the log synchronization is completed, its role will be changed to a service replica to provide services externally. And the service replica to be removed will stop providing services externally after not receiving messages from the leader within 3 heartbeat timeouts to ensure the security of the system. Among them, the heartbeat timeout is the period of the message sent by the Leader to other nodes, which is set to about 10 - 20ms.
[0031] 3) When a fault tolerance replica fails, the present invention does not handle it.
[0032] The method for writing operations on different replica sets in the present invention is as follows:
[0033] After the Leader receives a command from the client, according to the size relationship between K and This command is written to the replica in the form of a log entry in different ways. Among them, according to the size relationship between K and The write operations are divided into two categories:
[0034] 1) When The Leader only synchronously writes to the service replica set.
[0035] 2) When The Leader, in addition to synchronously sending to the service replica set, also needs to select replicas with high scores from the fault tolerance replica set and synchronously send write messages.
[0036] The present invention can reduce the load pressure on the Leader by randomly selecting L shadow replicas from the service replicas to be responsible for asynchronously sending log entries to the fault tolerance replicas in batches. The shadow replicas are certain service replicas, responsible for stripping the function of synchronizing log entries to the fault tolerance replicas from the Leader and asynchronously sending log entries to the fault tolerance replicas in batches.
[0037] The present invention provides an accelerated reading method, and the process is as follows:
[0038] The read operations of M clients can be sent to the Leader and the service replica set, directly obtaining the latest written data from the data replicas to meet the replica reading requirements of linear consistency, and accelerating the log entry search process by establishing a cuckoo filter for the log entries. At the same time, the multi-threaded parallel reading technology is used to shorten the latency of the read operation.
[0039] The cuckoo filter consists of an array. The key is hashed through a hash function and written to the corresponding position. When looking up, it is determined whether this log entry exists based on the hashed position, thus avoiding linear scanning. Two cuckoo filters with three hash functions are established for uncommitted and committed log entries respectively to accelerate the process of searching for log entries. Since there is no dependency relationship in the scanning of the three types of log entries, multiple threads can be enabled to perform parallel scanning to accelerate the reading process.
[0040] The present invention defines a copy with log currency and designs a local linear consistency reading method based on the Raft algorithm. This method optimizes the consensus protocol in a distributed database system, uses different copy sets to solve the single-point bottleneck problem in the consensus protocol, and reduces the request load on the leader. In addition, the present invention can also provide local reading operations on the copies and meet the requirements of linear consistency at the same time. Finally, a fault tolerance guarantee is provided by dynamically adjusting the copy set strategy to avoid affecting the overall performance of the system due to the failure of a certain server and reduce the latency of write operations.
[0041] The present invention has the following beneficial technical effects and remarkable technical progress compared with the prior art:
[0042] 1) It gives play to the characteristic of log currency, enables the copy to provide local linear consistency reading results, reduces the communication overhead with the leader, and alleviates the single-point bottleneck problem of the leader. At the same time, the multi-copy reading scheme also greatly improves the overall throughput of the system.
[0043] 2) Through the dynamic set adjustment strategy, faults can be detected and resolved in a timely manner, thus ensuring that the system has a low write latency.
[0044] 3. By using the optimization technology combining the cuckoo filter and parallel reading, the sequential scanning is converted into multi-thread parallel scanning, reducing the latency of read operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is the architecture diagram of the present invention;
[0046] Figure 2 It is the linear consistency reading method on the copy of the prior art;
[0047] Figure 3 It is the flow chart of copy set division of the present invention;
[0048] Figure 4 It is the flow chart of service copy fault handling of the present invention;
[0049] Figure 5 Schematic diagram of the log writing process of the present invention;
[0050] Figure 6 Flow chart of the log reading program of the present invention. Specific implementation method
[0051] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and in conjunction with specific embodiments. It should be noted that, except for the specifically mentioned content below, the processes, conditions, etc. for implementing the present invention are all general knowledge and common general knowledge in the art, and the present invention has no particularly restricted content.
[0052] Embodiment 1
[0053] Refer to Figure 1 , the architecture of the linear consistency reading method on the replicas of the present invention is as Figure 1 shown. The replicas mainly include three modules: a consensus module, log entries, and a state machine. The consensus module is mainly responsible for sending and receiving messages. The present invention divides the replicas other than the leader in the distributed system into two sets, namely a dynamic service set and a dynamic fault-tolerant set. The replicas in the dynamic service replica set all have the attribute of log currency, can respond to the read requests of clients, directly provide locally linear-consistent read results, and accelerate the retrieval speed through the combination of the cuckoo filter and parallel reading. At the same time, if a dynamic service replica fails, it can be replaced by a fault-tolerant replica to automatically perform fault tolerance processing.
[0054] Refer to Figure 3 , the process of dividing the replica sets of the present invention is as Figure 3 shown. The method specifically includes the following processing steps:
[0055] Step 201: After the leader successfully completes the election, it sends a Separate RPC message to other replicas to collect the current performance parameters of each replica, including the storage space size, log entry length, and disk write speed.
[0056] Step 202: The leader scores the collected information according to the following formula (a) and obtains a final ranking.
[0057]
[0058] Among them, S space , S entry , S velocity are the corresponding measurement units, and W space , W entry , W velocity are the corresponding weights; p space, p entry and p velocity are the remaining storage space size, the log entry length, and the disk write speed, respectively.
[0059] Step 203: The leader will select the top K replicas, and send a Change Status RPC to convert the role to Service Follower.
[0060] Step 204: The Leader will synchronize the log by writing empty log entries, forcing all replicas in the dynamic service set to have the latest log entries.
[0061] To minimize the load pressure on the Leader as much as possible, the present invention will randomly select L shadow replicas, which are responsible for asynchronously sending log entries to the fault-tolerant replicas in batches.
[0062] Refer to Figure 4 , the service replica failure handling process is as Figure 4 shown, and the method specifically includes the following processing steps:
[0063] Step 301: The Leader will wait for the time of 5 heartbeat timeouts. If no message from the Service Follower is received within this time, it means that this replica may have failed or there is a network partition. Otherwise, end the failure handling process.
[0064] Step 302 and Step 303: These two steps can be carried out in parallel. When a Service Follower fails, the highest-rated Tolerance Follower replica needs to be used for replacement. Therefore, on the one hand, the state of the failed replica is switched to Tolerance Follower; on the other hand, the log entries of the highest-rated Tolerance Follower need to be supplemented for subsequent state switching.
[0065] Step 304: After the above process ends, switch the highest-rated Tolerance Follower replica to Service Follower, and end the failure handling process.
[0066] The service replica will stop providing services externally after not receiving a message from the leader within 3 heartbeat timeouts to ensure the security of the system. Among them, the heartbeat timeout time is the cycle of the Leader sending messages to other nodes, and the default setting is about 10 - 20 ms.
[0067] Refer to Figure 5 , the present invention is based on the number K of dynamic service replicas and the majority of the system According to the size relationship, two writing strategies are adopted.
[0068] Refer to Figure 5 a. When , the Leader only synchronously writes to the service replica set (as shown by W1 in Figure 5 ). In this case, it is necessary to wait for all service replicas to successfully write before completing the write operation to ensure the linear consistency of subsequent read operations.
[0069] Refer to Figure 5 b. When , in addition to synchronously sending to the service replica set (as shown by W1 in Figure 5 ), the Leader also needs to select a replica with a high score from the fault-tolerant replica set (as shown by W3 in Figure 5 b) and synchronously send the write message.
[0070] In addition, for the replicas in the dynamic fault-tolerant set, they can asynchronously obtain the written log entries from the service replicas (as shown by W2 in Figure 5 ), which can reduce the load pressure on the leader and improve the throughput of the system.
[0071] The specific process of the read operation of the present invention: Two cuckoo filters are set for the uncommitted log entries and the committed but not applied log entries on each replica respectively, and their function is to speed up the search. For each log entry, the present invention uses three hash functions to map the target key to an array and perform a +1 operation on each bit. When a replica needs to determine whether the key to be searched exists in the log entry, it first hashes the key and searches the constructed cuckoo filter. If each position matches, it proves that the target value may exist. At this time, it is necessary to scan the corresponding log entry from back to front to find this value. Of course, there is also a possibility of false positives, that is, although the cuckoo filter shows existence, the value actually does not exist. In this case, it is necessary to judge the result by scanning, but it will not affect the correctness of the system. If the target key is not found in the cuckoo filter, it proves that the target value must not exist in this section of the log and needs to continue the search.
[0072] Refer to Figure 6 , the read operation of the present invention can be specifically divided into three parts, and it can simultaneously scan the written log entries, the committed but not applied to the state machine log entries, and the results in the state machine. Once the position of the latest written target key is found, the result can be immediately returned after judgment.
[0073] The first part is to scan the uncommitted log entries (such as R1 in Figure 1 , Figure 6As shown in 501, check whether the read key is in the currently uncommitted but already written log entries. First, hash the key, and then scan the cuckoo filter 1 to determine whether there is a hit (as shown in Figure 6 504). If there is a hit, search for the data from the back to the front in the current log entry. If the target key is found, wait for the log entry to be committed and then return the read result. If there is no hit or the target key is not found due to the false positive factor of the cuckoo filter, proceed to the second part.
[0074] Second part: Scan the log entries that have been committed but not yet applied to the state machine (as shown in Figure 1 R2 in Figure 6 502). Hash the key and scan the cuckoo filter 2 in parallel to determine whether there is a hit (as shown in Figure 6 505). Since these log entries are already determined not to be changed, if the corresponding key has a hit, search for the data from the back to the front and return the result. If there is no hit or the target key is not found, proceed to the third part.
[0075] Third part: Directly scan the state machine and read the data from it (as shown in Figure 1 R3 in Figure 6 503), and then return the final result.
[0076] The protection scope of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, the changes and advantages that can be conceived by those skilled in the art are included in the present invention, and the appended claims are used as the protection scope.
Claims
1. A method for local linear consistency reading on replicas based on the Raft algorithm, characterized in that, a method of optimizing the consensus protocol in a distributed database system is adopted. By periodically detecting, the replica with the best performance is dynamically selected to achieve low-latency and high-throughput local linear consistency reading on replicas. The method specifically includes: Step S01: Selection of nodes After the cluster starts, all servers are initialized to the state of Tolerance Follower. According to the Follower role in the Raft algorithm, they are divided into service replicas and fault-tolerant replicas, and a node is selected as the Leader for the cluster; Step S02: Partitioning of replicas The Leader periodically sends Separate RPC messages to other replicas to obtain the performance and status information on each node, scores and sorts the nodes through a weighted evaluation method, and for the top K nodes with the highest scores, through the ChangeStatus RPC message, completes the conversion of the role status from Tolerance Follower to Service Follower, so as to partition the replicas; Step S03: Fault tolerance processing 1) When the Leader fails, a new Leader is reselected through the leader election strategy; 2) When a service replica fails, first remove it from the service replica set and convert it to a fault-tolerant replica, then select the replica with the highest score from other fault-tolerant replicas. After the log entries are supplemented completely, convert it to a service replica and add it to the service replica set; 3) When a fault-tolerant replica fails, no fault tolerance processing is performed; Step S04: Selection of the best replica After the Leader receives a command from the client, according to the size relationship with , this command is written to the replica in the form of a log entry, and another L shadow replicas are randomly selected from the service replicas to be responsible for asynchronously sending the log entries to the fault-tolerant replicas in batches; Step S05: Replica consistency reading The read operations of M clients are sent to the Leader and the service replica set, and the latest written data is directly obtained from the data replicas, meeting the requirements of linear consistency replica reading. By establishing a cuckoo filter for the log entries to accelerate the search process of the log entries, and through multi-threaded parallel reading to shorten the latency of the read operation.
2. The method for local linear consistency reading on replicas based on the Raft algorithm according to claim 1, characterized in that, the service replica in step S01 has the attribute of log currency, synchronously receives the logs from the leader, and can provide local linear consistency reading results externally; the fault-tolerant replicas are responsible for providing the fault tolerance guarantee of the system. All service replicas form a dynamic service set, and all fault-tolerant replicas form a dynamic fault tolerance set; the service replica has the latest written log entries. Under the global clock, the latest written information can be read by subsequent read operations, that is, the read data result is the latest globally.
3. The method for local linear consistency reading on replicas based on the Raft algorithm according to claim 1, characterized in that, The Leader in step S02 periodically sends Separate RPC messages to other replicas, collects the remaining storage space size, log entry length, and disk write speed of the server, and performs a weighted evaluation score using the following formula (a): Among them, S space , S entry and S velocity are the corresponding measurement units respectively; W space , W entry , W velocity are the corresponding weights respectively; p space , p entry and p velocity are the remaining storage space size, the log entry length, and the disk write speed respectively.
4. The method for local linear consistent reading on replicas based on the Raft algorithm according to claim 1, wherein, the Change Status RPC message in step S02 carries both the missing log entries and the new role status.
5. The method for local linear consistent reading on replicas based on the Raft algorithm according to claim 1, wherein, the specific fault tolerance handling after the failure of the service replica in step S03 is as follows: The Leader will remove it from the service replica set after waiting for 5 heartbeat timeouts. During the waiting process, the Leader will select a replica with the highest score from the dynamic fault tolerance set to complete the log entry supplementation work. After the log synchronization is completed, its role will be changed to a service replica to provide services externally; And the service replica to be removed will stop providing services externally after not receiving a message from the leader within 3 heartbeat timeouts to ensure the security of the system; the heartbeat timeout is the period of the message sent by the Leader to other nodes, which is set to 10 - 20ms.
6. The method for local linear consistent reading on replicas based on the Raft algorithm according to claim 1, wherein, the log entries in step S04 include three different states, namely: 1) the state of being written but not committed; 2) the state of being committed but not applied; 3) the state of being applied to the state machine; the log entries in the first two states are used to determine data by looking up the log entries, and the latter is through querying the state machine.
7. The method for local linear consistent reading on replicas based on the Raft algorithm according to claim 1, wherein, According to the size relationship between K and , write this command in the form of a log entry to the replica. The writing operation is divided into the following two categories: 1) When Leader only synchronously writes to the service replica set; 2) When occurs, in addition to synchronously sending to the service replica set, the Leader also needs to select a replica with a high score from the fault-tolerant replica set and synchronously send the write message.
8. The method for local linear consistent reading on replicas based on the Raft algorithm according to claim 1, wherein, the shadow replicas in step S04 are several service replicas, responsible for separating the function of synchronizing log entries to the fault tolerance replicas from the Leader and asynchronously sending log entries to the fault tolerance replicas in batches.
9. The method for local linear consistent reading on replicas based on the Raft algorithm according to claim 1, wherein, the cuckoo filter in step S05 consists of an array, hashes the key through a hash function and writes it to the corresponding position. When looking up, it determines whether this log entry exists by the hashed position. Two cuckoo filters with three hash functions are established for uncommitted and committed log entries respectively, and multiple threads are started to scan in parallel to accelerate the reading process.