Method and system for improving writing efficiency of distributed database based on Raft protocol

CN120010787APending Publication Date: 2025-05-16上海沄熹科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510139258.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-16

Smart Images

  • Figure CN120010787A_ABST
    Figure CN120010787A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for improving the writing efficiency of a distributed database based on a Raft protocol, belongs to the technical field of databases, unifies a raftlog and a data load during data writing, and is implemented as follows: when the raftlog is written into a disk, the data load in the raftlog is extracted and independently written into the disk; the RAFT and the storage engine can access the same data load by combining a write-in disk of the RAFT log with an actual data write-in disk; and combining a write-in disk of the raftlog with a write-in disk of the WAL, so that the WAL and the raftlog share one data load. Aiming at the problem that excessive resources are occupied when the raftlog is written into the disk, the disk writing mechanism of the raftlog is optimized, the load of the raftlog is reduced, the writing frequency of the data load is reduced, the occupation of the disk space is reduced, and the writing efficiency of the distributed database is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of database technology, and in particular to a method and system for improving the writing efficiency of a distributed database based on the Raft protocol. Background Art

[0002] The current Internet is huge in scale, and a large amount of new data is generated every moment. How to efficiently and securely store massive data has become an important challenge. Distributed databases, with their high scalability, perfectly meet the needs of massive data storage. Through multiple servers and multiple nodes, they can effectively store the ever-expanding data.

[0003] However, write rate and data security have always been issues that need to be considered in distributed databases. The collaboration of multiple servers and multiple nodes makes the write logic of distributed databases more complex than that of single-machine databases, and is more affected by physical factors such as network conditions and distance between nodes. Data security also requires distributed databases to introduce distributed transactions, consistency protocols and other mechanisms, which not only require more complex logic, but also occupy more resources, slowing down the speed of actual data writing.

[0004] To this end, distributed databases need to consider more aspects, reduce redundant consumption during data writing, and improve the efficiency of actual writing, while not sacrificing data security and service availability.

[0005] At present, the commonly used consistency algorithms for distributed databases are the raft algorithm and the paxos algorithm, which is similar to raft. In the Raft algorithm, data consistency is guaranteed through the consensus of raftlog. The raftlog contains the actual data load, and the raftlog will be written to the disk before the actual data to ensure that the data is not lost. This effectively ensures data security, but it will add additional data, resulting in more data actually written to the disk; at the same time, the writing of raftlog to the disk precedes the completion of raft consensus and the writing of actual data to the disk, which will also reduce the data writing speed. Because the IO operation itself is slow, the reduction in writing speed here will also be more obvious. Summary of the invention

[0006] The technical task of the present invention is to address the above shortcomings and provide a method and system for improving the writing efficiency of a distributed database based on the Raft protocol. In order to address the problem that raftlog writing to disk takes up too many resources, the disk writing mechanism of raftlog is optimized to improve the writing efficiency of the distributed database.

[0007] The technical solution adopted by the present invention to solve its technical problem is:

[0008] A method to improve the writing efficiency of a distributed database based on the Raft protocol, unifying the data load when writing raftlog and data, and implementing it in the following way:

[0009] When writing raftlog to disk, the data payload is extracted and written to disk separately. By combining the writing of raftlog to disk with the writing of actual data to disk, raft and storage engine can access the same data payload. Combining the writing of raftlog to disk with the writing of WAL to disk allows WAL and raftlog to share the same data payload.

[0010] Furthermore, after the raft consensus succeeds, the data payload in the WAL is read and written to the disk; after the raft consensus fails, the corresponding raftlog and WAL are deleted.

[0011] Furthermore, when the leader needs to obtain the raftlog from the storage and send it to the follower, it obtains the data payload from the corresponding WAL, encapsulates it into a raft message and sends it to the follower.

[0012] Furthermore, the steps of combining the writing of raftlog to disk with the writing of WAL to disk include:

[0013] 1.1) Separate the data payload in raftlog;

[0014] 1.2) Write the data payload to disk by writing WAL;

[0015] 1.3) Record the corresponding KEY or label in raftlog to associate WAL;

[0016] 1.4) Write the raftlog stripped of the data payload to disk.

[0017] When a distributed database writes data, it converts the write request into a raftlog through the raft protocol and passes it to each node to initiate consensus. After each node receives it, it needs to write the raftlog to disk to prevent data loss.

[0018] Before writing raftlog to disk, separate the write requests from it;

[0019] The KEY or label pointing to this data write request is saved in the raftlog after stripping the data payload;

[0020] Then, the data payload in the write request is written to disk using the data write logic;

[0021] Finally, the raftlog stripped of data payload is written to disk;

[0022] The data write request is a request initiated by the client or the system to write data to the database, including a description of the write action and the data payload to be written; the KEY or tag is an ID that uniquely identifies a write request; the raftlog is a data block used in the raft protocol to transfer data and achieve consensus between nodes; the data write logic is the logic of writing data to disk through the storage engine, which may or may not include WAL.

[0023] Furthermore, the method of using raftlog to write data is as follows:

[0024] 2.1) Read the KEY or label from the raftlog;

[0025] 2.2) Find the corresponding WAL by KEY or label;

[0026] 2.3) Get the data payload from WAL;

[0027] 2.4) Write the data to disk.

[0028] Furthermore, the method to cancel data writing when raft consensus fails is as follows:

[0029] 3.1) Read KEY or label from raftlog;

[0030] 3.2) Find the corresponding WAL by KEY or label;

[0031] 3.3) Delete WAL;

[0032] 3.4) Delete raftlog.

[0033] Furthermore, the method of obtaining data from disk, encapsulating a complete raft message and sending it to the lagging follower is as follows:

[0034] 4.1) Read the KEY or label from the raftlog;

[0035] 4.2) Find the corresponding WAL by KEY or label;

[0036] 4.3) Get data payload from WAL;

[0037] 4.4) Encapsulate the data payload and raftlog into the raft message;

[0038] 4.5) Send the raft message to the follower.

[0039] The present invention also claims a system for improving the writing efficiency of a distributed database based on the Raft protocol, wherein:

[0040] Reduce the number of times data is written to disk by unifying the data load in raftlog with the data load during actual writing.

[0041] Associate the data payload in the raftlog with the actual data payload through labels, keys, or references;

[0042] By writing different status values ​​at different stages, the data status can be marked as uncommitted, committed, rolled back, etc.

[0043] And associate the data payload and the corresponding status through the label or key or reference;

[0044] The disk data payload is written when the raftlog is written to the disk, and the status is modified when the disk is actually written. This ultimately reduces the number of times and amount of data actually written to the disk, and improves data writing efficiency.

[0045] The system specifically improves the writing efficiency of the distributed database based on the Raft protocol through the above method.

[0046] The present invention also claims a device for improving the writing efficiency of a distributed database based on the Raft protocol, comprising: at least one memory and at least one processor;

[0047] The at least one memory is used to store a machine-readable program;

[0048] The at least one processor is used to call the machine-readable program to implement the above method.

[0049] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which can implement the above method when executed by a processor.

[0050] Compared with the prior art, the method and system of the present invention for improving the writing efficiency of a distributed database based on the Raft protocol have the following beneficial effects:

[0051] The present invention effectively reduces the performance loss caused by writing raftlog to disk in the raft protocol, reduces the load of raftlog, reduces the number of write times of data payload, reduces the occupancy of disk space, and has a significant effect on improving the writing speed of distributed databases and reducing the occupancy of disk space. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a diagram showing the writing process of a distributed database after combining the data load of raftlog and WAL provided by an embodiment of the present invention;

[0053] Figure 2 The diagram is a process diagram of obtaining data from a disk and encapsulating it into a raft message and sending it to a follower, provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The present invention will be further described below in conjunction with specific embodiments.

[0055] Regarding the issues raised in the background technology, in fact, the data payload occupies the largest space in the raftlog. If the data payload in the raftlog can be reduced or even cleared, the IO resource usage can be greatly reduced, the time consumed by writing the raftlog to the disk can be reduced, and the overall data writing efficiency can be improved.

[0056] Therefore, an embodiment of the present invention provides a method for improving the writing efficiency of a distributed database based on the Raft protocol, combining the writing of raftlog with WAL to disk and unifying the data load in raftlog and WAL.

[0057] Combine the writing of raftlog to disk with the writing of WAL to disk as follows:

[0058] 1.1. Separate the data load in raftlog;

[0059] 1.2. Write the data payload to disk by writing WAL;

[0060] 1.3. Record the corresponding KEY or label in raftlog to associate WAL;

[0061] 1.4. Write the raftlog stripped of data payload to disk.

[0062] The method of using raftlog to write data is as follows:

[0063] 2.1. Read KEY or label from raftlog;

[0064] 2.2. Find the corresponding WAL by KEY or label;

[0065] 2.3. Get data payload from WAL;

[0066] 2.4. Write data to disk.

[0067] The method to cancel data writing when raft consensus fails is as follows:

[0068] 3.1. Read KEY or label from raftlog;

[0069] 3.2. Find the corresponding WAL by KEY or label;

[0070] 3.3. Delete WAL;

[0071] 3.4. Delete raftlog.

[0072] The method to obtain data from disk and encapsulate a complete raft message to send to the lagging follower is as follows:

[0073] 4.1. Read KEY or label from raftlog;

[0074] 4.2. Find the corresponding WAL by KEY or label;

[0075] 4.3. Get data payload from WAL;

[0076] 4.4) Encapsulate the data payload and raftlog into the raft message;

[0077] 4.5. Send the raft message to the follower.

[0078] This method unifies the data load when raftlog and data are written, such as Figure 1 As shown, the specific implementation process of this method is as follows:

[0079] When writing raftlog to disk, the data payload is extracted and written to disk separately. By combining the writing of raftlog to disk with the writing of actual data to disk, raft and storage engine can access the same data payload. Combining the writing of raftlog to disk with the writing of WAL to disk allows WAL and raftlog to share the same data payload.

[0080] After the raft consensus succeeds, the data payload in the WAL is read and written to the disk; after the raft consensus fails, the corresponding raftlog and WAL are deleted.

[0081] When the leader needs to obtain the raftlog from the storage and send it to the follower, it obtains the data payload from the corresponding WAL, encapsulates it into a raft message and sends it to the follower.

[0082] The specific process of unifying the data load when writing raftlog to disk is as follows:

[0083] First, when the distributed database executes data writing, the write request is converted into raftlog through the raft protocol and passed to each node to initiate consensus. After each node receives it, it needs to write raftlog to disk to prevent data loss.

[0084] Secondly, before writing raftlog to disk, it separates the write requests;

[0085] Then, the KEY or label pointing to this data write request is saved in the raftlog stripped of the data payload;

[0086] Then, the data payload in the write request is written to disk using the data write logic;

[0087] Finally, the raftlog stripped of data payload is written to disk;

[0088] The data write request is a request initiated by the client or the system to write data to the database, including a description of the write action and the data payload to be written; the KEY or tag is an ID that uniquely identifies a write request; the raftlog is a data block used in the raft protocol to transfer data and achieve consensus between nodes; the data write logic is the logic of writing data to disk through the storage engine, which may or may not include WAL.

[0089] After the raft consensus succeeds, the specific process of using raftlog to write data to disk is as follows:

[0090] First, get the KEY or label of the write request in the raftlog;

[0091] Then, obtain the data payload in the corresponding WAL through the KEY or label;

[0092] Finally, the data is written to the disk.

[0093] After the raft consensus fails, the specific process of deleting the raftlog is as follows:

[0094] First, get the KEY or label of the write request in the raftlog;

[0095] Then, delete the corresponding WAL by KEY or tag;

[0096] Finally, delete the raftlog.

[0097] like Figure 2 As shown in the figure, the specific process of sending a raft message containing a data payload to a follower is:

[0098] First, get the KEY or label of the write request in the raftlog;

[0099] Then, obtain the data payload in the corresponding WAL through the KEY or label;

[0100] Finally, the data payload and raftlog are encapsulated into a raft message and sent to the follower.

[0101] This method improves the writing efficiency of the distributed database based on the raft protocol and reduces the resource usage of raftlog. It reduces the number of disk writes during data writing and improves the writing speed. At the same time, it also effectively reduces the disk space usage. It also retains the reliability of raftlog and WAL and ensures data security. This method has clear logic, strong operability, and is easy to promote.

[0102] The embodiment of the present invention further provides a system for improving the writing efficiency of a distributed database based on the Raft protocol, the system:

[0103] Reduce the number of times data is written to disk by unifying the data load in raftlog with the data load during actual writing.

[0104] Associate the data payload in the raftlog with the actual data payload through labels, keys, or references;

[0105] By writing different status values ​​at different stages, the data status can be marked as uncommitted, committed, rolled back, etc.

[0106] And associate the data payload and the corresponding status through the label or key or reference;

[0107] The disk data payload is written when the raftlog is written to the disk, and the status is modified when the disk is actually written. This ultimately reduces the number of times and amount of data actually written to the disk, and improves data writing efficiency.

[0108] The system specifically improves the writing efficiency of a distributed database based on the Raft protocol through the method for improving the writing efficiency of a distributed database based on the Raft protocol described in the above embodiment.

[0109] When writing raftlog to disk, the data payload is extracted and written to disk separately. By combining the writing of raftlog to disk with the writing of actual data to disk, raft and storage engine can access the same data payload. Combining the writing of raftlog to disk with the writing of WAL to disk allows WAL and raftlog to share the same data payload.

[0110] After the raft consensus succeeds, the data payload in the WAL is read and written to the disk; after the raft consensus fails, the corresponding raftlog and WAL are deleted.

[0111] When the leader needs to obtain the raftlog from the storage and send it to the follower, it obtains the data payload from the corresponding WAL, encapsulates it into a raft message and sends it to the follower.

[0112] The specific process of unifying the data load when writing raftlog to disk is as follows:

[0113] First, when the distributed database executes data writing, the write request is converted into raftlog through the raft protocol and passed to each node to initiate consensus. After each node receives it, it needs to write raftlog to disk to prevent data loss.

[0114] Secondly, before writing raftlog to disk, it separates the write requests;

[0115] Then, the KEY or label pointing to this data write request is saved in the raftlog stripped of the data payload;

[0116] Then, the data payload in the write request is written to disk using the data write logic;

[0117] Finally, the raftlog stripped of data payload is written to disk;

[0118] The data write request is a request initiated by the client or the system to write data to the database, including a description of the write action and the data payload to be written; the KEY or tag is an ID that uniquely identifies a write request; the raftlog is a data block used in the raft protocol to transfer data and achieve consensus between nodes; the data write logic is the logic of writing data to disk through the storage engine, which may or may not include WAL.

[0119] After the raft consensus succeeds, the specific process of using raftlog to write data to disk is as follows:

[0120] First, get the KEY or label of the write request in the raftlog;

[0121] Then, obtain the data payload in the corresponding WAL through the KEY or label;

[0122] Finally, the data is written to the disk.

[0123] After the raft consensus fails, the specific process of deleting the raftlog is as follows:

[0124] First, get the KEY or label of the write request in the raftlog;

[0125] Then, delete the corresponding WAL by KEY or tag;

[0126] Finally, delete the raftlog.

[0127] The specific process of sending a raft message containing data payload to a follower is as follows:

[0128] First, get the KEY or label of the write request in the raftlog;

[0129] Then, obtain the data payload in the corresponding WAL through the KEY or label;

[0130] Finally, the data payload and raftlog are encapsulated into a raft message and sent to the follower.

[0131] An embodiment of the present invention further provides a device for improving the write efficiency of a distributed database based on the Raft protocol, comprising: at least one memory and at least one processor;

[0132] The at least one memory is used to store a machine-readable program;

[0133] The at least one processor is used to call the machine-readable program to implement the method for improving the writing efficiency of a distributed database based on the Raft protocol as described in the above embodiment.

[0134] The embodiment of the present invention further provides a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by the processor, the method for improving the write efficiency of a distributed database based on the Raft protocol described in the above embodiment is implemented. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0135] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0136] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.

[0137] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0138] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0139] The present invention is shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.

Claims

1. A method for improving the writing efficiency of a distributed database based on the Raft protocol, characterized in that: Unify the data load of raftlog and data writing, and implement it as follows: When writing raftlog to disk, the data payload is extracted and written to disk separately. By combining the writing of raftlog to disk with the writing of actual data to disk, raft and storage engine can access the same data payload. Combining the writing of raftlog to disk with the writing of WAL to disk allows WAL and raftlog to share the same data payload.

2. According to claim 1, a method for improving the writing efficiency of a distributed database based on the Raft protocol is characterized in that: After the raft consensus succeeds, the data payload in the WAL is read and written to the disk; after the raft consensus fails, the corresponding raftlog and WAL are deleted.

3. According to a method for improving the writing efficiency of a distributed database based on the Raft protocol according to claim 1, it is characterized in that: When the leader needs to obtain the raftlog from the storage and send it to the follower, it obtains the data payload from the corresponding WAL, encapsulates it into a raft message and sends it to the follower.

4. A method for improving the write efficiency of a distributed database based on the Raft protocol according to claim 1, 2 or 3, characterized in that: The steps of combining the writing of raftlog to disk with the writing of WAL to disk include: 1.1) Separate the data payload in raftlog; 1.2) Write the data payload to disk by writing WAL; 1.3) Record the corresponding KEY or label in raftlog to associate WAL; 1.4) Write the raftlog stripped of the data payload to disk.

5. According to claim 4, a method for improving the writing efficiency of a distributed database based on the Raft protocol is characterized in that: The method of using raftlog to write data is as follows: 2.1) Read the KEY or label from the raftlog; 2.2) Find the corresponding WAL by KEY or label; 2.3) Get the data payload from WAL; 2.4) Write the data to disk.

6. A method for improving the write efficiency of a distributed database based on the Raft protocol according to claim 5, characterized in that: The method to cancel data writing when raft consensus fails is as follows: 3.1) Read KEY or label from raftlog; 3.2) Find the corresponding WAL by KEY or label; 3.3) Delete WAL; 3.4) Delete raftlog.

7. A method for improving the writing efficiency of a distributed database based on the Raft protocol according to claim 1, characterized in that: The method to obtain data from disk and encapsulate a complete raft message to send to the lagging follower is as follows: 4.1) Read the KEY or label from the raftlog; 4.2) Find the corresponding WAL by KEY or label; 4.3) Get data payload from WAL; 4.4) Encapsulate the data payload and raftlog into the raft message; 4.5) Send the raft message to the follower.

8. A system for improving the writing efficiency of a distributed database based on the Raft protocol, characterized in that: Reduce the number of times data is written to disk by unifying the data load in raftlog with the data load during actual writing. Associate the data payload in the raftlog with the actual data payload through labels, keys, or references; By writing different status values ​​at different stages, the data status can be marked as uncommitted, committed, or rolled back. And associate the data payload and the corresponding state through the label or KEY or reference; Write the disk data payload when the raftlog is written to disk, and modify the state when it is actually written to disk; The system specifically improves the writing efficiency of a distributed database based on the Raft protocol through the method described in any one of claims 1 to 7.

9. A device for improving the writing efficiency of a distributed database based on the Raft protocol, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.

10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, can implement the method according to any one of claims 1 to 7.