Data processing method and system
By using the ‘majority write’ method in the replication group and allowing the client to continue to provide services when the server fails, the problems of high write operation delay and system unavailability in the prior art are solved, and efficient data processing and fault tolerance are achieved.
Patent Information
- Application Number
- WO2025107984P0
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-05-30
AI Technical Summary
When the server uses the master-slip replication protocol, the paxos protocol or the raft protocol for data reading and writing, there are problems such as high write delay and low read-write operation throughput, and when the master node fails over to the slave node, the system cannot provide services to the client.
A data processing method and system are proposed to write data by using the ‘majority write’ method in the replication group, and allowing the client to continue to provide services when the server fails, reducing operation delays and improving data processing efficiency. The specific implementation includes a client broadcasting write commands, the server in the copy group writes data when the conditions are met, and sends confirmation information to the client, and the client broadcasts and submits the command after receiving confirmation from more than half of the servers.
It effectively reduces write operation delay, improves data processing efficiency, and maintains system availability in a few server failures, solving the write amplification problem.
Smart Images

Figure CN2024127011_30052025_PF_FP_ABST
Abstract
Description
Data processing method and system
[0001] This application claims priority to the Chinese patent application filed with the China National Intellectual Property Administration on November 21, 2023, with application number 202311557722.2 and application name “A Data Processing Method and System”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of data replication, and in particular to a data processing method and system. Background Art
[0003] When the server uses a master-slave replication protocol for data reading and writing, there are issues with high write latency and low read and write throughput. Furthermore, during a master node failover to a slave node, the system will be unable to provide services to clients. When the server uses the Paxos or Raft protocols for data reading and writing, the globally unique write order must be determined by replicating the state machine log before data replication, leading to write amplification.
[0004] How to achieve a certain degree of fault tolerance while reducing operation latency and solving write amplification when implementing data replication is an urgent technical problem that needs to be solved.
[0005] Summary of the Invention
[0006] The present application discloses a data processing method and system, which can achieve data replication while continuing to provide services to the client when a small number of server failures occur, and can also reduce operation delays and improve data processing efficiency.
[0007] In the first aspect, the present application provides a data processing method, which is applied to a data processing system, which includes a first client and a replication group, wherein the local maximum term identifier of more than half of the servers in the replication group is the first identifier, and the first identifier indicates that the replication group has reached a consensus on the first server in the replication group as the first role. The method includes: the first client broadcasts a first write command, and the first write command includes the first identifier, the first content of the first object to be written, and the first version number to be written; when the local maximum term identifier of the server in the replication group is the first identifier and the maximum version number corresponding to the local first object is less than the first version number, the server writes the first content and the first version number locally, and sends a first message to the first client, and the first message indicates agreement to the first write command; when the first client receives the first message sent by more than half of the servers in the replication group, including the first server, the first client broadcasts a commit command, and the commit command indicates that the status of updating the first content corresponding to the first version number is submitted.
[0008] Here, the server side can be a network-side device with data processing capabilities. The network-side device can be, for example, a server deployed on the network side, or a component in the server (the component can be, for example, a chip, an integrated circuit, etc.). The network-side device can be deployed in a cloud environment or an edge environment, and is not specifically limited here.
[0009] Here, the client can be a terminal device, which can be, for example, a user device (mobile phone, computer, tablet computer, PDA, desktop, headphones, speakers, wearable device, vehicle-mounted device, virtual reality device, augmented reality device, etc.), smart home device, smart transportation equipment, smart manufacturing equipment (such as robots, industrial equipment, smart logistics, smart factories, etc.), or a component within the terminal device (such as a chip or integrated circuit).
[0010] Here, the first version number is associated with the first content object of the first object. The first version number identifies the write of the first content of the first object. It is understood that different objects to be written will carry different version numbers in the write command. For the same object, different content will be written, and the version number carried in the write command will also be different.
[0011] Here, when applied to a file storage system, the first object can be a basic data unit in the file storage system, which can be represented by a file ID and offset. When applied to object storage, the first object is the object's key; when applied to block storage, the first object can be the block's ID.
[0012] Exemplarily, the first role is a leader or a semi-leader. If the first role is a leader, the first server is the server that stores the latest version of data for all objects in the replication group. If the first role is a semi-leader, the first server may, for example, be the server that stores the latest version of data for some objects in the replication group. In a replication group, a server can only have one leader and one semi-leader, and they cannot exist simultaneously.
[0013] Correspondingly, when the first role is a leader, the replication state of the replication group is leader-based quorum replication, which can be recorded as the first state or the replication state corresponding to the leader; when the first role is a semi-leader, the replication state of the replication group is semi-leader quorum replication, which can be recorded as the second state or the replication state corresponding to the semi-leader.
[0014] In this application, the slot ID where the consensus value is located in the local log of the server is called the term identifier. In other words, the slot ID where the consensus value is stored in the log is called the term identifier. Here, the slot ID is a positive integer. The larger the slot ID, the larger the term identifier. The slot ID can also be understood as an entry in the log. Each slot in the log stores the consensus value of an instance. The consensus value records the identifier of the server that is the leader or semi-leader that the current replication group has reached a consensus on and the corresponding replication status. The server that reaches this consensus in the replication group stores the consensus value locally in the same slot of the log. Here, the process in which the replication group reaches a consensus on a server as a leader or semi-leader is called an instance.
[0015] It can be understood that if the server has participated in multiple elections for leaders or semi-leaders and a consensus has been reached in each election, the local log of the server also stores the consensus values of multiple instances. Further, assuming that these multiple instances include instance 1 and instance 2, where instance 2 occurs later than instance 1, the slot ID corresponding to the consensus value of instance 2 in the log is greater than the slot ID corresponding to the consensus value of instance 1 in the log, that is, the term identifier corresponding to instance 2 is greater than the term identifier corresponding to instance 1, then the current local maximum term identifier of the server is the term identifier corresponding to instance 2. In this application, the local maximum term identifier of the server can also be referred to as the latest local term identifier of the server.
[0016] As you can understand, each successful election corresponds to a new term identifier. When the replication group reaches consensus on the same server as the same role at different times, the corresponding term identifier is different.
[0017] For example, after the server writes the first content and the first version number, the status of the first content of the first object corresponding to the first version number is "Preparing," meaning that the first content of the current first object can be rolled back or committed. When the status of the first content of the first object corresponding to the first version number is "Committed," the first content of the current first object is readable.
[0018] In the above method, the write process uses a "majority write" method to write the content of the object to the replication group, that is, the client broadcasts a write command, and each server in the replication group can directly respond to the client after receiving the write command after locally judging whether the write conditions are met. When the client receives a reply from the majority (that is, more than half) of the servers in the replication group indicating that they agree to write, the client determines that the write can be completed. Compared with the use of the master-slave replication protocol, the existing paxos protocol or the raft protocol for write operations, this method can reduce the delay of write operations and improve data processing efficiency. In addition, compared with the master-slave replication protocol, the present application also has a certain degree of fault tolerance. For example, when less than half of the servers in the replication group fail, it can still continue to provide services to the client. Compared with the existing paxos protocol or raft protocol, the present application does not need to determine a globally unique write order before writing data, so it solves the write amplification problem caused by object writing.
[0019] Optionally, the number of servers included in the replication group is 2f+1, where f is a positive integer. In this case, "more than half" in the replication group can be understood as greater than f, or a majority. This provides a certain degree of fault tolerance.
[0020] Optionally, before the first client broadcasts the first write command, the method further includes: the first client obtains the first identifier from the replication group. Exemplarily, the "acquisition" can be performed each time a write is required, or periodically, or based on a trigger condition. In this way, the term identifier obtained by the first client from the replication group can be the latest term identifier stored locally by the majority of servers in the replication group, so that the first client can read and write objects in the correct manner.
[0021] Exemplarily, the method further includes: when the local maximum term identifier of the server in the replication group is the first identifier and the maximum version number corresponding to the local first object is greater than or equal to the first version number, a second message is sent to the first client, the second message indicates rejection of the first write command, and the second message includes the maximum version number corresponding to the local first object of the server; when the first client receives the second message sent by more than half of the servers in the replication group, the first client broadcasts a write command, the write command including the first identifier, the first content of the first object to be written, and the second version number to be written, the second version number being greater than the maximum version number corresponding to the first object in the second message received by the first client. That is, when the client receives the second message sent by the majority of the servers in the replication group indicating rejection of the first write command, the client knows that the write has failed, and the client can redetermine the version number based on the received second message, and re-initiate the write command using the version number to request the replication group to write the first content of the first object again.
[0022] Optionally, the first role is a leader, the data processing system also includes a second client, and the method also includes: the second client sends a read command to the first server, the read command is used to request reading the first object, and the read command includes the first identifier; when the local maximum term identifier of the first server is the first identifier, the latest content of the first object is sent to the second client.
[0023] Exemplarily, the latest content of the first object may be the above-mentioned first content, or may be other content of the first object written after the first content of the first object.
[0024] In the above implementation method, the first role is the leader. The client knows that the replication group has reached a consensus on the first server as the leader through the obtained first identifier. In this case, the client uses the "leader read" mode to directly request the first server as the leader to read the target object.
[0025] Exemplarily, the first server sends the latest content of the first object to the second client, including: after the first server updates the status of the first content corresponding to the first version number to submitted based on the commit command, the first server sends the latest content of the first object to the second client. That is, when the server receives a read command requesting to read the first object, if the content of the first object locally on the server is currently in the above-mentioned "prepared" state and has not yet been in the "submitted" state, the server needs to wait until the content of the first object is in the "submitted" state before responding to the read command. In this way, the client can avoid reading expired data.
[0026] Optionally, the first role is a leader, the data processing system also includes a second client, and the method also includes: the second client sends a read command to the first server, the read command includes a second identifier different from the first identifier; the first server replies to reject the read command and sends the first identifier to the second client.
[0027] It can be understood that when there is no failure on the first server, the first identifier is still the latest term identifier (or maximum term identifier) stored locally by the majority of servers in the replication group. In this case, the second client may not have queried a new term identifier from the replication group for a long time. Therefore, the read command sent by the second client carries the second identifier that was most recently used by the second client in history. Since the second identifier has expired, the read command sent by the second client will be rejected by the first server, thereby avoiding the second client from reading expired data.
[0028] Optionally, the first role is a leader, the data processing system further includes a third client, and the method further includes: the third client sending a read command to the first server; and when any of the following conditions is met, the first server responds to reject the read command sent by the third client:
[0029] Before receiving the read command sent by the third client, the first server fails to retry the heartbeat communication within a preset time period; or
[0030] Before receiving the read command sent by the third client, the first server determines that the first identifier is invalid based on the heartbeat information sent by the second server in the replication group, and the heartbeat information includes the local maximum term identifier of the second server and the local maximum term identifier of the second server is greater than the first identifier.
[0031] Here, the failure of the first server to retry the heartbeat communication within the preset time period can be understood as, for example, the failure of the first server cannot be recovered. Taking a successful heartbeat communication as an example, the first server broadcasts a heartbeat message once and receives heartbeat messages sent by more than half of the servers in the replication group. The first server determines that this heartbeat is successful. Assuming that the server sends a heartbeat message once every interval of time T, and the server is allowed to try a maximum of M times, the above-mentioned preset time period can be M*T, which can also be called the maximum time allowed for the server to retry the heartbeat communication.
[0032] For example, when any of the above conditions is met, the first server can also mark itself as invalid locally. When the first server is in an invalid state, the first server rejects read commands and write commands from any client, that is, the first client does not provide services to the outside.
[0033] When the above implementation method is implemented, when the failure of the server acting as the leader cannot be recovered or the server acting as the leader determines that the local maximum term identifier has expired (or is invalid), the server rejects the read command from the client to prevent the client from reading expired content of the object.
[0034] Optionally, the first role is the leader, and less than half of the servers in the replication group fail, excluding the first server. This way, if a few servers in the replication group (excluding the first server, which is the leader) fail, the read and write processes in leader mode are not affected, and the replication group can continue to provide services to clients.
[0035] Optionally, the first role is a leader, and in the event of a failure of the first server, the local maximum term identifier of more than half of the servers in the replication group is a third identifier, and the third identifier is greater than the first identifier. The third identifier indicates that the replication group has reached a consensus that the third server in the replication group serves as the second role, and the second role is a semi-leader. The method also includes: within the waiting period from the time the third server serves as the second role, if the server in the replication group whose local maximum term identifier is the third identifier receives a second write command, the server replies with a rejection of the second write command; wherein, the waiting period is greater than the maximum period allowed for the first server to retry heartbeat communication.
[0036] Here, the maximum time duration allowed for the first server to retry heartbeat communication can refer to the above description of the maximum time duration allowed for the server to retry heartbeat communication, which will not be repeated here.
[0037] By implementing the above-mentioned implementation method, in the event of a failure of the first server, by limiting the servers in the replication group that have reached a consensus on the third server as the semi-leader, if a write command is received from the client within the waiting period starting from the local maximum term identifier of the third identifier, it must be rejected, and the write service will not be started until the waiting period has passed. The waiting period is sufficient for the first server as the leader to discover that it has failed. In this way, when new content of the target object is written to the replication group, the client can avoid reading expired content of the target object from the first server using an expired first identifier.
[0038] Optionally, the method further includes: when the third server determines that it is the server that currently stores the latest version of data for all objects in the replication group, the third server broadcasts an election request, wherein the election request is used to request that the third server be elected as the leader. In this manner, the third server, acting as a semi-leader, may also initiate a new election to request that the third server be elected as the leader.
[0039] Optionally, the first role is a semi-leader, and when the local maximum term identifier of the server in the replication group is the first identifier and the maximum version number corresponding to the local first object is less than the first version number, the first content and the first version number are written, including: if the server in the replication group receives the first write command after the waiting time since the first server acts as the first role, the local maximum term identifier of the server is the first identifier and the maximum version number corresponding to the local first object is less than the first version number, the first content and the first version number are written; wherein, the waiting time is greater than the maximum time allowed for the server as the leader to retry heartbeat communication.
[0040] In the above implementation method, the first role is a semi-leader. By limiting the servers in the replication group that reach a consensus on the first server as the semi-leader, they must wait for the local maximum term identifier to become the first identifier before they can start the write service. The waiting time is sufficient for the server as the leader to discover that it has failed. Therefore, when new content of the target object is written to the replication group, when the client uses an expired term identifier to read the target object from the replication group, the client is prevented from reading expired content of the target object.
[0041] Optionally, the data processing system also includes a second client, and the method also includes: the second client broadcasts a read command, the read command is used to request reading of the first object, and the read command includes the first identifier; when the local maximum term identifier of the server in the replication group is the first identifier, the read result is sent to the second client, and the read result includes a second version number and the second content of the first object corresponding to the second version number, and the second version number is the maximum version number corresponding to the first object locally on the server; when the second client receives the read result sent by more than half of the servers in the replication group, it obtains the second content of the first object corresponding to the maximum version number in the received read result.
[0042] In the implementation described above, when the first role is a semi-leader, the client obtains the first identifier to know that the replication group has reached a consensus on the first server as the semi-leader. In this case, the client uses the "majority read" mode to broadcast a read command to read the target object. Since the object content is written using the "majority write" mode, the servers that participated in writing the latest content of the target object and the servers that fed back the read results must have an intersection (i.e., there is at least one identical server), thereby ensuring that the client will definitely read the latest content of the target object. Compared to the master-slave replication protocol, the Paxos protocol, or the Raft protocol, this ensures that the read process will not be affected when a few servers fail.
[0043] In the second aspect, the present application provides a data processing method, which is applied to a client, and the method includes: broadcasting a write command, the write command including a first identifier, the first content of the first object to be written, and the first version number to be written, the first identifier being the local maximum term identifier of more than half of the servers in the replication group, and the first identifier indicating that the replication group has reached a consensus on the first server in the replication group as the first role; when the client receives the first information sent by more than half of the servers in the replication group, broadcasting a commit command, wherein the first information indicates agreement to the write command, and the commit command indicates that the status of updating the first content corresponding to the first version number is submitted.
[0044] Exemplarily, the first role is a leader or a semi-leader.
[0045] Here, the client, replication group, first role, etc. can refer to the relevant description of the corresponding content of the first aspect above, and will not be repeated here.
[0046] In the above method, the client writes the content of the object into the replication group using a "majority write" method, that is, the client broadcasts a write command, and when the client receives a reply from the majority (that is, more than half) of the servers in the replication group indicating that they agree to write, the client determines that the write can be completed. Compared with the use of the master-slave replication protocol, the existing paxos protocol or the raft protocol for write operations, this method can reduce the delay of write operations and improve data processing efficiency. In addition, compared with the master-slave replication protocol, the present application also has a certain degree of fault tolerance. For example, when less than half of the servers in the replication group fail, the write process initiated by the client is not affected. Compared with the existing paxos protocol or raft protocol, the present application does not need to determine a globally unique write order before writing data, so it solves the write amplification problem caused by object writing.
[0047] Optionally, before the broadcast write command, the method further includes: obtaining the first identifier from the replication group. Here, the description of "obtaining" can refer to the description of the corresponding content of the first aspect. In this way, the term identifier that the client can obtain from the replication group is the latest term identifier locally stored by the majority of servers in the replication group, so that the first client can read and write objects in the correct manner.
[0048] Optionally, the first role is a semi-leader, and the method further includes: broadcasting a read command, the read command being used to request reading of the first object, the read command including the first identifier; receiving a read result sent by a server in the replication group, the read result including a second version number and a second content of the first object corresponding to the second version number, the second version number being the maximum version number corresponding to the first object locally on the server; when the client receives the read result sent by more than half of the servers in the replication group, obtaining the second content of the first object corresponding to the maximum version number from the received read result.
[0049] When the above implementation method is implemented, when the first server acts as a semi-leader, the client uses the "majority read" method to read the target object from the replication group. Since the server participating in writing the latest content of the target object and the server that feeds back the read result must have an intersection (that is, there is at least one identical server), it can be ensured that the client will definitely read the latest content of the target object.
[0050] Optionally, the first role is a leader, and the method further includes: sending a read command to the first server, the read command being used to request reading of the first object, the read command including the first identifier; and receiving the latest content of the first object sent by the first server.
[0051] When the above implementation method is implemented, when the first server acts as a leader, after the client obtains the first identifier in a timely manner, the client uses the "leader read" mode to directly read the target object from the first server, which reduces the read latency and is conducive to improving data processing efficiency.
[0052] Optionally, the first role is a leader, and the method further includes: obtaining a second identifier from the replication group, the second identifier being the local maximum term identifier of more than half of the servers currently in the replication group, the second identifier indicating that the replication group has reached a consensus on the second server in the replication group as the second role, the second role is a semi-leader, and the second identifier is greater than the first identifier; and performing a data read operation or a data write operation according to the second identifier.
[0053] By implementing the above implementation method, the client knows that the current second server of the replication group is the semi-leader through the second identifier obtained in time to reach a consensus. When the client has a writing or reading requirement, the term identifier carried in the write command or read command sent is the second identifier.
[0054] Optionally, the method also includes: sending a read command to the first server, the read command being used to request reading of the first object, the read command including the first identifier; obtaining the second identifier from the replication group includes: obtaining the second identifier from the replication group when receiving a reply from the first server rejecting the read command.
[0055] It can be seen that the client sends a read command to the first server as the leader based on the first identifier. When the client receives the information from the first server indicating that the read command is rejected, it will trigger the client to obtain the latest term identifier from the replication group.
[0056] Exemplarily, after broadcasting the first write command, the method further includes: receiving a second message sent by the server in the replication group, the second message indicating rejection of the first write command, the second message including the maximum version number corresponding to the first object locally on the server; when the client receives the second message sent by more than half of the servers in the replication group, broadcasting a second write command, the second write command including a first identifier, the first content of the first object to be written, and a third version number to be written, the third version number being greater than the maximum version number corresponding to the first object in the second message received by the first client.
[0057] That is to say, when the client receives the second information sent by the majority of servers in the replication group indicating the rejection of the first write command, the client knows that the write has failed. The client can redetermine the version number based on the received second information and use the version number to re-initiate the write command to request the replication group to write the first content of the first object again.
[0058] In the third aspect, the present application provides a data processing system, which includes a first client and a replication group including multiple servers, wherein the local maximum term identifier of more than half of the servers in the replication group is a first identifier, and the first identifier indicates that the replication group has reached a consensus on the first server in the replication group as the first role. The first client is used to broadcast a first write command, and the first write command includes the first identifier, the first content of the first object to be written, and the first version number to be written; the server in the replication group is used to write the first content and the first version number when the local maximum term identifier is the first identifier and the maximum version number corresponding to the local first object is less than the first version number, and send a first message to the first client, wherein the first message indicates agreement to the first write command; the first client is used to broadcast a commit command when receiving the first message sent by more than half of the servers in the replication group, including the first server, and the commit command indicates that the status of updating the first content corresponding to the first version number is submitted.
[0059] Optionally, the number of servers included in the replication group is 2f+1, where f is a positive integer.
[0060] Optionally, before the first client is used to broadcast the first write command, the first client is further used to obtain the first identifier from the replication group.
[0061] Exemplarily, the server in the replication group is further used to send a second message to the first client when the local maximum term identifier is the first identifier and the maximum version number corresponding to the local first object is greater than or equal to the first version number, the second message indicating rejection of the first write command, the second message including the maximum version number corresponding to the local first object of the server; the first client is used to broadcast a write command when receiving the second message sent by more than half of the servers in the replication group, the write command including the first identifier, the first content of the first object to be written and the second version number to be written, the second version number being greater than the maximum version number corresponding to the first object in the second message received by the first client.
[0062] Optionally, the first role is a leader, and the system also includes a second client, which is used to send a read command to the first server, and the read command is used to request to read the first object, and the read command includes the first identifier; the first server is used to send the latest content of the first object to the second client when the local maximum term identifier is the first identifier.
[0063] Furthermore, the first server is specifically used to: after updating the status of the first content corresponding to the first version number to submitted based on the submission command, if the local maximum term identifier is the first identifier, send the latest content of the first object to the second client.
[0064] Optionally, the first role is a leader, and the system also includes a second client, which is used to send a read command to the first server, and the read command includes a second identifier different from the first identifier; the first server is also used to reply to reject the read command and send the first identifier to the second client.
[0065] Optionally, the first role is a leader, and the system further includes a third client, wherein the third client is configured to send a read command to the first server; and the first server is further configured to reply to reject the read command sent by the third client when any of the following conditions is met:
[0066] Before receiving the read command sent by the third client, the first server fails to retry the heartbeat communication within a preset time period; or
[0067] Before receiving the read command sent by the third client, the first server determines that the first identifier is invalid based on the heartbeat information sent by the second server in the replication group, and the heartbeat information includes the local maximum term identifier of the second server and the local maximum term identifier of the second server is greater than the first identifier.
[0068] Optionally, the first role is a leader, less than half of the servers in the replication group are faulty, and the faulty servers do not include the first server.
[0069] Optionally, the first role is a leader. In the event of a failure of the first server, the local maximum term identifier of more than half of the servers in the replication group is a third identifier, and the third identifier is greater than the first identifier. The third identifier indicates that the replication group has reached a consensus that the third server in the replication group serves as the second role. The second role is a semi-leader. Within the waiting period from the time the third server serves as the second role, if the server in the replication group whose local maximum term identifier is the third identifier receives a second write command, the server is also used to reply to reject the second write command; wherein, the waiting period is greater than the maximum period allowed for the first server to retry heartbeat communication.
[0070] Optionally, when the third server determines that it is the server that currently stores the latest version data of all objects in the replication group, the third server is further used to broadcast an election request, which is used to request the election of the third server as a leader.
[0071] Optionally, the first role is a semi-leader. If the server in the replication group receives the first write command after the waiting time since the first server took on the role of the first role, the server is used to write the first content and the first version number when the local maximum term identifier is the first identifier and the maximum version number corresponding to the local first object is less than the first version number; wherein the waiting time is greater than the maximum time allowed for the server as the leader to retry heartbeat communication.
[0072] Optionally, the system also includes a second client, which is used to broadcast a read command, wherein the read command is used to request reading of the first object, and the read command includes the first identifier; the server in the replication group is used to send a read result to the second client when the local maximum term identifier is the first identifier, and the read result includes a second version number and the second content of the first object corresponding to the second version number, and the second version number is the maximum version number corresponding to the first object locally on the server; the second client is used to obtain the second content of the first object corresponding to the maximum version number in the received read result when receiving the read result sent by more than half of the servers in the replication group.
[0073] In a fourth aspect, the present application provides a device for data processing, which is a client or is included in the client, and the device includes: a sending unit, used to broadcast a write command, the write command including a first identifier, the first content of the first object to be written, and the first version number to be written, the first identifier being the local maximum term identifier of more than half of the servers in the replication group, and the first identifier indicating that the replication group has reached a consensus on the first server in the replication group as the first role; the sending unit is also used to broadcast a commit command when the receiving unit in the device receives the first information sent by more than half of the servers in the replication group, wherein the first information indicates agreement to the write command, and the commit command indicates that the status of updating the first content corresponding to the first version number is submitted.
[0074] Optionally, the receiving unit is further configured to obtain the first identifier from the replication group.
[0075] Optionally, the first role is a semi-leader: the sending unit is also used to broadcast a read command, the read command is used to request reading of the first object, and the read command includes the first identifier; the receiving unit is used to receive the read result sent by the server in the replication group, the read result includes a second version number and the second content of the first object corresponding to the second version number, and the second version number is the maximum version number corresponding to the first object locally on the server; when the receiving unit receives the read result sent by more than half of the servers in the replication group, the processing unit in the device is used to obtain the second content of the first object corresponding to the maximum version number from the received read result.
[0076] Optionally, the first role is a leader, and the sending unit is further used to send a read command to the first server, where the read command is used to request reading of the first object, and the read command includes the first identifier; the receiving unit is further used to receive the latest content of the first object sent by the first server.
[0077] Optionally, the first role is a leader, and the receiving unit is further used to obtain a second identifier from the replication group, where the second identifier is the local maximum term identifier of more than half of the servers currently in the replication group, and the second identifier indicates that the replication group has reached a consensus on the second server in the replication group as the second role, the second role is a semi-leader, and the second identifier is greater than the first identifier; the processing unit in the device is used to perform data read or write operations based on the second identifier.
[0078] Optionally, the sending unit is also used to send a read command to the first server, where the read command is used to request reading of the first object, and the read command includes the first identifier; the receiving unit is specifically used to obtain the second identifier from the replication group when a reply from the first server rejecting the read command is received.
[0079] Exemplarily, the receiving unit is further used to receive second information sent by the server in the replication group, the second information indicating rejection of the first write command, the second information including the maximum version number corresponding to the first object locally on the server; when the receiving unit receives the second information sent by more than half of the server in the replication group, the sending unit is further used to broadcast a second write command, the second write command including a first identifier, the first content of the first object to be written, and a third version number to be written, the third version number being greater than the maximum version number corresponding to the first object in the second information received by the first client.
[0080] In a fifth aspect, the present application provides a device for data processing, which includes a processor and a memory, wherein the memory is used to store program instructions; the processor calls the program instructions in the memory so that the device executes the method in the second aspect or any possible implementation of the second aspect.
[0081] In a sixth aspect, the present application provides a computer-readable storage medium comprising computer instructions, which, when executed by a processor, implement the method in the above-mentioned second aspect or any possible implementation of the second aspect; and when the computer instructions are executed by the above-mentioned data processing system, implement the method in the above-mentioned first aspect or any possible implementation of the first aspect.
[0082] In the seventh aspect, the present application provides a computer program product, which, when executed by a processor, implements the method in the above-mentioned second aspect or any possible embodiment of the second aspect; or, when the computer program product is executed by a data processing system, implements the method in the above-mentioned first aspect or any possible embodiment of the first aspect.
[0083] Exemplarily, the computer program product may be a software installation package. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] FIG1 is a schematic diagram of the architecture of a data processing system provided in an embodiment of the present application;
[0085] FIG2 is a flow chart of a data writing method provided in an embodiment of the present application;
[0086] FIG3 is a flow chart of a data reading method provided in an embodiment of the present application;
[0087] FIG4 is a flow chart of another data reading method provided in an embodiment of the present application;
[0088] FIG5 is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0089] FIG6A is a schematic diagram of a process of electing a semi-leader in a replication group running the Paxos protocol according to an embodiment of the present application;
[0090] FIG6B is a schematic diagram of another process of electing a semi-leader by running the Paxos protocol in a replication group according to an embodiment of the present application;
[0091] FIG7 is a schematic diagram of a log for storing consensus values locally on each server in a replication group provided by an embodiment of the present application;
[0092] FIG8 is a schematic structural diagram of a device for data processing provided in an embodiment of the present application;
[0093] FIG9 is a schematic structural diagram of another apparatus for data processing provided in an embodiment of the present application;
[0094] FIG10 is a schematic structural diagram of a communication device provided in this embodiment of the present application. DETAILED DESCRIPTION
[0095] It should be noted that the prefixes such as "first" and "second" used in this application are only for distinguishing different description objects, and do not have any limiting effect on the position, order, priority, quantity or content of the described objects. For example, if the described object is a "field", then the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields", and "first" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the described object is a "level", then the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of described objects is not limited by the prefix and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the described object is a "device," then the "first device" and the "second device" can be the same device, the same type of device, or different types of devices. For another example, if the described object is "information," then the "first information" and the "second information" can be information of the same content or information of different contents. In short, the use of prefixes to distinguish the described objects in the embodiments of this application does not constitute a limitation on the described objects. For the description of the described objects, please refer to the description in the context of the claims or embodiments, and the use of such prefixes should not constitute an unnecessary limitation.
[0096] It should be noted that the descriptions used in the embodiments of the present application, such as "at least one of a1, a2, ..., and an" and the like, include any one of a1, a2, ..., and an existing alone, and any combination of any multiple of a1, a2, ..., and an, each of which can exist alone. For example, the description "at least one of a, b, and c" includes a alone, b alone, c alone, a combination of a and b, a combination of a and c, a combination of b and c, or a combination of ab and c.
[0097] To facilitate understanding, the following first introduces relevant terms that may be involved in the embodiments of this application.
[0098] (1) Hybrid logical clock
[0099] The Hybrid Logical Clock (HLC) is a clock algorithm used in distributed systems to provide ordered timestamps for events. Event-based time can be used to determine the order of events in a distributed system.
[0100] The HLC clock algorithm is based on a combination of physical and logical clocks. The physical clock is a real-world clock based on physical time. All nodes in a distributed system have access to the physical clock and provide a globally synchronized logical clock. Each node participating in the distributed system maintains its own logical clock counter, which increments when an event occurs, assigning a logical timestamp to each event. Therefore, the logical clock can be understood as an incrementing counter.
[0101] In an embodiment of the present application, the version number carried in the write command can be generated by a hybrid logical clock. The version number is used to distinguish different versions of a data object.
[0102] (2) Broadcast
[0103] Broadcasting is a one-to-many communication method. It allows a broadcast source to send information to all receivers within its broadcast communication domain. Each receiver within the broadcast communication domain is identified by a broadcast address. When a broadcast source sends information to this broadcast address, all receivers within the broadcast communication domain can receive and process the information.
[0104] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0105] Referring to FIG1 , FIG1 is a schematic diagram of the architecture of a data processing system provided in an embodiment of the present application. The system can implement data read and / or write operations, enabling clients to access services provided by servers with high availability. As shown in FIG1 , the system includes a replication group and at least one client, wherein the replication group includes multiple servers, and the client communicates with the servers in the replication group via a data center network, either wired or wirelessly, without specific limitation herein.
[0106] Each server in the replication group is deployed with distributed service logic and a replication protocol. The distributed service logic is used to process user requests from the client according to service requirements, and the replication protocol can synchronize the relevant operations of the distributed service logic to the other servers in the replication group. Here, the replication protocol can be understood as the data processing method provided in the following embodiment of the present application, which can be integrated into distributed storage or distributed service software as program code, or deployed on the server in a combination of software and hardware. No specific limitation is made here.
[0107] It can be understood that a replication group is a collection of multiple servers with data replication capabilities. The total number of servers in a replication group is n, where n = 2f + 1, and f is a positive integer. For data objects assigned to each replication group (hereinafter referred to as objects), the data processing method provided by this application can achieve high-availability access to the data objects (for example, when fewer than f servers in the replication group fail, read and write operations on the data objects can still be performed).
[0108] For example, each server in a replication group can play one of three roles at any given time: leader, semi-leader, or follower. The server's role can switch at different times according to pre-set rules. At any given time, a server in a replication group can only play one leader or semi-leader role, and they cannot exist simultaneously. However, a server in the replication group can play multiple followers.
[0109] For example, suppose the replication group determines, through consensus, that Server 1 is the leader. In this case, Server 1 is the server that stores the latest versions of all objects in the replication group. For another example, suppose the replication group determines, through consensus, that Server 2 is the semi-leader. In this case, Server 2 could be the server that stores the latest versions of some objects in the replication group.
[0110] In an embodiment of the present application, there are two replication states in the replication group, wherein the first state is leader-based quorum replication (leader-based quorum replication), and the second state is semi-leader quorum replication (semi-leader quorum replication). Each server in the replication group locally maintains a log to store the consensus value of an instance of the consensus protocol (such as paxos), which records the identifier of the server that is the leader or semi-leader that the current replication group has reached consensus on and the corresponding replication state. The server that reaches this consensus in the replication group stores the consensus value locally in the same slot of the log. In an embodiment of the present application, the log is only used to store the consensus value, and the slot ID in the log where the consensus value is stored is called the term identifier. Among them, the slot ID is a positive integer, and the larger the slot ID, the larger the term identifier. In some possible embodiments, the slot ID can also be understood as an entry in the log. Here, the process of the replication group reaching a consensus on a certain server as a leader or semi-leader is called an instance.
[0111] It can be understood that the local log of the server may store the consensus values of multiple instances, including instance 1 and instance 2. Among them, instance 2 occurs later than instance 1, and the slot ID corresponding to the consensus value of instance 2 in the log is greater than the slot ID corresponding to the consensus value of instance 1 in the log, that is, the term identifier corresponding to instance 2 is greater than the term identifier corresponding to instance 1, then the current local maximum term identifier of the server is the term identifier corresponding to instance 2. Here, the local maximum term identifier of the server can also be called the local latest term identifier.
[0112] Taking the replication group performing a consensus operation as an example, suppose that server 1 in the replication group initiates a consensus vote request to elect server 1 as the leader, and after this consensus vote the replication group reaches a consensus on electing server 1 as the leader (that is, more than half of the servers in the replication group voted in favor). Finally, the servers that have reached consensus in the replication group store the consensus value of this instance (that is, the result of reaching consensus) in the slot corresponding to the log. Then, the local maximum term identifier of the current majority of servers in the replication group is the slot ID (or log ID) of the consensus value in the log. The consensus value corresponding to the term identifier records the current replication status of the replication group and the identifier of server 1 as the leader corresponding to the replication status.
[0113] As you can understand, each election generates a term identifier, which increments with each new election. A successful election means that more than half of the servers in the replication group have stored the consensus value of the instance corresponding to the election and the term identifier corresponding to that consensus value.
[0114] Taking the election of the leader as an example, suppose there are 3 servers in the replication group, namely server 1, server 2 and server 3. Suppose server 1 initiates a consensus vote request to elect server 1 as the leader, and both server 1 and server 2 voted in favor. Since more than half of the servers in the replication group voted in favor, the replication group currently reaches a consensus on server 1 as the leader, where the term identifier corresponding to server 1 is identifier 1. Then identifier 1 is the local maximum term identifier of more than half of the servers in the replication group (i.e., server 1 and server 2), and identifier 1 indicates that the replication group has determined server 1 in the replication group as the leader through consensus operation.
[0115] Here, the embodiment of the present application does not limit the number of replication groups. Figure 1 only shows one replication group. In some possible embodiments, the data processing system may also include multiple replication groups. Each replication group can refer to the description of the replication group in Figure 1 and will not be repeated here.
[0116] Here, the server side can be a network-side device with data processing capabilities. The network-side device can be, for example, a server deployed on the network side, or a component in the server (the component can be, for example, a chip, an integrated circuit, etc.). The network-side device can be deployed in a cloud environment or an edge environment, and is not specifically limited here.
[0117] Here, the client can be a terminal device, which can be, for example, a user device (mobile phone, computer, tablet computer, PDA, desktop, headphones, speakers, wearable device, vehicle-mounted device, virtual reality device, augmented reality device, etc.), smart home device (such as TV, sweeping robot, smart desk lamp, audio system, smart lighting system, electrical control system, home background music, home theater system, intercom system, video surveillance, etc.), smart transportation equipment (such as cars, ships, drones, trains, vans, trucks, etc.), smart manufacturing equipment (such as robots, industrial equipment, smart logistics, smart factories, etc.), or a component in the terminal device (such as a chip or integrated circuit).
[0118] It is understood that both the client and server mentioned above are nodes. The server mentioned above is a name for a device with data processing capabilities. In some application scenarios or certain network types, the server may also be called a replication node.
[0119] The data processing system shown in Figure 1 can be applied to various network types, for example, one or more of the following network types: long term evolution (LTE) network, fifth generation mobile communication technology (5G), wireless local area network (for example, Wi-Fi), Bluetooth (BT), Zigbee, or vehicular short-range wireless communication network, etc.
[0120] It should be noted that Figure 1 is merely an exemplary architecture diagram and does not limit the number of network elements included in the system shown in Figure 1. Although not shown in Figure 1, Figure 1 may also include other functional entities in addition to the functional entities shown in Figure 1. Furthermore, the methods provided in the embodiments of the present application can be applied to the data processing system shown in Figure 1. Of course, the methods provided in the embodiments of the present application can also be applied to other data processing systems, and the embodiments of the present application are not limited in this regard.
[0121] Based on the system architecture shown in Figure 1, for example, in order to ensure that the client can read the latest data (or latest content) of the object, any client can first query the consensus value of the latest instance of the replication group and the term identifier corresponding to the consensus value before reading or writing a replication group. The client can read or write based on the obtained maximum term identifier, for example, the maximum term identifier is carried in the read command or write command sent.
[0122] In a specific implementation, the client queries the consensus value of the latest instance stored locally on each server in the replication group and the term identifier corresponding to the consensus value, and determines the consensus value of the same latest instance on the majority of servers (that is, more than half of the servers in the replication group) and the term identifier corresponding to the consensus value from the replies of each server in the replication group. The term identifier is the maximum term identifier local to more than half of the servers in the replication group. The client can verify the maximum term identifier obtained when reading / writing the replication group. For example, when the local maximum term identifier of a server in the replication group is not equal to the term identifier carried by the write command sent by the client, the server rejects the write command sent by the client. In addition, the client can know the replication status of the current replication group and the identifier of the target server as the leader or semi-leader corresponding to the replication status through the consensus value of the latest instance obtained, so that the client can select the corresponding read / write method in a targeted manner.
[0123] In some possible embodiments, a server serving as a leader in a replication group may fail, and a new round of elections may be initiated in the replication group through consensus operations to determine a server serving as a semi-leader. This means that compared to the server elected as the leader in the replication group in the previous round, the consensus value of the same latest instance on most servers in the replication group and the term identifier corresponding to the consensus value have changed. Therefore, the client can periodically query the replication group for the consensus value of the latest instance and the term identifier corresponding to the consensus value. In this way, any client can read the latest content of the object after the latest content of the object is written to the replication group.
[0124] Exemplarily, the client's replication group query period can be user-set or system-default. The length of the client's replication group query period is not limited. As an example, the client's replication group query period is less than the maximum duration allowed for the server to retry heartbeat communications. In this way, when the server serving as the leader in the replication group fails and a new server serving as the semi-leader is elected, the client can promptly obtain the replication group's current and latest term identifier.
[0125] In another implementation, the client can also query the replication group for the latest instance's consensus value and the term identifier corresponding to that consensus value without periodically querying the replication group. For example, after the client first queries the replication group to obtain the latest instance's consensus value and the term identifier corresponding to that consensus value, it can query the replication group again only when a trigger condition is met. For example, the trigger condition can be that the client requests to read the target object from the server that is the leader in the replication group and is rejected by the server.
[0126] See Figure 2, which is a flow chart of a data writing method provided in an embodiment of the present application. This method can be applied to the data processing system shown in Figure 1. The data processing system includes at least a client 1 and a replication group, wherein the replication group includes multiple servers. The method includes but is not limited to the following steps:
[0127] S201: Client 1 broadcasts write command 1, which includes a first identifier, the first content of the first object to be written, and the first version number to be written. The first identifier is the local maximum term identifier of more than half of the servers in the replication group. The first identifier indicates that the replication group has reached a consensus on server 1 in the replication group as the first role.
[0128] Here, when applied to a file storage system, the first object can be a basic data unit in the file storage system, which can be represented by a file ID and offset. When applied to object storage, the first object is the object's key; when applied to block storage, the first object can be the block's ID.
[0129] For example, before client 1 broadcasts write command 1, client 1 obtains the first identifier from the replication group. Here, the process of obtaining the first identifier can refer to the description of the corresponding content above, and will not be repeated here.
[0130] Here, the number of servers included in the replication group is 2f+1, where f is a positive integer.
[0131] It can be understood that the first identifier is also the latest local term identifier of more than half of the servers in the replication group, which means that when server 1 initiates a consensus vote to request the election of server 1 as the first role, more than half of the servers in the replication group (including server 1) voted in favor, that is, the replication group reached a consensus on server 1 as the first role through the consensus operation.
[0132] In the embodiment of the present application, the first role is a leader or a semi-leader.
[0133] Exemplarily, the first role is the leader, and server 1 is the server that stores the latest version data of all objects in the replication group; when the first role is a semi-leader, server 1 can be, for example, the server that stores the latest version data of some objects in the replication group.
[0134] Furthermore, if the first role is a leader, the current replication state of the replication group is the first state; or if the first role is a semi-leader, the current replication state of the replication group is the second state. Here, the first state and the second state can refer to the description of the corresponding content in Figure 1.
[0135] Here, the first version number corresponds to the first content of the first object. The first version number identifies the write of the first content of the first object. It can be understood that the objects to be written are different, and the version numbers carried by the write command are different. For the same object, the content written is different, and the version numbers carried by the write command are also different. In this way, when a server in the replication group stores content corresponding to multiple version numbers of the first object, by comparing the first version number with the version number of the local first object on the server, it can be determined whether to write the first content of the first object corresponding to the first version number, and writing expired content of the first object can be avoided. In addition, by carrying the version number in the write command, concurrency conflicts caused by different clients writing to the same object can also be resolved.
[0136] For example, the first version number can be generated by client 1 using a hybrid logical clock (HLC) algorithm. The HLC algorithm consists of a high-order physical clock and a low-order logical clock. For details, see the aforementioned description of the term "hybrid logical clock." For example, when client 1 generates write command 1, the first version number is generated based on the current physical clock and the incremented logical clock.
[0137] S202: When the local maximum term identifier of the server in the replication group is the first identifier and the maximum version number corresponding to the local first object is less than the first version number, the server writes the first content and the first version number of the first object and sends the first information to the client 1, where the first information indicates that the write command 1 is agreed.
[0138] In the embodiment of the present application, since the difference in the first role may affect the writing of the content of the object in the replication group, the following describes the following situations:
[0139] Case 1: The first role is the leader
[0140] In one implementation, the first role is the leader, that is, server 1 acts as the leader, and when server 1 does not fail, the server in the replication group can determine whether to write the first content and the first version number of the first object based on write command 1 by checking the first identifier (that is, determining whether the local maximum term identifier is the first identifier) and comparing the maximum version number of the local first object with the first version number when receiving the write command 1.
[0141] Taking server 1 as an example, the local maximum term identifier of server 1 is the first identifier. If the maximum version number corresponding to the local first object of server 1 is smaller than the first version number, it means that server 1 has not written the content of the first object corresponding to the first version number in history. Therefore, server 1 will write the above-mentioned first content and first version number based on the received write command 1.
[0142] Taking server 2 in the replication group as an example, assuming that server 2 receives write command 1, if the local maximum term identifier of server 2 is the first identifier, it means that server 2 participated in the election of server 1 as the leader and server 2 voted in favor. If the maximum version number corresponding to the first object local to server 2 is less than the first version number, server 2 will write the above-mentioned first content and first version number based on the received write command 1.
[0143] Case 2: The first role is a semi-leader
[0144] In another implementation, when the local maximum term identifier of the server in the replication group is the first identifier and the maximum version number corresponding to the local first object is less than the first version number, the first content and the first version number of the first object are written, including: if the server in the replication group receives a write command 1 after the waiting time since server 1 acts as a semi-leader, the server's local maximum term identifier is the first identifier and the maximum version number corresponding to the local first object is less than the first version number, the first content and the first version number of the first object are written, wherein the waiting time is greater than the maximum time allowed for the server as the leader to retry heartbeat communication.
[0145] Exemplarily, it is assumed that each server in the replication group can broadcast a heartbeat message to other servers once every fixed time interval (for example, time interval T) to exchange their respective local maximum term identifiers. In each heartbeat communication, if a server receives heartbeat messages sent by more than half of the servers in the replication group, the server determines that the heartbeat is successful; if a server does not receive heartbeat messages sent by more than half of the servers in the replication group, the server determines that the heartbeat has failed. Therefore, assuming that a heartbeat failure occurs in a server, if the server is allowed to attempt heartbeats up to M times, the maximum time allowed for the server to retry heartbeat communications is M*T, so the maximum time allowed for the server as the leader to retry heartbeat communications is also M*T, where M is an integer greater than 1.
[0146] It is understood that the server serving as the leader can detect whether it has failed by broadcasting a heartbeat. For example, if the server serving as the leader fails to heartbeat M times in a row, the server serving as the leader can mark itself as failed. In the embodiment of the present application, if a server serving as the leader is marked as failed, the server will reject read commands from any client until the server's heartbeat succeeds, and then resume the read service.
[0147] Here, the first role is a semi-leader, which means that server 1 in the replication group detects that the server that serves as the leader has failed, and server 1 initiates a new election so that the replication group reaches a consensus on server 1 as the second role, and the second role is a semi-leader. The replication state of the replication group switches from the above-mentioned first state corresponding to the original leader to the above-mentioned second state corresponding to the current semi-leader. The achievement of this consensus also generates a new term identifier (i.e., the first identifier). For the client, the client may not be able to query the replication group in time to obtain the first identifier, so the client cannot perceive that the replication group has currently switched from leader to semi-leader. In order to avoid the client using an expired term identifier (i.e., the server that originally served as the leader in the replication group The term identifier corresponding to the server) reads the expired content of an object. Through the above-mentioned condition "If the server in the replication group receives the write command 1 after the waiting time since server 1 became the semi-leader", the server in the replication group that reaches a consensus on server 1 as the semi-leader is restricted to the waiting time from the local maximum term identifier to the first identifier before starting the write service. The waiting time is enough for the server as the leader to discover that it is invalid. Therefore, when new data of the target object is written to the replication group, when the client uses the expired term identifier to read the target object from the replication group, the server that has marked itself as invalid will reject the client's reading of the target object, thereby avoiding the client reading the expired content of the target object.
[0148] In some possible embodiments, no matter the first role is a leader or a semi-leader, the server may reject the write command 1.
[0149] In one specific implementation, the first role is a leader, or the first role is a semi-leader and the server in the replication group receives a write command 1 after a waiting period from when server 1 became a semi-leader. In this case, if the local maximum term identifier of the server is the first identifier but the local maximum version number of the server is greater than or equal to the first version number, the server sends a second message to client 1, and the second message indicates that the write command 1 is rejected. The second message includes the maximum version number corresponding to the first object local to the server.
[0150] In one specific implementation, the first role is the leader, or the first role is the semi-leader and the server in the replication group receives the write command 1 after the waiting time since the server 1 became the semi-leader. In this case, if the local maximum term identifier of the server is not the first identifier, it means that the server did not participate in the election of server 1 as the leader initiated by server 1, and it also means that the server does not belong to the server in the replication group that has reached a consensus on server 1 as the leader. In this case, the server sends a third message to the client 1, and the third message indicates the rejection of the write command 1. The third message includes the current local maximum term identifier of the server.
[0151] In another specific implementation, the first role is a semi-leader, but the time interval between the time when server 1 in the replication group becomes the semi-leader and the time when server 1 receives the write command 1 is less than the maximum time allowed for the server as the leader to retry heartbeat communication, then the server replies to client 1 to reject the write command 1.
[0152] In an embodiment of the present application, after the server writes the first content and the first version of the first object, it also records that the status of the first content of the first object corresponding to the first version number is "preparing." For example, server 1 records the status of the first content of the first object corresponding to the first version number as "preparing," meaning that the first content of the current first object can be rolled back or committed.
[0153] In some possible embodiments, the server may receive multiple write commands for the same object from different clients. Assume that the term identifiers carried by these multiple write commands are all the first identifiers, but the contents of the first objects to be written are different and the version numbers to be written are different. In this case, if the server first receives the write command corresponding to the largest version number among these multiple write commands and executes the write based on this write command, the server will reject other write commands among these multiple write commands. In this way, the concurrency conflict caused by different clients writing to the same object is resolved through the version number.
[0154] S203: When client 1 receives the first information sent by more than half of the servers in the replication group including server 1, client 1 broadcasts a commit command, where the commit command indicates that the status of updating the first content corresponding to the first version number is submitted.
[0155] In an embodiment of the present application, when client 1 receives a reply from the majority of servers (including server 1) in the replication group accepting the write, client 1 determines that the write indicated by write command 1 can be successful, so client 1 broadcasts a commit command. Here, client 1 can end the write process without waiting for the server in the replication group to reply to the commit command. Accordingly, when the server in the replication group receives the commit command, the server in the replication group that has written the first version number and the first content of the first object updates the status of the first content of the first object corresponding to the first version to submitted.
[0156] For example, if the server 1 records the status of the first content of the first object corresponding to the first version number as "submitted", it means that the first content of the current first object is readable.
[0157] In some possible embodiments, client 1 does not receive the first message sent by more than half of the servers in the replication group, including server 1, but client 1 receives the above-mentioned second message sent by more than half of the servers in the replication group. The second message can refer to the description of "second message" in S202 above. In this case, client 1 determines that the write indicated by write command 1 is unsuccessful, and client 1 broadcasts write command 2. Write command 2 includes a first identifier, the first content of the first object to be written, and a second version number to be written, where the second version number is greater than the maximum version number corresponding to the first object in the second message received by client 1. It can be understood that the second version number is greater than the first version number.
[0158] It can be seen that when implementing the embodiments of the present application, no matter whether the replication group is currently in the replication state corresponding to the leader or the replication state corresponding to the semi-leader, the write process uses the "majority write" method to write the content of the object to the replication group, that is, the client broadcasts the write command, and each server in the replication group can directly respond to the client after receiving the write command after locally judging whether the write conditions are met. When the client receives the reply from the majority of the servers in the replication group that accept the write, the client determines that the write can be completed. Compared with the use of the master-slave replication protocol, the existing paxos protocol or the raft protocol for write operations, this method can reduce the delay of the write operation. In addition, compared with the master-slave replication protocol, the present application also has a certain degree of fault tolerance. For example, when less than half of the servers in the replication group fail, it can continue to provide services to the client. Compared with the existing paxos protocol or raft protocol, the present application does not need to determine the globally unique write order before writing data, so it solves the write amplification problem caused by object writing.
[0159] Based on the embodiment of Figure 2, it is assumed that the first role is the leader, that is, the server 1 in the replication group serves as the leader. When there is no failure in the server 1, the reading process adopts the "leader read" mode. For details, please refer to the description of the embodiment of Figure 3.
[0160] See Figure 3, which is a flow chart of a data reading method provided in an embodiment of the present application. This method can be applied to the data processing system shown in Figure 1. Compared to the data processing system described in the embodiment of Figure 2, this data processing system also includes client 2, where client 2 is different from client 1, and server 1 in the replication group serves as the leader. The method includes but is not limited to the following steps:
[0161] S301: Client 2 sends a read command to server 1 in the replication group. The read command includes a target term identifier and is used to request to read a first object.
[0162] Here, the target term identifier is obtained by client 2 from the replication group query before client 2 sends the read command. The process of obtaining the target term identifier can refer to the above description of obtaining the term identifier, which will not be repeated here.
[0163] S302: The server 1 compares the target term identifier with the first identifier to see if they are the same.
[0164] Exemplarily, when the target term identifier is the same as the first identifier, the server 1 executes S303 ; when the target term identifier is different from the first identifier, the server 1 executes S304 .
[0165] Here, the target term identifier and the first identifier may differ. This can mean that the target term identifier is less than the first identifier, indicating that the target term identifier has expired. Since the client acquired the target term identifier, the replication group has reached a consensus on Server 1 as the leader through a new election. The maximum term identifier currently localized by the majority of servers in the replication group is the first identifier. In this case, the client requeries the replication group to obtain the first identifier and sends a read command to Server 1 based on the first identifier.
[0166] S303: Server 1 sends the latest content of the first object to client 2.
[0167] Here, when the target term identifier is equal to the first identifier, the server 1 sends the latest content of the first object to the client 2 .
[0168] Exemplarily, if the first content and the first version number of the above-mentioned first object are successfully written into the replication group based on the embodiment of Figure 2, and before the client 1 reads the first object from the server 1, the replication group does not write other contents of the first object, then the latest content of the first object sent by the server 1 to the client 2 is the first content.
[0169] In this case, server 1 sends the first content of the first object to client 2, including: after server 1 updates the status of the first content of the first object corresponding to the first version number to submitted based on the submit command, server 1 sends the first content of the first object to client 2. Here, the submit command can be referred to the description of the corresponding content in S203 in the embodiment of Figure 2 above, and will not be repeated here.
[0170] Exemplarily, if the first content and the first version number of the above-mentioned first object are successfully written into the replication group based on the embodiment of Figure 2, and before the client 1 reads the first object from the server 1, for the first object, only other clients successfully write the third content of the first object and the corresponding version number into the replication group, then the latest content of the first object sent by the server 1 to the client 2 is the third content.
[0171] S304: The server 1 responds to the client 2 with a rejection of the read command.
[0172] In one implementation, when the target term identifier is smaller than the first identifier, the server 1 replies to the client 2 with a rejection of the read command.
[0173] In some possible embodiments, the server 1 may also reply to the client 2 with a read rejection command when any of the following conditions are met:
[0174] Condition 1: Before server 1 receives the read command sent by client 2, server 1 fails to retry heartbeat communication within the preset time period;
[0175] Condition 2: Before server 1 receives the read command sent by client 2, server 1 determines that the first identifier is invalid based on the heartbeat information sent by server 2 in the replication group. The heartbeat information sent by server 2 includes the local maximum term identifier of server 2 and the local maximum term identifier of server 2 is greater than the first identifier.
[0176] For condition 1, the preset duration is, for example, the maximum duration allowed for the server to retry heartbeat communication. That is, when server 1 detects that it has failed to retry heartbeat communication, server 1 refuses to provide read service to the client. In this case, when server 1 detects that it has failed to retry heartbeat communication, server 1 can locally mark itself as invalid. It can be understood that if server 1 fails to retry heartbeat communication within the preset duration, it means that the failure of server 1 is very likely to be unrecoverable. Server 1 promptly marks itself as invalid and refuses to provide read service to the outside world, which can prevent the client from reading expired data.
[0177] For condition 2, it is possible that after server 1 fails, it attempts heartbeat communication successfully within a preset time period, but during this period, a server in the replication group (such as server 2) detects a failure of server 1 (for example, the heartbeat information of server 1 is not received) and initiates a new election for server 2 as the semi-leader. The replication group reaches a consensus on server 2 as the semi-leader and generates a new term identifier. Since server 1 successfully attempts heartbeat communication, server 1 can receive the heartbeat information sent by server 2. Based on the heartbeat information, server 1 determines that the local first identifier has failed. In this case, server 1 can locally mark itself as invalid. In this way, server 1 promptly marks itself as invalid and refuses to provide read services to the outside world, which can prevent the client from reading expired data.
[0178] Under the above condition 2, in the above S301, when the server 1 compares whether the target term identifier is the same as the first identifier, it may happen that the target term identifier is greater than the first identifier, or the target term identifier is equal to the first identifier, or the target term identifier is less than the first identifier. In this case, the server 1 replies to the client 2 to reject the read command.
[0179] In some possible embodiments, the first role is the leader. After server 1 in the replication group fails, the local maximum term identifier of more than half of the servers in the replication group is the second identifier. The second identifier is greater than the first identifier. The second identifier indicates that the replication group has reached a consensus on server 2 as the semi-leader. In this scenario:
[0180] In one implementation, if, before reading the first object, client 2 obtains the second identifier in a timely manner through the above-mentioned periodic query of the replication group mechanism and knows that the current replication status of the replication group is the replication status corresponding to the semi-leader, then the reading process will switch from the "leader read" mode to the "majority read" mode. The process of client 2 using the "majority read" mode to read the first object can be referred to the description of the embodiment of Figure 4 below, and will not be repeated here.
[0181] In another implementation, if client 2 fails to obtain the second identifier from the replication group in time before reading the first object, that is, client 2 does not perceive that the replication status of the replication group has changed, client 2 still sends a read command 1 to server 1. Read command 1 is used to request to read the first object. The target term identifier carried by read command 1 is the first identifier. Accordingly, when client 2 receives a reply from server 1 rejecting read command 1, client 2 is triggered to obtain the second identifier from the replication group, so that client 2 can read the first object in "majority read" mode based on the second identifier.
[0182] It can be seen that when the server that serves as the leader in the replication group is fault-free, since this server is the server that stores the latest version data of all objects in the replication group, the client uses the "leader read" mode to directly request the leader to read the target object, and can obtain the latest content of the target object.
[0183] Based on the embodiment of Figure 2, it is assumed that the first role is the leader, that is, server 1 in the replication group serves as the leader. In the event of a failure of server 1, the local maximum term identifier of more than half of the servers in the replication group is the second identifier, and the second identifier is greater than the first identifier. The second identifier indicates that the replication group has reached a consensus on server 2 as the semi-leader. In this case, the read process adopts the "majority read" mode. For details, please refer to the description of the embodiment of Figure 4.
[0184] See Figure 4, which is a flow chart of another data reading method provided in an embodiment of the present application. This method can be applied to the data processing system shown in Figure 1. The data processing system includes client 2, and server 2 in the replication group serves as a semi-leader. The method includes but is not limited to the following steps:
[0185] S401: Client 2 broadcasts a read command, where the read command is used to request to read a first object and carries a second identifier.
[0186] For example, before client 2 broadcasts a read command, client 2 queries the replication group to obtain the second identifier and determines that server 2 is the semi-leader in the replication group. Therefore, client 2 knows that the replication group's current replication state corresponds to the semi-leader state, so client 2 uses the "majority read" mode to read the object.
[0187] S402: When the local maximum term identifier of the server in the replication group is the second identifier, the server sends a read result to client 2. The read result includes the target version number and the second content of the first object corresponding to the target version number. The target version number is the maximum version number corresponding to the local first object of the server.
[0188] It can be understood that the local maximum term identifier of the server in the replication group is the second identifier, which means that the server participated in the election of server 2 as the semi-leader and voted in favor. The server belongs to the server in the replication group that has reached a consensus on server 2 as the semi-leader.
[0189] In some possible embodiments, when the local maximum term identifier of the server in the replication group is not the second identifier, the server replies to the client 2 to reject the read command.
[0190] See Figure 5, which is a schematic diagram of an application scenario provided by an embodiment of the present application. In Figure 5, the number of servers in the replication group is 5, namely server 1, server 2, server 3, server 4 and server 5, among which server 2 serves as a semi-leader. The rectangular boxes listed behind each server in Figure 5 show the local object writing status of the server. Taking the content "2_4, x←5" of the last rectangular box of server 2 as an example, "2_4, x←5" means that the content "5" of object x corresponding to version number "4" is written when the term identifier is "2".
[0191] As can be seen from Figure 5, the local maximum term identifier of server 1 is 1, and the local maximum term identifiers of server 2, server 3, server 4, and server 5 are all 2. It can also be seen from Figure 5 that during the period when the term identifier was 1, two writes occurred, namely, writing "x=3" and "y=2" respectively; during the period when the term identifier was 2, two writes occurred, namely, writing "x=4" and "y=5" respectively. In addition, it can be seen from Figure 5 that the latest content of object x currently stored by server 1 is 3 and the latest content of object y is 2, the latest content of object x currently stored locally by server 2 and server 3 is 5 and the latest content of object y is 2, server 4 currently only stores object x and the latest content of object x is 5, and server 5 currently stores the latest content of object x as 4 and the latest content of object y as 2. In addition, it can be seen from Figure 5 that each time the content of an object is written, more than half of the servers in the replication group successfully write the content of the object locally.
[0192] In Figure 5, client 2 first broadcasts a read command requesting to read object x (the first object mentioned above). The second identifier carried in the read command is 2. After receiving the read command, the servers in the replication group compare the local maximum term identifier with the second identifier. The responses from each server in the replication group are as follows:
[0193] 1) The local maximum term identifier of server 1 is 1, which is smaller than the second identifier, so server 1 replies to client 2 rejecting the read command;
[0194] 2) Server 2's local maximum term identifier is 2 and the local maximum version number of object x is 4. Therefore, Server 2 sends a read result to Client 2, which carries a target version number of 4 and the content of x corresponding to the target version number of 5.
[0195] 3) Server 3's local maximum term identifier is 2, and the local maximum version number of object x is 4. Therefore, server 3 sends a read result to client 2, which carries a target version number of 4 and the content of x corresponding to the target version number is 5.
[0196] 4) Server 4's local maximum term identifier is 2, and the local maximum version number of object x is 4. Therefore, server 4 sends a read result to client 2, which carries a target version number of 4 and the content of x corresponding to the target version number of 5.
[0197] 5) The local maximum term identifier of server 5 is 2 and the maximum version number corresponding to the local object x is 3, so server 3 sends a read result to client 2, which carries a target version number of 3 and the content of x corresponding to the target version number is 4.
[0198] Here, FIG5 is only used as an example. FIG5 is only for clearly showing the writing status of the local objects on the server side in the replication group, and does not limit the storage form and storage content of the local data on the server side to only those shown in FIG5 .
[0199] S403: When client 2 receives read results sent by more than half of the servers in the replication group, it obtains the second content of the first object corresponding to the largest version number in the received read results.
[0200] Based on the description in Figure 5 in S402, it can be seen that client 2 receives read results sent by more than half of the servers in the replication group, and the largest target version number in the read results received by client 2 is 4 and the content of object x corresponding to version number "4" is "5". Therefore, the largest version number obtained by client 2 from the received read results is 4 and the content of object x corresponding to this version number is "5".
[0201] It can be seen from the embodiment of Figure 2 above that no matter whether the replication group is currently in the replication state corresponding to the leader or the replication state corresponding to the semi-leader, the write process uses the "majority write" method to write the content of the object to the replication group. Therefore, when the replication group reaches a consensus on a certain server as the semi-leader, the client reads the target object from the replication group through the "majority read" method. Since the server that participates in writing the latest content of the target object and the server that feeds back the read result must have an intersection (that is, there is at least one identical server), it can ensure that the client will definitely read the latest content of the target object.
[0202] The following is a specific example to illustrate the process of electing a semi-leader in a replication group.
[0203] In one implementation, server 1 in the replication group acts as the leader and the corresponding term identifier is identifier 1. The servers in the replication group sense whether other servers in the replication group are online through the above-mentioned heartbeat detection mechanism. Assuming that server 2 in the replication group does not receive the heartbeat information sent by server 1, server 2 believes that server 1 has failed, so server 2 initiates an election request. The election request is used to request the election of server 2 as a semi-leader. If more than half of the servers in the replication group agree to the election request, the replication group reaches a consensus on server 2 as a semi-leader and the corresponding term identifier is identifier 2, where identifier 2 is greater than identifier 1. Here, server 2 can be any currently online server in the replication group except server 1, where "currently online" means that the heartbeat communication of server 2 is normal.
[0204] See Figure 6A, which is a schematic diagram of the process of electing a semi-leader in a replication group running the paxos protocol provided by an embodiment of the present application. In Figure 6A, it is assumed that server 2 has not detected the heartbeat information of server 1, and the largest slot ID that is not empty (i.e., stores a consensus value) in the log maintained locally by server 2 is 1, i.e., the term identifier is 1, and the consensus value stored in the slot corresponding to the term identifier 1 in the log records the identifier of server 1 and the replication state is the first state corresponding to the leader, so server 2 creates a new slot ID "2" in the local log, in stage one: broadcast a proposal, which carries the term identifier (i.e., slot ID "2") and the proposal number of the proposal. It can be seen from Figure 6A that more than half of the servers in the replication group (i.e., server 1, server 3, server 4, and server 5) agree with the proposal and none of them agree. Return the consensus value; enter stage two: Server 2 broadcasts a consensus value A, which records the identifier of Server 2 and the second state corresponding to the semi-leader. As shown in Figure 6A, more than half of the servers in the replication group agree to accept the consensus value A; enter stage three: Server 2 broadcasts the proposal submission information and stores the consensus value A in the slot ID "2" of the local log. After receiving the proposal submission information, Server 3, Server 4, and Server 5 all store the consensus value A in the slot ID "2" of the local log and reply to Server 2. It can be seen that the local maximum term identifier of Server 1, Server 3, Server 4, and Server 5 in the replication group is "2".
[0205] In some possible embodiments, in the above-mentioned stage one, if a server responds to the proposal of server 2 and returns a consensus value B, server 2 stores the consensus value B in the slot ID "2" of the local log, creates a new slot ID "3" in the log, and rebroadcasts a new proposal. The term identifier carried by the proposal is slot ID "3" and the proposal also carries a new proposal number. If more than half of the servers in the replication group agree with the proposal and none of them return a consensus value, server 2 can broadcast the consensus value A.
[0206] Here, the description of FIG6A is only an exemplary brief description, not a complete description of the Paxos protocol, and should not limit the description of the process of electing a semi-leader by running the Paxos protocol in the replication group.
[0207] For example, under the Paxos protocol, the starting time of the waiting period in the aforementioned "after the waiting period from when the server becomes the semi-leader" can be the end time of stage 3 in Figure 6A. In this way, there is no need to modify the existing consensus protocol (e.g., the Paxos protocol). By simply constraining the servers in the replication group that have reached consensus on Server 2 as the semi-leader to provide write services after the waiting period from when Server 2 becomes the semi-leader, clients can be prevented from reading expired content of the target object using expired term identifiers.
[0208] In some possible embodiments, the existing paxos protocol can also be modified, that is, the waiting time is set in the election process. Referring to Figure 6B, Figure 6B is a schematic diagram of the process of electing a semi-leader by running the paxos protocol in another replication group provided by an embodiment of the present application. The difference compared to Figure 6A is that after the end of stage two, the server 2 enters stage three after a waiting time. In this way, when the above-mentioned term identifier "2" takes effect after the end of stage three, the server 1 has marked itself as invalid. When other clients use the new term identifier "2" to write new content of the target object to the replication group, even if the client fails to obtain the term identifier "2" from the replication group in time, the client cannot read data from the server 1 using the expired term identifier "1", thereby avoiding the client from reading expired data.
[0209] It can be understood that in the scenario shown in FIG6B, after completing the third stage shown in FIG6B, the server in the replication group that has reached a consensus on server 2 as the semi-leader in this election can provide write services to the outside world, that is, when the server receives a write command from the client for the first object, if the term identifier carried by the write command is the local maximum term identifier of the server and the version number carried by the write command is greater than the maximum version number corresponding to the local first object of the server, then the server writes the content of the first object carried by the write command and the version number carried by the write command. In this case, for the timing when the server in the replication group can accept writes, there is no need to limit the waiting time since server 2 in the replication group serves as the semi-leader before the write command received can be processed.
[0210] In conjunction with the description of Figures 5 and 6A above, when the replication group reaches a consensus on server 2 as the semi-leader, the replication group completes the switch from leader to semi-leader. At this time, the logs maintained locally by each server in the replication group can be shown in Figure 7. Figure 7 is a schematic diagram of a log for storing consensus values locally by each server in a replication group provided by an embodiment of the present application. Since in Figure 5, server 1 only wrote the content of the object locally under term identifier "1", while other servers in the replication group except server 1 wrote the content of the object under term identifier "1" and term identifier "2", so in Figure 7, the local log of server 1 only has the log entry of slot ID "1" and the slot ID "1" stores the consensus value, which records the identification of server 1 and the first state corresponding to replication status 1 being the leader; and the local logs of servers 2, server 3, server 4 and server 5 in the replication group all contain two log entries, slot ID "1" and slot ID "2", among which the consensus value stored in slot ID "1" records the identification of server 1 and the first state corresponding to replication status 1 being the leader, and the consensus value stored in slot ID "2" records the identification of server 2 and the second state corresponding to replication status 2 being the semi-leader.
[0211] In some possible embodiments, after server 2 acts as a semi-leader, server 2 can also communicate with other clients in the replication group to interactively check which data objects it lacks and the version numbers corresponding to the latest content of the data objects. In this way, server 2 can obtain the latest content of the currently missing data objects and the version numbers corresponding to the latest content of the data objects from other servers in the replication group. When server 2 determines that it is the server that currently stores the latest version of data of all objects, server 2 can broadcast an election request, which is used to request the election of server 2 as the leader. After the replication group reaches a consensus on server 2 as the leader through the election, the replication status of the replication group is switched from the second status corresponding to the semi-leader to the first status corresponding to the semi-leader.
[0212] For example, as can be seen from Figure 5, all objects currently stored in the replication group include object x and object y. Server 2 determines that it currently stores the latest version of data of all objects through interaction with other online servers in the replication group. Therefore, server 2 can initiate a new election request to request that server 2 be elected as the leader.
[0213] In some possible embodiments, in addition to the situation where the server that serves as the leader in the above-mentioned replication group goes offline due to a failure, causing the replication group to run a consensus protocol, when a server is added to the replication group, causing the number of servers in the replication group to change, the replication group also needs to run a consensus protocol so that each server has a consistent view of the current server in the replication group, thereby ensuring that the provided services can be executed correctly and consistently.
[0214] 8 is a schematic diagram of a data processing apparatus according to an embodiment of the present invention. The data processing apparatus 30 includes a sending unit 310 and a receiving unit 312. The data processing apparatus 30 can be implemented in hardware, software, or a combination of hardware and software.
[0215] As an example, the device 30 for data processing may be any of the above-mentioned clients or included in the client.
[0216] Exemplarily, when executing the write process, the sending unit 310 is used to broadcast a write command, which includes a first identifier, the first content of the first object to be written, and the first version number to be written. The first identifier is the local maximum term identifier of more than half of the servers in the replication group, and the first identifier indicates that the replication group has reached a consensus on the first server in the replication group as the first role; the sending unit 310 is also used to broadcast a submission command when the receiving unit 312 receives the first information sent by more than half of the servers in the replication group, wherein the first information indicates agreement to the write command, and the submission command indicates that the status of the first content corresponding to the first version number is updated to be submitted.
[0217] The apparatus 30 for data processing may be used to implement the method on the client 1 side described in the embodiment of Figure 2. In the embodiment of Figure 2, the sending unit 310 may be used to execute S201 and S203, and the receiving unit 312 may be used to execute S203.
[0218] In some possible embodiments, the data processing apparatus 30 may also perform a read process. In this case, the data processing apparatus 30 may also be used to implement the method on the client side 2 described in the embodiment of FIG. 3 or the method on the client side 2 described in the embodiment of FIG. 4 . For example, in the embodiment of FIG. 3 , the sending unit 310 may be used to perform S301 , and the receiving unit 312 may be used to perform S303 and S304 . For another example, in the embodiment of FIG. 4 , the sending unit 310 may be used to perform S401 , and the receiving unit 312 may be used to perform S402 and S403 .
[0219] 9 is a schematic diagram of the structure of another apparatus for data processing provided in an embodiment of the present application. The apparatus 40 for data processing includes a receiving unit 410, a processing unit 412, and a sending unit 414. The apparatus 40 can be implemented in hardware, software, or a combination of hardware and software.
[0220] Exemplarily, when executing the write process, the receiving unit 410 is used to receive a write command, which is transmitted in the form of a broadcast. The write command includes a first identifier, the first content of the first object to be written, and the first version number to be written. The first identifier indicates that the replication group has reached a consensus on the first server in the replication group as the first role; the processing unit 412 is used to write the first content and the first version number when it is determined that the local maximum term identifier is the first identifier and the local maximum version number corresponding to the first object is less than the first version number, and send the first information through the sending unit 414. The first information indicates agreement to the write command.
[0221] The apparatus 40 for data processing can be used to implement the server-side method described in the embodiment of Figure 2. In the embodiment of Figure 2, the receiving unit 410 can be used to execute S201 and S203, the processing unit 412 can be used to execute S202, and the sending unit 414 can be used to execute S202.
[0222] In some possible embodiments, the data processing device 40 can also perform a read process. In this case, the data processing device 40 can also be used to implement the method on the server 1 side described in the embodiment of Figure 3 or the method on the server side described in the embodiment of Figure 4. For example, in the embodiment of Figure 3, the receiving unit 410 can be used to perform S301, the processing unit 412 can be used to perform S302, and the sending unit 414 can be used to perform S303 and S304. For another example, in the embodiment of Figure 4, the receiving unit 410 can be used to perform S401, and the processing unit 412 and the sending unit 414 can be used to perform S402.
[0223] One or more of the various units in the embodiments shown in FIG8 or FIG9 above may be implemented in software, hardware, firmware, or a combination thereof. The software or firmware includes, but is not limited to, computer program instructions or codes, and may be executed by a hardware processor. The hardware includes, but is not limited to, various integrated circuits, such as a central processing unit (CPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC).
[0224] It should be understood that the division of the various units in the above devices (e.g., device 30 for data processing or device 40 for data processing) is merely a division of logical functions. In actual implementation, they may be fully or partially integrated into a single physical entity, or they may be physically separated. In addition, the units in the device may be implemented in the form of a processor calling software; for example, the device includes a processor, the processor is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the various units of the device, wherein the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory within the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits, and the functions of some or all of the units can be realized by designing the hardware circuits. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC), which realizes the functions of some or all of the above units by designing the logical relationship of the components in the circuit. For another example, in another implementation, the hardware circuit can be implemented by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by configuring the configuration file, thereby realizing the functions of some or all of the above units. All units of the above devices can be implemented in the form of software called by the processor, or in the form of hardware circuits, or in part by software called by the processor, and the rest by hardware circuits.
[0225] In an embodiment of the present application, a processor is a circuit with a signal processing capability. In one implementation, the processor can be a circuit with instruction reading and execution capability, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationship of a hardware circuit. The logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by a processor as an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the processor loads a configuration document to implement the process of hardware circuit configuration, which can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc.
[0226] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0227] In addition, the various units in the above devices can be fully or partially integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a system-on-a-chip (SOC). The SOC may include at least one processor for implementing any of the above methods or implementing the functions of the various units of the device. The type of the at least one processor can be different, for example, including a CPU and FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0228] Referring to Figure 10 , Figure 10 is a schematic diagram of the structure of a communication device provided in an embodiment of the present application. As shown in Figure 10 , communication device 50 includes: a processor 501, a communication interface 502, a memory 503, and a bus 504. Processor 501, memory 503, and communication interface 502 communicate with each other via bus 504. It should be understood that this application does not limit the number of processors and memories in communication device 50.
[0229] In one implementation, the communication device 50 may be the aforementioned server, which may be a network-side device with data processing capabilities. The network-side device may be, for example, a server deployed on the network side, or a component in the server (the component may be, for example, a chip, an integrated circuit, etc.). The network-side device may be deployed in a cloud environment or an edge environment, and is not specifically limited here.
[0230] In another implementation, the communication device 50 may be the aforementioned client. The client may be a terminal device, which may be, for example, a user device (a mobile phone, computer, tablet computer, PDA, desktop computer, headset, speaker, wearable device, vehicle-mounted device, virtual reality device, augmented reality device, etc.), a smart home device (such as a television, a sweeping robot, a smart desk lamp, a sound system, a smart lighting system, an appliance control system, home background music, a home theater system, an intercom system, a video surveillance system, etc.), an intelligent transportation device (such as a car, a ship, a drone, a train, a van, a truck, etc.), an intelligent manufacturing device (such as a robot, industrial equipment, intelligent logistics, an intelligent factory, etc.), or a component within the terminal device (such as a chip or integrated circuit).
[0231] Bus 504 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG10 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 504 may include a path for transmitting information between various components of communication device 50 (e.g., memory 503, processor 501, and communication interface 502).
[0232] The processor 501 can refer to the relevant description of the processor in the above embodiment, which will not be repeated here.
[0233] Memory 503 is used to provide storage space for storing data such as the operating system and computer programs. Memory 503 can be one or a combination of random access memory (RAM), erasable programmable read-only memory (EPROM), read-only memory (ROM), or compact disc read-only memory (CD-ROM). Memory 503 can exist independently or be integrated into processor 501.
[0234] The communication interface 502 can be used to provide information input or output for the processor 501. Alternatively, the communication interface 502 can be used to receive data transmitted externally and / or transmit data externally. It can be a wired link interface such as an Ethernet cable, or a wireless link interface (such as Wi-Fi, Bluetooth, general wireless transmission, etc.). Alternatively, the communication interface 502 can also include a transmitter (such as a radio frequency transmitter, antenna, etc.) or a receiver coupled to the interface.
[0235] The processor 501 in the communication device 50 is used to read the computer program stored in the memory 503 to execute the aforementioned method, such as the method described in FIG. 2 , FIG. 3 or FIG. 4 .
[0236] In one possible design, the communication device 50 may be one or more modules in an execution entity (e.g., client 1) that executes the method shown in FIG. 2 , and the processor 501 may be configured to read one or more computer programs stored in a memory to perform the following operations:
[0237] A write command is broadcasted by the sending unit 310, where the write command includes a first identifier, first content of a first object to be written, and a first version number to be written. The first identifier is a local maximum term identifier of more than half of the servers in the replication group, and the first identifier indicates that the replication group has reached a consensus on the first server in the replication group as the first role.
[0238] When the receiving unit 312 receives the first information sent by more than half of the servers in the replication group, the submit command is broadcast through the sending unit 310, wherein the first information indicates that the write command is agreed, and the submit command indicates that the status of the first content corresponding to the first version number is updated as submitted.
[0239] In one possible design, the communication device 50 may be one or more modules in an execution entity (e.g., a server) that executes the method shown in FIG. 2 , and the processor 501 may be configured to read one or more computer programs stored in a memory to perform the following operations:
[0240] A write command is received by the receiving unit 410, where the write command is transmitted in a broadcast format. The write command includes a first identifier, first content of a first object to be written, and a first version number to be written. The first identifier indicates that the replication group has reached a consensus on the first server in the replication group as the first role.
[0241] When it is determined that the local maximum term identifier is the first identifier and the local maximum version number corresponding to the first object is less than the first version number, the first content and the first version number are written, and the first information is sent through the sending unit 414, and the first information indicates that the write command is approved.
[0242] In the embodiments described above, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant descriptions of other embodiments. In addition, in the various embodiments of this application, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between the various embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0243] It should be noted that, those skilled in the art can see that all or part of the steps in the various methods of the above embodiments can be completed by a program to instruct relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0244] The technical solution of the present application may essentially or contribute to the part or all or part of the technical solution in the form of a software product. The computer program product is stored in a storage medium and includes a number of instructions for enabling a device (which may be a personal computer, a server, or a network device, a robot, a single-chip microcomputer, a chip, a robot, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
Claims
1. A data processing method, characterized in that: The method is applied to a data processing system, the data processing system comprising a first client and a replication group, wherein more than half of the servers in the replication group have a local maximum term identifier as a first identifier, and the first identifier indicates that the replication group has reached a consensus that the first server in the replication group serves as a first role. The method comprises: The first client broadcasts a first write command, where the first write command includes the first identifier, the first content of the first object to be written, and the first version number to be written; When the local maximum term identifier of the server in the replication group is the first identifier and the local maximum version number corresponding to the first object is less than the first version number, the server writes the first content and the first version number locally and sends first information to the first client, wherein the first information indicates that the first write command is agreed; When the first client receives the first information sent by more than half of the servers in the replication group including the first server, the first client broadcasts a submission command, where the submission command indicates that the status of updating the first content corresponding to the first version number is submitted.
2. The method according to claim 1, characterized in that The number of servers included in the replication group is 2f+1, where f is a positive integer.
3. The method according to claim 1 or 2, characterized in that: Before the first client broadcasts the first write command, the method further includes: The first client obtains the first identifier from the replication group.
4. The method according to any one of claims 1 to 3, characterized in that: The first role is a leader, the data processing system further includes a second client, and the method further includes: The second client sends a read command to the first server, where the read command is used to request to read the first object, and the read command includes the first identifier; When the local maximum term identifier of the first server is the first identifier, the first server sends the latest content of the first object to the second client.
5. The method according to any one of claims 1 to 3, characterized in that: The first role is a leader, the data processing system further includes a second client, and the method further includes: The second client sends a read command to the first server, where the read command includes a second identifier different from the first identifier; The first server responds by rejecting the read command and sends the first identifier to the second client.
6. The method according to any one of claims 1 to 5, characterized in that: The first role is a leader, the data processing system further includes a third client, and the method further includes: The third client sends a read command to the first server; When any of the following conditions is met, the first server responds to reject the read command sent by the third client: Before receiving the read command sent by the third client, the first server fails to retry the heartbeat communication within a preset time period; or Before receiving the read command sent by the third client, the first server determines that the first identifier is invalid based on the heartbeat information sent by the second server in the replication group, and the heartbeat information includes the local maximum term identifier of the second server and the local maximum term identifier of the second server is greater than the first identifier.
7. The method according to any one of claims 1 to 6, characterized in that: The first role is a leader, less than half of the servers in the replication group are faulty, and the faulty servers do not include the first server.
8. The method according to any one of claims 1 to 6, characterized in that: The first role is a leader, and in the case of a failure of the first server, the local maximum term identifier of more than half of the servers in the replication group is a third identifier, the third identifier is greater than the first identifier, and the third identifier indicates that the replication group has reached a consensus that the third server in the replication group serves as the second role, and the second role is a semi-leader, and the method further includes: Within the waiting time from when the third server takes on the second role, if the server whose local maximum term identifier in the replication group is the third identifier receives a second write command, the server replies with a rejection of the second write command; The waiting time is greater than the maximum time allowed for the first server to retry heartbeat communication.
9. The method according to claim 8, characterized in that The method further comprises: When the third server determines that it is the server that currently stores the latest version data of all objects in the replication group, the third server broadcasts an election request, where the election request is used to request that the third server be elected as a leader.
10. The method according to any one of claims 1 to 3, characterized in that: When the first role is a semi-leader, the local maximum term identifier of the server in the replication group is the first identifier and the local maximum version number corresponding to the first object is smaller than the first version number, writing the first content and the first version number includes: If the server in the replication group receives the first write command after the waiting time since the first server took on the first role, and the server's local maximum term identifier is the first identifier and the local maximum version number corresponding to the first object is less than the first version number, the first content and the first version number are written; The waiting time is greater than the maximum time allowed for the server as a leader to retry the heartbeat communication.
11. The method according to claim 10, characterized in that The data processing system further includes a second client, and the method further includes: The second client broadcasts a read command, where the read command is used to request to read the first object, and the read command includes the first identifier; When the local maximum term identifier of the server in the replication group is the first identifier, sending a read result to the second client, the read result including a second version number and a second content of the first object corresponding to the second version number, the second version number being the maximum version number corresponding to the local first object of the server; When the second client receives the read results sent by more than half of the servers in the replication group, it obtains the second content of the first object corresponding to the largest version number in the received read results.
12. A data processing method, characterized in that: The method is applied to a client, and the method comprises: Broadcasting a write command, wherein the write command includes a first identifier, a first content of a first object to be written, and a first version number to be written, wherein the first identifier is a local maximum term identifier of more than half of the servers in the replication group, and the first identifier indicates that the replication group has reached a consensus that the first server in the replication group serves as the first role; When the client receives the first information sent by more than half of the servers in the replication group, a commit command is broadcast, wherein the first information indicates that the write command is agreed, and the commit command indicates that the status of updating the first content corresponding to the first version number is submitted.
13. The method according to claim 12, characterized in that Before the broadcast write command, the method further includes: The first identifier is obtained from the replication group.
14. The method according to claim 12 or 13, characterized in that The first role is a semi-leader, and the method further comprises: broadcasting a read command, where the read command is used to request to read the first object, and the read command includes the first identifier; receiving a read result sent by a server in the replication group, the read result comprising a second version number and second content of the first object corresponding to the second version number, the second version number being the maximum version number corresponding to the first object locally on the server; When the client receives the read results sent by more than half of the servers in the replication group, the second content of the first object corresponding to the largest version number is obtained from the received read results.
15. The method according to claim 12 or 13, characterized in that The first role is a leader, and the method further includes: Sending a read command to the first server, where the read command is used to request to read the first object, and the read command includes the first identifier; Receive the latest content of the first object sent by the first server.
16. The method according to claim 12 or 13, characterized in that The first role is a leader, and the method further includes: Acquire a second identifier from the replication group, where the second identifier is the local maximum term identifier of more than half of the servers in the replication group, the second identifier indicates that the replication group has reached a consensus that the second server in the replication group serves as the second role, the second role is a semi-leader, and the second identifier is greater than the first identifier; A data read operation or data write operation is performed according to the second identifier.
17. The method according to claim 16, characterized in that The method further includes: sending a read command to the first server, the read command being used to request to read the first object, the read command including the first identifier; The obtaining a second identifier from the replication group includes: In the case of receiving a reply from the first server rejecting the read command, obtaining the second identifier from the replication group.
18. A data processing system, characterized in that: The system includes a first client and a replication group, wherein more than half of the servers in the replication group have a local maximum term identifier as a first identifier, and the first identifier indicates that the replication group has reached a consensus that the first server in the replication group serves as the first role, The first client is used to broadcast a first write command, where the first write command includes the first identifier, the first content of the first object to be written, and the first version number to be written; The server in the replication group is configured to write the first content and the first version number when the local maximum term identifier is the first identifier and the local maximum version number corresponding to the first object is smaller than the first version number, and send first information to the first client, wherein the first information indicates that the first write command is agreed to; The first client is used to broadcast a submission command when receiving the first information sent by more than half of the servers in the replication group including the first server, wherein the submission command indicates that the status of updating the first content corresponding to the first version number is submitted.
19. The system according to claim 18, characterized in that The number of servers included in the replication group is 2f+1, where f is a positive integer.
20. The system according to claim 18 or 19, characterized in that Before the first client broadcasts the first write command, The first client is further used to obtain the first identifier from the replication group.
21. The system according to any one of claims 18 to 20, characterized in that: The first role is a leader, and the system further includes a second client. The second client is used to send a read command to the first server, where the read command is used to request to read the first object, and the read command includes the first identifier; The first server is configured to send the latest content of the first object to the second client when the local maximum term identifier is the first identifier.
22. The system according to any one of claims 18 to 20, characterized in that: The first role is a leader, and the system further includes a second client. The second client is used to send a read command to the first server, where the read command includes a second identifier different from the first identifier; The first server is further configured to reply to reject the read command and send the first identifier to the second client.
23. The system according to any one of claims 18 to 22, characterized in that: The first role is a leader, and the system further includes a third client. The third client is used to send a read command to the first server; When any of the following conditions is met, the first server is further configured to reply to reject the read command sent by the third client: Before receiving the read command sent by the third client, the first server fails to retry the heartbeat communication within a preset time period; or Before receiving the read command sent by the third client, the first server determines that the first identifier is invalid based on the heartbeat information sent by the second server in the replication group, and the heartbeat information includes the local maximum term identifier of the second server and the local maximum term identifier of the second server is greater than the first identifier.
24. The system according to any one of claims 18 to 23, characterized in that: The first role is a leader, less than half of the servers in the replication group are faulty, and the faulty servers do not include the first server.
25. The system according to any one of claims 18 to 23, characterized in that: The first role is a leader. When the first server fails, the local maximum term identifier of more than half of the servers in the replication group is a third identifier, the third identifier is greater than the first identifier, and the third identifier indicates that the replication group has reached a consensus that the third server in the replication group serves as the second role, and the second role is a semi-leader. Within the waiting time from when the third server takes on the second role, if the server whose local maximum term identifier in the replication group is the third identifier receives a second write command, the server is further used to reply to reject the second write command; The waiting time is greater than the maximum time allowed for the first server to retry heartbeat communication.
26. The system according to claim 25, characterized in that When the third server determines that it is the server that currently stores the latest version data of all objects in the replication group, the third server is further used to broadcast an election request, where the election request is used to request the election of the third server as a leader.
27. The system according to any one of claims 18 to 20, characterized in that: The first role is a semi-leader, If a server in the replication group receives the first write command after a waiting period from when the first server acts as the first role, the server is configured to write the first content and the first version number when the local maximum term identifier is the first identifier and the local maximum version number corresponding to the first object is less than the first version number; The waiting time is greater than the maximum time allowed for the server as a leader to retry the heartbeat communication.
28. The system according to claim 27, characterized in that The system further includes a second client, The second client is used to broadcast a read command, where the read command is used to request to read the first object, and the read command includes the first identifier; The server in the replication group is used to send a read result to the second client when the local maximum term identifier is the first identifier, the read result including a second version number and the second content of the first object corresponding to the second version number, and the second version number is the maximum version number corresponding to the first object locally on the server; The second client is used to obtain the second content of the first object corresponding to the largest version number in the received read results when receiving the read results sent by more than half of the servers in the replication group.
29. A device for data processing, characterized in that: The device is a client or is included in the client, and the device includes: A sending unit, configured to broadcast a write command, wherein the write command includes a first identifier, a first content of a first object to be written, and a first version number to be written, wherein the first identifier is a local maximum term identifier of more than half of the servers in the replication group, and the first identifier indicates that the replication group has reached a consensus on the first server in the replication group as the first role; The sending unit is also used to broadcast a submission command when the receiving unit in the device receives first information sent by more than half of the servers in the replication group, wherein the first information indicates agreement with the write command, and the submission command indicates that the status of updating the first content corresponding to the first version number is submitted.
30. The device according to claim 29, characterized in that The receiving unit is further configured to obtain the first identifier from the replication group.
31. The device according to claim 29 or 30, characterized in that The first role is a semi-leader, The sending unit is further used to broadcast a read command, where the read command is used to request to read the first object, and the read command includes the first identifier; The receiving unit is configured to receive a read result sent by a server in the replication group, wherein the read result includes a second version number and a second content of the first object corresponding to the second version number, wherein the second version number is the maximum version number corresponding to the first object locally on the server; When the receiving unit receives the read results sent by more than half of the servers in the replication group, the processing unit in the device is used to obtain the second content of the first object corresponding to the largest version number from the received read results.
32. The device according to claim 29 or 30, characterized in that The first role is the leader, The sending unit is further used to send a read command to the first server, where the read command is used to request to read the first object, and the read command includes the first identifier; The receiving unit is further configured to receive the latest content of the first object sent by the first server.
33. The device according to claim 29 or 30, characterized in that The first role is the leader, The receiving unit is further used to obtain a second identifier from the replication group, the second identifier being a local maximum term identifier of more than half of the servers in the replication group, the second identifier indicating that the replication group has reached a consensus on the second server in the replication group as a second role, the second role being a semi-leader, and the second identifier being greater than the first identifier; The processing unit in the device is used to perform a data read operation or a data write operation according to the second identifier.
34. The device according to claim 33, characterized in that The sending unit is further used to send a read command to the first server, where the read command is used to request to read the first object, and the read command includes the first identifier; The receiving unit is specifically configured to: upon receiving a reply from the first server rejecting the read command, obtain the second identifier from the replication group.
35. A device for data processing, characterized in that: The device comprises a processor and a memory, wherein the memory stores computer program instructions, and the processor executes the computer program instructions to enable the device to perform the method according to any one of claims 12 to 17.
36. A computer-readable storage medium, characterized in that: The method comprises computer instructions, which, when executed by a processor, implement the method according to any one of claims 12 to 17; or, when executed by a data processing system, implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Data processing method and system
CN120029791A
Data processing method and system, computer equipment and storage medium
CN111368002A
Data processing method and device and electronic equipment
CN114244859A
Consensus method with high consensus efficiency and distributed system
CN115102967A
In-order fault tolerant consensus logs for replicated services
US20220286378A1