Distributed cluster updating method and apparatus, device, and medium
By identifying node role labels and implementing automated update strategies, the problems of business interruption and inefficiency in distributed cluster updates are solved, achieving efficient and reliable cluster updates and ensuring high availability and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU OVERSEAS KANGBAZI NETWORK TECHNOLOGY CO LTD
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-31
AI Technical Summary
Existing distributed cluster update methods suffer from problems such as service interruption due to full shutdown updates and low efficiency and error-proneness of manual updates on a per-machine basis, failing to meet high availability requirements.
By identifying the role tags of nodes to be updated and implementing differentiated update strategies, the leader node is updated last, and other nodes are updated first. The version replacement process is automated to avoid manual intervention. By leveraging the backward compatibility of high-version nodes, the cluster always has stable nodes providing services during the update process.
It enables efficient and reliable updates to distributed clusters, avoiding business interruptions and configuration errors, ensuring high availability of core business systems and robustness of the update process, and improving update efficiency and reliability.
Smart Images

Figure REF-OBJ-1782121624548-000002 
Figure REF-OBJ-1782121624548-000003 
Figure REF-OBJ-1782121624548-000004
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a distributed cluster update method and the corresponding apparatus, computer equipment, and computer-readable storage medium. Background Technology
[0002] ZooKeeper, as an open-source distributed coordination service, is typically deployed in clusters to provide highly available configuration management, naming services, and distributed synchronization capabilities. With business growth and security vulnerability patching, version updates are often required for distributed clusters integrated with ZooKeeper. Existing update methods primarily include full-shutdown updates and manual updates on each node. Full-shutdown updates require stopping all nodes in the entire cluster, replacing all nodes with the new version, and then restarting. While this method is simple in logic, it causes the entire cluster to completely stop service during the update window, resulting in prolonged business interruption and failing to meet the high availability requirements of core business systems. Manual updates on each node involve operations personnel logging into each node in the cluster, manually stopping the service process, replacing runtime files, and then restarting the service. While this method can keep some nodes running, the entire process is highly dependent on manual operation, inefficient, and prone to update failures in multi-node cluster environments due to missed steps or configuration errors. Summary of the Invention
[0003] The primary objective of this application is to solve at least one of the above-mentioned problems by providing a distributed cluster update method and corresponding apparatus, computer equipment, and computer-readable storage medium.
[0004] To achieve the various objectives of this application, the following technical solution is adopted: A distributed cluster update method provided for one of the purposes of this application includes the following steps: In response to a distributed cluster update event, determine the node to be updated, its current runtime version, the target runtime version, and its role tag. If the current runtime version is detected to be lower than the target runtime version, determine whether the node to be updated belongs to the leader node based on the role tag; If the node to be updated belongs to the leader node, then check whether the current runtime version of all other nodes in the distributed cluster, excluding the node to be updated, is the target runtime version. If the current runtime version of all other nodes is the target runtime version, then the current runtime version of the node to be updated will be updated to the target runtime version.
[0005] On the other hand, a distributed cluster update apparatus provided to meet one of the purposes of this application includes an event response module, a node analysis module, a version analysis module, and a version update module. The event response module is used to respond to a distributed cluster update event and determine the node to be updated corresponding to the event, its current runtime version, the target runtime version, and its role label. The node analysis module is used to determine whether the node to be updated is a leader node based on its role label if the current runtime version is detected to be lower than the target runtime version. The version analysis module is used to detect whether the current runtime version of all other nodes in the distributed cluster, excluding the node to be updated, is the target runtime version if the node to be updated is a leader node. The version update module is used to update the current runtime version of the node to be updated to the target runtime version if the current runtime version of all other nodes is the target runtime version.
[0006] On another front, a computer device provided for one of the purposes of this application includes a central processing unit and a memory, wherein the central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the distributed cluster update method of this application.
[0007] In another aspect, a computer-readable storage medium is provided to suit another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implementing a distributed cluster update method, which, when invoked by a computer, performs the steps included in the method. Attached Figure Description
[0008] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart illustrating a typical embodiment of the distributed cluster update method of this application; Figure 2 This is a schematic diagram of the distributed cluster update device of this application; Figure 3 This is a schematic diagram of the structure of a computer device used in this application. Detailed Implementation
[0009] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0010] Those skilled in the art will understand that, unless explicitly stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in the specification of this application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0011] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0012] Those skilled in the art will understand that the terms "client," "terminal," and "terminal device" as used herein include both devices that receive wireless signals, devices that only possess wireless signal receiver capabilities without transmission capabilities, and devices with receiving and transmitting hardware, devices that have receiving and transmitting hardware capable of bidirectional communication over a bidirectional communication link. Such devices may include: cellular or other communication devices such as personal computers or tablets, having single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service) that can combine voice, data processing, fax, and / or data communication capabilities; PDAs (Personal Digital Assistants) that may include radio frequency receivers, pagers, internet / intranet access, web browsers, notebooks, calendars, and / or GPS (Global Positioning System) receivers; and conventional laptops and / or handheld computers or other devices that have and / or include radio frequency receivers. As used herein, "client," "terminal," and "terminal device" can be portable, transportable, installed in a means of transportation (air, sea, and / or land), or suitable and / or configured to operate locally and / or in a distributed manner, operating in any other location on Earth and / or in space. "Client," "terminal," and "terminal device" as used herein can also be a communication terminal, an internet access terminal, or a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback capabilities, or a smart TV, set-top box, etc.
[0013] The hardware referred to by the names "server," "client," and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann architecture, such as a central processing unit (including an arithmetic logic unit and a control unit), memory, input devices, and output devices. The computer program is stored in its memory, and the central processing unit loads the program stored in the secondary storage into the main memory to run it, execute the instructions in the program, and interact with the input and output devices to complete specific functions.
[0014] It should be noted that the concept of "server" used in this application can also be extended to apply to server clusters. Based on network deployment principles as understood by those skilled in the art, each server should be a logical division. Physically, these servers can be independent of each other but accessible through interfaces, or they can be integrated into a single physical computer or a computer cluster. Those skilled in the art should understand this flexibility and should not use it to constrain the implementation of the network deployment method described in this application.
[0015] One or more of the technical features of this application, unless explicitly specified herein, can be deployed on a server and accessed by a client remotely calling the online service interface provided by the server, or can be directly deployed and run on a client for access.
[0016] Unless otherwise specified, all data involved in this application may be stored remotely on a server or on a local terminal device, as long as it is suitable for use by the technical solution of this application.
[0017] Those skilled in the art will understand that although the various methods in this application are described based on the same concept and thus present commonality among them, they can be performed independently unless otherwise specified. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept; therefore, concepts expressed in the same way, as well as concepts that are appropriately changed for convenience but are expressed differently, should be understood equivalently.
[0018] Unless otherwise expressly stated, the various embodiments disclosed in this application can be combined in a cross-cutting manner to flexibly construct new embodiments, as long as such combination does not depart from the inventive spirit of this application and can meet the needs of the prior art or solve a certain deficiency in the prior art. Those skilled in the art should be aware of such modifications.
[0019] The distributed cluster update method of this application can be programmed into a computer program product and deployed on a client or server. For example, in the exemplary application scenario of this application, it can be deployed on servers in the fields of content delivery network (CDN), video platform, live streaming platform, e-commerce platform, etc., so that the method can be executed by human-computer interaction with the process of the computer program product through a graphical user interface by accessing the interface opened after the computer program product is run.
[0020] Please see Figure 1 In some embodiments of the distributed cluster update method of this application, one embodiment includes the following steps: Step S1100: Respond to the distributed cluster update event and determine the node to be updated, its current runtime version, target runtime version, and role tag corresponding to the event.
[0021] This application targets a distributed ZooKeeper cluster containing multiple nodes, categorized by their roles as leader, follower, and observer. Distributed cluster update events are triggered by users (typically operations personnel) through an operations management platform or command-line interface. When triggered, the user can provide information about the node in the ZooKeeper distributed cluster requiring the version update, typically including the node's IP address and port or hostname, such as the IPv4 address 192.168.1.10 and the default SSH port 22. Alternatively, the user can instruct the update to begin without explicitly specifying a node. The event is then triggered by scheduling a node from the queue of nodes to be updated according to a pre-defined update order. This node is then designated as the node to be updated for this event. This update order typically places follower nodes before the leader node to ensure that the leader node is updated only in the final stage of the rolling update process. Furthermore, under the rolling update strategy adopted in this application, in order to ensure that more than half of the nodes in the distributed cluster are always online to provide normal coordination services during the update period, the update process is carried out on a node-by-node basis. Therefore, in a single update event, one node is designated as the node to be updated. Those skilled in the art can further configure the node update order as needed based on the disclosure herein.
[0022] Furthermore, following the user's trigger, a server independent of the distributed cluster obtains and parses the node information provided by the user or the dequeued node information to determine the node to be updated in this update operation. The node to be updated refers to the physical server or virtual machine instance in the ZooKeeper distributed cluster that actually performs the version update operation. Subsequently, the server directly checks its local storage directory to see if the target version data packet for the current version needs to be updated exists. If the target version data packet does not exist locally, it usually indicates that the node to be updated is the first node in the distributed cluster to update this version. The server directly pushes the update script to the node to be updated. When the node executes the script, it triggers the corresponding third-party download. If the target version data packet exists locally, the server sends the target version data packet and its hash data (used to verify the integrity of the data packet) along with the update script to the node to be updated, thereby preventing nodes that are not the first to update from repeatedly downloading the target version data packet.
[0023] After receiving the update script, the node to be updated responds to the distributed cluster update event and executes the update script to determine its current runtime library version and the target runtime library version. The current runtime library version refers to the specific version number of the Zookeeper software currently running on the node to be updated, and the target runtime library version refers to the specific version number of the Zookeeper software that the node to be updated expects to update to. There are two specific implementation methods for the node to be updated to obtain the current runtime library version. In one embodiment, the update script reads the version identifier file in the Zookeeper installation directory, reads a text file named "version.txt" or "build.txt", and extracts the specific version string from it. In another embodiment, the update script calls Zookeeper's built-in four-letter command to send the "envi" command to the service port of the node to be updated, obtains the environment configuration output information of the Zookeeper process, and extracts the value corresponding to "zookeeper.version" from the output information using regular expressions. In specific implementations, the update script parses the update task configuration file issued by the server and reads the preset target version number from it.
[0024] Role labels refer to the specific role an updated node assumes in the ZAB distributed consensus protocol of a ZooKeeper distributed cluster. These roles include leader, follower, and observer. Functionally, the leader handles all write requests in the distributed cluster, initiates transaction proposals, coordinates data synchronization, and manages transaction logs, serving as the core provider of cluster coordination services. Followers handle client read requests, forward write requests to the leader, and participate in leader election voting and transaction proposal confirmation. Observers also handle read requests and synchronize data with the leader, but do not participate in leader election voting; their primary purpose is to extend the cluster's read capacity without affecting election performance. There are two specific implementations for an updated node to obtain its role label: In one embodiment, the update script reads state variables from the ZooKeeper process memory or reads state files from the local data directory, combines this with the ZAB protocol's election state determination logic, calculates and returns the role identifier. Specific implementation details will be provided in subsequent embodiments; this step will not be discussed further here. In another embodiment, the update script sends the "stat" or "srvr" command to the client port of the node to be updated, which has an established communication connection, by invoking Zookeeper's built-in four-letter command to obtain the current running status output information of the node to be updated. The update script parses the output information and extracts key fields that represent the node's role. If the extracted key field is "Mode: leader", the role label of the node to be updated is determined to be leader; if the extracted key field is "Mode: follower", the role label of the node to be updated is determined to be follower; if the extracted key field is "Mode: observer", the role label of the node to be updated is determined to be observer.
[0025] Step S1200: If the current runtime version is detected to be lower than the target runtime version, determine whether the node to be updated belongs to the leader node based on the role tag.
[0026] If the current runtime library version is detected to be lower than the target runtime library version, it means that the node to be updated has not been updated to the target runtime library version. The update script then continues to execute to accurately determine the node's identity. During this process, the update script parses the current runtime library version and the target runtime library version into strings conforming to semantic versioning specifications, extracts the major version number, minor version number, and revision number, and compares them digit by digit. When the numerical combination of the current runtime library version is significantly less than the numerical combination of the target runtime library version, it is confirmed that the current runtime library version is lower than the target runtime library version. Then, based on the role label, it is determined whether the node to be updated belongs to the leader node. In one embodiment, if the update script confirms that the role label is a leader role, it directly determines that the node to be updated belongs to the leader node; otherwise, if the role label is a follower role or an observer role, it directly determines that the node to be updated does not belong to the leader node.
[0027] Step S1300: If the node to be updated belongs to the leader node, then check whether the current runtime version of all other nodes in the distributed cluster, excluding the node to be updated, is the target runtime version.
[0028] Under the premise that the node to be updated assumes the role of leader, a comprehensive check of the version status of the remaining nodes in the entire cluster is carried out. The purpose is to strictly ensure that the leader node is the last object to be updated in the cluster. When it is updated, all follower nodes and observer nodes have already completed their version updates, thereby eliminating the risk of incompatibility of distributed consensus protocols or data interfaces caused by the leader node's version being higher than that of other nodes.
[0029] The other nodes mentioned above are every node that makes up the ZooKeeper distributed cluster, excluding the node currently identified as the leader and awaiting update, encompassing all follower and observer nodes. "Current runtime version" refers to the specific version identifier of the ZooKeeper software running on these nodes, such as "3.4.14"; "Target runtime version" refers to the unified version identifier expected to be achieved in this update operation, such as "3.5.6". The essence of the detection is to verify, one by one, whether the current runtime version value of each non-leader node in the cluster is completely consistent with the target runtime version value. The specific detection process can be implemented through various embodiments, including but not limited to: in one embodiment, the detection is carried out using the communication links already established within the ZooKeeper cluster. Since the node awaiting update is still running online as the leader, it maintains effective data synchronization and heartbeat connections with all follower and observer nodes. The update script can send lightweight custom probe commands sequentially to each other node in the cluster via the leader node process, or directly reuse existing commands in the ZooKeeper management protocol. For example, the command "envi" is sent to the service port of each node to receive the environment configuration information returned by each node, from which the version number corresponding to the "zookeeper.version" field is parsed. Then, the parsed version number string is compared with the target runtime library version string. When the version information of every node in the cluster except itself is successfully obtained and each version number is equal to the target runtime library version, the current runtime library version of all other nodes is determined to be the target runtime library version.
[0030] In another embodiment, a Secure Shell protocol is used to connect to other nodes for local queries. The update script runs on the node to be updated, but with the help of pre-configured passwordless authentication information, it remotely logs into the operating system environment of each other node in the cluster via the Secure Shell protocol. After logging in, a pre-written local version query command is executed on each remote node. Specifically, this can be done by reading the version identifier file in the ZooKeeper installation directory, such as accessing " / usr / local / zookeeper / version.txt" and directly extracting the version number string recorded in the file; or by calling the command-line tool included with ZooKeeper and executing the command "bin / zkServer.sh status" to extract the version field from the output information. After obtaining the version number of each node, it is compared with the target runtime version. Similarly, the check is considered successful only when the version numbers obtained from all remote queries are the same as the target runtime version.
[0031] In another embodiment, detection is implemented using node metadata stored in an independent cluster management service or configuration center. During distributed cluster deployment, the current runtime version information of each node is automatically reported and persistently stored in a centralized metadata repository or the backend database of the operations and maintenance management platform. The update script can directly send a query request to this management service to retrieve the latest runtime version field of all nodes in the cluster whose role tag is not the leader. The management service returns a mapping list of node identifiers and version numbers. The update script iterates through this list, checking whether the version number of each non-leader node matches the target runtime version. This method does not require establishing real-time connections for each node, making it highly efficient, but it relies on the version information in the management service being synchronized with the actual deployment status in real time. Therefore, an additional round of sampling verification can be added: one or two nodes are randomly selected from the list, and their reported version authenticity is cross-verified using the aforementioned four-word command or remote login method to enhance the reliability of the detection.
[0032] Regardless of the implementation method used, if it is detected that the current runtime version of any other node is not equal to the target runtime version, it indicates that some nodes in the cluster have not yet completed the update. At this point, the update process will not continue to perform the version replacement operation on the current leader node. Instead, the update script will generate a clear detection failure notification, indicating the identifier of the non-compliant node and its current version number, and suspend or terminate the current update task for that leader node. Operations personnel can prioritize processing those non-updated nodes based on this notification. After all non-leader nodes have been updated, the update process for the leader node will be retried, thereby ensuring the rigor of the entire rolling update sequence and the consistency of cluster services.
[0033] Step S1400: If the current runtime version of all other nodes is the target runtime version, then update the current runtime version of the node to be updated to the target runtime version.
[0034] When the node to be updated is the leader node and all other nodes in the distributed cluster have been updated to the target runtime version, the node to be updated is updated. During this process, first, the ZooKeeper service process on the node to be updated is stopped.
[0035] Subsequently, the node to be updated obtains the target version data package corresponding to the target runtime version. The technical background for this step is as follows: at the beginning of responding to a distributed cluster update event, the server identifies the node to be updated as the target for this update and immediately checks whether the target version data package exists in its local storage. Since the leader node is the last node scheduled for update in the cluster, the server may have previously sent the data package to the last follower or observer node and then deleted it from its local storage due to a storage space cleanup policy. Furthermore, the scheme disclosed in this application allows nodes in the cluster to run different versions at different times without forcing version locking. Therefore, the existence of the data package on the server's local storage is not a definitive event when scheduling the leader node update. When the data package exists on the server's local storage, the server pushes the data package and its hash data to the designated cache directory of the node to be updated through this event response process; if the data package does not exist on the server's local storage, this push will not occur.
[0036] Therefore, when the node to be updated executes the update script after service shutdown, it first actively checks whether the target version data package has been pre-received and stored in its local designated cache directory. If the data package already exists locally, meaning it received the data package normally pushed by the server without deletion, it is directly read and used; if the data package is not found locally, it indicates that the server did not push it before, and the update script triggers the node to download the complete target version data package from the preset code repository or the official release address. The code repository is a network-accessible database that can store runtime library versions customized by the distributed cluster's operations and maintenance personnel based on the officially released runtime library version. This data package completely replaces the runtime library files of the node to be updated, and reads the personalized configuration file saved for the node in the preset cluster backup data, restoring it to the new version installation directory, ensuring that the node can integrate into the cluster with the correct cluster member identity and communication parameters after restarting. After the node restarts, it compares the difference between its own and the current leader node's transaction identifier, and chooses incremental or full snapshot mode to synchronize data, making up for the state changes that occurred during the downtime. Once data synchronization is complete and the node is marked as available, it registers with the leader and officially rejoins the distributed cluster, thereby restoring full service capabilities with the target runtime version. This demonstrates that the distributed cluster continues to provide stable and reliable service even during leader node updates.
[0037] It's easy to understand that the version update process for non-leader nodes is similar to that for leader nodes, except that there's no need for downtime and the election of a new leader node. Furthermore, non-leader nodes don't need to wait to become the last node in the distributed cluster to update to that version; they can update directly.
[0038] As can be seen from the above embodiment of this application, the technical solution of this application has many advantages, including but not limited to the following aspects: This application effectively solves the technical problems of long-term business interruption caused by full shutdown updates and low efficiency and error-proneness of manual configuration updates on a node-by-node basis by identifying the role tags of the nodes to be updated and implementing differentiated update strategies.
[0039] Upon responding to a distributed cluster update event, the system checks if the current runtime version of the node to be updated is lower than the target version. Subsequent processes are only triggered when an update is required, avoiding invalid operations. By further identifying node role labels, an update order control logic centered on the leader node is constructed. When the node to be updated is the leader, it is mandatory to complete the version updates of all other nodes in the cluster first, and only then is the leader node processed. This ensures that throughout the update process, a stable leader node always exists in the distributed cluster to provide coordination services, preventing cluster instability or service interruptions due to frequent changes or offline status of the leader node. This achieves true rolling updates and guarantees the high availability of core business systems. Simultaneously, it fundamentally avoids the incompatibility between a high-version leader node and a low-version follower node simultaneously in the cluster. Since the leader node is the core provider of cluster coordination services, if its runtime version is updated to a higher version prematurely, the adopted new distributed coordination protocol or data service interface may not be understood and interacted with by other nodes in the cluster that have not yet updated to the lower version, which could easily lead to service anomalies throughout the entire cluster. To address this, all follower nodes are forced to update before the leader node. This leverages the backward compatibility of higher-version nodes with lower-version leaders, ensuring that the current leader node can effectively collaborate with any other node within a mutually compatible protocol framework throughout the update cycle. This eliminates the possibility of some nodes being isolated due to version incompatibility.
[0040] Furthermore, the entire update process is automatically driven and executed, from node role determination to version consistency detection, and finally to version replacement. The entire process requires no manual intervention, fundamentally eliminating the risk of update failure due to missed steps, configuration errors, or operation delays in the traditional manual update method, and significantly improving the efficiency and reliability of distributed cluster updates.
[0041] In a further embodiment, step S1200, determining whether the node to be updated belongs to the leader node based on the role tag, includes the following steps: Step S1210: Send role detection requests to all nodes in the distributed cluster except the node to be updated, and obtain the cluster leader information returned by these nodes.
[0042] During the operation of a distributed cluster, the role status of nodes dynamically switches due to anomalies such as network partitions, node restarts, or election timeouts. If the determination of whether a node is a leader node relies solely on the role label acquired locally by the node to be updated, in transitional states such as network split-brain or before global consensus has been reached, inconsistencies may arise between the node's locally perceived role and the actual consensus state of the cluster. This leads to incorrect node identity determinations and violates the security policy that the leader node must be the last to be updated. To address the technical problem of misidentification caused by biases in the local state perception of a single node and to ensure the absolute accuracy of leader node identity determination, this further embodiment introduces a cross-validation mechanism based on a global cluster perspective. This mechanism obtains cluster leader information from a global perspective by sending probe requests to other nodes and performs double verification by combining this information with the node's local role label.
[0043] Based on this consideration, the first step is to send role detection requests to all nodes in the distributed cluster except the node to be updated, and obtain the cluster leader information returned by these nodes. The role detection request is used to query the identity of the cluster leader currently recognized by the requesting node. The cluster leader information is the identifier of the leader node currently recognized by the corresponding node, including the Internet Protocol address, hostname, or unique node identifier. In one embodiment, the update script, based on the four-word command mechanism built into the distributed coordination service, sends status query commands to the client service ports of all other nodes via Transmission Control Protocol (TCP) connections, such as sending the "stat" command or "srvr" command to each node. The text response returned by the requesting node contains information about the current cluster's operating mode and the specific identifier of the leader node. The update script parses the leader field in the returned text and extracts the Internet Protocol address and port of the leader node. In another embodiment, the update script, based on the underlying communication mechanism of the distributed consensus protocol, sends a custom leader query message through the cluster's internal election port. After receiving the query message, the protocol stacks of other nodes return the current term number and the unique identifier of the leader node corresponding to that term. In another embodiment, the update script sends a query request through the application programming interface of the operation and maintenance management platform. The operation and maintenance management platform reads and returns the cluster leader information reported by each node from its persistent cluster state metadata database. This information includes the unique identifier of the current leader recognized by each node. Regardless of the specific implementation method, the update script aggregates all the collected cluster leader information into a set for subsequent comparison.
[0044] Step S1220: Determine whether the role label represents the leader role, and whether all cluster leader information points to the node to be updated.
[0045] The role tag is a local role identifier obtained by the node to be updated, with a value of one of three: "leader," "follower," or "observer." Representing the leader role means the role tag value equals the preset leader standard identifier, i.e., the string "leader." Pointing to the node to be updated means the node identifier contained in the cluster leader information returned by other nodes completely matches the identifier of the node to be updated itself. This identifier can be an Internet Protocol address, hostname, or node unique identifier. When determining whether a role tag represents a leader role, the update script reads the role tag string obtained in the previous steps and performs an exact match with the preset leader standard string "leader" to determine if the role tag is a leader identifier. When determining whether all cluster leader information points to the node to be updated, the update script iterates through all other cluster leader information returned by other nodes. In one embodiment, the update script extracts the Internet Protocol address (including the corresponding node's IP address and port) from each piece of cluster leader information and performs a string-by-string comparison with the Internet Protocol address bound to the local network interface of the node to be updated. In another embodiment, the update script extracts the unique node identifier from each cluster leader message, which is the integer identifier defined by the myid file in the ZooKeeper configuration. It then compares this integer identifier with the unique node identifier recorded in the myid file in the local data directory of the node to be updated. Only when the cluster leader information returned by every other node is completely consistent with the identifier of the node to be updated is it determined that all cluster leader messages point to the node to be updated. If any node returns a null value, an identifier pointing to another node, or an unknown state for its cluster leader information, it is determined that not all messages point to the node to be updated.
[0046] Step S1230: If the role label represents the leader role and all cluster leader information points to the node to be updated, then determine that the node to be updated belongs to the leader node; otherwise, determine that the node to be updated does not belong to the leader node.
[0047] After completing the above judgments, when both of the above conditions are met simultaneously—that is, the node to be updated locally considers itself the leader, and all other nodes in the cluster unanimously recognize the node to be updated as the sole leader of the current cluster—the update script determines that the node to be updated belongs to the leader node, thereby triggering the subsequent strict version check and delayed update process for the leader node, i.e., checking whether all other nodes in the cluster have been updated to the target runtime version. When either of the above two conditions is not met, for example, the role label shows as follower, or the role label shows as leader but at least one other node returns cluster leader information pointing to another node, it indicates that a network split brain or inconsistent election state has occurred in the cluster, and the update script determines that the node to be updated does not belong to the leader node. At this time, the update script treats the node to be updated as an ordinary non-leader node and directly executes the normal non-leader node update process, updating the node's current runtime version to the target runtime version, without waiting for other nodes to complete the update.
[0048] In this embodiment, by introducing a cross-validation mechanism from a cluster-wide perspective, the risk of misjudgment caused by biases in the local state perception of a single node is completely eliminated. When distributed systems face complex scenarios such as network jitter, partitioning, or election transition periods, the role state cached locally by a node may lag behind the actual consensus state of the cluster. This embodiment mandates that the local role label of the node to be updated be absolutely consistent with the cluster leader information reported by all other nodes in the cluster before confirming its leader status. This dual-verification mechanism effectively prevents concurrent update conflicts caused by multiple nodes simultaneously claiming to be the leader due to split-brain phenomena, ensuring that the core security strategy of the leader node being the last to update is strictly and accurately executed. This significantly improves the robustness and data consistency of the rolling update process while guaranteeing the high availability of the distributed cluster.
[0049] In a further embodiment, step S1400, updating the current runtime version of the node to be updated to the target runtime version, includes the following steps: Step S1410: Stop running the node to be updated, and obtain the target version data package and its hash data corresponding to the target runtime version.
[0050] After triggering the version replacement operation on the leader node, this further embodiment elaborates on the specific implementation process of the replacement operation to ensure that the new version of the runtime is deployed securely and completely on the node to be updated, and reintegrated into the distributed cluster with the correct cluster identity configuration, thereby fully restoring its service capabilities.
[0051] First, the ZooKeeper service process on the node to be updated is stopped. In one specific embodiment, the update script calls the "zkServer.sh stop" command included in the ZooKeeper installation directory to send a termination signal to the running ZooKeeper process, waiting for the process to completely exit and release the port and file locks it occupies. In another specific embodiment, the update script executes the service stop command through an operating system-level process management tool, such as systemctl or supervisor, which is responsible for sending a termination signal to the ZooKeeper process and monitoring its exit status.
[0052] After the service process has completely stopped, the operation of retrieving the target version data package and its hash data corresponding to the target runtime library version is executed. The target version data package is an archive compressed package containing all executable programs, runtime dependencies, default configuration templates, and other necessary files for the new version of ZooKeeper, usually provided in tar.gz or zip format. The hash data is an integrity verification benchmark value associated with the target version data package, a digest string calculated from the data package using a preset hash algorithm, such as a 64-bit hexadecimal string generated using the SHA-256 algorithm.
[0053] The specific implementation method for obtaining the target version data packet includes: at the beginning of a distributed cluster update event, a server independent of the cluster, after determining the node to be updated, checks whether the target version data packet is cached in its local storage directory. The condition for the server to store this data packet locally is that the data packet has been used in previous update operations on other nodes and has not yet been cleaned up. The leader node, as the last scheduled update target in the cluster, may have its data packet deleted from its local storage after the server pushes it to the penultimate node due to a storage space cleanup policy. Therefore, when the server locally possesses the data packet, it pushes the data packet and its corresponding hash data to the designated cache directory of the node to be updated; when the server locally does not possess the data packet, this push will not occur.
[0054] When the node to be updated executes the update script after the service process stops, it first actively queries its local designated cache directory to see if the target version data package has been pre-received and stored. In one embodiment, the update script checks whether there is an archive file with a fixed naming rule corresponding to the target runtime library version in the cache directory. If the data package already exists locally, it directly reads the data package file and its accompanying hash data. If the data package is not found locally, it indicates that the server did not push it previously due to the lack of local caching. In this case, the update script triggers the node to download the complete target version data package corresponding to the target runtime library version from the preset official release address, and simultaneously download the corresponding hash data from the same source address. At this point, the target version data package and its hash data are located in the local file system of the node to be updated.
[0055] Step S1420: Determine whether the hash value of the data packet corresponding to the target version data packet is consistent with the source hash value in the hash data.
[0056] After performing the stop and get operations, the update script checks whether the hash value of the data packet corresponding to the target version data packet matches the source hash value in the hash data. The data packet hash value is a digest string recalculated by the node to be updated after receiving the target version data packet, using the same one-way hash function as the sender. The source hash value is the original digest string calculated by the sender and recorded in the hash data. To determine if they match, the update script calls the operating system's built-in verification tool or a third-party encryption library to fully read the target version data packet and calculate its hash value. In one embodiment, the update script uses a secure hash algorithm to calculate the data packet hash value of the target version data packet and compares the calculated hexadecimal string with the source hash value character by character. If all characters are identical, they are considered to match. In another embodiment, the update script uses a message digest algorithm to calculate the data packet hash value, converts it to Base64 encoding, and performs a binary-level equivalence check with the source hash value, which is also Base64 encoded in the hash data. If there is any difference between the data packet hash value and the source hash value, they are considered inconsistent, and the update script immediately terminates the update process and triggers a data re-download mechanism.
[0057] Step S1430: If the data packet hash value is consistent with the source hash value, then update the version configuration data corresponding to the current runtime version of the node to be updated based on the target version data packet, and read the configuration file of the node to be updated from the preset cluster backup data and configure it for the node.
[0058] After the data integrity verification passes, a substantive replacement operation on the node runtime library files is performed. First, the update script decompresses the target version data package to a temporary directory. The decompressed temporary directory contains all the constituent files of the new version of ZooKeeper, specifically including executable scripts (such as zkServer.sh, zkCli.sh), Java archive files (such as zookeeper-version.jar), native dependency libraries (such as libzookeeper*.so), and version identifier files (such as version.txt), etc.
[0059] Subsequently, the update script completely replaces the existing files in the ZooKeeper installation root directory on the node to be updated with the new version files in the temporary directory. In one embodiment, the update script first moves all the current old version files in the ZooKeeper installation directory to a timestamped backup directory, and then recursively copies all the new version files in the temporary directory to the corresponding path in the installation directory. This backup-then-overwrite approach provides the ability to roll back immediately. In another embodiment, for hot-swappable components such as Java archive files, the update script directly uses file system-level copy operations to overwrite the old version files with the same name with the new version files.
[0060] After replacing the runtime files, the update script reads the personalized configuration file saved for the node to be updated from the preset cluster backup data and restores it to the corresponding location in the new version's installation directory. The cluster backup data is generated before responding to a distributed cluster update event. The update management system traverses each node in the distributed cluster, reading the configuration file used by each node when running under the old runtime version, i.e., the zoo.cfg main configuration file. Based on the unique identifier of the node to be updated, such as its Internet Protocol address or myid value, the update script precisely extracts the main configuration file belonging to that node from the cluster backup data, copies these files to the corresponding conf subdirectory in the new installation directory, overwriting the general default configuration included with the target version's data package.
[0061] Step S1440: Re-add the node to be updated to the distributed cluster and run it again, and determine that the current runtime version of the node to be updated after running is the target runtime version.
[0062] After completing the runtime library replacement and configuration restoration, the node to be updated is restarted and rejoined to the distributed cluster. Specifically, the update script calls the "zkServer.sh start" command in the ZooKeeper installation directory, or the service is triggered to start via the operating system's process management tools. After the node process starts, it loads all executable programs and dependent libraries of the target runtime library version from the newly installed directory. During the node startup process, it reads the configuration file restored from the cluster backup data, obtains the address list of other members in the cluster, and attempts to establish a connection with the cluster.
[0063] After a node starts, it enters the transaction data synchronization phase. Once synchronization is complete, the node status is updated and it officially joins the cluster. Specifically, when the update script or the node's own monitoring process detects that the node's status identifier has changed to an available state, it confirms that the node has completed data synchronization and is capable of providing coordination services. Subsequently, the node awaiting update proactively sends a heartbeat registration request to the current leader node. This heartbeat request contains the node's unique identifier and the identifier of the latest processed transaction. Upon receiving this heartbeat registration request, the current leader node re-adds it to the cluster member list. The node awaiting update then begins receiving read and write requests from clients or participating in voting to officially restore all services using the target runtime version.
[0064] In this embodiment, by forcibly performing an integrity check on the target version data package after the node stops running, it is ensured that the runtime library files written to the disk are absolutely consistent with the release source. At the same time, by accurately restoring the node's personalized configuration file from the preset cluster backup data, the risk of the new version's default configuration overwriting the node's original cluster topology and identity is effectively shielded.
[0065] In a further embodiment, step S1440, re-adding the node to be updated to the distributed cluster and re-running, includes the following steps: Step S1441: Obtain the latest transaction identifier of the current leader node in the distributed cluster, and the latest transaction identifier of the node to be updated stored locally.
[0066] In typical distributed cluster update scenarios involving downtime, service can be restored by incrementally synchronizing new transactions generated during the downtime after the node restarts. However, in practice, risks exist such as corrupted local transaction logs, disk anomalies leading to transaction loss, or excessive downtime causing incremental logs to be cleared by the cluster. If only incremental synchronization is relied upon, any missing local transactions will cause an irreversible fork in the node's state and the cluster's state. To address the issue that incremental synchronization alone cannot handle local transaction anomalies and ensure the absolute integrity of transaction data on updated nodes, this further embodiment introduces a dynamic synchronization and fallback mechanism based on transaction differences. Incremental synchronization serves as a conventional and efficient recovery method, while full snapshot synchronization acts as a fallback mechanism to handle transaction anomalies or excessive lag.
[0067] Based on this consideration, the update script is executed first to obtain the latest transaction identifier of the current leader node in the distributed cluster, as well as the latest transaction identifier of the node to be updated stored locally. The current leader node refers to the core node responsible for handling write requests and coordinating data synchronization, elected by the distributed cluster through an election mechanism during the node's downtime for updates. If the node to be updated is a leader node, then the current leader node in this step is the leader node re-elected during the node's downtime for updates; otherwise, the current leader node in this step is the already elected leader node in the distributed cluster. The latest transaction identifier of the current leader node in the distributed cluster is used to uniquely and monotonically increase the global sequence number of the last state change operation in the distributed cluster. In ZooKeeper's ZAB protocol, it is called zxid, a 64-bit long integer value composed of a term number and a transaction sequence number. The latest transaction identifier stored locally on the node to be updated refers to the sequence number corresponding to the last transaction log successfully persisted to the local disk before the node's downtime. In one embodiment, the update script sends the "srvr" command to the client service port of the current leader node through the management interface of the distributed coordination service, parses the "Zxid" field in the returned message, and obtains the latest transaction identifier of the current leader node. Simultaneously, it reads the transaction log file in the local data directory of the node to be updated. This file is typically prefixed with "log.", and the last complete transaction record is read byte-by-byte, from which the zxid field is parsed as the latest transaction identifier stored locally on the node to be updated. In another embodiment, the update script remotely logs into the current leader node via the Secure Shell protocol, executes a transaction log parsing tool in its local data directory, reads the last record of the transaction log file, extracts its zxid value as the latest transaction identifier of the current leader node, and scans the local transaction log directory of the node to be updated to obtain the name of the last log file sorted by timestamp. The sequence number within the file is then parsed using a regular expression as the latest transaction identifier stored locally on the node to be updated.
[0068] Step S1442: Determine the transaction difference between the latest transaction identifier of the current leader node and the latest transaction identifier stored locally on the node to be updated.
[0069] The transaction difference refers to the absolute numerical difference between the latest transaction identifier of the current leader node and the latest transaction identifier stored locally on the node to be updated. This difference quantifies the number of normal state changes missed by the node to be updated during downtime and is a key indicator for measuring the health and completeness of the local transaction log of the node to be updated. In one embodiment, the update script converts both the latest transaction identifier of the current leader node and the latest transaction identifier stored locally on the node to be updated into long integer values, performs a subtraction operation, takes the absolute value, and assigns the result to the transaction difference variable. For example, if the latest transaction identifier of the current leader node is 0x30000005 and the latest transaction identifier stored locally on the node to be updated is 0x30000001, then the difference is 4, indicating that the node to be updated is missing 4 committed transactions.
[0070] Step S1443: If the transaction difference is less than or equal to the preset synchronization threshold, the node to be updated is triggered to perform incremental transaction data synchronization with the current leader node, and after the incremental transaction data synchronization is completed, the node status identifier of the node to be updated is updated to the available state.
[0071] The preset synchronization threshold is a critical value pre-configured by the system to define the boundary between regular incremental synchronization and fallback synchronization for abnormal situations. It is typically set to an integer value, such as 1000, based on the normal downtime for cluster updates and the transaction generation rate. Incremental transaction data synchronization means that the current leader node only sends the newly added transaction log corresponding to the transaction difference to the nodes to be updated, which then replay it locally to catch up on the status. Node status identifiers are state variables used in process memory to characterize the service readiness of the current node, and their values include initialization status, synchronization status, and available status. This branch represents an efficient recovery path under normal circumstances, indicating that the local transaction log of the node to be updated is intact and only a small number of missing transactions need to be caught up. In one embodiment, when the transaction difference is determined to be less than or equal to the preset synchronization threshold, the node to be updated initiates a protocol-defined follower recovery request to the current leader node, carrying its latest locally stored transaction identifier. After receiving the request, the current leader node locates the first transaction record with a transaction identifier greater than this in its transaction log, and from that position, sends all subsequent committed transaction records one by one in a streaming manner to the nodes to be updated. The node to be updated receives these transaction records and applies them one by one to its in-memory database and local transaction log in their original commit order. After completion, it changes the node status flag in memory from synchronized to available. In another embodiment, the leader node packages the missing transaction records into multiple fixed-size data blocks and transmits them concurrently. The node to be updated receives the data blocks, verifies the sequence number continuity, and writes them to its local disk. After all data blocks have been written, it calls the state machine interface to update the node status flag to available.
[0072] Step S1444: If the transaction difference is greater than the preset synchronization threshold, the current leader node is triggered to send full snapshot data and incremental transaction data generated after the full snapshot data is generated to the node to be updated. After the node to be updated loads the full snapshot data and incremental transaction data, the node status identifier of the node to be updated is updated to the available state.
[0073] Full snapshot data refers to the static data file generated by the current leader node after serializing its complete state machine data in memory at a specific transaction marker. Incremental transaction data generated after the full snapshot data is generated refers to the transaction logs newly generated and committed by the cluster during and after the generation of the full snapshot data. This branch serves as a fallback mechanism to handle transaction anomalies. When the difference exceeds a threshold, it indicates that the node to be updated has experienced anomalies such as local transaction loss, log corruption, or excessive lag before or during the update. In this case, local incremental synchronization is unreliable, and the existing data with local anomalies must be discarded, and the state baseline must be reset through a full snapshot. In one embodiment, when the transaction difference is determined to be greater than a preset synchronization threshold, the node to be updated initiates a snapshot synchronization request to the current leader node. Upon receiving the request, the current leader node immediately triggers a full serialization operation of the memory data, serializing its current complete database state into a binary snapshot file and recording the baseline transaction marker corresponding to the snapshot generation time. The leader node sends this full snapshot file to the node to be updated, along with the newly cached incremental transaction data generated after the snapshot. The node to be updated first clears its local old state data, deserializes and loads the full snapshot file into memory, establishing a complete baseline data state. It then continues to receive and replay incremental transaction data, catching up on all transactions newly committed after the snapshot was generated. Finally, it updates the node status flag to an available state. In another embodiment, the leader node does not generate snapshots in real time. Instead, it directly reads the latest snapshot file generated in its local data directory and sends this existing snapshot file, along with the newly generated incremental transaction log file, to the node to be updated. The node to be updated assembles and loads the complete snapshot file on its local disk, simultaneously establishing an incremental log receiving pipeline to receive and apply incremental transaction data in real time. After all data has been applied, it modifies the node status flag in its local configuration file to an available state.
[0074] Step S1445: When the node status is detected as available, the node to be updated is triggered to send a heartbeat registration request to the current leader node, so that the current leader node adds the node to be updated to the distributed cluster and restarts.
[0075] After data synchronization is completed and node status identifiers are updated, when a node status identifier is detected to be in an available state, the node to be updated is triggered to send a heartbeat registration request to the current leader node, so that the current leader node can add the node to be updated to the distributed cluster and restart operation. The heartbeat registration request is a protocol message sent by the node to be updated to the current leader node after data synchronization is complete and it has service capabilities, to declare its own survival and request to be re-integrated into the cluster routing and election system. In one embodiment, the update script continuously reads the node status identifier in the memory of the node to be updated through a polling mechanism, or queries the node's status by calling the "stat" command, checking whether the "Mode" field in the returned text has changed from "syncing" to "follower" or "observer". When an available state is detected, the underlying network communication module is invoked to send a heartbeat registration request containing the unique identifier of the node to be updated and the latest transaction identifier to the election port of the current leader node. Upon receiving the request, the current leader node verifies whether the transaction identifier in the request matches the known transaction progress. It confirms that the node to be updated has been fully synchronized and the data is consistent. Then, it adds the node to the list of active nodes and begins forwarding client read and write requests and subsequent new transactions to it, enabling the node to be updated to restart with the target runtime version.
[0076] In this embodiment, a dual data recovery mechanism combining regular incremental synchronization and full snapshot synchronization as a fallback for exceptions is constructed to completely resolve the potential for inconsistencies in transaction data after node updates. Under normal circumstances, incremental synchronization enables efficient and rapid node recovery, minimizing the impact of downtime updates on business operations. However, when faced with anomalies such as local transaction loss or log corruption occurring before or during the update, a full snapshot synchronization is automatically triggered as a fallback mechanism, forcibly resetting the node state and updating the latest data. This not only ensures recovery efficiency but also fundamentally guarantees that the transaction data state of each node in the distributed cluster is absolutely consistent with that of the cluster leader node after an update and restart, greatly improving the cluster's strong data consistency and disaster recovery capabilities under complex fault scenarios.
[0077] In a further embodiment, before step S1441, obtaining the latest transaction identifier of the current leader node in the distributed cluster, the following steps are included: Step S2400: Respond to the sandbox verification event, obtain the sandbox distributed cluster in the sandbox environment, and determine the equivalent node corresponding to the node to be updated in the sandbox distributed cluster. The sandbox distributed cluster is a replicated deployment version of the distributed cluster.
[0078] To address the risk of production cluster failure due to incompatibility between the new runtime version and existing data or cluster protocols before rejoining the node to be updated to the production cluster and performing data synchronization, this further embodiment introduces a pre-emptive sandbox verification mechanism. Specifically, although the target runtime version has been developed and verified, unforeseen compatibility issues may still arise under specific production cluster data formats, configuration parameter combinations, and actual interaction scenarios with other nodes. Directly updating the production node and attempting to synchronize it could trigger such issues, leading to repeated node crashes, data corruption, and even contaminating other normal nodes in the cluster through protocol interactions. Therefore, before irreversibly advancing the update operation to data synchronization and joining the production cluster, a full-scale simulation is conducted in a sandbox environment completely identical to the production environment to pre-expose and diagnose compatibility risks, thereby determining whether to allow the update process to continue.
[0079] This further embodiment is executed by the aforementioned server, independent of the distributed cluster. After the node to be updated performs version replacement and configuration restoration, its update script does not directly enter the data synchronization process. Instead, it sends a sandbox verification request to the server to trigger a sandbox verification event and waits for the server's response. Upon receiving the request, the server responds to the event and obtains the sandbox distributed cluster in the sandbox environment.
[0080] To ensure absolute consistency between the sandbox environment and the production environment at the time of this verification, the server completed a full replication operation of the distributed cluster in the production environment and performed an initial equivalent backup of the replicated cluster in the sandbox environment before responding to this distributed cluster update event. The full replication operation refers to the server completely copying the operating system image, all runtime library files in the ZooKeeper installation directory, complete transaction logs and snapshot files in the data directory, and all configuration files of each node in the production distributed cluster to the sandbox environment using snapshot or synchronization tools, generating a sandbox distributed cluster with the exact same topology, membership, and persistent data as the production cluster. The initial equivalent backup refers to accurately deploying the data and configurations obtained from the full replication to the corresponding nodes in the sandbox distributed cluster before the sandbox distributed cluster starts for the first time, ensuring that it is in a ready-to-start state completely consistent with the production distributed cluster at the time of replication.
[0081] In responding to this distributed cluster update event, the server controls the execution of the same node update process in the sandbox environment as in the production environment. Specifically, in the production environment, the server performs version replacement and configuration restoration on the nodes to be updated in the distributed cluster according to the aforementioned steps. Simultaneously, in the sandbox environment, the server performs the exact same update operation on the equivalent nodes with the same cluster membership in the sandbox distributed cluster, using the exact same target version data package and the exact same configuration restoration logic. Because the source node states in both environments are guaranteed to be consistent by the aforementioned full replication operation, and the subsequent update operations are completely isomorphic, the equivalent nodes in the sandbox environment and the nodes to be updated in the production environment are in the exact same state during this step—running the same target runtime version, loading the same personalized configuration, and possessing the same amount of local transaction data. Under this state consistency, a mapping relationship can be established between the two corresponding equivalent nodes in the distributed clusters of the two environments.
[0082] Based on the aforementioned state consistency, the server can obtain the sandboxed distributed cluster in its current state within the sandbox environment, and query the aforementioned mapping relationship to determine the equivalent node corresponding to the node to be updated within the sandboxed distributed cluster. This sandboxed distributed cluster is the mirror cluster constructed through full replication and synchronized update operations, completely consistent with the current state of the distributed cluster in the production environment. The equivalent node is the node in this sandboxed distributed cluster that has the exact same cluster membership, configuration, and updated status as the node to be updated in the production environment.
[0083] Step S2410: Obtain preset simulated read requests and simulated write requests to perform service verification on equivalent nodes in the sandbox distributed cluster and obtain the verification results.
[0084] Next, the server obtains preset simulated read and write requests to perform service verification on the equivalent node in the sandbox distributed cluster and obtain the verification results. The preset simulated read and write requests are a set of operation instructions manually constructed by ZooKeeper cluster operations engineers, covering the core coordination service functions of the ZooKeeper cluster. Simulated read requests include operations such as querying node status, reading configuration data under a specified path, and obtaining a list of child nodes. Simulated write requests include operations such as creating new nodes, updating existing node data, deleting nodes, and setting access control lists. The server sends these simulated read and write requests to the sandbox distributed cluster through an automated verification framework to verify whether the equivalent node, after applying the updated script of the target runtime version, can correctly handle client requests and perform normal distributed consensus interaction with other nodes in the sandbox distributed cluster. During the service verification execution, the server monitors the process status, memory usage, network throughput, and request response latency of the equivalent node in real time, and collects the response status codes and response data returned by the equivalent node. Finally, these monitoring data and response results are aggregated to generate the verification results.
[0085] During the process of obtaining preset simulated read and write requests to verify services on equivalent nodes in the sandbox distributed cluster, the server sends these requests to the client access endpoint of the sandbox distributed cluster. Upon receiving the requests, the sandbox distributed cluster coordinates the processing internally through its distributed consensus protocol. Specifically, if the request is a simulated read request, the receiving sandbox node queries its local in-memory database and returns the result; if the request is a simulated write request, the receiving sandbox node forwards it to the leader node in the sandbox distributed cluster. The leader node then initiates a transaction proposal, coordinates log replication by follower nodes, commits the transaction, and finally returns the processing result. During this internal coordination process, the server performs in-depth service monitoring of the equivalent nodes. If the equivalent node is a leader node, the server monitors its throughput for handling write requests, the latency of transaction proposals, and its heartbeat synchronization status with other nodes; if the equivalent node is a follower node, the server monitors its response time for forwarding write requests, the processing efficiency of local read requests, and the correctness of log replay. The server monitors the process status, memory usage, and network input / output throughput of the equivalent nodes in real time, and collects the global response status codes and response data returned by the sandbox distributed cluster for these simulated requests. Finally, these monitoring indicators and response results are summarized to obtain the verification results.
[0086] Step S2420: If the verification result indicates that the equivalent node is running normally in the sandbox distributed cluster, then execute the step of obtaining the latest transaction identifier of the current leader node in the distributed cluster; otherwise, skip all subsequent steps, confirm that the current runtime version of the node to be updated cannot be updated to the target runtime version, and end the update of the node to be updated.
[0087] In determining whether the verification result indicates that the equivalent node is running normally in the sandbox distributed cluster, the server performs deep analysis and rule matching on the output verification result. The server checks whether all global response status codes in the verification result are success indicators, whether the verification response data is completely consistent with the expected result, and confirms that the equivalent node has not experienced process crashes, memory overflows, network disconnections, or triggered abnormal elections during the service verification period. If the verification result meets all the above conditions, indicating that the equivalent node is running normally in the sandbox distributed cluster, the server returns a verification success confirmation instruction to the node to be updated. After receiving the confirmation instruction, the update script on the node to be updated continues to execute the subsequent steps of obtaining the latest transaction identifier of the current leader node in the distributed cluster. If the verification result contains a response status code of failure, mismatched response data, or an abnormal phenomenon in the equivalent node, indicating that the equivalent node cannot run normally in the sandbox distributed cluster, the server returns a verification failure interception instruction to the node to be updated. After receiving the interception instruction, the update script on the node to be updated immediately skips all subsequent steps, confirms that the current runtime version of the node to be updated cannot be updated to the target runtime version, and forcibly terminates the operation of updating the node to be updated.
[0088] In this embodiment, a sandbox distributed cluster that is absolutely isomorphic to the production environment in terms of data state, topology, and update progress is constructed by performing a full replication and initial equivalent backup of the production environment before responding to a distributed cluster update event, and by synchronously executing the same node update process in both environments during the response. This dual-track parallel pre-verification mechanism follows the real working principle of distributed systems, injecting simulated requests into the sandbox cluster and having it coordinate and process them internally. Without interfering with the business continuity of the production environment, it accurately monitors the actual performance of equivalent nodes in cluster consensus, data synchronization, and request processing. When verification passes, the update process in the production environment is smoothly advanced; when verification fails, the update operation is immediately intercepted and terminated. This fundamentally eliminates catastrophic consequences such as distributed cluster data corruption, cluster split-brain, or service interruption caused by blind updates, greatly improving the security and robustness of distributed cluster version updates.
[0089] In a further embodiment, after step S2420, ending the update of the node to be updated, the following steps are included: Step S2421: Obtain the runtime error logs generated by the equivalent node during the service verification process, as well as the distributed consensus protocol messages exchanged between the equivalent node and other nodes in the sandbox distributed cluster.
[0090] When the sandbox verification process determines that the node to be updated cannot be successfully updated to the target runtime version, the entire update process has been safely terminated. However, operations and maintenance personnel are faced with a "failed update" conclusion without knowing the root cause. This forces them to manually check massive amounts of logs and protocol messages, resulting in extremely low efficiency. Therefore, this further embodiment aims to solve this technical problem by automatically collecting and analyzing key diagnostic data after an update failure, and accurately determining the specific cause of the failure based on preset judgment logic. Ultimately, it generates a diagnostic report that operations and maintenance personnel can directly use, thereby improving the efficiency and accuracy of troubleshooting.
[0091] After the update operation on the node to be updated is completed, the server immediately initiates the fault diagnosis process. First, it retrieves the runtime error logs generated by the equivalent node during the service verification process, as well as the distributed consensus protocol messages exchanged between the equivalent node and other nodes in the sandbox distributed cluster.
[0092] The runtime error log refers to the text file (e.g., zookeeper.out or logs / zookeeper.log) on the local file system, containing error, warning, and exception stack traces, output by the ZooKeeper service process to the equivalent node in the sandbox environment during the entire service verification process, including startup, data loading, and handling simulated read / write requests. There are several specific implementations for obtaining this log: In one implementation, a monitoring agent deployed on the equivalent node reads all log files in the local log directory after service verification, filters entries containing keywords such as "ERROR," "FATAL," and "Exception," and packages them for return to the diagnostic server. In another implementation, after verification failure, the diagnostic server remotely logs into the equivalent node's operating system via Secure Shell (SSH), directly accesses the ZooKeeper log output directory (e.g., / var / log / zookeeper / ), and uses command-line tools (such as grep and tail) to extract exception-related log content. By reading the already written log files, no additional operations on the process status are required after a failure, thus ensuring the originality and integrity of the fault scene data.
[0093] Distributed consensus protocol messages refer to the binary data packets actually transmitted in the network link when equivalent nodes communicate with other nodes (especially leader nodes and follower nodes) within a sandbox distributed cluster based on the ZAB (ZooKeeper Atomic Broadcast) protocol. These messages are used to understand the internal working mechanisms of cluster members, such as state negotiation, data replication, and leader election. There are several ways to obtain these protocol messages: In a preferred embodiment, when building the sandbox environment, a network packet capture tool (such as tcpdump) is deployed in the host machine or container network namespace where the equivalent node is located. The filtering condition is set to capture all TCP traffic between the equivalent node's IP address and the dedicated election port (default 3888) and data synchronization port (default 2888) of other nodes in the cluster. After the service verification is completed, the captured .pcap file is provided to the diagnostic server. Since the packet capture tool is started and continuously running when deployed in the sandbox environment, its collection process does not depend on any process restart or state change after a failure, and can fully reflect the real network interaction process during the entire verification period. In another embodiment, if the ZooKeeper nodes in the sandbox environment have enabled audit or debug logging during deployment via startup parameters (such as -Dzookeeper.audit.enable=true) or configuration files, the diagnostic server can directly read these log files to extract key message types and parameters of inter-node interactions. However, it should be noted that enabling debug logging increases disk I / O and storage overhead and may affect performance. Therefore, it is recommended to pre-deploy network packet capture tools in the sandbox environment to ensure comprehensive diagnostic data collection and non-intrusiveness to the original fault scene.
[0094] Step S2422: Extract abnormal feature information from the runtime error log and perform protocol parsing based on the distributed consensus protocol message to obtain the parsing result.
[0095] Anomaly signature information is key structured data extracted from unstructured runtime error log text through methods such as pattern matching, representing specific fault types. For example, from an exception message like "java.io.StreamCorruptedException: invalid stream header," "deserialization exception" can be extracted as an anomaly signature. Specifically, an exception pattern library can be predefined, where each item contains a regular expression (e.g., .*StreamCorruptedException.*) and a corresponding feature label (e.g., "deserialization exception"). The server scans the error log line by line, matching the log content against the regular expressions in the library; once a match is found, the corresponding feature label is extracted. Another example is that if the log contains "java.io.IOException: CRC check failed," the feature "cyclic redundancy check failed" (CRC check failed) can be extracted.
[0096] Protocol parsing refers to the process of decoding and analyzing distributed consensus protocol messages in binary or text form layer by layer. Its purpose is to understand the negotiation process between nodes and reveal the nature of conflicts. In one embodiment, the server has a built-in ZAB protocol parsing engine. For captured binary messages, it first reassembles the data stream according to the TCP protocol, and then, according to the ZAB protocol, parses out key fields such as message type (e.g., proposal, confirmation, commit), node ID, epoch number, and transaction identifier (zxid). For text messages output in debug log format, the parsing engine uses regular expressions to match the text and extract the same key fields. Ultimately, the parsing result is a series of structured records, each containing a timestamp, source node, target node, message type, and the specific parameter values parsed.
[0097] Step S2423: If the abnormal feature information indicates that the equivalent node triggers a deserialization exception when parsing the local full snapshot file, or if the equivalent node fails to trigger cyclic redundancy check when reading the local incremental transaction log, then it is determined that the reason for the inability to update is that the data storage format of the target runtime version is incompatible with the existing data in the local data directory of the node to be updated.
[0098] Deserialization exceptions typically occur after the equivalent node starts and attempts to load a full snapshot file (e.g., a file named snapshot.xxxxx) from the local data directory. This is because newer code versions have modified the serialization / deserialization format of the snapshot data (e.g., class definitions, field order, or data types), making it impossible to correctly restore the old binary snapshot file into a data object in memory. It should be noted that while ZooKeeper's official releases generally maintain backward compatibility with snapshot file formats, such incompatibility issues can still arise when there are significant version differences (e.g., across multiple major versions) or when using customized versions.
[0099] Cyclic redundancy check (CRUD) failure occurs when an equivalent node attempts to replay the local incremental transaction log (i.e., the transaction log file, such as log.xxxxx) to catch up on pending transactions. The new version's code's log file read verification logic determines that the log file cannot pass the integrity check. In actual production environments, such failures are more often caused by physical disk damage, file system anomalies leading to incomplete log files, or the old version's transaction log containing data records deemed non-compliant under the new version's verification framework, causing the new version's code to fail during verification. Therefore, if either of these two characteristics is recognized by the server, the root cause of the problem is pointed to the incompatibility between the "data storage format" and the "existing data." For example, if the new version changes the node data in the snapshot from Java's native serialization method to another serialization format, reading the old format snapshot file will inevitably trigger a deserialization exception.
[0100] Step S2424: If the parsing result indicates that there is a term number conflict between the equivalent node and the leader node in the sandbox distributed cluster, or indicates that the minimum committed transaction identifier that the leader node in the sandbox distributed cluster needs to synchronize is greater than the maximum global transaction identifier stored locally by the equivalent node, and the leader node does not support log truncation operation, then it is determined that the reason for the inability to update is that the distributed consensus protocol of the target runtime version is incompatible with the distributed consensus protocol of the unupdated node in the sandbox distributed cluster.
[0101] Regarding term number conflicts: In the ZAB protocol, the term number (epoch) is a monotonically increasing integer used to identify the term of the leader node, preventing the previous leader from continuing to cause trouble. If the analysis results show that when the equivalent node (as a follower) initiates a connection to the leader of the sandbox cluster, its recorded current term number (obtained from its local transaction log) contradicts the term number claimed by the leader node. For example, the equivalent node believes it is currently in term 3, while the leader node (possibly elected from an older version of the code) claims it is in term 2. In this case, according to the ZAB protocol's operating mechanism, if the acceptedEpoch (i.e., term 3) carried by the equivalent node is greater than the leader node's currentEpoch (i.e., term 2), the leader node should update its currentEpoch to the new value and re-initiate an election to attempt to unify the cluster to the new term. However, if the new version of the protocol running on the equivalent node has stricter validation logic during the election process than the old version (for example, requiring followers to include an election payload in a specific format to recognize a new term), then both sides will repeatedly trigger elections and fail to enter a normal synchronization state, causing the cluster to fall into election oscillation, and the recovery process of the equivalent node will never be completed. The parsing results will clearly capture this conflict phenomenon; for example, the acceptedEpoch field in the FOLLOWERINFO message sent by the follower is 3, while the currentEpoch field in the LEADERINFO message replied by the leader is 2.
[0102] For nodes where the smallest committed transaction ID is greater than the equivalent node's local largest global transaction ID, and the leader does not support log truncation, when a node recovers data, the leader typically determines what type of synchronization data to send to the follower based on the local largest committed transaction ID (lastZxid) carried in the FOLLOWERINFO message sent by the follower. Specifically, if the follower's lastZxid falls within the valid range of the leader's local transaction log, incremental logs are sent in DIFF mode; if the follower lags too far behind, and its lastZxid is less than the leader's local smallest zxid of a transaction log not covered by a snapshot, the leader typically needs to send a full snapshot in SNAP mode to the follower; if the follower's lastZxid is greater than the leader's local lastZxid, log truncation is performed in TRUNC mode. If the leader node determines that the equivalent node is significantly behind based on the old version protocol (i.e., lastZxid is less than the leader's local minimum valid transaction log zxid), but the new version follower requests a synchronization mode combination during synchronization negotiation that the old version leader cannot provide or that is not recognized by the old version protocol—for example, the new version requires the old version leader to perform log truncation under certain specific conditions before synchronization can begin—and the old version leader's ZAB protocol implementation lacks or has inconsistencies in its handling logic for this scenario, this will lead to synchronization negotiation failure. The parsing results will capture the leader responding with an error message indicating "unable to process" after receiving the synchronization request from the equivalent node, or directly disconnecting, thus clearly indicating the root cause of the failure due to protocol-level incompatibility.
[0103] Step S2425: Generate a version update diagnostic report based on the reason for the inability to update and return it.
[0104] After identifying the specific type of incompatibility, the server structures and integrates key information from the entire diagnostic process to generate a human-readable version update diagnostic report. This report may include the following sections: a conclusion section, clearly stating whether the "failure to update" is due to "incompatible data storage formats" or "incompatible distributed consensus protocols"; an evidence section, listing specific log snippets (such as the stack trace of the line that triggered the deserialization exception) and protocol parsing results (such as message details showing term number conflicts); and a recommendation section, providing solutions for different causes, such as recommending using data migration tools for format conversion for data incompatibility, and recommending deploying compatibility patches on some nodes in the cluster for protocol incompatibility. Once generated, the report can be returned to the user who triggered the distributed cluster update event via email, push notifications, or directly displayed on the graphical user interface of the operations and maintenance platform.
[0105] In this embodiment, when sandbox verification fails, two deep-seated technical faults—"incompatible data storage formats" and "incompatible distributed consensus protocols"—are automatically and accurately located, achieving a leap from "knowing only the failure" to "knowing the root cause of the failure." The entire process is automated, requiring no manual intervention. By collecting runtime error logs and network protocol messages, extracting abnormal features, and analyzing protocol interaction details, a structured diagnostic report is generated based on pre-set, hard-sense judgment logic that accurately characterizes specific faults. This allows maintenance personnel to take direct remedial actions. This significantly improves operational efficiency, and the accuracy and certainty of the diagnostic results effectively avoid secondary risks caused by misjudgments, significantly enhancing the controllability and robustness of the distributed cluster version update process.
[0106] In a further embodiment, after step S1400, updating the current runtime version of the node to be updated to the target runtime version, the following steps are included: Step S1500: Respond to the version rollback event, obtain the node to be rolled back corresponding to the event, and the preset cluster backup data. The cluster backup data is constructed by obtaining the full configuration data of each node in the distributed cluster under the old runtime version before responding to the distributed cluster update event.
[0107] A version rollback event is an emergency processing signal triggered under specific conditions after the current runtime version of the node to be updated has been updated to the target runtime version. The triggering conditions for this event can be diverse, aiming to quickly respond to various production anomalies. In one embodiment, when operations personnel discover on the monitoring platform that the service performance indicators of the updated node, such as average request response latency or transactions per second, exceed preset business thresholds for multiple consecutive sampling periods, they can manually trigger the event by clicking the "One-Click Rollback" button on the graphical user interface of the operations management platform. In another embodiment, the updated node itself has automatic health check capabilities. Its deployed health check agent continuously calls local four-letter commands (such as srvr or mntr). Once it detects an abnormal status reported by the process itself, memory usage exceeding a dangerous level, or critical error information output, it automatically generates and sends a version rollback event notification to the server. This event information explicitly includes the unique identifier of the node that needs to perform the rollback operation. This identifier can be its Internet Protocol address, hostname, or its unique server identifier in the cluster; this node is the node to be rolled back. The old runtime version refers to the old runtime version that each node in the distributed cluster was running before updating to the target runtime version.
[0108] The server responds to the version rollback event and retrieves pre-set cluster backup data. This cluster backup data, as revealed in the preceding steps, is a dataset constructed by fully collecting and encapsulating the critical configuration data of each node in the distributed cluster before the formal start of the update process in response to this distributed cluster update event. Specifically, constructing this cluster backup data can be achieved by the server automatically traversing the entire node list of the distributed cluster before the update event is triggered, remotely logging into each node via Secure Shell (SSH) to read the full configuration data actually used under the old runtime version. The full configuration data does not refer to a single main configuration file (such as zoo.cfg), but rather includes the complete set of configuration information that ensures node identity, behavior, and communication. In one embodiment, the configuration data set includes: a core configuration file, zoo.cfg, which defines parameters such as the client port, a list of cluster member server addresses, data and log storage directory paths, and snapshot and transaction log trigger thresholds; a node identity file, myid, typically located in the data directory, containing a unique integer used to identify itself in the cluster's configured server list; and a ZooKeeper Java Virtual Machine startup parameter file, java.env or zookeeper-env.sh, which sets key runtime environment parameters such as heap memory size and garbage collector type. In another embodiment, if a node uses dynamic configuration (such as reconfig), these dynamic configuration items stored in ZooKeeper's own data tree are also exported and recorded. The server further packages this full configuration data, indexed by node identifiers and obtained from each node, into a structured dataset and stores it in the server's own reliable non-volatile storage; this is the preset cluster backup data. When a rollback event occurs, the server can directly and accurately retrieve this dataset from local storage.
[0109] Step S1510: Based on the cluster backup data, revert the current runtime version of the node to be rolled back to the old runtime version.
[0110] After obtaining the cluster backup data, the server initiates an automated rollback operation for the node to be rolled back. The core objective of this process is to completely restore all runtime files, dependent libraries, and configuration files on the node to be rolled back to the old runtime version state before the update began, and then rejoin the cluster with the old version.
[0111] First, stop the service processes of the target runtime version running on the node to be rolled back. In one embodiment, the server remotely executes a stop script in the ZooKeeper installation directory on the node to be rolled back via SSH, such as executing the command `bin / zkServer.sh stop`, and waits for the process to exit safely and release the port it occupies. In another embodiment, if the node service is managed by a process management tool such as systemd or supervisor, the server calls the corresponding tool's service stop command, such as `systemctl stop zookeeper.service`.
[0112] After the service process stops, the runtime library files are replaced. The server needs to obtain the old version data package corresponding to the old runtime library version. This old version data package is an archive file containing all executable programs and dependent libraries of the old version of ZooKeeper. The source of this data package can be implemented in different ways. In one embodiment, the server maintains its own version repository, storing archive packages of all versions used in the cluster's history. The server directly retrieves and reads the corresponding old version data package from the local version repository based on the old version number reported by the node to be rolled back before the update. In another embodiment, during the preparation phase before the first update of the cluster, the update script has already packaged the old version runtime library currently running on the cluster as part of a backup and stored it in a designated directory on the server. In this case, the server can directly read the backup archive package.
[0113] Next, the server pushes the acquired old version data package to the temporary directory of the node to be rolled back via a secure file transfer channel (such as SCP or SFTP protocol). Then, the decompression and replacement commands are remotely executed on the node. In one embodiment, all current new version files in the ZooKeeper installation directory on the node are first cleared or moved to a backup directory with a "_rollback" suffix to ensure a clean and traceable replacement environment. Then, all files of the decompressed old version runtime library are recursively copied to the ZooKeeper installation root directory. In another, more efficient embodiment, if the node adopts a "backup first, overwrite later" strategy during the update, i.e., the old version files are still stored in a timestamped local backup directory, the server can directly send commands to use operating system commands (such as rsync or cp) to overwrite all files in the local backup directory back to the installation directory, without needing to remotely transfer the entire old version data package.
[0114] After rolling back the runtime files, a crucial step is restoring the configuration files. The server precisely extracts the full configuration data corresponding to the unique identifier of the node to be rolled back from the preset cluster backup data. Based on this configuration data, the server remotely performs a configuration file rewrite operation on the node to be rolled back. Specifically, in one embodiment, the server completely overwrites the content of the zoo.cfg file in the backup with the file of the same name in the ZooKeeper configuration directory of the node to be rolled back; it writes the content of the myid file in the backup to the myid file in the node's data directory; and it also overwrites the content of environment-related files such as java.env. This process strictly ensures that after the node is restored, all configuration parameters, such as its cluster membership, communication address with other nodes, and local data path, are completely consistent with those before the update, eliminating any risk that the node will not be able to integrate into the cluster due to configuration residues or mismatches.
[0115] Finally, restart the node to be rolled back and add it to the cluster. The server starts the old version of the ZooKeeper process by remotely executing a startup script (such as zkServer.sh start) or calling the process management service. After the node process starts, it will contact other members in the cluster according to the restored configuration file. Since the full snapshot and incremental transaction log in the node's local data directory are still the existing data before its shutdown and rollback (this data still exists after the update and can be parsed by the old version of the code), it can use this local data to perform regular data recovery and synchronization with the current cluster leader node, catching up on the few transactions missed during the failure and rollback. Once the data synchronization is complete, the node is restored to normal operation as the old runtime version and provides services.
[0116] In this embodiment, a complete and reliable two-way version change closed loop is constructed by deeply coupling the one-time update operation with the always-on version rollback operation. The key is that a precise old version configuration snapshot of each node—i.e., cluster backup data—is captured and persistently stored at the beginning of the update. When a rollback is triggered, this mechanism is executed fully automatically. Using this original configuration snapshot, not only are runtime library files accurately replaced with the old version, but the most critical configuration files that determine a node's position in the cluster are precisely restored "as is." This fundamentally avoids the configuration errors that are easily caused by manual rollback, achieving true "one-click restoration." The entire process is fast, accurate, and requires no manual intervention, greatly reducing business interruption time caused by new version issues. It provides a solid last line of defense for large-scale, routine, and automated updates of distributed clusters, significantly improving the overall resilience and operational reliability of the distributed system.
[0117] Please see Figure 2This application provides a distributed cluster update apparatus to meet one of its objectives. It is a functional embodiment of the distributed cluster update method of this application. The apparatus includes an event response module 1100, a node analysis module 1200, a version analysis module 1300, and a version update module 1400. The event response module 1100 is used to respond to distributed cluster update events and determine the node to be updated, its current runtime version, the target runtime version, and its role label. The node analysis module 1200 is used to determine whether the node to be updated is a leader node based on its role label if the current runtime version is lower than the target runtime version. The version analysis module 1300 is used to detect whether the current runtime version of all other nodes in the distributed cluster, excluding the node to be updated, is the target runtime version if the current runtime version of all other nodes is the target runtime version. The version update module 1400 is used to update the current runtime version of the node to be updated to the target runtime version if the current runtime version of all other nodes is the target runtime version.
[0118] In a further embodiment, the node analysis module 1200 includes: an information acquisition submodule, used to send role detection requests to all nodes in the distributed cluster except the node to be updated, and acquire the cluster leader information returned by these nodes; a node judgment submodule, used to determine whether the role label represents a leader role, and whether all cluster leader information points to the node to be updated; and a node parsing submodule, used to determine that the node to be updated belongs to the leader node if the role label represents a leader role and all cluster leader information points to the node to be updated; otherwise, determine that the node to be updated does not belong to the leader node.
[0119] In a further embodiment, the version update module 1400 includes: a data acquisition submodule, used to stop the node to be updated and acquire the target version data packet and its hash data corresponding to the target runtime version; a data packet verification submodule, used to determine whether the hash value of the data packet corresponding to the target version data packet is consistent with the source hash value in the hash data; a version update submodule, used to update the version configuration data corresponding to the current runtime version of the node to be updated based on the target version data packet if the hash value of the data packet is consistent with the source hash value, and read the configuration file of the node to be updated from the preset cluster backup data and configure it for the node; and a node reload submodule, used to re-add the node to be updated to the distributed cluster and run it again, and determine that the current runtime version of the node to be updated after running is the target runtime version.
[0120] In a further embodiment, the node reload submodule includes: an identifier acquisition unit, used to acquire the latest transaction identifier of the current leader node in the distributed cluster, and the latest transaction identifier stored locally by the node to be updated; a difference determination unit, used to determine the transaction difference between the latest transaction identifier of the current leader node and the latest transaction identifier stored locally by the node to be updated; an incremental synchronization unit, used to trigger the node to be updated to synchronize incremental transaction data with the current leader node if the transaction difference is less than or equal to a preset synchronization threshold, and update the node status identifier of the node to be updated to an available state after the incremental transaction data synchronization is completed; a full synchronization unit, used to trigger the current leader node to send full snapshot data and incremental transaction data generated after the full snapshot data is generated to the node to be updated if the transaction difference is greater than the preset synchronization threshold, and update the node status identifier of the node to be updated to an available state after the node to be updated loads the full snapshot data and incremental transaction data; and a node reload unit, used to trigger the node to be updated to send a heartbeat registration request to the current leader node when the node status identifier is detected to be available, so that the current leader node adds the node to be updated to the distributed cluster and restarts it.
[0121] In a further embodiment, before the identifier acquisition unit, the system includes: a sandbox verification unit, used to respond to a sandbox verification event, acquire the sandbox distributed cluster in the sandbox environment, and determine the equivalent node corresponding to the node to be updated in the sandbox distributed cluster, wherein the sandbox distributed cluster is a replicated deployment version of the distributed cluster; a node verification unit, used to acquire preset simulated read requests and simulated write requests to perform service verification on the equivalent node in the sandbox distributed cluster and obtain the verification result; and a route execution unit, used to execute the step of acquiring the latest transaction identifier of the current leader node in the distributed cluster if the verification result indicates that the equivalent node is running normally in the sandbox distributed cluster; otherwise, skip all subsequent steps, confirm that the current runtime version of the node to be updated cannot be updated to the target runtime version, and end the update of the node to be updated.
[0122] In a further embodiment, after the split execution unit, there is: an information acquisition unit, used to acquire runtime error logs generated by the equivalent node during service verification, and distributed consensus protocol messages interacting between the equivalent node and other nodes in the sandbox distributed cluster; an information parsing unit, used to extract abnormal feature information from the runtime error logs and perform protocol parsing based on the distributed consensus protocol messages to obtain the parsing result; and a first attribution unit, used to determine that the reason for the inability to update is the target runtime library if the abnormal feature information indicates that the equivalent node triggered a deserialization exception when parsing the local full snapshot file, or indicates that the equivalent node triggered a cyclic redundancy check failure when reading the local incremental transaction log. The data storage format of the version is incompatible with the existing data in the local data directory of the node to be updated; the second attribution unit is used to determine that the reason for the inability to update is that the distributed consensus protocol of the target runtime version is incompatible with the distributed consensus protocol of the unupdated nodes in the sandbox distributed cluster if the parsing result indicates that there is a term number conflict between the equivalent node and the leader node in the sandbox distributed cluster, or indicates that the minimum committed transaction identifier required to be synchronized by the leader node in the sandbox distributed cluster is greater than the maximum global transaction identifier stored locally by the equivalent node, and the leader node does not support log truncation operation; the report return unit is used to generate a version update diagnostic report based on the reason for the inability to update and return it.
[0123] In a further embodiment, after the version update module 1400, there is a node rollback module, used to respond to a version rollback event, obtain the node to be rolled back corresponding to the event, and preset cluster backup data. The cluster backup data is constructed by obtaining the full configuration data of each node in the distributed cluster running under the old runtime version before responding to the distributed cluster update event. The version rollback module is used to roll back the current runtime version of the node to be rolled back to the old runtime version based on the cluster backup data.
[0124] To address the aforementioned technical problems, embodiments of this application also provide computer equipment. For example... Figure 3 The diagram shows the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, data, and computer-readable instructions. The data may store a sequence of control information. When the computer-readable instructions are executed by the processor, they enable the processor to implement a distributed cluster update method. The processor of the computer device provides computing and control capabilities, supporting the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When these computer-readable instructions are executed by the processor, they enable the processor to execute the distributed cluster update method of this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0125] In this embodiment, the processor is used to execute... Figure 2 The system contains the specific functions of each module and its sub-modules. The memory stores the program code and various data required to execute these modules or sub-modules. The network interface is used for data transmission between user terminals and the server. In this embodiment, the memory stores the program code and data required to execute all modules / sub-modules in the distributed cluster update device of this application. The server can call the server's program code and data to execute the functions of all sub-modules.
[0126] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the distributed cluster update method of any embodiment of this application.
[0127] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM).
[0128] In summary, this application can efficiently and accurately update distributed clusters.
[0129] Those skilled in the art will understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in this application can be alternated, modified, combined, or deleted. Furthermore, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted. Furthermore, steps, measures, and solutions in the prior art that are similar to those disclosed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted.
[0130] The above are only some embodiments of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for distributed cluster update, the method comprising: Includes the following steps: In response to a distributed cluster update event, determine the node to be updated, its current runtime version, the target runtime version, and its role tag. If the current runtime version is detected to be lower than the target runtime version, then it is determined whether the node to be updated belongs to the leader node based on the role tag; If the node to be updated belongs to the leader node, then check whether the current runtime version of all other nodes in the distributed cluster, excluding the node to be updated, is the target runtime version. If the current runtime version of all other nodes is the target runtime version, then the current runtime version of the node to be updated will be updated to the target runtime version.
2. The distributed cluster update method according to claim 1, characterized in that, Determining whether the node to be updated belongs to the leader node based on the role tag includes the following steps: Send role detection requests to all nodes in the distributed cluster except for the node to be updated, and obtain the cluster leader information returned by these nodes; Determine whether the role label represents a leader role, and whether all cluster leader information points to the node to be updated; If the role label represents the leader role, and all cluster leader information points to the node to be updated, then the node to be updated is determined to be a leader node; otherwise, the node to be updated is determined not to be a leader node.
3. The distributed cluster update method according to claim 1, characterized in that, Updating the current runtime library version of the node to be updated to the target runtime library version includes the following steps: Stop running the node to be updated, and obtain the target version data package and its hash data corresponding to the target runtime version; Determine whether the hash value of the data packet corresponding to the target version data packet is consistent with the source hash value in the hash data; If the hash value of the data packet is consistent with the hash value of the source, then the version configuration data corresponding to the current runtime version of the node to be updated is updated based on the target version data packet, and the configuration file of the node to be updated is read from the preset cluster backup data and configured for the node. The node to be updated is rejoined to the distributed cluster and run again, and the current runtime library version of the node to be updated after running is determined to be the target runtime library version.
4. The distributed cluster update method according to claim 3, characterized in that, The process of re-adding the node to be updated to the distributed cluster and restarting the process includes the following steps: Obtain the latest transaction identifier of the current leader node in the distributed cluster, and the latest transaction identifier stored locally on the node to be updated; Determine the transaction difference between the latest transaction identifier of the current leader node and the latest transaction identifier stored locally on the node to be updated; If the transaction difference is less than or equal to the preset synchronization threshold, the node to be updated is triggered to perform incremental transaction data synchronization with the current leader node, and after the incremental transaction data synchronization is completed, the node status identifier of the node to be updated is updated to the available state. If the transaction difference is greater than the preset synchronization threshold, the current leader node is triggered to send full snapshot data and incremental transaction data generated after the full snapshot data is generated to the node to be updated. After the node to be updated loads the full snapshot data and incremental transaction data, the node status identifier of the node to be updated is updated to the available state. When the node status is detected to be in the available state, the node to be updated is triggered to send a heartbeat registration request to the current leader node, so that the current leader node adds the node to be updated to the distributed cluster and restarts it.
5. The distributed cluster update method according to claim 4, characterized in that, Before obtaining the latest transaction identifier of the current leader node in the distributed cluster, the following steps are included: In response to the sandbox verification event, obtain the sandbox distributed cluster in the sandbox environment, and determine the equivalent node corresponding to the node to be updated in the sandbox distributed cluster. The sandbox distributed cluster is a replicated deployment version of the distributed cluster. Obtain preset simulated read requests and simulated write requests to perform service verification on equivalent nodes in the sandbox distributed cluster, and obtain the verification results; If the verification result indicates that the equivalent node is running normally in the sandbox distributed cluster, then the step of obtaining the latest transaction identifier of the current leader node in the distributed cluster is executed; Otherwise, skip all subsequent steps, confirm that the current runtime version of the node to be updated cannot be updated to the target runtime version, and end the update of the node to be updated.
6. The distributed cluster update method according to claim 5, characterized in that, After finishing the update of the node to be updated, the following steps are included: Obtain the runtime error logs generated by the equivalent node during the service verification process, as well as the distributed consensus protocol messages exchanged between the equivalent node and other nodes in the sandbox distributed cluster; Extract the abnormal feature information from the runtime error log, and perform protocol parsing based on the distributed consensus protocol message to obtain the parsing result; If the abnormal feature information indicates that the equivalent node triggers a deserialization exception when parsing the local full snapshot file, or indicates that the equivalent node triggers a cyclic redundancy check failure when reading the local incremental transaction log, then it is determined that the reason for the inability to update is that the data storage format of the target runtime version is incompatible with the existing data in the local data directory of the node to be updated. If the parsing result indicates that there is a term number conflict between the equivalent node and the leader node in the sandbox distributed cluster, or indicates that the minimum committed transaction identifier that the leader node in the sandbox distributed cluster needs to synchronize is greater than the maximum global transaction identifier stored locally by the equivalent node, and the leader node does not support log truncation operation, then it is determined that the reason for the inability to update is that the distributed consensus protocol of the target runtime version is incompatible with the distributed consensus protocol of the unupdated node in the sandbox distributed cluster. A version update diagnostic report is generated and returned based on the reasons for the inability to update.
7. The distributed cluster update method according to claim 1, characterized in that, After updating the current runtime version of the node to be updated to the target runtime version, the following steps are included: In response to a version rollback event, the corresponding node to be rolled back and the preset cluster backup data are obtained. The cluster backup data is constructed by obtaining the full configuration data of each node in the distributed cluster under the old runtime version before responding to the distributed cluster update event. Based on the cluster backup data, the current runtime library version of the node to be rolled back will be rolled back to the old runtime library version.
8. A distributed cluster update device, characterized in that, include: The event response module is used to respond to distributed cluster update events and determine the node to be updated, its current runtime version, the target runtime version, and the role tag corresponding to the event. The node analysis module is used to determine whether the node to be updated belongs to the leader node based on the role tag if the current runtime version is detected to be lower than the target runtime version. The version analysis module is used to detect whether the current runtime version of all nodes in the distributed cluster other than the node to be updated is the target runtime version if the node to be updated belongs to the leader node. The version update module is used to update the current runtime version of the node to be updated to the target runtime version if the current runtime version of all other nodes is the target runtime version.
9. A computer device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 7, which, when invoked by a computer, executes the steps included in the corresponding method.