An ATS primary and backup center hot backup redundancy system based on an optimized bully algorithm

CN122607395APending Publication Date: 2026-08-21TIANJIN JINHANG INTELLIGENT CONTROL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610336389.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-19
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

①不支持“设备级云化”,无法实现OCC和BOCC设备交叉服务运营

Benefits of technology

本发明提出一种基于优化Bully算法实现的ATS主备中心热备冗余系统,本发明的关键点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122607395A_ABST
    Figure CN122607395A_ABST
Patent Text Reader

Abstract

The application relates to an ATS master and standby center hot standby redundancy system based on an optimized Bully algorithm, and belongs to the field of rail transit. OCC and BOCC contain application server A / B machines, database server A / B machines, interface server A / B machines and dispatching workstations; wherein the hot standby devices are application servers, database servers and interface servers; the server A / B machines of the OCC and the server A / B machines of the BOCC form a cluster, all the devices in the cluster perform UDP communication, and the communication structure is one-to-many, that is, each device communicates with all the devices in the cluster. The application supports OCC and BOCC device cross operation, improves signal system center business service availability, solves the problem that when double master centers appear, signal system ATS related businesses may be confused, reduces the workload of OCC and BOCC hot standby switching personnel, reduces the possible failure rate and improves the system fault tolerance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of rail transit, specifically relating to an ATS primary / backup hot standby redundancy system based on an optimized Bully algorithm. Background Technology

[0002] Currently, some subway lines in China typically have two control centers: a primary control center (OCC) and a backup control center (BOCC). An ATS OCC generally includes central application server A / B machines, database server A / B machines, central interface server A / B machines, communication front-end server A / B machines, and several dispatch workstations. An ATS BOCC may include central application server A / B machines, database server A / B machines, central interface server A / B machines, communication front-end server A / B machines, and several dispatch workstations, or it may only contain a single server.

[0003] In the initial stages of proposing the primary and backup center configuration, the establishment of the BOCC (Built-in-Card Center) mainly considered scenarios of OCC (Operational Control Center) equipment failure and disaster recovery, so that the BOCC could ensure subway operation in the event that the OCC is completely unavailable (such as OCC power outage, fire, etc.). With the gradual popularization of fully automated subway operation systems (FAO), train operation command is centralized in the control center, and the role of the BOCC in subway operation has become even more important.

[0004] The existing implementations most similar to this invention are: (1) Hot standby redundancy management system for main and backup control centers of rail transit This solution adds hardware controllers to both the OCC and BOCC to enable primary / backup switching between them. Essentially, it connects the controllers to the power supplies of the OCC and BOCC's egress switches, using the controllers to control the switches' power supply to allow either the OCC or BOCC to access the mains network, thus indirectly achieving the goal of switching between the OCC and BOCC.

[0005] (2) Switching scheme between the control center and the backup control center of the fully automatic operation system Two solutions are proposed. One is to switch over the OCC and BOCC as a whole, but the operator needs to develop a targeted switching process, clarify the switching timing, and the operation organization and maintenance response strategy when applying it. The other is to make the switching of the signal system's OCC and BOCC independent of the switching of external systems, which can improve flexibility and reduce the scope of impact to a certain extent.

[0006] Disadvantages of existing technical solutions: ① It does not support "device-level cloudification" and cannot achieve cross-service operation of OCC and BOCC devices.

[0007] ②The switching process between OCC and BOCC takes a long time, which cannot guarantee the continuity of ATS services.

[0008] ③Adding extra equipment to achieve hot standby for OCC and BOCC increases the potential points of equipment failure.

[0009] ④ OCC / BOCC only supports fixed device configurations, such as only supporting 2x2 application servers and not supporting flexible configurations such as 2x1, making it difficult to expand the device cluster.

[0010] ⑤ Central cluster expansion is not supported.

[0011] ⑥ There may be two main data centers, which could cause confusion in ATS-related services and lead to misalignment of ATS planning adjustments.

[0012] ⑦ It does not support extreme scenarios where the server is in OCC and the scheduling workstation is in BOCC.

[0013] ⑧ When switching between OCC and BOCC, there is a high level of personnel involvement and a high error rate.

[0014] ⑨ The method of enabling OCC and BOCC to access the backbone network by controlling the power supply of the switch cannot guarantee uninterrupted continuous operation due to the long startup time of the switch. Summary of the Invention

[0015] (a) Technical problems to be solved The technical problem to be solved by this invention is how to provide an ATS primary and backup hot standby redundancy system based on an optimized Bully algorithm, so as to solve the above-mentioned problems of primary and backup redundancy between OCC and BOCC in the rail transit field.

[0016] (II) Technical Solution To address the aforementioned technical problems, this invention proposes an ATS primary / standby center hot standby redundancy system based on an optimized Bully algorithm. This system includes: OCC and BOCC. The equipment included in OCC and BOCC consists of application server A / B machines, database server A / B machines, interface server A / B machines, and scheduling workstations; among them, the hot standby equipment includes application servers, database servers, and interface servers; the combination of server equipment in OCC and BOCC forms a computing resource platform to provide computing support for business operations. The application servers A / B of OCC and the application servers A / B of BOCC form a cluster. The database servers A / B of OCC and the database servers A / B of BOCC form a cluster. The interface servers A / B of OCC and the interface servers A / B of BOCC form a cluster. Internal communication within the cluster is achieved through an internal network communication topology. All devices within the cluster communicate via UDP. The communication structure is one-to-many, meaning that each device communicates with all other devices in the cluster.

[0017] (III) Beneficial Effects This invention proposes an ATS primary / standby hot standby redundancy system based on an optimized Bully algorithm. The key points of this invention are: (1) OCC and BOCC are “device-level cloudification”, supporting cross-service operation of OCC and BOCC devices.

[0018] (2) Master-slave state decision-making method based on optimized Bully algorithm.

[0019] (3) Hot standby management method for master-slave centers based on optimized Bully algorithm.

[0020] Advantages of this invention: (1) Achieve “device-level cloudification” for OCC and BOCC, with high scalability. Support cross-operation of OCC and BOCC equipment to improve the availability of signal system center business services.

[0021] (2) To resolve the problem of potential confusion in ATS-related services due to the emergence of dual master centers.

[0022] (3) Reduce the workload of personnel involved in OCC and BOCC hot standby switching, reduce the possible error rate, and improve the system fault tolerance.

[0023] (4) No additional equipment is required to achieve hot standby for OCC and BOCC, reducing the probability of failure. Attached Figure Description

[0024] Figure 1 This is a diagram of the OCC and BOCC device architecture of the present invention; Figure 2 This is a block diagram of the hot standby equipment nodes; Figure 3 This is a diagram of the internal network topology of the application server. Figure 4 This is an overall framework diagram of the present invention; Figure 5 Flowchart for initializing host election; Figure 6 Example diagram of the election and promotion process; Figure 7 A schematic diagram of the master / standby decision state machine. Detailed Implementation

[0025] To make the objectives, contents, and advantages of the present invention clearer, the specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples.

[0026] Terminology Explanation: OCC: Operation Control Center (Primary Control Center) BOCC: Backup Operation Control Center DDOCC: Dynamic Debugging Operation Control Center XXOCC: XX Operation Control Center (An unexpectedly added control center) Bully: A Distributed Election Algorithm FAO: Fully Automatic Operation System ATS: Automatic Train Supervision Hot standby is a technical term in the IT and system operations field, referring to keeping redundant devices or services in a ready state in real time during system operation so that they can immediately take over work when the main system fails, ensuring business continuity and high availability.

[0027] Main machine: The machine that provides business services to the outside world. For example, if there are two OCC application servers, the machine that provides services to the outside world is defined as the main machine, and the redundant / ready machine is defined as the standby machine.

[0028] Primary group: The group that provides external services is the primary group. For example, the OCC application server and the BOCC application server are two groups. The service is the primary group, and the redundant / ready state is defined as the backup group.

[0029] Primary group host: The host in the primary group that provides services to the outside world.

[0030] Device cluster: A collection of devices consisting of multiple groups of servers that provide a specified computing service to the outside world.

[0031] Device-level cloudification: Multiple device clusters are combined into a resource platform, which provides a set of computing services to the outside world.

[0032] Challenge lifecycle: refers to the valid duration for a challenge sent by node A. Typically, a challenge has a valid duration of N seconds, within which the challenged node must process the challenge.

[0033] The purpose of this invention: ① Achieve “device-level cloudification” for OCC and BOCC, support cross-operation of OCC and BOCC equipment, and improve the availability of signal system center business services.

[0034] ②Reduce the switching time between OCC and BOCC to ensure the continuity of ATS services in the signaling system.

[0035] ③ To resolve the potential confusion in ATS-related services caused by the emergence of dual master centers.

[0036] ④ Improve the flexibility of OCC and BOCC equipment configuration ⑤ Improve the scalability of the central cluster ⑥ Reduce the workload of personnel involved in OCC and BOCC hot standby switching, reduce the possible error rate, and improve the system fault tolerance.

[0037] To address the problems in the background art, the present invention provides the following solutions: OCC and BOCC server devices cannot be "cloudified at the device level," thus failing to meet the requirements for cross-operation. This solution can meet the requirements for cross-operation. OCC and BOCC servers cannot be cross-operated, which cannot maximize the availability of the ATS control center. This solution maximizes the availability of the ATS control center by enabling cross-operation of the servers. As the OCC and BOCC server clusters expand, maintaining consistency across servers becomes increasingly difficult. This solution addresses this by employing a distributed consensus algorithm based on optimized Bully to manage the server cluster state, thereby maximizing server state consistency. This invention uses an optimized Bully algorithm to implement a hot standby management solution for control center clusters, which is the first of its kind in the industry. Some OCC and BOCC primary / standby redundancy functions require additional equipment, which increases the potential for failure. This solution does not rely on additional equipment, reducing potential failure points. Expanding OCC and BOCC servers at the device level or center level is difficult. This solution provides differentiated management for OCC and BOCC server clusters, enabling hot standby redundancy management regardless of whether the number of a certain type of server device is added or reduced in OCC or BOCC, or whether additional DDOCC or XXOCC is added at the center level.

[0038] 1. Equipment Architecture Introduction The OCC and BOCC device architectures designed in this scheme are as follows: Figure 1 As shown: OCC and BOCC comprise equipment including application server A / B machines, database server A / B machines, interface server A / B machines, and scheduling workstations. Hot standby equipment includes application servers, database servers, and interface servers. The server equipment combination of OCC and BOCC forms a computing resource platform, providing computing support for business operations. Hot standby nodes include... Figure 2 As shown.

[0039] Taking application servers as an example, OCC's application server A / B machines and BOCC's application server A / B machines form a cluster, and the internal network communication topology is as follows: Figure 3 As shown, this topology enables internal communication within the cluster. All devices within the cluster communicate via UDP, and the communication structure is one-to-many, meaning each device communicates with all other devices in the cluster.

[0040] Other node clusters (interface servers, database servers) of OCC and BOCC are handled in the same way as application server clusters.

[0041] 2. Implementation Scheme Introduction Solution implementation architecture as follows Figure 4 As shown, the hot standby function of the OCC and BOCC hot standby devices is mainly implemented by the following modules, each of which implements logic processing through a state machine.

[0042] Primary / Secondary Management Module: Manages the service status of OCC and BOCC server clusters and the status of primary and secondary machines within the cluster; uses an optimized Bully algorithm to decide which primary or secondary machine to use.

[0043] Heartbeat Management Module: Manages distributed heartbeat status.

[0044] Local State Management Module: Manages local state, mainly network state and process state.

[0045] Command line operation module: Allows you to view the status of the primary and backup centers and issue manual switchover commands.

[0046] Command control module: Responds to commands.

[0047] Front-end communication module: Communication between the management server and the scheduling workstation.

[0048] 2.1 Heartbeat Management Module This module is used for communication between devices. It adopts a distributed design, and the heartbeat state machines of each device send and receive their own state information to understand the state of other devices. It also sends heartbeat information to the primary / backup management module to make primary / backup status decisions.

[0049] 2.2 Local Status Management Module The local state management module is used to periodically monitor the local state, including the state of the network and other applications, and report to the heartbeat state machine.

[0050] 2.3 Primary / Secondary Management Module 1) Introduction to the Bully Algorithm The Bully algorithm is a classic leader election algorithm for distributed systems, primarily used to quickly elect a new coordinating node after a node failure. Its core components include the following: Inter-node communication: Each node can communicate with other nodes through message passing.

[0051] Election process: When a node discovers that the current leader is unavailable, it will initiate a challenge.

[0052] The node that initiates the challenge sends a challenge message to all nodes with a higher identifier (ID).

[0053] If a node receives a challenge message and its ID is higher than the sender's, it will respond with an OK message and continue sending challenge messages to nodes with higher IDs.

[0054] If the node that initiated the challenge does not receive a response from any node with a higher ID, it will declare itself the new leader and send a challenge success message to all other nodes.

[0055] Leader confirms: Once a node is elected as the leader, it sends a challenge success message to all other nodes.

[0056] After receiving the success message, other nodes will confirm the new leader and update their own status.

[0057] The core idea of ​​the Bully algorithm is to use the unique identifier of a node (usually an integer ID) to determine who should become the leader. Nodes with higher IDs have higher priority and are more likely to be elected as leaders. This mechanism ensures the efficiency and determinism of the election process.

[0058] Advantages: Simple to implement, fast convergence. Disadvantages: High-ID node failures can lead to frequent elections. The message complexity is O(n²) (worst case). Depends on the fact that the node IDs are globally unique and ordered. 2) Optimize the Bully algorithm for specific application scenarios. To meet the needs of business scenarios, the following modifications and optimizations were made to the traditional Bully algorithm: In the Bully algorithm, the leader is defined as the master state in the business scenario. In this design, the OCC and BOCC application servers (4 in total) form a cluster. The OCC application servers (2) and BOCC application servers (2) are each defined as a group, GroupA and GroupB respectively. Within a group, machines are hot-standby, and the master state device is defined as the host. Between groups, there is hot standby, and the master state group is defined as the master group.

[0059] In business scenarios, the integer ID of the Bully algorithm is defined by business status (network status, process status) and dynamic random numbers. This changes the traditional fixed ID of Bully, solves the drawbacks of fixed priority, and addresses the problem of frequent elections caused by the failure of high-ID nodes. The calculation logic is as follows: ID in, ID: An 8-byte unsigned integer, except for the following bits which are defined, the rest are reserved; Network status, bits 26-31 of the ID; Process status, bits 8-25 of the ID; R : A real random number, consisting of bits 0-7 of the ID; the calculation method will be explained in a later chapter.

[0060] Random number definition: Abnormal nodes have a random number of 0; naturally generated random numbers are 1-98; the initial random number for the master host is 99; the target host's random number is 100 during manual switching to ensure victory in decision-making; when in the master state, the state machine's random number is 127, which can suppress challenges from all nodes. Naturally generated random numbers 1-98 are real random numbers; 0, 100, and 127 are nominal random numbers assigned by the system.

[0061] Real random numbers The generation algorithm is determined by the initial random number value 'a', the final random number value 'b', the divisor (divisor), and the remainder (remainder). The calculation formula is as follows: in, : is the basic random number generation function; Start value a, end value b: These are configuration values ​​that are consistent across machines in the cluster and are used for basic random number generation. divisor: the number of machines in the cluster; Remainder: This is the machine ID, which is unique within the cluster.

[0062] This avoids the situation where random values ​​overlap among machines within the cluster.

[0063] In business scenarios, there are 4 OCC and BOCC application servers (configurable in total). The main host is the primary service device, and the rest are backup devices, i.e., 1 primary and 3 backups. There is no high complexity problem.

[0064] Configure a specified master group host, and the election process is as follows: Figure 5 As shown: Machine 1 is configured as the primary host, with its initial random number set to 99. Machines without initialized primary host configurations generate random numbers ranging from 1 to 98 using an algorithm. Upon receiving a challenge from another machine, Machine 1 compares its network status, process status, and the size of its random number. If the high-order bits of the ID value (network status, process status) are consistent, Machine 1 initializes its primary host configuration with the largest random number, sends a suppression signal to the other machines, and completes the initial host election process.

[0065] If the initialization host machine is connected to the cluster when the cluster already has a host, the initialization host will remain in standby mode to ensure that the existing host state of the system remains unchanged.

[0066] Example of a leader election process without initialization of a specified host. Assume there are 4 machines, and the challenge process is as follows: Figure 6 As shown: When a new round of elections is initiated, each machine generates a corresponding ID based on its own state; After learning the ID status of other machine nodes, Machine 1 initiates a challenge. All other machine nodes determine that Machine 1's ID value is superior and take no action. If Machine 1 is not suppressed within the challenge's lifecycle, the challenge is successful. If suppressed, the challenge ends after comparing with the succession conditions. Upon receiving the challenge, Machine 4 determines that it is superior to Machine 1, sends a suppression message, then sends a victory packet. If no suppression occurs, it becomes the successor.

[0067] When the challenge is initiated, each server generates a corresponding ID according to the method for generating integer IDs, and compares it with the machines in the cluster to achieve the function of election suppression or selection of surrender.

[0068] The optimal machine is selected as the master machine through the above process, and the suppressed machine is set as the backup machine.

[0069] Advantages of modifying the Bully algorithm: There is no situation where a high-ID node failure would lead to frequent elections. There is no high complexity problem Solving the problem of ensuring that node IDs are globally unique and ordered. 3) Design of primary and backup decision state machines The primary / standby management module is implemented based on a state machine, such as... Figure 7 As shown, this is a master-slave decision state machine. The Bully algorithm is used to implement the logical interaction between the machines and finally decide the master-slave state. Table 1 shows the event definition of the state machine; Table 2 shows the response logic of each state of the state machine receiving each event.

[0070] Table 1 Definition of Primary and Backup Decision Events Event Name meaning EventStartElection Initiate decision-making operation command event EventChallenge Send Challenge Message Event EventSuppress Send suppression message event EventSurrender Send surrender message event EventVictory Send victory message event EventElectionTimer Waiting for feedback timeout event EventNetRecoverTimer Network recovery timeout event EventReturnToBack Immediate backup event EventSwitchToMainCommand Master Ascension Command Event EventKeepPriorityOutTime Maintain priority timeout events EventShowStatusInGroup Display group status events EventShowALLNodesStatus Display all node status events EventHeartbeat Send heartbeat status events periodically OPEventHeartbeat Heartbeat status events are periodically forwarded to the primary and backup operation state machines. EventSendALLStatusToPanel Send status updates to the front end periodically EventPeriodSendALLStatusToPanel Periodically send status to the front end EventSendALLStatusToPanelNow Send status to the front end when the status changes. EventTransToBackReply Feedback on backup reduction operation results EventTransToMainReply Feedback on the success of the master upgrade operation EventSendStatusRequestReply Display all node status events Table 2 Event handling methods for each state 2.4 Command Line Operation Module Used to issue commands to the command-line operation state machine to view the status of the primary and backup centers and to manually switch over.

[0071] 2.5 Command Control Module The command control module is mainly used to respond to commands from the command line operation module. After obtaining the information required for the command from the primary and backup decision state machines, it feeds it back to the command line operation module.

[0072] Key points of this invention: (1) OCC and BOCC are “device-level cloudification”, supporting cross-service operation of OCC and BOCC devices.

[0073] (2) Master-slave state decision-making method based on optimized Bully algorithm.

[0074] (3) Hot standby management method for master-slave centers based on optimized Bully algorithm.

[0075] Advantages of this invention: (1) Achieve “device-level cloudification” for OCC and BOCC, with high scalability. Support cross-operation of OCC and BOCC equipment to improve the availability of signal system center business services.

[0076] (2) To resolve the problem of potential confusion in ATS-related services due to the emergence of dual master centers.

[0077] (3) Reduce the workload of personnel involved in OCC and BOCC hot standby switching, reduce the possible error rate, and improve the system fault tolerance.

[0078] (4) No additional equipment is required to achieve hot standby for OCC and BOCC, reducing the probability of failure.

[0079] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An ATS primary / standby hot standby redundancy system based on an optimized Bully algorithm, characterized in that, The system includes: OCC and BOCC; The equipment included in OCC and BOCC consists of application server A / B machines, database server A / B machines, interface server A / B machines, and scheduling workstations; among them, the hot standby equipment includes application servers, database servers, and interface servers; the combination of server equipment in OCC and BOCC forms a computing resource platform to provide computing support for business operations. The application servers A / B of OCC and the application servers A / B of BOCC form a cluster. The database servers A / B of OCC and the database servers A / B of BOCC form a cluster. The interface servers A / B of OCC and the interface servers A / B of BOCC form a cluster. Internal communication within the cluster is achieved through an internal network communication topology. All devices within the cluster communicate via UDP. The communication structure is one-to-many, meaning that each device communicates with all other devices in the cluster.

2. The ATS primary / standby hot standby redundancy system based on the optimized Bully algorithm as described in claim 1, characterized in that, The hot standby function of the hot standby device is jointly implemented by the following modules, each of which uses a state machine to perform logical processing. Primary / Secondary Management Module: Used to manage the service status of OCC and BOCC server clusters and the status of primary and secondary machines within the cluster; Use the optimized Bully algorithm to decide the host or master group; Heartbeat Management Module: Used to manage distributed heartbeat status; Local state management module: Used to manage local state, including network state and process state; Command line operation module: Used to view the status of the primary and backup centers and issue manual switchover commands; Command control module: used to respond to commands; Front-end communication module: Used to manage communication between the server and the scheduling workstation.

3. The ATS primary / standby hot standby redundancy system based on the optimized Bully algorithm as described in claim 2, characterized in that, The heartbeat management module is used for communication between devices. It adopts a distributed design, and the heartbeat state machines of each device send and receive their own state information to understand the state of other devices. They also send heartbeat information to the primary / backup management module to make primary / backup status decisions.

4. The ATS primary / standby hot standby redundancy system based on the optimized Bully algorithm as described in claim 2, characterized in that, The local status management module is used to periodically monitor the local status, including network and application status, and report to the heartbeat status machine; the command line operation module is used to issue commands to the command line operation status machine to view the status of the primary and backup centers and to manually switch over. The command control module is used to respond to commands from the command line operation module. After obtaining the information required for the command from the primary and backup decision state machines, it feeds it back to the command line operation module.

5. The ATS primary / standby hot standby redundancy system based on the optimized Bully algorithm as described in any one of claims 1-4, characterized in that, The primary / standby management module has been optimized to meet the needs of specific business scenarios. In the business scenario, the leader of the Bully algorithm is defined as the master state; there are a total of 4 servers in OCC and BOCC, forming a cluster; among them, the 2 servers of OCC and the 2 servers of BOCC are each defined as a group, namely GroupA and GroupB; within a group, the machine is hot standby, and the master state device is defined as the host; between groups, the master state group is hot standby, and the master group is defined as the master group. In business scenarios, the integer ID of the Bully algorithm is defined by business status and dynamic random numbers, and the calculation logic is as follows: ID in, ID: An 8-byte unsigned integer, except for the following bits which are defined, the rest are reserved; Network status, bits 26-31 of the ID; Process status, bits 8-25 of the ID; R : A real random number, consisting of bits 0-7 of the ID; In business scenarios, among the four servers in OCC and BOCC, the primary host is the main service device, and the rest are backup devices, i.e., 1 primary and 3 backups.

6. The ATS primary / standby hot standby redundancy system based on the optimized Bully algorithm as described in claim 5, characterized in that, Random number definition: Abnormal nodes have a random number of 0; naturally generated random numbers are 1-98; the initial random number for the main host is 99; the random number for the target host during manual switching is 100 to ensure victory in decision-making; when in the main state, the random number for the state machine is 127, which can suppress the challenges of all nodes; naturally generated random numbers 1-98 are real random numbers. 0, 100, and 127 are nominal random numbers assigned by the system; Real random numbers The generation algorithm is determined by the initial random number value 'a', the final random number value 'b', the divisor, and the remainder; the calculation formula is as follows: in, : is the basic random number generation function; Start value a, end value b: These are configuration values ​​that are consistent across machines in the cluster and are used for basic random number generation. divisor: the number of machines in the cluster; Remainder: This is the machine ID, which is unique within the cluster.

7. The ATS primary / standby hot standby redundancy system based on the optimized Bully algorithm as described in claim 6, characterized in that, Configure a specified master group host, and the election process includes: Machine 1 is configured as the main host, and its initial random number is 99. Machines without initialized main host configurations generate random numbers ranging from 1 to 98 according to the algorithm. After receiving challenges from other machines, Machine 1 compares the network status, process status, and the size of the random number. If the high-order bits of the ID value, i.e., the network status and process status, are consistent, then the initialized main host configuration random number is the largest, and a suppression signal is sent to other machines, completing the initial host election function. If the initialization host machine is connected to the cluster when the cluster already has a host, the initialization host will remain in standby mode to ensure that the existing host state of the system remains unchanged.

8. The ATS primary / standby hot standby redundancy system based on the optimized Bully algorithm as described in claim 6, characterized in that, Example of a leader election process without initialization of a designated host, assuming there are 4 machines, the challenge process includes: When a new round of elections is initiated, each machine generates a corresponding ID based on its own state; After learning the ID status of other machine nodes, Machine 1 initiates a challenge; all other machine nodes determine that Machine 1's ID value is better than their own and do not take any action; if Machine 1 is not suppressed within the challenge lifecycle, the challenge is successful; if it is suppressed, the challenge ends after comparing the promotion conditions; after receiving the challenge, Machine 4 determines that it is better than Machine 1, sends a suppression message, then sends a victory packet; if no suppression occurs, it becomes the master. When the challenge is initiated, each server generates a corresponding ID according to the method for generating integer IDs, and compares it with the machines in the cluster to realize the function of election suppression or selection of surrender. The optimal machine is selected as the master machine through the above process, and the suppressed machine is set as the backup machine.

9. The ATS primary / standby hot standby redundancy system based on the optimized Bully algorithm as described in claim 6, characterized in that, The primary / standby management module is implemented based on a state machine, serving as the primary / standby decision state machine. The Bully algorithm is used to implement logical interactions between the machines, ultimately determining the primary / standby state. The state machine events are defined as follows: Table 1 Definition of Primary and Backup Decision Events 。 10. The ATS primary / standby hot standby redundancy system based on the optimized Bully algorithm as described in claim 9, characterized in that, The response logic for each state of the primary and backup decision state machines to each event is as follows: Table 2 Event handling methods for each state 。