Main and backup control methods and devices

By utilizing the computing channel status table and self-test results of the master control resource module in the cloud secure computing platform, combined with periodic broadcast messages, master-slave switching control was achieved, solving the problem that traditional methods could not determine the master-slave relationship, and ensuring the reliability and flexibility of the system.

CN121441680BActive Publication Date: 2026-07-31BEIJING HOLLYSYS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING HOLLYSYS
Filing Date
2025-12-31
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In cloud security computing platforms, traditional relay interlocking mechanisms cannot effectively determine the primary and backup relationship, resulting in the inability to perform primary and backup control. Especially in resource pooling and dynamic reconfiguration environments with scalable modules, any idle module may become a third backup host, which violates the original intention of the elastic architecture of the cloud platform.

Method used

By using the periodic status table and self-test results of the computing channel of the master control resource module, it is determined whether the current master control resource module is faulty. It then communicates with other modules through periodic broadcast messages to achieve master-slave switching control, forming a decentralized, highly reliable collective decision-making system.

Benefits of technology

Effective master-slave control is implemented in the cloud security computing platform, avoiding the situation where two masters output security data at the same time, thus ensuring the reliability and flexibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121441680B_ABST
    Figure CN121441680B_ABST
Patent Text Reader

Abstract

This invention provides a primary / standby control method and apparatus. The method determines whether the current primary control resource module is faulty based on the system periodic status table and self-test results of the computing channels in the current primary control resource module; and determines whether other primary control resource modules are faulty based on the periodic broadcast messages received by the computing channels in the current primary control resource module and the system periodic status table. If any primary control resource module fails, primary / standby switching control is performed according to the primary / standby type of the primary control resource module. This invention does not rely on any external dedicated equipment, but rather makes each primary control resource module in the cloud platform an equal participant in the decision-making network. Through the system periodic status table and periodic broadcast messages, a decentralized, highly reliable collective decision-making system is formed, achieving effective primary / standby control even in a cloud secure computing platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of secure computer platform technology, and in particular to a primary / backup control method and apparatus. Background Technology

[0002] In the field of rail transit, the safety computer platform of the train control system is usually designed with a relay interlock mechanism to determine the unique master-slave relationship between the A system and the B system, which effectively avoids the situation where two master systems output safety data at the same time.

[0003] With the continuous and in-depth exploration of cloud technology in the research field of next-generation train operation control systems, the application of cloud technology in the train operation control systems of the new generation of rail transit will surely become a technological breakthrough of great significance in the future rail transit field.

[0004] However, in cloud security platforms with resource pooling, scalable module numbers, and dynamic reconfiguration, any idle module can become a third warm standby host, dynamically forming a new primary / standby unit with surviving modules. Pre-installing physical relays for all possible primary / standby combinations is impractical and violates the original intention of the elastic architecture of cloud platforms. In other words, the traditional method of using interlocking relays to determine the primary / standby relationship can no longer be effectively applied in cloud security computing platforms, thus leading to the inability to perform effective primary / standby control. Summary of the Invention

[0005] In view of this, the present invention provides a primary / backup control method and apparatus to solve the problem that primary / backup control cannot be performed in a cloud security computing platform.

[0006] The first aspect of the present invention provides a primary / standby control method, comprising:

[0007] Based on the system periodic status table of the computing channels in the current master control resource module and the self-test results, it is determined whether the current master control resource module is faulty; wherein, the current master control resource module periodically sends periodic broadcast messages to the system and receives periodic broadcast messages sent by other master control resource modules; each master control resource module includes two computing channels, and each computing channel periodically updates its corresponding system periodic status table;

[0008] Based on the periodic broadcast messages received by the computing channel in the current master control resource module and the system periodic status table, determine whether there are other master control resource modules malfunctioning;

[0009] If any master control resource module fails, master-slave switching control will be completed according to the master-slave type of the master control resource module.

[0010] Optionally, any two main control resource modules running security services constitute a primary / standby subsystem; the primary / standby type of any main control resource module in the primary / standby subsystem is determined by its unique number.

[0011] Optionally, the current master control resource module periodically sends periodic broadcast messages to the system, including:

[0012] The first computing channel in the current main control resource module sends periodic broadcast messages to the system through the first local area network and the first control local area network bus;

[0013] The second computing channel in the current master control resource module sends periodic broadcast messages to the system through the second local area network and the second control local area network bus.

[0014] Optionally, the periodic broadcast message includes module information of the computing channel, the operating status of the computing channel, and perception information of other main control resource modules. Before the current main control resource module periodically sends the periodic broadcast message to the system, it also includes:

[0015] For each computing channel in the current master control resource module, obtain the module information of the computing channel; wherein, the module information of the computing channel includes the location number of the computing channel, the computing channel number, and the current system role;

[0016] For each computing channel in the current master control resource module, a functional safety self-test is performed on the computing channel to obtain the self-test result, and the self-test result is sent to the adjacent channel in each running cycle; the adjacent channel is another computing channel in the current master control resource module.

[0017] Based on the self-test results of all the computing channels, determine the current operating status of the main control resource module;

[0018] Based on the received periodic broadcast messages, the status of other master control resource modules is perceived, and the perception information of other master control resource modules is obtained.

[0019] Optionally, the step of sensing the status of other master control resource modules based on the received periodic broadcast messages to obtain sensing information of other master control resource modules includes:

[0020] Receive a first target periodic broadcast message; wherein the first target periodic broadcast message is received by the first computing channel through the first local area network and the first control local area network bus; the first target periodic broadcast message is a periodic broadcast message sent by the third computing channel; the third computing channel is any computing channel under the same local area network and the same control local area network bus as the first computing channel;

[0021] The first target periodic broadcast message is transmitted to the second computing channel, and the second target periodic broadcast message transmitted by the second computing channel is received simultaneously.

[0022] Based on the first target periodic broadcast message and the second target periodic broadcast message, the perception information of the master control resource module to which the third computing channel belongs is determined.

[0023] Optionally, determining whether the current main control resource module is faulty based on the system cycle status table of the computing channel in the current main control resource module and the self-test results includes:

[0024] A functional safety self-test is performed on the computing channel cycle in the current main control resource module to obtain the first fault result;

[0025] The system periodic status table corresponding to the computing channel in the current main control resource module is periodically checked, and the second fault result is determined based on the online status of all main control resource modules in the system periodic status table.

[0026] The third fault result is determined based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current master control resource module, the identifier set of the fault detection module, and the number of fault detection modules; wherein, the fault detection module is the master control resource module that detected the fault in the current master control resource module.

[0027] The fourth fault result is determined based on the number of fault-aware modules in the system cycle status table corresponding to the computing channel in the current main control resource module and the number of currently online main control resource modules.

[0028] Based on the first fault result, the second fault result, the third fault result, and the fourth fault result, it is determined whether the current main control resource module is faulty.

[0029] Optionally, determining whether there are other main control resource modules malfunctioning based on the periodic broadcast messages received by the computing channel in the current main control resource module and the system periodic status table includes:

[0030] The fifth fault result is determined based on the running status field in the periodic broadcast message received by the computing channel in the current master control resource module;

[0031] The sixth fault result is determined based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current main control resource module and the number of checkpoints in the online main control resource module;

[0032] Based on the fifth and sixth fault results, determine whether there are other main control resource modules that are faulty.

[0033] Optionally, if any main control resource module fails, a primary / standby switchover control is performed based on the primary / standby type of the main control resource module, including:

[0034] If the faulty master control resource module belongs to a master-slave type that is a standby system, then determine whether the target master control resource module can operate normally; wherein, the target master control resource module is another master control resource module in the master-slave subsystem to which the faulty master control resource module belongs;

[0035] If the target master control resource module can operate normally, continue to output security data using the target master control resource module;

[0036] If the primary / standby type of the faulty master control resource module is primary, then the primary / standby type of the target master control resource module is changed to primary, and the target master control resource module continues to output security data.

[0037] Optionally, after continuing to output security data using the target master control resource module, the method further includes:

[0038] The main control resource module in the third backup unit is used to replace the main control resource module of the main backup subsystem to which the faulty main control resource module belongs. If at least one main control resource module in the main control unit is not running security services, then the main control unit is used as the third backup unit. The main control unit consists of two consecutive main control resource modules.

[0039] A second aspect of the present invention provides a primary / standby control device, comprising:

[0040] The first analysis unit is used to determine whether the current master control resource module is faulty based on the system periodic status table of the computing channels in the current master control resource module and the self-test results; wherein, the current master control resource module periodically sends periodic broadcast messages to the system and receives periodic broadcast messages sent by other master control resource modules; each master control resource module includes two computing channels, and each computing channel periodically updates its corresponding system periodic status table;

[0041] The second analysis unit is used to determine whether there are other main control resource modules malfunctioning based on the periodic broadcast messages received by the computing channel in the current main control resource module and the system periodic status table.

[0042] The switching control module is used to perform master-slave switching control based on the master-slave type of the master control resource module if any master control resource module fails.

[0043] Optionally, any two main control resource modules running security services constitute a primary / standby subsystem; the primary / standby type of any main control resource module in the primary / standby subsystem is determined by its unique number.

[0044] Optionally, the first computing channel in the current master control resource module sends periodic broadcast messages to the system through the first local area network and the first control local area network bus; the second computing channel in the current master control resource module sends periodic broadcast messages to the system through the second local area network and the second control local area network bus.

[0045] Optionally, the periodic broadcast message includes module information of the computing channel, the operating status of the computing channel, and sensing information of other main control resource modules. The main / standby control device further includes:

[0046] The module information acquisition unit is used to acquire module information for each computing channel in the current master control resource module; wherein, the module information of the computing channel includes the location number of the computing channel, the computing channel number, and the current system role;

[0047] The self-test unit is used to perform a functional safety self-test on each computing channel in the current master control resource module, obtain the self-test result, and send the self-test result to the adjacent channel in each running cycle; the adjacent channel is another computing channel in the current master control resource module.

[0048] The running status determination unit is used to determine the running status of the current main control resource module based on the self-test results of all the computing channels;

[0049] The sensing unit is used to sense the status of other master control resource modules based on the received periodic broadcast messages, and obtain the sensing information of other master control resource modules.

[0050] Optionally, the sensing unit includes:

[0051] A receiving unit is configured to receive a first target periodic broadcast message; wherein the first target periodic broadcast message is received by the first computing channel through a first local area network and a first control local area network bus; the first target periodic broadcast message is a periodic broadcast message sent by a third computing channel; the third computing channel is any computing channel under the same local area network and the same control local area network bus as the first computing channel;

[0052] The transmission unit is configured to transmit the first target periodic broadcast message to the second computing channel, and simultaneously receive the second target periodic broadcast message transmitted by the second computing channel;

[0053] The sensing subunit is used to determine the sensing information of the main control resource module to which the third computing channel belongs based on the first target periodic broadcast message and the second target periodic broadcast message.

[0054] Optionally, the first analysis unit includes:

[0055] The first fault determination subunit is used to perform a functional safety self-check on the computing channel cycle in the current main control resource module and obtain the first fault result.

[0056] The second fault determination subunit is used to periodically check the system periodic status table corresponding to the computing channel in the current main control resource module, and determine the second fault result based on the online status of all main control resource modules in the system periodic status table.

[0057] The third fault determination subunit is used to determine the third fault result based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current main control resource module, the identifier set of the fault perception module, and the number of fault perception modules; wherein, the fault perception module is the main control resource module that detects the fault of the current main control resource module;

[0058] The fourth fault determination subunit is used to determine the fourth fault result based on the number of fault perception modules in the system cycle status table corresponding to the computing channel in the current main control resource module and the number of currently online main control resource modules.

[0059] The first analysis subunit is used to determine whether the current main control resource module is faulty based on the first fault result, the second fault result, the third fault result, and the fourth fault result.

[0060] Optionally, the second analysis unit includes:

[0061] The fifth fault determination subunit is used to determine the fifth fault result based on the running status field in the periodic broadcast message received by the computing channel in the current main control resource module.

[0062] The sixth fault determination subunit is used to determine the sixth fault result based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current main control resource module and the number of checkpoints in the online main control resource module.

[0063] The second analysis subunit is used to determine whether there are other main control resource modules that are faulty, based on the fifth fault result and the sixth fault result.

[0064] Optionally, the switching control unit includes:

[0065] The operation status determination subunit is used to determine whether the target main control resource module can operate normally if the main control resource module to which the faulty main control resource module belongs is a standby system; wherein, the target main control resource module is another main control resource module in the main standby subsystem to which the faulty main control resource module belongs;

[0066] The first switching control subunit is used to continue outputting security data using the target master control resource module if the target master control resource module can operate normally.

[0067] The second switching control subunit is used to switch the primary / standby type of the target primary control resource module to primary if the primary / standby type of the faulty primary control resource module is primary, and to continue outputting security data using the target primary control resource module.

[0068] Optionally, the main / standby control device further includes:

[0069] A replacement unit is used to replace the main control resource module of the main backup subsystem to which the faulty main control resource module belongs with a main control resource module in the third backup unit. If at least one main control resource module in the main control unit is not running security services, then the main control unit is used as the third backup unit. The main control unit consists of two consecutive main control resource modules.

[0070] A third aspect of the present invention provides an electronic device, comprising:

[0071] One or more processors;

[0072] A storage device on which one or more programs are stored;

[0073] When the one or more programs are executed by the one or more processors, the one or more processors implement the primary / standby control method as described in any one of the first aspects.

[0074] A fourth aspect of the present invention provides a computer storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the primary / backup control method as described in any one of the first aspects.

[0075] As can be seen from the above scheme, the present invention provides a primary / standby control method and apparatus. This method determines whether the current primary control resource module is faulty based on the system periodic status table and self-test results of the computing channels in the current primary control resource module; and determines whether other primary control resource modules are faulty based on the periodic broadcast messages received by the computing channels in the current primary control resource module and the system periodic status table. If any primary control resource module fails, primary / standby switching control is completed according to the primary / standby type of the primary control resource module. This invention does not rely on any external dedicated equipment, but rather makes each primary control resource module in the cloud platform an equal participant in the decision-making network. Through the system periodic status table and periodic broadcast messages, a decentralized, highly reliable collective decision-making system is formed, achieving the goal of effective primary / standby control even in a cloud secure computing platform. Attached Figure Description

[0076] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0077] Figure 1 A schematic diagram of a cloud security computing platform environment provided in an embodiment of the present invention;

[0078] Figure 2 A flowchart of a primary / backup control method provided in another embodiment of the present invention;

[0079] Figure 3 A flowchart of a primary / backup control method provided in another embodiment of the present invention;

[0080] Figure 4 A schematic diagram of a CAN bus broadcast message format provided for another embodiment of the present invention;

[0081] Figure 5 A schematic diagram illustrating a local area network broadcast message format according to another embodiment of the present invention;

[0082] Figure 6 A flowchart illustrating a method for sensing information of other master control resource modules, provided as another embodiment of the present invention;

[0083] Figure 7 A flowchart of a method for determining whether the current master control resource module is faulty, provided in another embodiment of the present invention;

[0084] Figure 8A flowchart of a method for determining whether other main control resource modules are faulty, provided as another embodiment of the present invention;

[0085] Figure 9 A schematic diagram of a primary / standby control device provided in another embodiment of the present invention;

[0086] Figure 10 This is a schematic diagram of an electronic device that implements a master / slave control method according to another embodiment of the present invention. Detailed Implementation

[0087] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0088] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0089] It should be noted that the concepts of "first" and "second" mentioned in this invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0090] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0091] First, the cloud security computing platform environment provided in this invention is described, such as... Figure 1 As shown, the cloud secure computing platform includes a scalable number of secure master control resource modules. Each master control resource module consists of two computing channels forming a "two-out-of-two" structure and function, and also has virtualization capabilities to run multiple applications. The orchestrator module and the vehicle communication module in the cloud secure computing platform are both master control resource modules, differing only in their business functions.

[0092] Each main control resource module of the cloud security computing platform has two independent channels connected to the platform's internal Ethernet switch, enabling local area network communication between the main control resource modules. Simultaneously, all security computer modules are also connected via an expansion bus, enabling bus communication between them. The expansion bus uses Controller Area Network (CAN) communication. After the cloud security platform starts running, all main control resource modules (including the configuration management module and communication management module) periodically send periodic broadcast messages to the system after initialization.

[0093] It should be noted that, Figure 1 The examples given are local area networks and CAN communication. In the actual application of this invention, networks with broadcast communication capabilities can achieve similar results.

[0094] Specifically, the first computing channel in the main control resource module (such as computing channel 1 in main control resource module 1) sends periodic broadcast messages to the system through the first local area network (LAN A switch) and the first control LAN bus (CAN bus A channel); the second computing channel in the main control resource module (such as computing channel 2 in main control resource module 1) sends periodic broadcast messages to the system through the second local area network (LAN B switch) and the second control LAN bus (CAN bus B channel).

[0095] In practical applications of this invention, any two main control resource modules running secure services can constitute a primary / backup subsystem with a "two-out-of-two" functional structure. In a primary / backup subsystem, the main control resource module operating in primary mode is called the primary system, and the main control resource module operating in standby mode is called the backup system. In practical applications of this invention, the primary / backup type can be determined based on the unique identifier of the main control resource module; this is not limited here.

[0096] Because the main control resource modules are deployed in a scalable number of device cages, each main control resource module is installed in a slot within the cage. Each slot is uniquely identified by its cage number and slot number. During operation, the main control resource module can obtain its own location number, i.e., its unique identifier, by reading hardware information.

[0097] One implementation method for determining the primary / standby type based on the unique identifier of the primary control resource module includes:

[0098] First, the unique ID of the main control resource module is calculated from the cage number and slot number. The calculation method is: Unique ID of main control resource module = Cage number × Maximum number of modules in the cage + Slot number. The cage number starts from 0, and the slot number starts from 1, therefore the position number starts from 1. When the unique ID of the main control resource module is odd, its role is configured as the primary system; when the unique ID of the main control resource module is even, its role is configured as the backup system.

[0099] Based on the aforementioned cloud security computing platform, this embodiment of the invention provides a primary / standby control method, such as... Figure 2 As shown, the specific steps include:

[0100] S201. Based on the system cycle status table of the computing channel in the current main control resource module and the self-test results, determine whether the current main control resource module is faulty.

[0101] Each computing channel periodically updates its corresponding system periodic status table.

[0102] It is important to emphasize that each computing channel independently maintains a system cycle status table. The entries in this table are used to record the status and sensing information of all modules in the system. The computing channel can receive the status and sensing information of all modules in the system (including the main control resource module, orchestrator module, and communication module), and update the system cycle status table based on all received information and the cycle detection method.

[0103] It should be noted that in this invention, perception information is provided on a computing channel basis, or on a main control resource module or application virtual machine basis, both of which can achieve similar effects, and no limitation is made here.

[0104] The main contents of the system cycle status table are shown in Table 1. It is understood that in actual applications, system information may also be included, but this is not relevant to the present invention and will not be described in detail here.

[0105] Table 1

[0106]

[0107] The update method for each entry in the system cycle state table for each computing channel can be as follows:

[0108] 1. D_id, S_id, and R are checked and updated based on the module information in the received periodic broadcast messages.

[0109] 2.S performs fault "OR" updates based on the module status in the periodic state.

[0110] 3. Before each periodic broadcast is sent, the CP value of all entries in the system periodic status table is incremented; whenever a periodic broadcast message from the module of that entry is received, the CP value is cleared to zero.

[0111] 4. Whenever a periodic broadcast message from this module is received, process the PI field in the broadcast message and parse out all D_ids whose values ​​correspond to the fault status. Update the LF_ids entry in the system periodic status table corresponding to the D_id by adding the D_id of the module that sent the broadcast message to LF_ids as a set.

[0112] 5. The value of N_Fault is the cardinality of the aforementioned LF_ids set, and N_Fault is updated whenever LF_ids is updated.

[0113] It should be noted that the periodic broadcast message includes module information of the computing channel, the operating status of the computing channel, and perception information of other main control resource modules. Therefore, in the actual application of this invention, before the current main control resource module periodically sends the periodic broadcast message to the system, such as... Figure 3 As shown, it also includes:

[0114] S301. For each computing channel in the current master control resource module, obtain the module information of the computing channel.

[0115] The module information for the computing channel includes the computing channel's location number, computing channel number, and current system role.

[0116] S302. For each computing channel in the current main control resource module, perform a functional safety self-test on the computing channel, obtain the self-test result, and send the self-test result to the adjacent channel in each running cycle.

[0117] The adjacent channel is another computing channel in the current master control resource module.

[0118] It should be noted that functional safety self-tests are hardware and software self-tests, including but not limited to RAM self-tests, cache self-tests (whether read and write operations are correct and whether tampering has occurred), CPU instruction self-tests, clock self-tests, timer self-tests, and system task scheduling self-tests (whether tasks are executed in sequence). There are no restrictions on these aspects.

[0119] S303. Based on the self-test results of all computing channels, determine the current operating status of the main control resource module.

[0120] Specifically, the operating status refers to the operating status of the two computing channels of the main control resource module. In accordance with functional safety requirements, each computing channel performs periodic functional safety self-checks, with the result being either normal or faulty. During each operating cycle, each of the two computing channels sends its own periodic self-check status. Before sending, the two computing channels exchange their respective periodic self-check statuses, performing a fault OR operation between their own status and the status of the adjacent channel. That is, if either status is faulty, the status to be sent is corrected to a faulty status. Therefore, the statuses sent by the two computing channels are consistent, both representing the status of the main control resource module.

[0121] S304. Based on the received periodic broadcast messages, perceive the status of other master control resource modules and obtain the perception information of other master control resource modules.

[0122] In the practical application of this invention, the broadcast message sent by each computing channel via the local area network also includes the sense information detected from other modules within the system, which are also in two states: normal and fault. To increase the amount of information about the internal system status for each master control resource module, the two computing channels maintain and send their own sense information without prior data transfer. This sense information field is represented by PI and uses several bytes, the number of bytes being configured based on the number of master control resource modules in the system. The index of the low-to-high bit position of the field uniquely maps to the location number of the master control resource module in the system, and the value of the corresponding bit represents the status of the corresponding module: 0 indicates a normal state, and 1 indicates a fault state. The data formats of the CAN bus and local area network broadcast messages are as follows: Figure 4 and Figure 5 As shown. CAN_id is the essential ID information for CAN bus communication, CRC is the cyclic redundancy check code for this broadcast message, and the remaining fields are the known fields mentioned above.

[0123] In practical applications of this invention, to simplify system logic and increase the timeliness of status updates, the broadcast message packets of the CAN bus are not split and reassembled. Therefore, the status content broadcast by the CAN bus only includes its own information and operating status, compressing the periodic status to 8 bytes. In contrast, the periodic status broadcast via Ethernet includes all information. However, the module information and operating status broadcast via Ethernet and CAN bus are consistent and redundant.

[0124] Optionally, in another embodiment of the present invention, one implementation of step S304 is as follows: Figure 6 As shown, it includes:

[0125] S601, Receive the first target periodic broadcast message.

[0126] The first target periodic broadcast message is received by the first computing channel through the first local area network and the first control local area network bus; the first target periodic broadcast message is a periodic broadcast message sent by the third computing channel; the third computing channel is any computing channel under the same local area network and the same control local area network bus as the first computing channel.

[0127] It should be noted that, since the first computing channel is connected to the first local area network and the first control local area network bus, each computing channel can receive two periodic broadcast messages from any computing channel under the same network and bus channel through its own connected local area network and CAN bus channel. Therefore, under normal conditions, each master control resource module can receive four periodic broadcast messages.

[0128] S602, Transmit the first target periodic broadcast message to the second computing channel, and simultaneously receive the second target periodic broadcast message transmitted by the second computing channel.

[0129] S603. Based on the first target periodic broadcast message and the second target periodic broadcast message, determine the perception information of the main control resource module to which the third computing channel belongs.

[0130] Specifically, the computation channel processes the S field in periodic broadcast messages using a "redundancy removal" method. That is, the computation channel can receive a maximum of four broadcast messages from any master control resource module in a given period. It checks the Seq field in these broadcast messages, comparing it to the Seq value of the last received broadcast message from that module. Broadcast messages with an increased Seq value within the tolerance range are considered valid, and the S field of redundant broadcast messages is ignored. (Since the sending end's mutual transmission status is consistent, this comparison is not necessary). Then, the computation channel processes the PI field in the periodic broadcast messages using a fault-based "OR" method. That is, the computation channel can receive a maximum of two broadcast messages containing the PI field from any master control resource module in a given period, and processes the PI field in each broadcast message using a fault-based "OR" method.

[0131] For example: If two broadcast messages from the same module are received at this time, assuming that the second bit of the PI field represents the perceived status of module 2, the value of this bit in the PI field of the first message is 0, indicating that module 2 is considered normal, and the value of this bit in the PI field of the second message is 1, indicating that module 2 is considered faulty, then this module will handle it according to "the module that sent the broadcast message considers module 2 to be faulty".

[0132] When any primary or backup subsystem in a cloud computing platform experiences a single-system failure, a crucial prerequisite for restoring dual-system operation using a third backup resource module is that the primary and backup subsystems can correctly identify their own or neighboring system failures. Relying solely on the primary control resource module's own judgment and communication between the primary and backup subsystems is insufficient to adequately prevent single-point-of-failure scenarios. Therefore, it is necessary to combine system awareness information from all resource modules in the system for decision-making. Thus, in another embodiment of the present invention, one implementation of step S201 is as follows... Figure 7 As shown, it includes:

[0133] S701. Perform a functional safety self-check on the computing channel cycle in the current main control resource module and obtain the first fault result.

[0134] Specifically, in accordance with functional safety requirements, each computing channel in the main control resource module will periodically perform self-checks on the hardware and operating environment of the computing channel. When a fault is actively detected, it will determine that it is faulty.

[0135] S702. Perform periodic checks on the system periodic status table corresponding to the computing channels in the current main control resource module, and determine the second fault result based on the online status of all main control resource modules in the system periodic status table.

[0136] Specifically, each computing channel of the main control resource module periodically checks the system's periodic status table. If all online main control resource modules in the table change to a fault state, then the module itself is considered to be faulty.

[0137] S703. Based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current main control resource module, the identifier set of the fault perception module, and the number of fault perception modules, determine the third fault result.

[0138] Among them, the fault perception module is the main control resource module that detects a fault in the current main control resource module.

[0139] Specifically, the main control resource module checks the periodic checkpoint count (CP) value of each module in each computing channel cycle. When the CP value exceeds the threshold, the module is determined to be in a fault state. If the identifier set (LF_ids) of the fault-aware module of the module is empty and the number of fault-aware modules (N_Fault) is 0 for 3 consecutive cycles, that is, no other module has detected the fault of the module, then the module is determined to be faulty.

[0140] S704. Determine the fourth fault result based on the number of fault perception modules in the system cycle status table corresponding to the computing channel in the current main control resource module and the number of currently online main control resource modules.

[0141] Specifically, each computing channel periodically checks the N_Fault value of its own module information. When the value is found to be greater than or equal to Nm / 2, it determines that it is faulty, where Nm is the number of online main control resource modules in the periodic status table.

[0142] It should be noted that, Figure 7 This example illustrates the sequence of steps S701 to S704. In actual application, the order of steps S701, S702, S703, and S704 is not restricted, and steps S701, S702, S703, and S704 can be executed simultaneously.

[0143] S705. Determine whether the current main control resource module is faulty based on the first fault result, the second fault result, the third fault result, and the fourth fault result.

[0144] Specifically, if any one of the first, second, third, or fourth fault results indicates that the current main control resource module is faulty, then the current main control resource module is confirmed to be faulty.

[0145] In the actual application of this invention, when the main control resource module determines that it is faulty, the main control resource module first shuts down the security output and calculation of the security application, and sets its own state to a fault state. After the fault occurs, it maintains a periodic broadcast message for 3 cycles, and then performs the operation shutdown process.

[0146] The three-cycle interval is a parameter used for implementation. After a fault, the fault information log and record need to be transmitted to the maintenance terminal through the maintenance interface, so at least three cycles are required to transmit the fault information. Simultaneously, the three-cycle broadcast also improves the reception of fault information by other modules in the system. Too many cycles are not recommended because broadcasting fault status would increase the fault handling burden on the module receiving the message; therefore, after maintaining a sufficient number of cycles to handle the fault, the system is shut down.

[0147] S202. Based on the periodic broadcast messages received by the computing channel in the current master control resource module and the system periodic status table, determine whether there are any other master control resource modules that are faulty.

[0148] In the practical application of this invention, when the main control resource module detects a fault in another module, it sets the operating status of that module in the periodic status table to fault. Simultaneously, it iterates through the LF_ids set of all modules, removes the D_id of that module from the LF_ids set of all modules, and decrements N_Fault by one. Then, when sending a periodic broadcast message, it updates the PI field of the broadcast message.

[0149] Optionally, in another embodiment of the present invention, one implementation of step S202 is as follows: Figure 8 As shown, it includes:

[0150] S801. Determine the fifth fault result based on the running status field in the periodic broadcast message received by the computing channel in the current master control resource module.

[0151] Specifically, if any computing channel receives a periodic broadcast message with a fault status in the R field, it determines that the master control resource module that sent the broadcast message is faulty.

[0152] S802. Based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current main control resource module and the number of checkpoints in the online main control resource module, determine the result of the sixth fault.

[0153] Specifically, the main control resource module checks the periodic status table for each computing channel cycle. If the CP value is greater than or equal to the threshold N1, and there is any online main control resource module with a check point of 0 in the periodic status table, then the main control resource module is determined to be faulty.

[0154] It should be noted that, Figure 8 This example illustrates the concept of executing step S801 first and then step S802. It is understood that step S802 can also be executed first and then second, or both steps S801 and S802 can be executed simultaneously.

[0155] S803. Based on the results of the fifth and sixth faults, determine whether there are other main control resource modules that are faulty.

[0156] Specifically, if either the fifth or sixth fault result indicates a fault in the main control resource module X, then the fault in the main control resource module X can be confirmed.

[0157] S203. If any main control resource module fails, the main / standby switchover control shall be completed according to the main control resource module’s main / standby type.

[0158] It should be noted that, Figure 2 This example illustrates the concept of executing step S201 first and then step S202. It is understood that step S202 can also be executed first and then step S202, or steps S201 and S202 can be executed simultaneously. As long as any main control resource module fails, the main / standby switching control is completed according to the main control resource module's main / standby type.

[0159] In the practical application of this invention, since the primary and backup subsystems include a primary system and a backup system, when one system detects a fault in another primary control resource module, it can know through the S_id information in the module information that the faulty module is a neighboring system of this primary and backup subsystem. Then, in each subsequent cycle, the N_Fault value information of the neighboring primary control resource module is checked. When the N_Fault value is found to be greater than or equal to Nm / 2, the neighboring primary control resource module is determined to be faulty.

[0160] Optionally, in another embodiment of the present invention, if the faulty master control resource module belongs to the master / standby type of standby system, and it is determined that the target master control resource module can operate normally, the target master control resource module continues to output security data. If the faulty master control resource module belongs to the master / standby type of master system, the master / standby type of the target master control resource module is converted to master system, and the target master control resource module continues to output security data.

[0161] The target master control resource module is another master control resource module in the master-slave subsystem to which the failed master control resource module belongs.

[0162] Understandably, if the primary / standby type of the faulty master control resource module belongs to the primary system, it is possible to try sending offline request confirmations to the primary system multiple times. When an offline confirmation is received from the original primary system or no response is received after a timeout, the primary / standby type of the target master control resource module is changed to the primary system, and the target master control resource module continues to output safe data.

[0163] Optionally, in another embodiment of the present invention, after the target master control resource module continues to output security data, the method further includes:

[0164] The main control resource module in the third backup unit is used to replace the main control resource module of the backup subsystem to which the faulty main control resource module belongs.

[0165] If at least one main control resource module in the main control unit is not running security services, the main control unit will be used as the third backup unit. The main control unit consists of two consecutive main control resource modules.

[0166] It should be noted that every two consecutively numbered master control resource modules constitute a master control unit. These two master control resource modules are respectively A-series and B-series modules. To meet safety requirements, the power supply for A-series and B-series modules in the master control unit is independent. The orchestrator module and vehicle communication module in the cloud security computing platform have the same structure as the master control unit. If at least one master control resource module in the master control unit is not running security services, the master control unit will be used as the third backup unit. During operation, if a master control resource module in the primary and backup subsystems fails, the redundant master control resource modules will operate as a single system. At this time, the A-series and B-series modules in the third backup unit can replace the primary or backup system of the primary and backup subsystems under certain conditions, so that the failed primary and backup subsystems can again form a "two-for-two" functional structure.

[0167] In this invention, the master control resource module not only uses the traditional periodic self-test method to identify its own faults, but also makes full use of the system-wide perception information to carry out its own fault identification. This approach indirectly improves the reliability of status confirmation between the master and backup systems, which is the key difference between this invention and existing technologies.

[0168] It is important to emphasize that each master control resource module provides perception information about other master control resource modules in the system, as well as its own status, but it does not conduct a voting operation for the failure of any master control resource module. Instead, the relevant primary and backup subsystems make autonomous decisions based on the equally shared perception information in the system. This constitutes a significant difference between this invention and existing technologies. Furthermore, the operating status provided by each master control resource module remains consistent. However, when perceiving other master control resource modules in the system, each master control resource module's two computing channels independently provide perception information. In this way, the system's perception information is provided with a richer amount of information, thereby enhancing the robustness of module fault identification.

[0169] It should be noted that the method in this invention can be applied to any cloud security computing platform that adopts a functional safety architecture (dual-machine hot standby, "two-out-of-two", "two-out-of-three"), such as in the fields of urban rail transit and aerospace, and is not limited here.

[0170] As can be seen from the above scheme, the present invention provides a primary / standby control method. This method determines whether the current primary control resource module is faulty based on the system periodic status table and self-test results of the computing channels in the current primary control resource module; and determines whether other primary control resource modules are faulty based on the periodic broadcast messages received by the computing channels in the current primary control resource module and the system periodic status table. If any primary control resource module fails, primary / standby switching control is completed according to the primary / standby type of the primary control resource module. This invention does not rely on any external dedicated equipment, but rather makes each primary control resource module in the cloud platform an equal participant in the decision-making network. Through the system periodic status table and periodic broadcast messages, a decentralized, highly reliable collective decision-making system is formed, achieving the goal of effective primary / standby control even in a cloud secure computing platform.

[0171] Another embodiment of the present invention provides a primary / standby control device, such as... Figure 9 As shown, it specifically includes:

[0172] The first analysis unit 901 is used to determine whether the current main control resource module is faulty based on the system cycle status table of the computing channel in the current main control resource module and the self-test results.

[0173] The current master control resource module periodically sends periodic broadcast messages to the system and receives periodic broadcast messages from other master control resource modules. Each master control resource module includes two computing channels, and each computing channel periodically updates its corresponding system periodic status table.

[0174] Optionally, in another embodiment of the present invention, any two master control resource modules running security services constitute a master-slave subsystem; the master control resource module in any master-slave subsystem determines the master-slave type based on its unique number.

[0175] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, and will not be repeated here.

[0176] Optionally, in another embodiment of the present invention, the first computing channel in the current master control resource module sends periodic broadcast messages to the system through the first local area network and the first control local area network bus; the second computing channel in the current master control resource module sends periodic broadcast messages to the system through the second local area network and the second control local area network bus.

[0177] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, and will not be repeated here.

[0178] Optionally, in another embodiment of the present invention, the periodic broadcast message includes module information of the computing channel, the operating status of the computing channel, and sensing information of other main control resource modules. One implementation of the main / standby control device further includes:

[0179] The module information acquisition unit is used to acquire module information for each computing channel in the current master control resource module.

[0180] The module information for the computing channel includes the computing channel's location number, computing channel number, and current system role.

[0181] The self-test unit is used to perform functional safety self-tests on each computing channel in the current master control resource module, obtain the self-test results, and send the self-test results to the adjacent channel in each running cycle; the adjacent channel is another computing channel in the current master control resource module.

[0182] The running status determination unit is used to determine the current running status of the main control resource module based on the self-test results of all computing channels.

[0183] The sensing unit is used to sense the status of other master control resource modules based on the received periodic broadcast messages, and obtain the sensing information of other master control resource modules.

[0184] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, and will not be repeated here.

[0185] Optionally, in another embodiment of the present invention, one implementation of the sensing unit includes:

[0186] The receiving unit is used to receive the first target periodic broadcast message.

[0187] The first target periodic broadcast message is received by the first computing channel through the first local area network and the first control local area network bus; the first target periodic broadcast message is a periodic broadcast message sent by the third computing channel; the third computing channel is any computing channel under the same local area network and the same control local area network bus as the first computing channel.

[0188] The transmission unit is used to transmit the first target periodic broadcast message to the second computing channel, and at the same time receive the second target periodic broadcast message transmitted by the second computing channel.

[0189] The sensing subunit is used to determine the sensing information of the main control resource module to which the third computing channel belongs, based on the first target periodic broadcast message and the second target periodic broadcast message.

[0190] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, and will not be repeated here.

[0191] Optionally, in another embodiment of the present invention, one implementation of the first analysis unit 901 includes:

[0192] The first fault determination subunit is used to perform a functional safety self-check on the computing channel cycle in the current main control resource module and obtain the first fault result.

[0193] The second fault determination subunit is used to periodically check the system periodic status table corresponding to the computing channel in the current main control resource module, and determine the second fault result based on the online status of all main control resource modules in the system periodic status table.

[0194] The third fault determination subunit is used to determine the third fault result based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current main control resource module, the identifier set of the fault perception module, and the number of fault perception modules.

[0195] Among them, the fault perception module is the main control resource module that detects a fault in the current main control resource module.

[0196] The fourth fault determination subunit is used to determine the fourth fault result based on the number of fault-aware modules in the system cycle status table corresponding to the computing channel in the current main control resource module and the number of currently online main control resource modules.

[0197] The first analysis subunit is used to determine whether the current main control resource module is faulty based on the first fault result, the second fault result, the third fault result, and the fourth fault result.

[0198] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, and will not be repeated here.

[0199] The second analysis unit 902 is used to determine whether there are other main control resource modules malfunctioning based on the periodic broadcast messages received by the computing channel in the current main control resource module and the system periodic status table.

[0200] Optionally, in another embodiment of the present invention, one implementation of the second analysis unit 902 includes:

[0201] The fifth fault determination subunit is used to determine the result of the fifth fault based on the running status field in the periodic broadcast message received by the computing channel in the current main control resource module.

[0202] The sixth fault determination subunit is used to determine the result of the sixth fault based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current main control resource module and the number of checkpoints in the online main control resource module.

[0203] The second analysis subunit is used to determine whether there are other main control resource modules that have failed, based on the fifth and sixth fault results.

[0204] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, and will not be repeated here.

[0205] The switching control unit 903 is used to perform master-slave switching control based on the master-slave type of any master control resource module if any master control resource module fails.

[0206] For details on the specific operation of the units disclosed in the above embodiments of the present invention, please refer to the corresponding method embodiments, such as... Figure 2 As shown, it will not be elaborated further here.

[0207] Optionally, in another embodiment of the present invention, one implementation of the switching control unit 903 includes:

[0208] The operation status determination subunit is used to determine whether the target main control resource module can operate normally if the main control resource module that has failed belongs to the main standby type of the standby system.

[0209] The target master control resource module is another master control resource module in the master-slave subsystem to which the failed master control resource module belongs.

[0210] The first switching control subunit is used to continue outputting security data using the target master control resource module if the target master control resource module can operate normally.

[0211] The second switching control subunit is used to switch the target main control resource module's main / standby type to main if the main control resource module that has failed belongs to the main / standby type, and then continue to output safety data using the target main control resource module.

[0212] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, and will not be repeated here.

[0213] Optionally, in another embodiment of the present invention, one implementation of the main / standby control device further includes:

[0214] The replacement unit is used to replace the main control resource module of the main standby subsystem to which the faulty main control resource module belongs with the main control resource module in the third backup unit.

[0215] If at least one main control resource module in the main control unit is not running security services, the main control unit will be used as the third backup unit. The main control unit consists of two consecutive main control resource modules.

[0216] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, and will not be repeated here.

[0217] As can be seen from the above scheme, the present invention provides a primary / standby control device. Based on the system periodic status table and self-test results of the computing channels in the current primary control resource module, it determines whether the current primary control resource module is faulty; and based on the periodic broadcast messages received by the computing channels in the current primary control resource module and the system periodic status table, it determines whether other primary control resource modules are faulty. If any primary control resource module fails, primary / standby switching control is completed according to the primary / standby type of the primary control resource module. This invention does not rely on any external dedicated equipment, but rather makes each primary control resource module in the cloud platform an equal participant in the decision-making network. Through the system periodic status table and periodic broadcast messages, a decentralized, highly reliable collective decision-making system is formed, achieving the goal of effective primary / standby control even in a cloud secure computing platform.

[0218] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0219] Another embodiment of the present invention provides an electronic device, such as... Figure 10 As shown, it includes:

[0220] One or more processors 1001.

[0221] Storage device 1002, on which one or more programs are stored.

[0222] When the one or more programs are executed by the one or more processors 1001, the one or more processors 1001 implement the primary / backup control method as described in the above embodiments.

[0223] Another embodiment of the present invention provides a computer storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the primary / backup control method as described in the above embodiments.

[0224] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0225] It should be noted that the computer-readable medium described above in this invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0226] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0227] Another embodiment of the present invention provides a computer program product, which, when executed, is used to perform the above-described master / slave control method.

[0228] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, it performs the functions defined in the methods of the embodiments of the present invention.

[0229] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in this invention is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely exemplary forms for implementing the invention.

[0230] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of the invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0231] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention is not limited to the specific combination of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with technical features of the present invention (but not limited to) that have similar functions.

Claims

1. A method for controlling a primary and a backup, characterized by, include: Based on the system periodic status table of the computing channels in the current master control resource module and the self-test results, it is determined whether the current master control resource module is faulty; wherein, the current master control resource module periodically sends periodic broadcast messages to the system and receives periodic broadcast messages sent by other master control resource modules; each master control resource module includes two computing channels, and each computing channel periodically updates its corresponding system periodic status table; Based on the periodic broadcast messages received by the computing channel in the current master control resource module and the system periodic status table, determine whether there are other master control resource modules malfunctioning; If any main control resource module fails, the main / standby switchover control will be completed according to the main control resource module’s main / standby type. The step of determining whether the current main control resource module is faulty based on the system cycle status table of the computing channel in the current main control resource module and the self-test results includes: A functional safety self-test is performed on the computing channel cycle in the current main control resource module to obtain the first fault result; The system periodic status table corresponding to the computing channel in the current main control resource module is periodically checked, and the second fault result is determined based on the online status of all main control resource modules in the system periodic status table. The third fault result is determined based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current master control resource module, the identifier set of the fault detection module, and the number of fault detection modules; wherein, the fault detection module is the master control resource module that detected the fault in the current master control resource module. The fourth fault result is determined based on the number of fault-aware modules in the system cycle status table corresponding to the computing channel in the current main control resource module and the number of currently online main control resource modules. Based on the first fault result, the second fault result, the third fault result, and the fourth fault result, determine whether the current main control resource module is faulty; Any two main control resource modules running security services constitute a primary / standby subsystem; the primary / standby type of any main control resource module in the primary / standby subsystem is determined by its unique number. If any main control resource module fails, the main / standby switchover control is performed according to the main control resource module's main / standby type, including: If the faulty master control resource module belongs to a master-slave type that is a standby system, then determine whether the target master control resource module can operate normally; wherein, the target master control resource module is another master control resource module in the master-slave subsystem to which the faulty master control resource module belongs; If the target master control resource module can operate normally, continue to output security data using the target master control resource module; If the primary / standby type of the faulty master control resource module is primary, then the primary / standby type of the target master control resource module is changed to primary, and the target master control resource module continues to output security data.

2. The master-slave control method according to claim 1, wherein The current master control resource module periodically sends periodic broadcast messages to the system, including: The first computing channel in the current main control resource module sends periodic broadcast messages to the system through the first local area network and the first control local area network bus; The second computing channel in the current master control resource module sends periodic broadcast messages to the system through the second local area network and the second control local area network bus.

3. The master-slave control method according to claim 2, wherein The periodic broadcast message includes module information of the computing channel, the operating status of the computing channel, and perception information of other main control resource modules. Before the current main control resource module periodically sends the periodic broadcast message to the system, it also includes: For each computing channel in the current master control resource module, obtain the module information of the computing channel; wherein, the module information of the computing channel includes the location number of the computing channel, the computing channel number, and the current system role; For each computing channel in the current master control resource module, a functional safety self-test is performed on the computing channel to obtain the self-test result, and the self-test result is sent to the adjacent channel in each running cycle; the adjacent channel is another computing channel in the current master control resource module. Based on the self-test results of all the computing channels, determine the current operating status of the main control resource module; Based on the received periodic broadcast messages, the status of other master control resource modules is perceived, and the perception information of other master control resource modules is obtained.

4. The master-slave control method according to claim 3, wherein The process of sensing the status of other master control resource modules based on received periodic broadcast messages to obtain sensing information of other master control resource modules includes: Receive a first target periodic broadcast message; wherein the first target periodic broadcast message is received by the first computing channel through the first local area network and the first control local area network bus; the first target periodic broadcast message is a periodic broadcast message sent by the third computing channel; the third computing channel is any computing channel under the same local area network and the same control local area network bus as the first computing channel; The first target periodic broadcast message is transmitted to the second computing channel, and the second target periodic broadcast message transmitted by the second computing channel is received simultaneously. Based on the first target periodic broadcast message and the second target periodic broadcast message, the perception information of the master control resource module to which the third computing channel belongs is determined.

5. The master-slave control method according to claim 1, wherein The step of determining whether there are other main control resource modules malfunctioning based on the periodic broadcast messages received by the computing channel in the current main control resource module and the system periodic status table includes: The fifth fault result is determined based on the running status field in the periodic broadcast message received by the computing channel in the current master control resource module; The sixth fault result is determined based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current main control resource module and the number of checkpoints in the online main control resource module; Based on the fifth and sixth fault results, determine whether there are other main control resource modules that are faulty.

6. The master-slave control method according to claim 1, wherein After continuing to output security data using the target master control resource module, the method further includes: The main control resource module in the third backup unit is used to replace the main control resource module of the main backup subsystem to which the faulty main control resource module belongs. If at least one main control resource module in the main control unit is not running security services, then the main control unit is used as the third backup unit. The main control unit consists of two consecutive main control resource modules.

7. A master / standby control device characterized by comprising: include: The first analysis unit is used to determine whether the current master control resource module is faulty based on the system periodic status table of the computing channels in the current master control resource module and the self-test results. The current master control resource module periodically sends periodic broadcast messages to the system and receives periodic broadcast messages from other master control resource modules. Each master control resource module includes two computing channels, and each computing channel periodically updates its corresponding system periodic status table. Any two master control resource modules running security services constitute a master-slave subsystem. The master control resource module in any master-slave subsystem determines its master-slave type based on its unique number. The second analysis unit is used to determine whether there are other main control resource modules malfunctioning based on the periodic broadcast messages received by the computing channel in the current main control resource module and the system periodic status table. The switching control unit is used to perform master-slave switching control according to the master-slave type of any master control resource module if any master control resource module fails. The first analysis unit includes: The first fault determination subunit is used to perform a functional safety self-check on the computing channel cycle in the current main control resource module and obtain the first fault result. The second fault determination subunit is used to periodically check the system periodic status table corresponding to the computing channel in the current main control resource module, and determine the second fault result based on the online status of all main control resource modules in the system periodic status table. The third fault determination subunit is used to determine the third fault result based on the periodic checkpoint count value in the system periodic status table corresponding to the computing channel in the current main control resource module, the identifier set of the fault perception module, and the number of fault perception modules. Among them, the fault perception module is the main control resource module that detects a fault in the current main control resource module; The fourth fault determination subunit is used to determine the fourth fault result based on the number of fault perception modules in the system cycle status table corresponding to the computing channel in the current main control resource module and the number of currently online main control resource modules. The first analysis subunit is used to determine whether the current main control resource module is faulty based on the first fault result, the second fault result, the third fault result, and the fourth fault result. The switching control unit includes: The operation status determination subunit is used to determine whether the target main control resource module can operate normally if the main control resource module to which the faulty main control resource module belongs is a standby system; wherein, the target main control resource module is another main control resource module in the main standby subsystem to which the faulty main control resource module belongs; The first switching control subunit is used to continue outputting security data using the target master control resource module if the target master control resource module can operate normally. The second switching control subunit is used to switch the primary / standby type of the target primary control resource module to primary if the primary / standby type of the faulty primary control resource module is primary, and to continue outputting security data using the target primary control resource module.