A hierarchical strategy-based centralized fault monitoring and collaborative processing method and system
Patent Information
- Application Number
- CN202610864831.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-11
AI Technical Summary
[0006]本发明提供一种基于分层策略集中式故障监控与协同处理方法及系统,解决了当前自动驾驶系统的故障管理多采用分布式架构,各子系统独立处理本地故障,缺乏统一协同机制,导致系统级故障响应能力不足的问题
[0025] This invention provides a centralized fault monitoring and collaborative processing method and system based on a hierarchical strategy.
Smart Images

Figure CN122733641A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving, and in particular to a method and system for centralized fault monitoring and collaborative processing based on a hierarchical strategy. Background Technology
[0002] With the rapid development of autonomous driving technology, system complexity and safety requirements are increasing. Traditional autonomous driving systems typically employ a distributed architecture for fault management, with each functional subsystem handling local faults independently.
[0003] The actual deployment of large-scale autonomous driving systems involves huge safety risks, including the chain reaction of system-level failures and the failure of multiple subsystems in coordination. In addition, in actual operation, there are safety hazards such as response delays and inconsistent decisions when dealing with complex fault scenarios.
[0004] Traditional distributed fault management solutions have the following limitations: each subsystem can only make decisions based on local information and lacks the ability to coordinate processing from a global perspective; fault strategies are usually hard-coded in the program, making it difficult to adapt to different operating scenarios and fault modes.
[0005] Therefore, it is necessary to provide a hierarchical strategy-based centralized fault monitoring and collaborative processing method and system to solve the above-mentioned technical problems. Summary of the Invention
[0006] This invention provides a hierarchical strategy-based centralized fault monitoring and collaborative processing method and system, which solves the problem that current autonomous driving systems mostly adopt a distributed architecture for fault management, with each subsystem independently handling local faults and lacking a unified collaborative mechanism, resulting in insufficient system-level fault response capabilities.
[0007] To address the aforementioned technical problems, this invention provides a centralized fault monitoring and collaborative processing method based on a hierarchical strategy, comprising the following steps:
[0008] S101: System initialization, loading the hierarchical fault handling strategy; when the system starts, the hierarchical fault handling strategy is loaded from the external configuration file. The hierarchical fault handling strategy includes a local fast response strategy executed by each functional sub-node and a global collaborative strategy managed by the central node.
[0009] S102: Each functional sub-node performs local fault detection and response; each functional sub-node performs fault detection based on local policies, and immediately performs a fast response when a fault is detected, while encapsulating the fault status information into a standardized fault message format;
[0010] S103: The central node collects and updates the global fault status; the central node subscribes to and receives fault status messages from multiple functional sub-nodes, and uses a mutex lock mechanism to update the global fault status table in a thread-safe manner;
[0011] S104: The central node performs policy evaluation and generates collaborative instructions; the central node performs policy evaluation at fixed intervals based on the global fault status table and global policy, and generates a set of collaborative action instructions to be triggered.
[0012] S105: Issue collaborative instructions and execute system-level responses; the central node periodically broadcasts collaborative action instruction sets and global system fault status to all functional sub-nodes, and each functional sub-node executes the corresponding collaborative response actions upon receiving them.
[0013] Preferably, the external configuration file predefines all fault types and their attribute characteristics supported by the system, forming a fault list; the local policy and global policy support logical combinations of multiple condition types, including basic fault conditions, composite logical conditions, and system state conditions; each functional sub-node performs fault detection and rapid response based on the local policy, and encapsulates the fault status information into a standardized fault message format and reports it to the central node; the central node subscribes to and receives fault status messages from multiple functional sub-nodes; the central node uses a mutex lock mechanism to thread-safely update the received fault status of each functional sub-node to a unified global fault status table; the global fault status table uses a bit storage structure, achieving efficient storage and fast access to fault status through bit operations; the central node performs policy evaluation at fixed intervals based on the global fault status table and the global policy, generating a set of cooperative action instructions to be triggered; the central node periodically broadcasts the set of cooperative action instructions and the global system fault status to all functional sub-nodes.
[0014] Preferably, the global fault status table uses a bit storage structure, supports batch fault status operations based on bitmasks, and can atomically update the fault status corresponding to a specified functional sub-node simultaneously.
[0015] Preferably, the global fault status table adopts modular bitmap management, dividing the fault status storage area according to functional modules, with each module having an independent fault status storage space.
[0016] Preferably, the attribute characteristics of the fault include at least one of the following: fault identifier, fault name, fault description information, default severity level, and functional safety level.
[0017] Preferably, the basic fault condition is determined based on a single fault state, the composite logic condition combines multiple fault conditions through logical operators, and the system state condition is determined based on a numerical comparison of system operating parameters.
[0018] Preferably, the composite logic condition supports nested combinations of logical operators, including multi-level combinations of AND, OR, and NOT logical operations, and supports the order of logical operations based on priority.
[0019] Preferably, the global policy supports the configuration and management of policy priorities. When multiple policy conditions are met simultaneously, the central node executes the corresponding collaborative actions according to the preset priority order. The functional sub-nodes maintain independent local fault states and execute corresponding local policies. When a fault is detected, the functional sub-nodes first respond quickly based on the local policy and then encapsulate the fault state information into a standardized fault message format and report it to the central node.
[0020] Preferably, the central node periodically evaluates the global fault status table, generates and publishes a set of collaborative action instructions and the global fault status; after receiving the instructions, each functional sub-node executes the corresponding collaborative response action according to the instruction set.
[0021] Preferably, the centralized fault monitoring and collaborative processing system based on a hierarchical strategy-based centralized fault monitoring and collaborative processing method includes a global management unit deployed at the central node and local execution units deployed at each functional sub-node:
[0022] The global management unit includes a policy configuration module, a fault status collection module, a global status maintenance module, a policy evaluation module, and a fault information publishing module. The policy configuration module loads hierarchical fault handling policies based on a fault list from an external configuration file. The fault status collection module subscribes to and receives fault status messages reported by each local execution unit. The global status maintenance module maintains a unified global fault status table using a bit storage structure and ensures thread-safe update operations through a mutex lock mechanism. The policy evaluation module performs policy evaluation at fixed intervals and generates a set of cooperative action instructions based on the global fault status table and global policies. The fault information publishing module periodically broadcasts the set of cooperative action instructions and the system's global fault status to all functional sub-nodes.
[0023] The local execution unit includes a fault detection module, a local policy execution module, and a fault status reporting module. The local execution unit is configured to execute local rapid response strategies on child nodes, perform local actions when a fault is detected, and encapsulate the fault status information into a standardized fault message format for reporting. The fault detection module monitors the operating status of the local node in real time and identifies various faults. The local policy execution module executes the local rapid response strategy and receives and executes global collaborative instructions. The fault status reporting module encapsulates the fault status into a standard message format and reports it to the central node.
[0024] Compared with related technologies, the centralized fault monitoring and collaborative processing method and system based on a hierarchical strategy provided by this invention has the following beneficial effects:
[0025] This invention provides a centralized fault monitoring and collaborative processing method and system based on a hierarchical strategy. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating a centralized fault monitoring and collaborative processing method based on a hierarchical strategy provided by the present invention.
[0027] Figure 2 This invention provides an implementation diagram of a centralized fault monitoring and collaborative processing method based on a hierarchical strategy.
[0028] Figure 3 This invention provides a system architecture block diagram for a centralized fault monitoring and collaborative processing system based on a hierarchical strategy.
[0029] The diagram is labeled as follows: 310, Global Management Unit; 311, Policy Configuration Module; 312, Fault Status Collection Module; 313, Global Status Maintenance Module; 314, Policy Evaluation Module; 315, Fault Information Release Module; 320, Local Execution Unit; 321, Fault Detection Module; 322, Local Policy Execution Module; 323, Fault Status Reporting Module. Detailed Implementation
[0030] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0031] First Embodiment
[0032] Please refer to the following: Figure 1 and Figure 2 ,in, Figure 1 This is a flowchart illustrating a centralized fault monitoring and collaborative processing method based on a hierarchical strategy provided by the present invention. Figure 2 This invention provides an implementation diagram of a centralized fault monitoring and collaborative processing method based on a hierarchical strategy.
[0033] A centralized fault monitoring and collaborative processing method based on a hierarchical strategy includes: S101: System initialization, loading a hierarchical fault processing strategy; when the system starts, the hierarchical fault processing strategy is loaded from an external configuration file, wherein the hierarchical fault processing strategy includes a local fast response strategy executed by each functional sub-node and a global collaborative strategy managed by the central node.
[0034] S102: Each functional sub-node performs local fault detection and response; each functional sub-node performs fault detection based on local policies, and immediately performs a fast response when a fault is detected, while encapsulating the fault status information into a standardized fault message format;
[0035] S103: The central node collects and updates the global fault status; the central node subscribes to and receives fault status messages from multiple functional sub-nodes, and uses a mutex lock mechanism to update the global fault status table in a thread-safe manner;
[0036] S104: The central node performs policy evaluation and generates collaborative instructions; the central node performs policy evaluation at fixed intervals based on the global fault status table and global policy, and generates a set of collaborative action instructions to be triggered.
[0037] S105: Issue collaborative instructions and execute system-level responses; the central node periodically broadcasts collaborative action instruction sets and global system fault status to all functional sub-nodes, and each functional sub-node executes the corresponding collaborative response actions upon receiving them.
[0038] The external configuration file predefines all fault types and their attribute characteristics supported by the system, forming a fault list. The local and global policies support logical combinations of various condition types, including basic fault conditions, composite logical conditions, and system state conditions. Each functional sub-node performs fault detection and rapid response based on the local policy, and encapsulates the fault status information into a standardized fault message format and reports it to the central node. The central node subscribes to and receives fault status messages from multiple functional sub-nodes. The central node uses a mutex lock mechanism to update the received fault status of each functional sub-node to a unified global fault status table in a thread-safe manner. The global fault status table uses a bit storage structure to achieve efficient storage and fast access to fault status through bit operations. The central node performs policy evaluation at fixed intervals based on the global fault status table and the global policy, generating a set of cooperative action instructions to be triggered. The central node periodically broadcasts the set of cooperative action instructions and the global system fault status to all functional sub-nodes.
[0039] The global fault status table uses a bit storage structure, supports batch fault status operations based on bitmasks, and can atomically update the fault status corresponding to a specified functional sub-node simultaneously.
[0040] The global fault status table adopts modular bitmap management, and the fault status storage area is divided according to functional modules. Each module has an independent fault status storage space.
[0041] The attribute characteristics of the fault include at least one of the following: fault identifier, fault name, fault description information, default severity level, and functional safety level.
[0042] The basic fault conditions are judged based on a single fault state, the composite logic conditions combine multiple fault conditions through logical operators, and the system state conditions are judged based on the numerical comparison of system operating parameters.
[0043] The compound logic condition supports nested combinations of logical operators, including multi-level combinations of AND, OR, and NOT logical operations, and supports the order of logical operations based on priority.
[0044] The global policy supports the configuration and management of policy priorities. When multiple policy conditions are met at the same time, the central node executes the corresponding collaborative actions according to the preset priority order. The functional sub-nodes maintain independent local fault states and execute corresponding local policies. When a fault is detected, the functional sub-nodes first respond quickly based on the local policy, and at the same time encapsulate the fault state information into a standardized fault message format and report it to the central node.
[0045] The central node periodically evaluates the global fault status table, generates and publishes a set of collaborative action instructions and global fault status; after receiving the instructions, each functional sub-node executes the corresponding collaborative response action according to the instruction set.
[0046] When the system starts, it loads the hierarchical fault handling strategy from the external configuration file. The hierarchical fault handling strategy includes local fast response strategies executed by each functional sub-node and global collaborative strategies managed by the central node.
[0047] Each functional sub-node performs fault detection based on local policies. When a fault is detected, a rapid response is immediately executed, and the fault status information is encapsulated into a standardized fault message format.
[0048] The central node subscribes to and receives fault status messages from multiple functional child nodes, and uses a mutex lock mechanism to update the global fault status table in a thread-safe manner.
[0049] The central node performs policy evaluation at fixed intervals based on the global fault status table and global policy, and generates a set of collaborative action instructions to be triggered.
[0050] The central node periodically broadcasts a set of collaborative action instructions and the global fault status of the system to all functional sub-nodes. Each functional sub-node then executes the corresponding collaborative response action upon receiving the instructions.
[0051] The centralized fault monitoring and collaborative processing method based on a hierarchical strategy, as described in this application, achieves rapid fault detection, collaborative processing, and system-level safety assurance for autonomous driving systems through unified fault definitions, efficient bit storage structures, and flexible strategy configuration. Using this centralized fault management method significantly improves the system's response speed, resource utilization efficiency, and operational reliability.
[0052] The working principle of the hierarchical strategy-based centralized fault monitoring and collaborative processing method provided by this invention is as follows:
[0053] During operation, the central node and each functional sub-node are started first, and a communication connection is established according to the API interface specifications provided by the system. The policy configuration module of the central node reads the configuration file, which contains three main parts: the fault definition section enumerates all fault types and their attribute characteristics supported by the system; the local policy section defines the fast response rules of each sub-node; and the global policy section defines the rules for complex faults that need to be processed collaboratively.
[0054] During the execution of local fault detection and response, the fault detection logic is implemented based on the C++ programming language. After each child node detects a fault, the local policy executor immediately triggers a predefined response action. The message processing component encodes the fault status into a standard message format that includes a timestamp, node ID, and fault status bitmap.
[0055] During the update of the global fault state, the state collection module listens for fault message topics through the ROS2 subscription mechanism, and the global state maintenance module manages the fault state using a bit storage structure. Each fault corresponds to a specific bit in the bitmap. By acquiring a mutex lock, performing bit operations to update the state, and releasing the mutex lock in an atomic operation sequence, data consistency is ensured when multiple threads access the data concurrently.
[0056] During the policy evaluation process, the policy evaluation engine is triggered periodically at 100ms intervals. The evaluation process includes: creating a copy of the global fault status table, parsing the conditional expressions in the global policy, performing compound logic judgments and priority sorting, and finally generating a set of cooperative action instructions containing target nodes, instruction types and execution parameters.
[0057] Compared with related technologies, the centralized fault monitoring and collaborative processing method based on a hierarchical strategy provided by this invention has the following beneficial effects:
[0058] This invention provides a centralized fault monitoring and collaborative processing method based on a hierarchical strategy. By constructing a centralized fault monitoring node and a configurable hierarchical strategy mechanism, it achieves unified decision-making and collaborative response to distributed faults, thereby elevating local fault handling to the global system level. This system can execute optimal safety strategies when dealing with single-point and compound faults, significantly improving the overall robustness and functional safety level of the autonomous driving system. Through unified fault definitions, efficient bit storage structures, and flexible strategy configurations, it achieves rapid detection, collaborative processing, and system-level safety assurance of autonomous driving system faults. Using this centralized fault management method significantly improves the system's response speed, resource utilization efficiency, and operational reliability.
[0059] Second Embodiment
[0060] Please refer to the following: Figure 3 Based on the first embodiment of this application, which provides a hierarchical strategy-based centralized fault monitoring and collaborative processing method and system, the second embodiment of this application proposes another hierarchical strategy-based centralized fault monitoring and collaborative processing method and system. The second embodiment is merely a preferred embodiment of the first embodiment, and its implementation will not affect the individual implementation of the first embodiment.
[0061] Specifically, the second embodiment of this application provides a centralized fault monitoring and collaborative processing system based on a hierarchical strategy, which differs in that the centralized fault monitoring and collaborative processing system based on the hierarchical strategy includes a global management unit 310 deployed at the central node and local execution units 320 deployed at each functional sub-node.
[0062] The global management unit 310 includes a strategy configuration module 311, a fault status collection module 312, a global status maintenance module 313, a strategy evaluation module 314, and a fault information publishing module 315. The strategy configuration module 311 loads a hierarchical fault handling strategy based on a fault list from an external configuration file. The fault status collection module 312 subscribes to and receives fault status messages reported by each local execution unit. The global status maintenance module 313 maintains a unified global fault status table using a bit storage structure and ensures thread-safe update operations through a mutex lock mechanism. The strategy evaluation module 314 performs strategy evaluation at fixed intervals and generates a set of cooperative action instructions based on the global fault status table and global strategy. The fault information publishing module 315 periodically broadcasts the set of cooperative action instructions and the system's global fault status to all functional sub-nodes.
[0063] The local execution unit 320 includes a fault detection module 321, a local policy execution module 322, and a fault status reporting module 323. The local execution unit 320 is configured to execute a local fast response strategy on the child node, perform local actions when a fault is detected, and encapsulate the fault status information into a standardized fault message format for reporting. The fault detection module 321 monitors the running status of the node in real time and identifies various faults. The local policy execution module 322 executes the local fast response strategy and receives and executes global coordination instructions. The fault status reporting module 323 encapsulates the fault status into a standard message format and reports it to the central node.
[0064] For specific limitations regarding the centralized fault monitoring and collaborative processing system based on the hierarchical strategy, please refer to the limitations of the centralized fault monitoring and collaborative processing method based on the hierarchical strategy mentioned above, which will not be repeated here. Each module of the centralized fault monitoring and collaborative processing system based on the hierarchical strategy can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in the processor of the computer device in hardware form or independent of it, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0065] Compared with related technologies, the centralized fault monitoring and collaborative processing system based on a hierarchical strategy provided by this invention has the following beneficial effects:
[0066] This invention provides a system for centralized fault monitoring and collaborative processing based on a hierarchical strategy. Through the unified management of the central node and the distributed execution of each functional sub-node, a complete fault monitoring and collaborative processing system is constructed. The system adopts a bit storage structure to achieve efficient management of fault status, supports flexible fault processing logic through hierarchical strategy configuration, and uses a mutex lock mechanism to ensure data consistency in a multi-threaded environment. Ultimately, it realizes system-level collaborative processing and safety assurance of autonomous driving system faults.
[0067] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A centralized fault monitoring and collaborative processing method based on a hierarchical strategy, characterized in that, Includes the following steps: S101: System initialization, loading the hierarchical fault handling strategy; when the system starts, the hierarchical fault handling strategy is loaded from the external configuration file. The hierarchical fault handling strategy includes a local fast response strategy executed by each functional sub-node and a global collaborative strategy managed by the central node. S102: Each functional sub-node performs local fault detection and response; each functional sub-node performs fault detection based on local policies, and immediately performs a fast response when a fault is detected, while encapsulating the fault status information into a standardized fault message format; S103: The central node collects and updates the global fault status; the central node subscribes to and receives fault status messages from multiple functional sub-nodes, and uses a mutex lock mechanism to update the global fault status table in a thread-safe manner; S104: The central node performs policy evaluation and generates collaborative instructions; the central node performs policy evaluation at fixed intervals based on the global fault status table and global policy, and generates a set of collaborative action instructions to be triggered. S105: Issue collaborative instructions and execute system-level responses; the central node periodically broadcasts collaborative action instruction sets and global system fault status to all functional sub-nodes, and each functional sub-node executes the corresponding collaborative response actions upon receiving them.
2. The method for centralized fault monitoring and collaborative processing based on a hierarchical strategy according to claim 1, characterized in that, The external configuration file predefines all fault types and their attribute characteristics supported by the system, forming a fault list; the local and global policies support logical combinations of various condition types, including basic fault conditions, composite logical conditions, and system status conditions; each functional sub-node performs fault detection and rapid response based on the local policy, and encapsulates the fault status information into a standardized fault message format and reports it to the central node. The central node subscribes to and receives fault status messages from multiple functional sub-nodes; the central node uses a mutex lock mechanism to update the received fault status of each functional sub-node to a unified global fault status table in a thread-safe manner. The global fault status table adopts a bit storage structure, which enables efficient storage and fast access to fault status through bit operations. The central node performs policy evaluation based on the global fault status table and the global policy at a fixed period to generate a set of collaborative action instructions to be triggered. The central node broadcasts the set of collaborative action instructions and the global fault status of the system to all functional sub-nodes periodically.
3. The method for centralized fault monitoring and collaborative processing based on a hierarchical strategy according to claim 1, characterized in that, The global fault status table uses a bit storage structure, supports batch fault status operations based on bitmasks, and can atomically update the fault status corresponding to a specified functional sub-node simultaneously.
4. The method for centralized fault monitoring and collaborative processing based on a hierarchical strategy according to claim 1, characterized in that, The global fault status table adopts modular bitmap management, and the fault status storage area is divided according to functional modules. Each module has an independent fault status storage space.
5. The method for centralized fault monitoring and collaborative processing based on a hierarchical strategy according to claim 1, characterized in that, The attribute characteristics of the fault include at least one of the following: fault identifier, fault name, fault description information, default severity level, and functional safety level.
6. The method for centralized fault monitoring and collaborative processing based on a hierarchical strategy according to claim 2, characterized in that, The basic fault conditions are judged based on a single fault state, the composite logic conditions combine multiple fault conditions through logical operators, and the system state conditions are judged based on the numerical comparison of system operating parameters.
7. The method for centralized fault monitoring and collaborative processing based on a hierarchical strategy according to claim 2, characterized in that, The compound logic condition supports nested combinations of logical operators, including multi-level combinations of AND, OR, and NOT logical operations, and supports the order of logical operations based on priority.
8. The method for centralized fault monitoring and collaborative processing based on a hierarchical strategy according to claim 1, characterized in that, The global policy supports the configuration and management of policy priorities. When multiple policy conditions are met at the same time, the central node executes the corresponding collaborative actions according to the preset priority order. The functional sub-nodes maintain independent local fault states and execute corresponding local policies. When a fault is detected, the functional sub-nodes first respond quickly based on the local policy, and at the same time encapsulate the fault state information into a standardized fault message format and report it to the central node.
9. The method for centralized fault monitoring and collaborative processing based on a hierarchical strategy according to claim 1, characterized in that, The central node periodically evaluates the global fault status table, generates and publishes a set of collaborative action instructions and global fault status; after receiving the instructions, each functional sub-node executes the corresponding collaborative response action according to the instruction set.
10. A centralized fault monitoring and collaborative processing system for implementing the centralized fault monitoring and collaborative processing method based on a hierarchical strategy as described in any one of claims 1 to 9, characterized in that, This includes a global management unit deployed at the central node and local execution units deployed at each functional sub-node: The global management unit includes a policy configuration module, a fault status collection module, a global status maintenance module, a policy evaluation module, and a fault information publishing module; the policy configuration module is used to load hierarchical fault handling policies based on fault lists from external configuration files; The fault status collection module is used to subscribe to and receive fault status messages reported by each local execution unit; The global state maintenance module is used to maintain a unified global fault state table with a bit storage structure and to ensure thread-safe update operations through a mutex lock mechanism; the strategy evaluation module is used to perform strategy evaluation at fixed intervals and generate a set of cooperative action instructions based on the global fault state table and the global strategy; the fault information publishing module is used to publish the set of cooperative action instructions and the global fault state of the system to all functional sub-nodes in a periodic broadcast manner. The local execution unit includes a fault detection module, a local policy execution module, and a fault status reporting module. The local execution unit is configured to execute a local fast response policy on the child node, perform local actions when a fault is detected, and encapsulate the fault status information into a standardized fault message format for reporting. The fault detection module monitors the operating status of the node in real time and identifies various faults; the local policy execution module executes the local fast response policy and receives and executes global collaborative instructions. The fault status reporting module encapsulates the fault status into a standard message format and reports it to the central node.