Fault processing method and device, equipment and medium
By implementing remote start, automatic restart and hot standby switch in the ROS2 system, the problems of low node startup efficiency and inconvenient management are solved, the system stability and self-recovery capabilities are improved, and it is suitable for autonomous driving and industrial control.
Patent Information
- Application Number
- CN202510454994.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-01
AI Technical Summary
In complex environments, the robot operating system ROS2 has problems such as low node startup efficiency, inconvenient management, imperfect automatic restart mechanism and lack of hot standby switching capabilities, resulting in insufficient system stability and security.
By assigning robot operating system node tasks to each computing device, obtaining configuration files and startup parameters, using the remote startup module to realize remote startup and unified management of nodes, configuring daemon programs to monitor node status in real time, and automatically restarting or migrating node processes in the event of a failure, and using the hot standby switch module to automatically switch to adjacent devices when the device fails.
It improves the high availability and self-restoration capabilities of the ROS2 system, improves node management efficiency, and realizes a more stable and reliable computing platform, suitable for critical mission scenarios such as autonomous driving and industrial control.
Smart Images

Figure CN120407251A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of detection technology, and particularly to a fault handling method, device, equipment and medium. Background Art
[0002] The Robot Operating System (ROS2 system) runs in a complex environment and may face problems such as hardware failures, software anomalies, sensor failures, network interruptions, etc. Therefore, a fault tolerance mechanism is required to improve the stability and security of the robot operating system.
[0003] Currently, based on the robot operating system, the following main disadvantages exist:
[0004] The decentralized design of the robot operating system makes it impossible to efficiently implement the startup of distributed nodes. When multiple computing devices run ROS2 system tasks simultaneously, it is necessary to manually start or stop ROS2 nodes on each computing device respectively. This method is inefficient and the ROS2 nodes are not easy to manage;
[0005] The automated restart mechanism is imperfect: when a key ROS2 node crashes, it is necessary for manual intervention to check the log and perform a manual restart, and the system cannot implement the dynamic perception and autonomous recovery functions;
[0006] The lack of hot standby switching ability: The key ROS2 nodes lack a real-time status synchronization mechanism between the primary and standby instances. When the primary node crashes, the standby node cannot directly inherit the running status of the primary node and needs to be re-initialized, resulting in the interruption of the service continuity of the robot operating system. Summary of the Invention
[0007] This application provides a fault handling method, device, equipment and medium. The method includes allocating robot operating system node tasks to each computing device; obtaining the robot operating system node configuration file and the robot operating system node startup parameters; starting the robot operating system nodes of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters; in response to a failure of the robot operating system node process of any one of the computing devices, restarting the robot operating system node of the computing device through the daemon process program of the computing device; performing misaligned loop monitoring on several computing devices, and in response to a failure of any one of the computing devices, allocating the robot operating system node process on the failed computing device to an adjacent computing device and starting it. This application improves the high availability and self-recovery ability of the ROS2 system, and realizes a more stable and reliable ROS2 computing platform.
[0008] This application provides a fault handling method. The method is applied to a fault handling system, and the fault handling system includes several computing devices. The method includes:
[0009] Allocate robot operating system node tasks to each computing device;
[0010] Obtain the robot operating system node configuration file and the robot operating system node startup parameters;
[0011] Start the robot operating system nodes of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters;
[0012] In response to a failure of the robot operating system node process of any one computing device, restart the robot operating system node of the computing device through the daemon process program of the computing device;
[0013] Perform staggered loop monitoring on several computing devices. In response to a failure of any one computing device, allocate the robot operating system node process on the failed computing device to an adjacent computing device and start it.
[0014] The present application also provides a fault handling device, which includes:
[0015] An allocation module for allocating robot operating system node tasks to each computing device;
[0016] An acquisition module for obtaining the robot operating system node configuration file and the robot operating system node startup parameters;
[0017] A remote startup module for starting the robot operating system nodes of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters;
[0018] An automatic restart module for restarting the robot operating system node of a computing device through the daemon process program of the computing device in response to a failure of the robot operating system node process of any one computing device;
[0019] A hot standby switching module for performing staggered loop monitoring on several computing devices. In response to a failure of any one computing device, allocate the robot operating system node process on the failed computing device to an adjacent computing device and start it.
[0020] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of the fault handling method when executing the computer program, and the method includes:
[0021] Allocate robot operating system node tasks to each computing device;
[0022] Obtain the robot operating system node configuration file and the robot operating system node startup parameters;
[0023] Start the Robot Operating System nodes of each computing device according to the Robot Operating System node configuration file and the Robot Operating System node startup parameters;
[0024] In response to a failure of the Robot Operating System node process on any one of the computing devices, restart the Robot Operating System node of that computing device through the daemon process program of that computing device;
[0025] Perform staggered loop monitoring on several computing devices. In response to a failure of any one of the computing devices, allocate the Robot Operating System node process on the failed computing device to an adjacent computing device and start it.
[0026] This application also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the fault handling method are implemented. The method includes:
[0027] Allocate Robot Operating System node tasks to each computing device;
[0028] Obtain the Robot Operating System node configuration file and the Robot Operating System node startup parameters;
[0029] Start the Robot Operating System nodes of each computing device according to the Robot Operating System node configuration file and the Robot Operating System node startup parameters;
[0030] In response to a failure of the Robot Operating System node process on any one of the computing devices, restart the Robot Operating System node of that computing device through the daemon process program of that computing device;
[0031] Perform staggered loop monitoring on several computing devices. In response to a failure of any one of the computing devices, allocate the Robot Operating System node process on the failed computing device to an adjacent computing device and start it.
[0032] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the fault handling method are implemented. The method includes:
[0033] Allocate Robot Operating System node tasks to each computing device;
[0034] Obtain the Robot Operating System node configuration file and the Robot Operating System node startup parameters;
[0035] Start the Robot Operating System nodes of each computing device according to the Robot Operating System node configuration file and the Robot Operating System node startup parameters;
[0036] In response to a failure of the robot operating system node process of any computing device, the robot operating system node of the computing device is restarted through the daemon program of the computing device;
[0037] Perform staggered ring monitoring on several computing devices. In response to a failure of any computing device, the robot operating system node process on the faulty computing device is allocated to an adjacent computing device and started.
[0038] Through this application, since the method includes allocating robot operating system node tasks to each computing device; obtaining the robot operating system node configuration file and the robot operating system node startup parameters; starting the robot operating system node of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters; in response to a failure of the robot operating system node process of any computing device, the robot operating system node of the computing device is restarted through the daemon program of the computing device; perform staggered ring monitoring on several computing devices. In response to a failure of any computing device, the robot operating system node process on the faulty computing device is started on an adjacent computing device. Therefore, this application improves the high availability and self-recovery ability of the ROS2 system, improves the management efficiency of ROS2 nodes, is applicable to key mission scenarios such as autonomous driving and industrial control, and can simultaneously achieve a more stable and reliable ROS2 computing platform.
[0039] The technical solution of this application can remotely start and uniformly manage ROS2 nodes on multiple computing devices through the remote start module, avoiding the problem of manually starting nodes on each device one by one in the traditional ROS2 architecture, greatly improving the system deployment efficiency and maintenance convenience, and enhancing the system's self-recovery ability; the automatic restart module monitors the running status of the ROS2 process in real time. When it detects that a key node crashes abnormally, the system can automatically restart the node process, avoiding manual intervention, improving the system's reliability and running stability, and achieving efficient task migration and fault tolerance; the hot standby switching module can automatically allocate all ROS2 node processes on the faulty computing device to other adjacent normal devices when the computing device fails or crashes, realizing seamless migration, ensuring that the task is not interrupted, and enhancing the system's business continuity. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 The first flowchart of the fault handling method provided by the embodiment of the present application;
[0042] Figure 2 The second flowchart of the fault handling method provided by the embodiment of the present application;
[0043] Figure 3 The flowchart of node remote startup provided by the embodiment of the present application;
[0044] Figure 4 The flowchart of automatic restart provided by the embodiment of the present application;
[0045] Figure 5 The flowchart of hot backup switching provided by the embodiment of the present application;
[0046] Figure 6 The specific flowchart of the fault handling method provided by the embodiment of the present application;
[0047] Figure 7 The structural diagram of the fault handling device provided by the embodiment of the present application;
[0048] Figure 8 The exemplary system that can be used to implement the various embodiments described in the present application provided by the embodiment of the present application. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0050] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0051] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0052] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the fault handling method depends, the specific application environment architecture or specific hardware architecture will be described herein.
[0053] The ROS2 system is a robot operating system widely used in the fields of robotics and autonomous driving. It supports multiple computing platforms and can support distributed communication, allowing different components of the robot operating system to run on different computer devices. It has strong flexibility, reliability, and real-time performance.
[0054] A daemon process is a program that runs in the background without direct user interaction control. It is independent of the control terminal and periodically performs certain tasks or waits to handle certain events that occur. It is usually used for system management, service monitoring, task scheduling, etc. functions, and can automatically run after the system starts and automatically recover after the node crashes or exits abnormally.
[0055] The embodiments of this application provide a fault handling method, as Figure 1 shown. The method is applied to a fault handling system, and the fault handling system includes several computing devices. The method includes:
[0056] Allocating robot operating system node tasks to each computing device;
[0057] Obtaining the robot operating system node configuration file and the robot operating system node startup parameters;
[0058] Starting the robot operating system nodes of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters;
[0059] In response to a robot operating system node process failure of any one of the computing devices, restart the robot operating system node of the computing device through the daemon process of the computing device;
[0060] Perform misaligned loop monitoring on several computing devices. In response to a failure of any one of the computing devices, allocate the robot operating system node process on the failed computing device to an adjacent computing device and start it.
[0061] It can be understood that this application provides a fault handling method based on the ROS2 system to improve the fault tolerance and robustness of the ROS2 system.
[0062] To achieve the above object, the present application mainly relates to a node remote start module, an automatic restart module, and a hot standby switching module. The remote start of ROS2 nodes on each computing device and the caching of start node allocation information are realized through the node remote start module; the automatic restart module monitors the running status of ROS2 node processes on each computing device in real time. When the ROS2 node process crashes due to an exception, the automatic restart module reads the ROS2 node start command and restarts the ROS2 process; the hot standby switching module realizes the monitoring of the ROS2 node process status at the computing device level. When any computing device crashes abnormally, the hot standby switching module timely allocates all ROS2 node processes on the crashed device to other adjacent normally running computing devices, achieving high-efficiency fault tolerance.
[0063] An embodiment of the present application provides a fault handling method, which is applied to a fault handling system. The fault handling system includes several computing devices, as Figure 2 shown. The method includes:
[0064] Step S01, allocate robot operating system node tasks to each computing device.
[0065] Specifically, a ROS2 node is a basic component unit in the ROS2 system. Each ROS2 node performs specific tasks, such as sensing data processing, control algorithms, communication, etc.; ROS2 nodes can run on different computing devices for calculation, and ROS2 nodes can interact through topics, services, etc.; the ROS2 system itself does not support the remote start management of ROS2 nodes, and further design and optimization are required to achieve fast and efficient remote start management functions.
[0066] First, manually allocate ROS2 node tasks for different computing devices. Different ROS2 nodes perform tasks with different functions; for example, computing device A mainly performs autonomous driving perception functions, and computing device B mainly performs autonomous driving planning and control, positioning functions.
[0067] Step S011, evenly allocate robot operating system node tasks according to the processor usage rate of the computing device and / or the graphics processor usage rate of the computing device and / or the memory usage rate of the computing device;
[0068] Prioritize allocating robot operating system nodes with high-frequency communication to the same computing device or adjacent computing devices;
[0069] Allocate the robot operating system nodes of each computing device according to the functions of the robot operating system nodes;
[0070] Allocating the robot operating system nodes of each computing device according to the functions of the robot operating system nodes includes:
[0071] When the robot operating system node is a graphics processing-intensive node, allocate the robot operating system node to a computing device with a graphics card;
[0072] When the robot operating system node is a real-time control node, allocate the robot operating system node to a computing device with low latency and high clock;
[0073] When the robot operating system node is a sensor-driven node, allocate the robot operating system node to a computing device physically close to the sensor;
[0074] Deploy the key robot operating system nodes for high-frequency communication dispersedly among several computing devices.
[0075] Specifically, monitor the usage rate of the processor CPU / graphics processor GPU / memory of the computing device, and evenly allocate the robot operating system node tasks to avoid overloading a single computing device;
[0076] To increase communication efficiency, the ROS2 nodes for high-frequency communication are preferentially deployed on the same computing device or adjacent computing devices;
[0077] To increase hardware affinity, for example, allocate the GPU node to a computing device with a graphics card, and allocate the real-time control node to a low-latency device;
[0078] Deploy the key robot operating system nodes for high-frequency communication dispersedly among several computing devices to avoid single-point failures.
[0079] Through the optimization of allocating robot operating system node tasks to each computing device, the CPU / GPU load balance degree of the computing device is improved by 30 - 50%; the cross-computing device communication latency is reduced by 20 - 40%; the fault recovery time is shortened from the minute level to the second level (depending on the heartbeat detection interval).
[0080] Step S02, obtain the corresponding target computing device IP of each robot operating system node, the login username and password of each computing device, and the name of each robot operating system node;
[0081] Obtain the robot operating system node configuration file according to the corresponding target computing device IP of each robot operating system node, the login username and password of each computing device, and the name of each robot operating system node;
[0082] Obtain the startup program and task parameters of each robot operating system node;
[0083] Obtain the robot operating system node startup parameters according to the startup program and task parameters of each robot operating system node.
[0084] Specifically, a ROS2 node configuration file is generated, which contains the target device IPs of different ROS2 nodes, the login usernames and passwords of each computing device, the names of each ROS2 node, etc.
[0085] Generate ROS2 node startup parameters, which include the startup programs and task parameters of each ROS2 node.
[0086] Step S02: Obtain the robot operating system node configuration file and the robot operating system node startup parameters.
[0087] Step S03: Start the robot operating system nodes of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters.
[0088] Step S031: Generate the startup commands for each robot operating system node according to the robot operating system node configuration file and the robot operating system node startup parameters.
[0089] Map the startup commands of each robot operating system node to the corresponding computing device to generate the mapping relationship between the robot operating system node startup commands and the computing devices.
[0090] Start the robot operating system nodes of each computing device according to the startup commands of each robot operating system node through the Secure Shell protocol service and the mapping relationship between the robot operating system node startup commands and the computing devices.
[0091] Specifically, as Figure 3 shown, after the ROS2 node allocation is completed, it is necessary to implement the function of uniformly starting the ROS2 nodes on a certain computing device. Therefore, first, information such as each node startup program, task parameters, and node configuration file is structured to generate the startup commands for each robot operating system node. This startup command can be directly jumped to the corresponding device through the SSH service for startup and task execution. Through a unified remote startup module, the startup commands of each ROS2 node can be asynchronously sent to the corresponding computing device through the SSH (Secure Shell) service to execute tasks on the ROS2 nodes of each computing device.
[0092] Step S04: Store the startup commands of each robot operating system node in the network file system.
[0093] Build a shared service platform for the startup information of the robot operating system nodes of several computing devices through the network file system.
[0094] Specifically, in order to achieve global sharing of remote startup information, a startup information sharing service platform for multiple ROS2 nodes is built using the Network File System (NFS), and the startup information of multiple ROS2 nodes is shared, enabling different computing devices to obtain the allocation and execution startup information of ROS2 node tasks from each other.
[0095] Step S05, in response to a failure of the robot operating system node process on any computing device, the daemon process program on this computing device restarts the robot operating system node on this computing device.
[0096] Step S051, configure a corresponding daemon process program for each computing device, and monitor the process status of the robot operating system node on the corresponding computing device through the daemon process program;
[0097] In response to a failure of the robot operating system node process on any computing device, the daemon process program on this computing device obtains the startup command of the robot operating system node on this computing device from the network file system, and restarts the robot operating system node on this computing device through the startup command of the robot operating system node on this computing device.
[0098] Specifically, as Figure 4 shown, due to reasons within the ROS2 node itself or uncontrollable factors, the process of the ROS2 node itself may crash, and the crash of a certain ROS2 node process will affect the tasks of the entire system; therefore, in response to abnormal phenomena such as the crash of the ROS2 node, this application designs an automatic restart module for the ROS2 node, configures a daemon process program for each computing device, and each daemon process program monitors the process status of the ROS2 node on its corresponding device. When the process status of a certain ROS2 node changes, for example, switches from the running state to the crashed state, the daemon process program on the corresponding computing device will read the startup command of the ROS2 node process from the shared information NFS, re-execute the startup command of the ROS2 node, start and execute the ROS2 node, and achieve the normal recovery and operation of the crashed ROS2 node process.
[0099] Step S06, perform staggered loop monitoring on several computing devices. In response to a failure of any computing device, allocate the robot operating system node process on the failed computing device to an adjacent computing device and start it.
[0100] Step S061, perform staggered cyclic full-scale monitoring on several computing devices through the daemon process programs corresponding to each computing device;
[0101] In response to the failure of any computing device, the daemon process corresponding to the adjacent computing device obtains the robot operating system node startup command of the faulty computing device from the network file system, allocates the robot operating system node process on the faulty computing device to the adjacent computing device, and starts the robot operating system node on the adjacent computing device through the startup command of the robot operating system node on the faulty computing device.
[0102] Specifically, as Figure 5 shown, in addition to the process crash of the ROS2 node itself, there are still some uncontrollable factors that cause failures at the computing device level, that is, the computing device fails and the corresponding programs on the computing device cannot run properly; for this situation, this application sets up a hot standby switch daemon process to achieve the fault tolerance function at the device level, that is, a daemon process is running on each computing device to monitor the information of other computing devices in real time. When a certain computing device fails, the daemon processes of other computing devices can timely monitor that the status of the faulty device has changed, obtain the ROS2 node startup information in the crashed device from the shared information NFS in time, and start and execute the corresponding ROS2 node process on the adjacent normal computing device.
[0103] For example: when a certain computing device fails, the hot standby switch daemon process monitors the status of the corresponding computing device. The hot standby switch daemon process can read the ROS2 node startup information in the crashed device from the shared information NFS, send the ROS2 node startup information in the crashed device to the adjacent computing device, and execute the ROS2 node to achieve the distributed fault tolerance computing function at the device level.
[0104] Here, as Figure 6 shown, the embodiments of this application will be further described in detail in combination with the Autoware.Universe autonomous driving framework:
[0105] The Autoware.Universe autonomous driving framework is an open-source software framework for autonomous driving, which includes autonomous driving business modules such as control, perception, planning, and positioning, and also has a simulation test platform for developers and researchers to carry out autonomous driving software development and testing; Autoware.Universe can run on a distributed computing platform and needs to allocate multiple business modules and tasks of the framework on each computing device to achieve an efficient and stable distributed autonomous driving computing system; in this embodiment, the computing device platform used is mainly composed of 4 edge computing devices with an ARM architecture, and the specific information of the devices is shown in Table 1 below.
[0106] Table 1
[0107]
[0108] The business modules of the Autoware.Universe autonomous driving framework are mainly divided into modules such as sensors, perception, control, positioning, and planning. According to the Autoware.Universe business modules, each ROS2 node in the business modules is divided into devices, and different computing devices execute different ROS2 nodes. The specific ROS2 node division results of the devices are shown in Table 2 below. Different computing devices (Orin) are assigned different functional ROS2 nodes, and different Orins are assigned different IP addresses, and they can communicate with each other through the SSH service.
[0109] Table 2
[0110]
[0111]
[0112]
[0113] Based on the ROS2 node division results of Autoware.Universe, and based on the SSH service, the ROS2 node startup commands of different functional ROS2 nodes are regenerated and reformatted into commands that can be remotely started through the SSH service. The startup commands of each robot operating system node are remotely sent to other devices Orin through the device Orin-a, so as to start all ROS2 nodes. At the same time, the startup commands of all ROS2 nodes are stored in the NFS shared space to achieve global sharing of the startup information of all ROS2 nodes.
[0114] The running status of the ROS2 node processes on their respective computing devices Orin is monitored in real time. If a certain ROS2 node process in a certain computing device Orin crashes and the process is closed due to certain reasons, the daemon process will read the startup command of the ROS2 node process from the shared information NFS and re-execute the startup command to start and execute the ROS2 node process, so as to achieve the normal recovery operation of the crashed ROS2 node process.
[0115] For example: when the / map / map_container process in the computing device Orin-d is closed due to map loading problems, the daemon process on this computing device will monitor that the map / map_container process changes from the running state to the closed state, and then read the startup command of this process from the NFS shared space and re-execute the startup command to achieve the restart of the map / map_container process.
[0116] In addition to the process crash of the ROS2 nodes themselves, the computing device may also malfunction, causing the device to be unable to run the corresponding program properly, thus resulting in the abnormal operation of the functions of the business module. For example, the computing device Orin-d shuts down due to overheating, and all business modules on this device cannot run properly. Four devices Orin implement a ring-type device status monitoring, that is, device Orin-a monitors the status of device Orin-b, device Orin-b monitors the status of device Orin-c, device Orin-c monitors the status of device Orin-d, and device Orin-d monitors the status of device Orin-a. This ring-type device status monitoring realizes full-scale monitoring. When device Orin-d shuts down due to overheating, the daemon process program of device Orin-c can timely monitor the status information of device Orin-d, read all the ROS2 node process startup commands on device Orin-d from the NFS shared space, and re-execute all the ROS2 node process startup commands on device Orin-c to start the ROS2 node processes, realizing the fault tolerance function between devices.
[0117] Based on the above fault handling method, the high availability and self-recovery ability of the Autoware.universe system are improved.
[0118] In addition, the fault handling method also includes:
[0119] In response to a fault in the communication network between several computing devices, multiple network cards (such as Ethernet + WiFi) are configured for the computing devices, and the redundant design of the network interface is used to improve the reliability of communication;
[0120] and / or select a data distribution middleware protocol and standard that supports high fault tolerance between computing devices to improve the performance, reliability and flexibility of the robot system;
[0121] In response to inconsistent communication data between computing devices, rosbridge_suite (an important toolkit in ROS / ROS2) is used to achieve cross-computing device data synchronization;
[0122] and / or add a cyclic redundancy check field and a hash field to the computing device messages to verify the data between computing devices;
[0123] In response to tight computing device resources, the non-critical functions of the computing device are shut down;
[0124] and / or integrate Kubernetes and Nomad to implement containerized orchestration for ROS2 nodes.
[0125] During actual deployment, it is recommended to first test the communication configuration of the data distribution middleware protocol and standard (DDS, Data Distribution Service) in a small-scale environment, and then gradually expand the node scale. For critical mission systems, more refined state management can be achieved by combining special nodes (LifecycleNode) that implement node lifecycle management in the ROS2 system.
[0126] Through the above methods, the ROS2 system can maintain its core functions in scenarios such as device failures and network jitters. During actual deployment, the redundancy level and recovery strategy need to be adjusted according to specific applications.
[0127] Without departing from the technical solution of this application, several improvements and optimizations can be made to the method for handling faults provided in the embodiments of this application, and these improvements and optimizations should also be regarded as the protection scope of this application.
[0128] The beneficial effects brought by the technical solution provided in the embodiments of this application are as follows:
[0129] This application improves the high availability and self-recovery ability of the ROS2 system, enhances the management efficiency of ROS2 nodes, is applicable to critical mission scenarios such as autonomous driving and industrial control, and can simultaneously implement a more stable and reliable ROS2 computing platform.
[0130] Through the remote startup module in the technical solution of this application, remote startup and unified management of ROS2 nodes on multiple computing devices can be achieved, avoiding the problem of manually starting nodes on each device one by one in the traditional ROS2 architecture, greatly improving the system deployment efficiency and maintenance convenience, and enhancing the system's self-recovery ability; by the automatic restart module, the running status of ROS2 processes is monitored in real time. When it is detected that a critical node crashes abnormally, the system can automatically restart the node process, avoiding manual intervention, improving the system's reliability and running stability, and achieving efficient task migration and fault tolerance capabilities; through the hot standby switching module, when a computing device fails or crashes, all ROS2 node processes on this device can be automatically allocated to other adjacent normal devices, realizing seamless migration, ensuring that tasks are not interrupted, and enhancing the system's business continuity.
[0131] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0132] The embodiments of this application also provide a fault handling device, as Figure 7 shown, the device includes: an allocation module, an acquisition module, a remote startup module, a storage module, an automatic restart module, and a hot standby switching module.
[0133] In this embodiment, a distribution module is configured to distribute Robot Operating System node tasks to each computing device;
[0134] An acquisition module is configured to acquire a Robot Operating System node configuration file and Robot Operating System node startup parameters;
[0135] A remote startup module is configured to start the Robot Operating System nodes of each computing device according to the Robot Operating System node configuration file and the Robot Operating System node startup parameters;
[0136] An automatic restart module is configured to, in response to a failure of the Robot Operating System node process of any one computing device, restart the Robot Operating System node of the computing device through the daemon program of the computing device;
[0137] A hot standby switching module is configured to perform staggered loop monitoring on several computing devices, and in response to a failure of any one computing device, allocate the Robot Operating System node process on the failed computing device to an adjacent computing device and start it.
[0138] In this embodiment, an acquisition module is configured to acquire the corresponding target computing device IP of each Robot Operating System node, the login username and password of each computing device, and the name of each Robot Operating System node;
[0139] Obtain a Robot Operating System node configuration file according to the corresponding target computing device IP of each Robot Operating System node, the login username and password of each computing device, and the name of each Robot Operating System node;
[0140] Obtain the startup program and task parameters of each Robot Operating System node;
[0141] Obtain Robot Operating System node startup parameters according to the startup program and task parameters of each Robot Operating System node.
[0142] In one of the embodiments, a remote startup module is configured to generate startup commands for each Robot Operating System node according to the Robot Operating System node configuration file and the Robot Operating System node startup parameters;
[0143] Map the startup commands of each Robot Operating System node to the corresponding computing device to generate a mapping relationship between the Robot Operating System node startup commands and the computing device;
[0144] Start the Robot Operating System nodes of each computing device according to the startup commands of each Robot Operating System node through the Secure Shell Protocol service and the mapping relationship between the Robot Operating System node startup commands and the computing device.
[0145] In one embodiment, a storage module is configured to store the startup commands of each Robot Operating System (ROS) node in a Network File System (NFS).
[0146] An ROS node startup information sharing service platform for several computing devices is constructed through the NFS.
[0147] In one embodiment, an automatic restart module is configured to configure a corresponding daemon process program for each computing device, and monitor the process status of the ROS node of the corresponding computing device through the daemon process program.
[0148] In response to a failure of the ROS node process of any computing device, the daemon process program of the computing device obtains the startup command of the ROS node of the computing device from the NFS, and restarts the ROS node of the computing device through the startup command of the ROS node of the computing device.
[0149] In one embodiment, a hot standby switching module is configured to perform staggered cyclic full monitoring on several computing devices through the daemon process programs corresponding to the respective computing devices.
[0150] In response to a failure of any computing device, the daemon process program of the adjacent computing device obtains the startup command of the ROS node of the failed computing device from the NFS, allocates the ROS node process on the failed computing device to the adjacent computing device, and starts the ROS node on the adjacent computing device through the startup command of the ROS node on the failed computing device.
[0151] In one embodiment, an allocation module is configured to evenly allocate ROS node tasks according to the processor usage rate of the computing device and / or the graphics processor usage rate of the computing device and / or the memory usage rate of the computing device.
[0152] Prioritize allocating ROS nodes with high-frequency communication to the same computing device or adjacent computing devices.
[0153] Allocate the ROS nodes of each computing device according to the functions of the ROS nodes.
[0154] Allocating the ROS nodes of each computing device according to the functions of the ROS nodes includes:
[0155] When the ROS node is a graphics processing intensive node, allocate the ROS node to a computing device with a graphics card.
[0156] When the robot operating system node is a real-time control node, the robot operating system node is allocated to a computing device with low latency and high clock.
[0157] When the robot operating system node is a sensor driver node, the robot operating system node is allocated to a computing device physically close to the sensor.
[0158] The key robot operating system nodes for high-frequency communication are dispersed and deployed among several computing devices.
[0159] Specifically, this application designs a fault handling device based on the ROS2 system. This solution proposes a ROS2 system fault-tolerant computing daemon process, which can remotely start ROS2 nodes on multiple computing devices to achieve unified scheduling across devices; at the same time, it can monitor the process status of ROS2 nodes in real time, automatically detect the crash or abnormal exit of ROS2 nodes; dynamically migrate the tasks of ROS2 nodes, and automatically re-allocate the ROS2 node processes to healthy devices when a device crashes. Through three core modules of remote start, automatic restart, and hot standby switch, this application improves the high availability and self-recovery ability of the ROS2 system, is applicable to key task scenarios such as autonomous driving and industrial control, and realizes a more stable and reliable ROS2 computing platform.
[0160] The beneficial effects brought by the technical solution provided in the embodiments of this application are as follows:
[0161] This application improves the high availability and self-recovery ability of the ROS2 system, improves the management efficiency of ROS2 nodes, is applicable to key task scenarios such as autonomous driving and industrial control, and can also realize a more stable and reliable ROS2 computing platform.
[0162] Through the remote start module, the technical solution of this application can remotely start and uniformly manage ROS2 nodes on multiple computing devices, avoiding the problem of manually starting nodes on each device one by one in the traditional ROS2 architecture, greatly improving the system deployment efficiency and maintenance convenience, and enhancing the system's autonomous recovery ability; through the automatic restart module, it monitors the running status of the ROS2 process in real time. When it detects that a key node has an abnormal crash, the system can automatically restart the node process, avoiding manual intervention, improving the system reliability and running stability, and realizing efficient task migration and fault tolerance; through the hot standby switch module, when a computing device fails or crashes, it can automatically allocate all ROS2 node processes on this device to other adjacent normal devices to achieve seamless migration, ensure that the task is not interrupted, and improve the system's business continuity.
[0163] For the description of the features in the corresponding embodiments of the fault handling device, reference can be made to the relevant descriptions in the corresponding embodiments of the fault handling method, which will not be elaborated here one by one.
[0164] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in the embodiment of the fault handling method. The method includes:
[0165] Allocating a Robot Operating System (ROS) node task to each computing device;
[0166] Obtaining a ROS node configuration file and ROS node startup parameters;
[0167] Starting the ROS node of each computing device according to the ROS node configuration file and the ROS node startup parameters;
[0168] In response to a fault in the ROS node process of any one computing device, restarting the ROS node of the computing device through the daemon process program of the computing device;
[0169] Performing staggered loop monitoring on several computing devices. In response to a fault in any one computing device, allocating the ROS node process on the faulty computing device to an adjacent computing device and starting it.
[0170] As Figure 8 shown, an embodiment of the present application further provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. Wherein, the computer program is configured to execute the steps in the embodiment of the fault handling method when running. The method includes:
[0171] Allocating a ROS node task to each computing device;
[0172] Obtaining a ROS node configuration file and ROS node startup parameters;
[0173] Starting the ROS node of each computing device according to the ROS node configuration file and the ROS node startup parameters;
[0174] In response to a fault in the ROS node process of any one computing device, restarting the ROS node of the computing device through the daemon process program of the computing device;
[0175] Performing staggered loop monitoring on several computing devices. In response to a fault in any one computing device, allocating the ROS node process on the faulty computing device to an adjacent computing device and starting it.
[0176] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs.
[0177] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program. When the computer program is executed by a processor, it implements the steps in the embodiment of the fault handling method. The method includes:
[0178] Assign robot operating system node tasks to each computing device;
[0179] Obtain the robot operating system node configuration file and the robot operating system node startup parameters;
[0180] Start the robot operating system node of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters;
[0181] In response to a failure of the robot operating system node process of any one computing device, restart the robot operating system node of the computing device through the daemon program of the computing device;
[0182] Perform staggered loop monitoring on several computing devices. In response to a failure of any one computing device, allocate the robot operating system node process on the failed computing device to an adjacent computing device and start it.
[0183] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the steps in the embodiment of the fault handling method. The method includes:
[0184] Assign robot operating system node tasks to each computing device;
[0185] Obtain the robot operating system node configuration file and the robot operating system node startup parameters;
[0186] Start the robot operating system node of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters;
[0187] In response to a failure of the robot operating system node process of any one computing device, restart the robot operating system node of the computing device through the daemon program of the computing device;
[0188] Perform misaligned ring monitoring on several computing devices. In response to a failure of any one of the computing devices, allocate the robot operating system node process on the failed computing device to an adjacent computing device and start it.
[0189] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0190] The above has introduced in detail a fault handling method, device, equipment, and medium provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A fault handling method, characterized in that The method is applied to a fault handling system, which includes several computing devices. The method includes: Allocating robot operating system node tasks to each computing device; Obtaining the robot operating system node configuration file and the robot operating system node startup parameters; Starting the robot operating system nodes of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters; In response to a fault in the robot operating system node process of any one of the computing devices, restart the robot operating system node of that computing device through the daemon process program of that computing device; Perform staggered loop monitoring on the several computing devices. In response to a fault in any one of the computing devices, allocate the robot operating system node process on the faulty computing device to an adjacent computing device and start it.
2. The fault handling method according to claim 1, wherein Before obtaining the robot operating system node configuration file and the robot operating system node startup parameters, it includes: Obtaining the corresponding target computing device IP of each robot operating system node, the login username and password of each computing device, and the name of each robot operating system node; Obtaining the robot operating system node configuration file according to the corresponding target computing device IP of each robot operating system node, the login username and password of each computing device, and the name of each robot operating system node; Obtaining the startup program and task parameters of each robot operating system node; Obtaining the robot operating system node startup parameters according to the startup program and task parameters of each robot operating system node.
3. The fault handling method according to claim 1, characterized in that, The starting the robot operating system nodes of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters includes: Generating startup commands for each robot operating system node according to the robot operating system node configuration file and the robot operating system node startup parameters; Mapping the startup commands of each robot operating system node to the corresponding computing device to generate a mapping relationship between the robot operating system node startup commands and the computing devices; Starting the robot operating system nodes of each computing device according to the startup commands of each robot operating system node through the Secure Shell Protocol service and the mapping relationship between the robot operating system node startup commands and the computing devices.
4. The fault handling method according to claim 3, wherein After starting the robot operating system nodes of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters, it includes: Storing the startup commands of each robot operating system node in the network file system; Constructing a robot operating system node startup information sharing service platform for several computing devices through the network file system.
5. The fault handling method according to claim 4, characterized in that The restarting the robot operating system node of a computing device through the daemon process program of that computing device in response to a fault in the robot operating system node process of any one of the computing devices includes: Configure a corresponding daemon process program for each computing device, and monitor the process status of the robot operating system nodes of the corresponding computing device through the daemon process program; In response to a failure of the robot operating system node process of any computing device, the daemon process program of the computing device obtains the startup command of the robot operating system node of the computing device from the network file system, and restarts the robot operating system node of the computing device through the startup command of the robot operating system node of the computing device.
6. The fault handling method according to claim 5, characterized in that, The misaligned ring monitoring of the several computing devices, in response to a failure of any computing device, allocating the robot operating system node process on the failed computing device to an adjacent computing device and starting it, includes: Performing misaligned cyclic full-scale monitoring on the several computing devices through the daemon process programs corresponding to the respective computing devices; In response to a failure of any computing device, obtain the startup command of the robot operating system node of the failed computing device from the network file system through the daemon process program corresponding to the adjacent computing device, allocate the robot operating system node process on the failed computing device to the adjacent computing device, and start the robot operating system node on the adjacent computing device through the startup command of the robot operating system node on the failed computing device.
7. The fault handling method according to claim 1, characterized in that, The allocating robot operating system node tasks to each computing device includes: Balancing the allocation of robot operating system node tasks according to the processor usage rate of the computing device and / or the graphics processor usage rate of the computing device and / or the memory usage rate of the computing device; Prioritizing the allocation of robot operating system nodes with high-frequency communication on the same computing device or adjacent computing devices; Allocating the robot operating system nodes of each computing device according to the functions of the robot operating system nodes; The allocating the robot operating system nodes of each computing device according to the functions of the robot operating system nodes includes: When the robot operating system node is a graphics processing-intensive node, allocate the robot operating system node to a computing device with a graphics card; When the robot operating system node is a real-time control node, allocate the robot operating system node to a computing device with low latency and high clock; When the robot operating system node is a sensor-driven node, allocate the robot operating system node to a computing device physically close to the sensor; Disperse the deployment of critical robot operating system nodes with high-frequency communication among several computing devices.
8. A fault handling device, characterized in that, The device includes: An allocation module for allocating robot operating system node tasks to each computing device; An acquisition module for acquiring the robot operating system node configuration file and the robot operating system node startup parameters; A remote startup module for starting the robot operating system nodes of each computing device according to the robot operating system node configuration file and the robot operating system node startup parameters; An automatic restart module, which is used to respond to a failure of the robot operating system node process of any computing device, and then restart the robot operating system node of the computing device through the daemon program of the computing device; A hot standby switching module, which is used to perform staggered ring monitoring on the several computing devices, and in response to a failure of any computing device, allocate the robot operating system node process on the failed computing device to an adjacent computing device and start it.
9. An electronic device, characterized in that, Comprising: A memory, which is used to store computer programs; A processor, which is used to implement the steps of the fault handling method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the fault handling method according to any one of claims 1 to 7 when executed by a processor.