A watchdog monitoring method for multi-core and multi-thread
By binding watchdog threads on multiple cores in a multi-core multi-threaded system, the problem of restricted supervision function caused by single core binding is solved, supervision and load balancing between cores is achieved, system stability and resource utilization are improved, and failure recovery time is reduced.
Patent Information
- Application Number
- CN202510511633.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-23
AI Technical Summary
In multi-core and multi-threading systems, the way a single core binds watchdog threads leads to limited supervision functions, and abnormal situations cannot be discovered and handled in time. The complexity and resource utilization of multi-core and multi-threading environment are low, and the failure recovery time is long.
The watchdog thread is bound to multiple cores, and the multi-core watchdog thread peer supervision, master-slave multi-core watchdog thread and dynamic load balancing design is adopted to achieve supervision and load balancing between cores through shared memory or message queue communication.
It improves the stability and reliability of the system, reduces the failure recovery time, optimizes resource utilization, and enhances the system's fault tolerance and management simplicity.
Smart Images

Figure CN120029814B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of watchdog monitoring, and specifically, to a watchdog monitoring method for multi-core and multi-threaded systems. Background Art
[0002] With the rapid development of modern computing technology, multi-core and multi-threaded processors have become the mainstream computing platforms. This architecture allows the processor to execute multiple threads simultaneously, thus significantly improving the overall performance of the system. However, the multi-core and multi-threaded environment also brings complexity and challenges, especially in applications that require high reliability and security.
[0003] In a multi-core and multi-threaded system, the execution of critical task threads and the stability of the system are crucial. For this purpose, the concept of a watchdog thread is introduced. It is an independent thread responsible for monitoring the execution of critical task threads and the state of the system. When an abnormal situation is detected, the watchdog thread can quickly take corresponding measures, such as restarting the critical task thread or triggering a fault recovery mechanism, to ensure the stability and availability of the system.
[0004] However, there are certain limitations in the way of binding the watchdog thread to a single core. When the CPU utilization rate of this core reaches saturation, the supervision function of the watchdog thread will be affected, which may lead to the inability to detect and handle abnormal situations in a timely manner. This situation is particularly prominent in complex multi-core and multi-threaded systems because different cores may carry different task loads, resulting in more intense resource usage on some cores.
[0005] In the prior art, a single-threaded supervision method that only binds the watchdog thread to core 1 is often used to supervise the program execution state. However, when the CPU utilization rate of core 1 reaches saturation, the supervision function of the watchdog thread will be affected, which may lead to the inability to detect and handle abnormal situations in a timely manner.
[0006] In summary, the prior art has the following problems:
[0007] 1. Limitations of a single-core watchdog thread: There are limitations in the way of binding the watchdog thread to a single core. When the CPU utilization rate of this core reaches saturation, the supervision function of the watchdog thread will be affected, which may lead to the inability to detect and handle abnormal situations in a timely manner;
[0008] 2. Complexity of multi-core and multi-threaded systems: The multi-core and multi-threaded environment brings complex thread scheduling and task allocation problems. Especially in applications with critical task threads and high reliability requirements, an effective mechanism is needed to supervise and manage these tasks;
[0009] 3. Resource utilization rate and fault recovery time: In a multi-core and multi-threaded system, how to optimize resource utilization and reduce fault recovery time is an important technical issue, and the traditional single-core watchdog thread design has deficiencies in this regard. Summary of the Invention
[0010] To overcome the deficiencies of the prior art, the present invention provides a watchdog monitoring method for multi-core and multi-threaded systems, which solves problems such as low reliability existing in the prior art.
[0011] The technical solution adopted by the present invention to solve the above problems is:
[0012] A watchdog monitoring method for multi-core and multi-threaded systems, binding watchdog threads on N cores; where N≥2 and N is an integer.
[0013] As a preferred technical solution, an implementation method of peer supervision of multi-core watchdog threads is adopted: a watchdog thread is bound on each core, and each watchdog thread is responsible for monitoring the watchdog threads of a specified number of other cores and the critical task threads of other cores.
[0014] As a preferred technical solution, it includes the following steps:
[0015] A1, allocate a watchdog thread for each core and configure the monitoring targets of the watchdog thread; where the monitoring targets of the watchdog thread include the watchdog threads of other cores and the critical task threads of other cores;
[0016] A2, each watchdog thread regularly sends a heartbeat signal to its monitored targets and waits for a response;
[0017] A3, if the watchdog thread does not receive a response from the monitored target within a predetermined time, it is determined that the monitored target may have failed, and a corresponding fault recovery mechanism is triggered.
[0018] As a preferred technical solution, when adopting the implementation method of peer supervision of multi-core watchdog threads, the watchdog threads communicate with each other through shared memory or message queue.
[0019] As a preferred technical solution, an implementation method of master-slave multi-core watchdog threads is adopted: there is a main watchdog thread and multiple slave watchdog threads, the main watchdog thread is responsible for supervising all slave watchdog threads, the slave watchdog threads are respectively bound on different cores, and the slave watchdog threads are responsible for supervising the execution of the critical task threads of the cores where they are located.
[0020] As a preferred technical solution, it includes the following steps:
[0021] B1. Select a core as the running core of the main watchdog thread and bind slave watchdog threads to other cores respectively.
[0022] B2. The main watchdog thread configures and starts all slave watchdog threads and monitors the running status of all slave watchdog threads.
[0023] B3. The slave watchdog threads periodically send heartbeat signals to the main watchdog thread and report the status of the critical task threads they monitor.
[0024] B4. If the main watchdog thread detects an abnormality in a certain slave watchdog thread or critical task thread, take exception handling measures.
[0025] As a preferred technical solution, the exception handling measures include restarting the slave watchdog thread or triggering a fault recovery mechanism.
[0026] As a preferred technical solution, an implementation method of a multi-core watchdog thread with dynamic load balancing is adopted: the number and / or binding position of the watchdog threads can be dynamically adjusted according to the real-time load situation.
[0027] As a preferred technical solution, it includes the following steps:
[0028] C1. Preset a certain number of watchdog threads according to the number of cores and task requirements.
[0029] C2. Monitor the load situation of each core of the system in real time.
[0030] C3. When the load of a certain core is higher than the set value, migrate the watchdog thread on that core to a core with a load lower than the set value to run.
[0031] Among them, the monitoring target of the watchdog thread can be dynamically adjusted during the running process to ensure that the watchdog thread always supervises the execution of critical task threads.
[0032] As a preferred technical solution, the load situation index is the CPU usage rate or memory occupancy.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] 1. Improve the stability and reliability of the system: By binding watchdog threads to multiple cores, mutual supervision and the supervision of critical task threads are realized. Even if the load of a certain core is too high, the watchdog threads on other cores can still maintain the supervision of the system, ensuring the stability and reliability of the system.
[0035] 2. Reduce fault recovery time: When an abnormal situation is detected, the watchdog threads on multiple cores can quickly take corresponding measures, such as restarting critical task threads or triggering a fault recovery mechanism, thereby reducing the impact of faults on the system and shortening the fault recovery time.
[0036] 3. Optimize resource utilization: The design of multi-core watchdog threads can better utilize the computing resources of multi-core and multi-thread processors, improving overall performance. Through reasonable thread scheduling and task allocation, it can ensure that critical task threads receive sufficient resource support while maintaining supervision and management of other tasks.
[0037] 4. Improve system fault tolerance: The design of multi-core watchdog threads can enhance the system's fault tolerance. When a fault occurs in a certain core or thread, the watchdog threads on other cores can detect it in a timely manner and take measures to prevent the spread of the fault and reduce the risk of system crash.
[0038] 5. Simplify system design and management: By binding watchdog threads to multiple cores, the system design and management can be simplified. Developers can more flexibly configure and adjust the parameters and behaviors of watchdog threads to meet the requirements of different applications. At the same time, the design of multi-core watchdog threads also provides richer monitoring and debugging means, facilitating developers to conduct in-depth analysis and optimization of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 Schematic diagrams of implementation methods listed for specific embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0041] Embodiment 1
[0042] To overcome the limitations of the prior art and improve the security and stability of multi-core and multi-thread systems, the present invention proposes a design concept of binding watchdog threads to multiple cores, as Figure 1 shown, including multi-core watchdog thread peer supervision, master-slave multi-core watchdog threads, and multi-core watchdog threads with dynamic load balancing. This design allows watchdog threads to run on different cores to achieve mutual supervision and supervision of critical task threads (the main thread that runs normally in the application program cycle is the critical thread, and multiple cores may run multiple critical threads that run in cycles). In this way, even if the load on a certain core is too high, the watchdog threads on other cores can still maintain supervision of the system to ensure the stability and reliability of the system.
[0043] In summary, based on the high reliability and security requirements of the multi-core and multi-thread architecture, the present invention proposes a design solution of binding watchdog threads on multiple cores to supervise each other, aiming to improve the stability and reliability of the system, reduce the fault recovery time, and optimize the utilization of resources, providing a strong guarantee for the security of multi-core and multi-thread programs.
[0044] The present invention can be implemented in the following manner (taking two cores as an example):
[0045] 1. Multi-core deployment: Bind a watchdog thread on Core 2 as well to achieve mutual supervision with the watchdog thread on Core 1. This design can improve the fault tolerance and response speed of the system.
[0046] 2. Mutual supervision mechanism: The two watchdog threads are respectively located on Core 1 and Core 2. They not only supervise each other's states but also simultaneously supervise the execution of critical task threads. When any thread detects an abnormality, corresponding measures can be taken quickly, such as restarting the critical task thread or triggering the fault recovery mechanism.
[0047] 3. Supervision of critical task threads: The two watchdog threads are jointly responsible for supervising the execution of critical task threads to ensure that they run in a normal state. When an abnormality of the critical task thread is detected, the watchdog thread will trigger the fault recovery mechanism to ensure the stability and availability of the system.
[0048] The present invention has the following advantages:
[0049] 1. Improve system stability and reliability: By binding watchdog threads on multiple cores, mutual supervision and supervision of critical task threads are achieved. Even if the load of a certain core is too high, the watchdog threads on other cores can still maintain supervision of the system to ensure the stability and reliability of the system.
[0050] 2. Reduce the fault recovery time: When an abnormal situation is detected, the watchdog threads on multiple cores can quickly take corresponding measures, such as restarting the critical task thread or triggering the fault recovery mechanism, thereby reducing the impact of the fault on the system and reducing the fault recovery time.
[0051] 3. Optimize resource utilization: The design of multi-core watchdog threads can better utilize the computing resources of multi-core and multi-thread processors to improve the overall performance. Through reasonable thread scheduling and task allocation, it can ensure that critical task threads obtain sufficient resource support while maintaining supervision and management of other tasks.
[0052] 4. Improve system fault tolerance: The design of multi-core watchdog threads can improve the system's fault tolerance. When a certain core or thread fails, the watchdog threads on other cores can detect it in time and take measures to avoid the spread of the fault and reduce the risk of system crash.
[0053] 5. Simplify system design and management: By binding watchdog threads to multiple cores, the system design and management can be simplified. Developers can configure and adjust the parameters and behaviors of watchdog threads more flexibly to meet the requirements of different applications. At the same time, the design of multi-core watchdog threads also provides richer monitoring and debugging means, facilitating developers to conduct in-depth analysis and optimization of the system.
[0054] Example 2
[0055] Based on Example 1, this example provides a more refined implementation.
[0056] Peer supervision of multi-core watchdog threads:
[0057] In this implementation, a watchdog thread is bound to each core. These threads are in a peer-to-peer position and monitor the execution status of other cores and critical task threads. Each watchdog thread is responsible for monitoring a specified number of other cores and critical task threads to ensure that they run as expected.
[0058] Specific steps:
[0059] 1. When the application is initialized, allocate a watchdog thread to each core and configure their monitoring targets (other cores and critical task threads).
[0060] 2. Each watchdog thread periodically sends a heartbeat signal to its monitored targets and waits for a response.
[0061] 3. If a watchdog thread does not receive a response from its monitored target within a predetermined time, it is determined that the target may have failed, and the corresponding fault recovery mechanism is triggered.
[0062] Among them, the watchdog threads communicate with each other through shared memory or message queues to coordinate fault recovery and other system operations.
[0063] Advantages:
[0064] Each core has an independent watchdog thread, improving the reliability and fault tolerance of the system.
[0065] Peer supervision between cores is achieved, avoiding the problem of single-point failure.
[0066] Example 3
[0067] Based on Example 1, this example provides a more refined implementation.
[0068] Master-slave multi-core watchdog threads:
[0069] In this embodiment, there is a main watchdog thread and multiple slave watchdog threads. The main watchdog thread is responsible for coordinating and supervising all slave watchdog threads, while the slave watchdog threads are respectively bound to different cores and are responsible for supervising the execution of critical task threads.
[0070] Specific steps:
[0071] 1. When the application is initialized, select a core as the running core of the main watchdog thread and bind the slave watchdog threads to other cores respectively.
[0072] 2. The main watchdog thread is responsible for configuring and starting all slave watchdog threads and monitoring their running status.
[0073] 3. The slave watchdog threads regularly send heartbeat signals to the main watchdog thread and report the status of the critical task threads they monitor.
[0074] 4. If the main watchdog thread detects an abnormality in a certain slave watchdog thread or critical task thread, it will take corresponding measures, such as restarting the slave watchdog thread or triggering a fault recovery mechanism.
[0075] Advantages:
[0076] The master-slave structure is clear, which is convenient for management and maintenance.
[0077] The main watchdog thread can centrally monitor all slave watchdog threads and critical task threads, improving the monitoring efficiency.
[0078] Embodiment 4
[0079] Based on Embodiment 1, this embodiment provides a more refined implementation.
[0080] Multi-core watchdog threads with dynamic load balancing:
[0081] In this embodiment, the number and binding positions of the watchdog threads can be dynamically adjusted according to the real-time load conditions of the system. Through dynamic load balancing, it can be ensured that the watchdog threads always run on the cores with lighter loads, thereby improving the overall performance and stability of the system.
[0082] Specific steps:
[0083] 1. When the application is initialized, preset a certain number of watchdog threads according to the number of cores and task requirements.
[0084] 2. Monitor the load conditions of each core of the system in real time, including CPU usage, memory occupancy, etc.
[0085] 3. When the load of a certain core is too high, migrate the watchdog thread on that core to a core with lighter load to run.
[0086] Among them, during the running process, the monitoring targets of the watchdog threads can be dynamically adjusted to ensure that they always supervise the execution of the critical task threads.
[0087] Advantages:
[0088] Through dynamic load balancing, it can ensure that the watchdog threads are always running in the best state, improving the overall performance and stability of the system.
[0089] It is applicable to complex and changeable task environments and can adaptively adjust the system configuration to meet the requirements.
[0090] As described above, the present invention can be preferably implemented.
[0091] All the features disclosed in all the embodiments in this specification, or all the steps in the methods or processes implicitly disclosed, except for the mutually exclusive features and / or steps, can be combined and / or extended, replaced in any way.
[0092] As mentioned above, it is only a preferred embodiment of the present invention, and there is no any formal limitation to the present invention. According to the technical essence of the present invention, any simple modification, equivalent replacement and improvement made to the above embodiments within the spirit and principle of the present invention still belong to the protection scope of the technical solution of the present invention.
Claims
1. A watchdog monitoring method for multi-core and multi-thread, characterized in that, Bind watchdog threads to N cores; where N≥2 and N is an integer; The watchdog monitoring method includes: Peer supervision of multi-core watchdog threads: One watchdog thread is bound to each core, and each watchdog thread is responsible for monitoring the watchdog threads of a specified number of other cores and the critical task threads of other cores; Master-slave multi-core watchdog threads: There is one master watchdog thread and multiple slave watchdog threads. The master watchdog thread is responsible for supervising all slave watchdog threads. The slave watchdog threads are respectively bound to different cores, and the slave watchdog threads are responsible for supervising the execution of the critical task threads of the cores where they are located; Multi-core watchdog threads with dynamic load balancing: The number and / or binding location of watchdog threads can be dynamically adjusted according to the real-time load situation.
2. The watchdog monitoring method for multi-core and multi-thread according to claim 1, wherein It includes the following steps: A1. Assign a watchdog thread to each core and configure the monitoring targets of the watchdog thread; where the monitoring targets of the watchdog thread include the watchdog threads of other cores and the critical task threads of other cores; A2. Each watchdog thread periodically sends a heartbeat signal to its monitored targets and waits for a response; A3. If the watchdog thread does not receive a response from the monitored target within a predetermined time, it is determined that the monitored target may have failed, and a corresponding fault recovery mechanism is triggered.
3. A watchdog monitoring method for multi-core and multi-thread according to claim 1 or 2, characterized in that, When adopting the implementation method of peer supervision of multi-core watchdog threads, the watchdog threads communicate with each other through shared memory or message queue.
4. A watchdog monitoring method for multi-core and multi-thread according to claim 1, characterized in that It includes the following steps: B1. Select a core as the running core of the master watchdog thread and bind slave watchdog threads to other cores respectively; B2. The master watchdog thread configures and starts all slave watchdog threads and monitors the running status of all slave watchdog threads; B3. The slave watchdog threads periodically send heartbeat signals to the master watchdog thread and report the status of the critical task threads they monitor; B4. If the master watchdog thread detects an abnormality in a certain slave watchdog thread or critical task thread, abnormal handling measures are taken.
5. A watchdog monitoring method for multi-core and multi-thread according to claim 4, characterized in that, The abnormal handling measures include restarting the slave watchdog thread or triggering the fault recovery mechanism.
6. A watchdog monitoring method for multi-core and multi-thread according to claim 1, characterized in that, It includes the following steps: C1. Preset a certain number of watchdog threads according to the number of cores and task requirements; C2. Monitor the load conditions of each core of the system in real time; C3. When the load of a certain core is higher than the set value, migrate the watchdog thread on that core to a core with a load lower than the set value to run; Among them, the monitoring targets of the watchdog threads can be dynamically adjusted during the running process to ensure that the watchdog threads always supervise the execution of critical task threads.
7. A watchdog monitoring method for multi-core and multi-thread according to claim 1 or 6, characterized in that The load condition indicator is CPU usage or memory occupancy.
Citation Information
Patent Citations
Watchdog feeding method and system in Linux system
CN115658356A