Watchdog monitoring method for multi-core and multi-thread

By binding watchdog threads on multiple cores and adopting multi-core watchdog thread peer supervision or master-slave design, the problem of the supervision function of a single core watchdog thread being affected when the CPU usage is high is solved, achieving high reliability and rapid failure recovery of the system.

CN120029814AActive Publication Date: 2025-05-23成都交控轨道科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510511633.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In multi-core and multi-threading systems, when the CPU usage of a single core reaches saturation, the supervision function is affected, which may lead to the inability to detect and handle exceptions in a timely manner.

Method used

The watchdog thread is bound to multiple cores, and the multi-core watchdog thread peer supervision or the master-slave multi-core watchdog thread design is adopted to achieve mutual supervision and supervision of mission-critical threads through heartbeat signals and fault recovery mechanisms.

Benefits of technology

It improves the stability and reliability of the system, reduces the failure recovery time, optimizes resource utilization, improves the fault tolerance of the system, and simplifies system design and management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029814A_ABST
    Figure CN120029814A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of watchdog monitoring, and discloses a multi-core and multi-thread watchdog monitoring method, which comprises the following steps of: binding watchdog threads on N cores; wherein N is an integer greater than or equal to 2. According to the invention, the watchdog threads are bound on a plurality of cores, so that mutual supervision and supervision of key task threads are realized. And even if the load of a certain core is too high, the watchdog threads on other cores can still keep monitoring the system, so that the stability and the reliability of the system are ensured. The problems of low reliability and the like in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of watchdog monitoring, in particular to a watchdog monitoring method for multi-core and multi-threading. Background Art

[0002] With the rapid development of modern computing technology, multi-core multi-threaded processors have become the mainstream computing platform. This architecture allows the processor to execute multiple threads at the same time, significantly improving the overall performance of the system. However, the multi-core multi-threaded environment also brings complexity and challenges, especially in applications that require high reliability and security.

[0003] In a multi-core multi-threaded system, the execution of mission-critical threads and the stability of the system are crucial. To this end, the concept of a watchdog thread is introduced, which is an independent thread responsible for monitoring the execution of mission-critical threads and the status of the system. When an abnormal situation is detected, the watchdog thread can quickly take corresponding measures, such as restarting the mission-critical thread or triggering a fault recovery mechanism, to ensure the stability and availability of the system.

[0004] However, there are certain limitations to binding a watchdog thread on a single core. When the CPU usage of the core reaches saturation, the monitoring function of the watchdog thread will be affected, which may result in failure to detect and handle abnormal situations in a timely manner. This situation is particularly prominent in complex multi-core and multi-threaded systems, because different cores may carry different task loads, resulting in more intensive resource usage of some cores.

[0005] In the prior art, a single-threaded monitoring method is often used to monitor the program execution status, in which a watchdog thread is bound only to core 1. However, when the CPU usage of core 1 reaches saturation, the monitoring function of the watchdog thread will be affected, which may result in failure to detect and handle abnormal situations in a timely manner.

[0006] In summary, the prior art has the following problems: 1. Limitation of watchdog thread on a single core: There are limitations on how to bind a watchdog thread on a single core. When the CPU usage of the core reaches saturation, the monitoring function of the watchdog thread will be affected, which may result in failure to detect and handle abnormal situations in a timely manner. 2. Complexity of multi-core and multi-threaded systems: Multi-core and multi-threaded environments bring complex thread scheduling and task allocation issues, especially in mission-critical threads and applications with high reliability requirements, which require an effective mechanism to supervise and manage these tasks; 3. Resource utilization and fault recovery time: In a multi-core multi-threaded system, how to optimize resource utilization and reduce fault recovery time is an important technical issue. The traditional single-core watchdog thread design has shortcomings in this regard. Summary of the invention

[0007] In order to overcome the deficiencies of the prior art, the present invention provides a watchdog monitoring method for multi-core and multi-threading, which solves the problems of low reliability and the like in the prior art.

[0008] The technical solution adopted by the present invention to solve the above problems is: A watchdog monitoring method for multi-core and multi-threaded systems is disclosed, wherein watchdog threads are bound to N cores; wherein N≥2 and N is an integer.

[0009] As a preferred technical solution, a multi-core watchdog thread peer supervision implementation method is adopted: a watchdog thread is bound to each core, and each watchdog thread is responsible for monitoring a specified number of watchdog threads of other cores and key task threads of other cores.

[0010] As a preferred technical solution, the following steps are included: A1, assign a watchdog thread to each core and configure the monitoring target of the watchdog thread; the monitoring target of the watchdog thread includes the watchdog threads of other cores and the key task threads of other cores; A2, each watchdog thread periodically sends a heartbeat signal to the target it monitors and waits for a response; A3, if the watchdog thread does not receive a response from the monitoring target within a predetermined time, it is determined that the monitoring target may have failed, and a corresponding fault recovery mechanism is triggered.

[0011] As a preferred technical solution, when a multi-core watchdog thread peer supervision implementation method is adopted, the watchdog threads communicate with each other through a shared memory or a message queue.

[0012] As a preferred technical solution, a master-slave multi-core watchdog thread implementation method is adopted: there is a master watchdog thread and multiple slave watchdog threads, the master watchdog thread is responsible for supervising all slave watchdog threads, the slave watchdog threads are bound to different cores respectively, and the slave watchdog thread is responsible for supervising the execution of the critical task thread of the core where it is located.

[0013] As a preferred technical solution, the following steps are included: B1, select a core as the running core of the master watchdog thread, and bind slave watchdog threads to other cores respectively; B2, the master watchdog thread configures and starts all slave watchdog threads, and monitors the running status of all slave watchdog threads; B3, the slave watchdog thread periodically sends heartbeat signals to the master watchdog thread and reports the status of the critical task threads monitored by itself; B4, if the master watchdog thread detects that a slave watchdog thread or a critical task thread is abnormal, it takes exception handling measures.

[0014] As a preferred technical solution, the exception handling measures include restarting the slave watchdog thread or triggering a fault recovery mechanism.

[0015] As a preferred technical solution, a multi-core watchdog thread implementation method with dynamic load balancing is adopted: the number and / or binding position of the watchdog threads can be dynamically adjusted according to the real-time load situation.

[0016] As a preferred technical solution, the following steps are included: C1, preset a certain number of watchdog threads according to the number of cores and task requirements; C2, real-time monitoring of the load of each core of the system; C3, when the load of a core is higher than the set value, the watchdog thread on the core is migrated to the core with a load lower than the set value; Among them, the monitoring target of the watchdog thread can be dynamically adjusted during the operation to ensure that the watchdog thread always supervises the execution of the critical task thread.

[0017] As a preferred technical solution, the load condition indicator is CPU usage or memory occupancy.

[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. Improve system stability and reliability: By binding watchdog threads on multiple cores, mutual supervision and supervision of critical task threads are achieved. Even if the load of a core is too high, the watchdog threads on other cores can still maintain supervision of the system, ensuring the stability and reliability of the system.

[0019] 2. Reduce fault recovery time: When an abnormal situation is detected, the watchdog thread on the multi-core can quickly take corresponding measures, such as restarting critical task threads or triggering a fault recovery mechanism, thereby reducing the impact of the fault on the system and reducing the fault recovery time.

[0020] 3. Optimize resource utilization: The design of multi-core watchdog threads can better utilize the computing resources of multi-core multi-threaded processors and improve overall performance. Through reasonable thread scheduling and task allocation, it can ensure that critical task threads receive sufficient resource support while maintaining supervision and management of other tasks.

[0021] 4. Improve system fault tolerance: The design of multi-core watchdog threads can improve the system's fault tolerance. When a core or thread fails, the watchdog threads on other cores can detect and take measures in time to prevent the fault from spreading and reduce the risk of system crashes.

[0022] 5. Simplify system design and management: By binding watchdog threads on multiple cores, system design and management can be simplified. Developers can configure and adjust the parameters and behaviors of watchdog threads more flexibly to meet the needs of different applications. At the same time, the design of multi-core watchdog threads also provides more abundant monitoring and debugging methods, which is convenient for developers to conduct in-depth analysis and optimization of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 The following is a schematic diagram of an implementation method listed for a specific embodiment of the present invention. DETAILED DESCRIPTION

[0024] The present invention will be further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.

[0025] Example 1 In order to overcome the limitations of the prior art and improve the security and stability of multi-core multi-threaded systems, the present invention proposes a design idea of ​​binding watchdog threads on multiple cores, such as Figure 1 As shown in the figure, it includes peer supervision of multi-core watchdog threads, master-slave multi-core watchdog threads, and multi-core watchdog threads with dynamic load balancing. This design allows watchdog threads to run on different cores, achieving mutual supervision and supervision of key task threads (the main thread of the application running in a normal cycle is the key thread, and multiple cores may run multiple cycles of key threads). In this way, even if the load of a core is too high, the watchdog threads on other cores can still maintain supervision of the system to ensure the stability and reliability of the system.

[0026] To summarize, based on the high reliability and security requirements of a multi-core, multi-threaded architecture, the present invention proposes a design scheme for binding watchdog threads on multiple cores to supervise each other, aiming to improve the stability and reliability of the system, reduce fault recovery time, and optimize resource utilization, thereby providing strong protection for the security of multi-core, multi-threaded programs.

[0027] The present invention can be implemented in the following manner (taking two cores as an example): 1. Multi-core deployment: A watchdog thread is also bound to core 2 to implement mutual supervision with the watchdog thread on core 1. This design can improve the fault tolerance and response speed of the system.

[0028] 2. Mutual supervision mechanism: The two watchdog threads are located on core 1 and core 2 respectively. They not only supervise each other's status, but also supervise the execution of critical task threads. When any thread detects an abnormality, it can quickly take corresponding measures, such as restarting the critical task thread or triggering the fault recovery mechanism.

[0029] 3. Mission-critical thread supervision: Two watchdog threads are jointly responsible for supervising the execution of mission-critical threads to ensure that they run in a normal state. When an abnormality is detected in a mission-critical thread, the watchdog thread will trigger a fault recovery mechanism to ensure the stability and availability of the system.

[0030] The present invention has the following advantages: 1. Improve system stability and reliability: By binding watchdog threads on multiple cores, mutual supervision and supervision of critical task threads are achieved. Even if the load of a core is too high, the watchdog threads on other cores can still maintain supervision of the system, ensuring the stability and reliability of the system.

[0031] 2. Reduce fault recovery time: When an abnormal situation is detected, the watchdog thread on the multi-core can quickly take corresponding measures, such as restarting critical task threads or triggering fault recovery mechanisms, thereby reducing the impact of the fault on the system and reducing the fault recovery time.

[0032] 3. Optimize resource utilization: The design of multi-core watchdog threads can better utilize the computing resources of multi-core multi-threaded processors and improve overall performance. Through reasonable thread scheduling and task allocation, it can ensure that critical task threads receive sufficient resource support while maintaining supervision and management of other tasks.

[0033] 4. Improve system fault tolerance: The design of multi-core watchdog threads can improve the system's fault tolerance. When a core or thread fails, the watchdog threads on other cores can detect and take measures in time to prevent the fault from spreading and reduce the risk of system crashes.

[0034] 5. Simplify system design and management: By binding watchdog threads on multiple cores, system design and management can be simplified. Developers can configure and adjust the parameters and behaviors of watchdog threads more flexibly to meet the needs of different applications. At the same time, the design of multi-core watchdog threads also provides more abundant monitoring and debugging methods, which is convenient for developers to conduct in-depth analysis and optimization of the system.

[0035] Example 2 Based on Example 1, this example provides a more detailed implementation method.

[0036] Multi-core watchdog thread peer supervision: In this implementation, a watchdog thread is bound to each core, and these threads are in a peer position to supervise the execution status of other cores and critical task threads. Each watchdog thread is responsible for monitoring a specified number of other cores and critical task threads to ensure that they run as expected.

[0037] Specific steps: 1. When the application is initialized, a watchdog thread is assigned to each core and their monitoring targets (other cores and mission-critical threads) are configured.

[0038] 2. Each watchdog thread periodically sends a heartbeat signal to the target it monitors and waits for a response.

[0039] 3. If the watchdog thread does not receive a response from its monitored target within a predetermined time, it determines that the target may have failed and triggers the corresponding fault recovery mechanism.

[0040] Among them, watchdog threads communicate with each other through shared memory or message queues to coordinate fault recovery and other system operations.

[0041] advantage: Each core has an independent watchdog thread, which improves the reliability and fault tolerance of the system.

[0042] It achieves peer supervision between cores and avoids the problem of single point failure.

[0043] Example 3 Based on Example 1, this example provides a more detailed implementation method.

[0044] Master-slave multi-core watchdog thread: In this implementation, there is a master watchdog thread and multiple slave watchdog threads. The master watchdog thread is responsible for coordinating and supervising all slave watchdog threads, while the slave watchdog threads are bound to different cores and are responsible for supervising the execution of critical task threads.

[0045] Specific steps: 1. When the application is initialized, select a core as the running core of the main watchdog thread, and bind slave watchdog threads on other cores respectively.

[0046] 2. The master watchdog thread is responsible for configuring and starting all slave watchdog threads and monitoring their running status.

[0047] 3. The slave watchdog thread periodically sends heartbeat signals to the master watchdog thread and reports the status of the critical task threads it monitors.

[0048] 4. If the master watchdog thread detects an abnormality in a slave watchdog thread or a critical task thread, it will take corresponding measures, such as restarting the slave watchdog thread or triggering a fault recovery mechanism.

[0049] advantage: The master-slave structure is clear and easy to manage and maintain.

[0050] The master watchdog thread can centrally monitor all slave watchdog threads and key task threads, thus improving monitoring efficiency.

[0051] Example 4 Based on Example 1, this example provides a more detailed implementation method.

[0052] Dynamic load balancing of multi-core watchdog threads: In this implementation, the number and binding position of the watchdog threads can be dynamically adjusted according to the real-time load of the system. Through dynamic load balancing, it can be ensured that the watchdog thread always runs on the core with a lighter load, thereby improving the overall performance and stability of the system.

[0053] Specific steps: 1. When the application is initialized, a certain number of watchdog threads are preset according to the number of cores and task requirements.

[0054] 2. Real-time monitoring of the load of each core of the system, including CPU usage, memory usage, etc.

[0055] 3. When the load of a core is too high, the watchdog thread on the core is migrated to a core with a lighter load.

[0056] Among them, the monitoring targets of the watchdog threads can be dynamically adjusted during operation to ensure that they always supervise the execution of critical task threads.

[0057] advantage: Through dynamic load balancing, it can be ensured that the watchdog thread always runs in the best state, improving the overall performance and stability of the system.

[0058] It is suitable for complex and changing mission environments and can adaptively adjust system configuration to meet needs.

[0059] As described above, the present invention can be preferably implemented.

[0060] All features disclosed in all embodiments in this specification, or steps in all methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or expanded or replaced in any manner.

[0061] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. According to the technical essence of the present invention, within the spirit and principles of the present invention, any simple modification, equivalent replacement and improvement made to the above embodiment still falls within the protection scope of the technical solution of the present invention.

Claims

1. A watchdog monitoring method for multi-core and multi-threaded systems, characterized in that: Bind the watchdog thread to N cores; where N ≥ 2 and N is an integer.

2. A watchdog monitoring method for multi-core and multi-thread according to claim 1, characterized in that: The implementation method of multi-core watchdog thread peer supervision is adopted: a watchdog thread is bound to each core, and each watchdog thread is responsible for monitoring the watchdog threads of a specified number of other cores and the critical task threads of other cores.

3. A watchdog monitoring method for multi-core and multi-thread according to claim 2, characterized in that: The following steps are involved: A1, assign a watchdog thread to each core and configure the monitoring target of the watchdog thread; the monitoring target of the watchdog thread includes the watchdog threads of other cores and the key task threads of other cores; A2, each watchdog thread periodically sends a heartbeat signal to the target it monitors and waits for a response; A3, if the watchdog thread does not receive a response from the monitoring target within a predetermined time, it is determined that the monitoring target may have failed, and a corresponding fault recovery mechanism is triggered.

4. A watchdog monitoring method for multi-core and multi-thread according to claim 2 or 3, characterized in that: When the multi-core watchdog thread peer supervision is implemented, the watchdog threads communicate with each other through shared memory or message queues.

5. The watchdog monitoring method for multi-core and multi-thread according to claim 1, characterized in that: The implementation method of the master-slave multi-core watchdog thread is adopted: there is a master watchdog thread and multiple slave watchdog threads. The master watchdog thread is responsible for supervising all slave watchdog threads. The slave watchdog threads are bound to different cores respectively. The slave watchdog thread is responsible for supervising the execution of the key task thread of the core.

6. A watchdog monitoring method for multi-core and multi-thread according to claim 5, characterized in that: The following steps are involved: B1, select a core as the running core of the master watchdog thread, and bind slave watchdog threads to other cores respectively; B2, the master watchdog thread configures and starts all slave watchdog threads, and monitors the running status of all slave watchdog threads; B3, the slave watchdog thread periodically sends heartbeat signals to the master watchdog thread and reports the status of the critical task threads monitored by itself; B4, if the master watchdog thread detects that a slave watchdog thread or a critical task thread is abnormal, it takes exception handling measures.

7. A watchdog monitoring method for multi-core and multi-thread according to claim 6, characterized in that: Exception handling measures include restarting the watchdog thread or triggering the fault recovery mechanism.

8. The watchdog monitoring method for multi-core and multi-thread according to claim 1, characterized in that: The implementation method of multi-core watchdog threads with dynamic load balancing is as follows: the number and / or binding position of the watchdog threads can be dynamically adjusted according to the real-time load situation.

9. A watchdog monitoring method for multi-core and multi-thread according to claim 8, characterized in that: The following steps are involved: C1, preset a certain number of watchdog threads according to the number of cores and task requirements; C2, real-time monitoring of the load of each core of the system; C3, when the load of a core is higher than the set value, the watchdog thread on the core is migrated to the core with a load lower than the set value; Among them, the monitoring target of the watchdog thread can be dynamically adjusted during the operation to ensure that the watchdog thread always supervises the execution of the critical task thread.

10. A watchdog monitoring method for multi-core and multi-thread according to claim 8 or 9, characterized in that: The load condition indicator is CPU usage or memory usage.

Citation Information

Patent Citations

  • Watchdog feeding method and system in Linux system

    CN115658356A

  • Multi-level watchdog design method and device, equipment and storage medium

    CN116048861A

  • Optical sensor

    EP4428583A1

  • Electronic device for building training data of artificial intelligence model using capsule endoscopy images and method for building the same

    KR1020240041220A