Multilayer architecture function safety monitoring system of complex software system

Through the multi-layered functional safety monitoring system, the problem of insufficient monitoring of cross-domain faults in complex software systems is solved, self-recovery and stable operation are achieved, the system availability and functional safety are improved, and hardware costs are saved.

CN120704934AActive Publication Date: 2025-09-26AUTOCORE INTELLIGENT TECH (NANJING) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511150151.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-09-26
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing functional safety monitoring mechanisms lack effective monitoring strategies when dealing with failures in complex software systems, especially in cross-domain and cross-system architecture designs, resulting in a poor user experience. Existing solutions usually use a restart method to handle failures, affecting system stability.

Method used

A multi-layered functional safety monitoring system for complex software systems is designed, which includes a system operation monitor, a state management and switching module, a virtual watchdog, and a global time scheduling state monitoring module. Through multi-level supervision and processing measures, cross-domain software execution timing and logic monitoring is achieved, and self-recovery and flexible fault handling are supported.

Benefits of technology

It achieves self-recovery and stable operation of cross-domain faults in complex software systems, improves system availability, meets the requirements of different functional safety levels, saves hardware costs, and avoids system restarts caused by single faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704934A_ABST
    Figure CN120704934A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-layer architecture function safety monitoring system of a complex software system. The multi-layer architecture function safety monitoring system comprises a system operation monitor, a state management and switching module, a standby state management and switching module, a virtual watchdog and a global time scheduling state monitoring module. The system operation monitor comprises a state monitoring module, a state switching self-checking module, a fault processing measure module and a watchdog management module; the system operation monitor is used for monitoring faults of application layer software, middle layer software and an operation system, an external virtual watchdog is designed to supervise the faults, and the virtual global time scheduling state supervision module is used for supervising global time scheduling services. According to the method, the global supervision state combination provides various fault judgment modes, flexible supervision according to services is facilitated, and various processing measures are configured; the layered fault monitoring and processing guarantee that self-recovery is carried out under the condition that single-function software or middle-layer software fails, and normal operation of other software is not affected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a functional safety monitoring mechanism for a software system, and in particular to a multi-layered functional safety monitoring system for a complex software system. Background Art

[0002] With the increasing popularity of intelligent driving, smart cockpit functions and SOA architecture, automotive electrical and electronic systems (EE systems) have become increasingly complex, covering key areas such as power control, chassis management and autonomous driving.

[0003] With the trend toward multi-domain integration, high-performance central computing (HPC) platforms are crucial to achieving this goal. HPC utilizes one or more powerful processors or processor clusters, deploying a variety of operating systems based on varying application requirements and functional safety levels. This allows various functional modules to run concurrently on different operating systems, significantly improving resource utilization and system adaptability.

[0004] Any failure in these systems could pose a direct threat to passenger safety. Today, the automotive industry is undergoing a major shift from hardware-driven to software-driven. This change not only improves functional flexibility and user experience, but also places higher demands on software security and reliability.

[0005] However, existing functional safety monitoring mechanisms based on AUTOSAR or other similar frameworks can partially self-recover from application-layer software failures. Problems with mid-layer monitoring and state management handoff modules typically require a reboot, which can lead to a poor user experience in HPC environments. Furthermore, when it comes to cross-domain and cross-system architecture design, most current solutions on the market are limited to reporting fault information for a single domain or system, lacking effective cross-domain system monitoring strategies. Summary of the Invention

[0006] In order to address the deficiencies in the prior art, the present invention aims to provide a multi-layer architecture functional safety monitoring system for complex software systems.

[0007] To achieve the purpose of the present invention, the technical solution adopted by the present invention is: A multi-layered functional safety monitoring system for a complex software system includes a system operation monitor, a state management and switching module, a standby state management and switching module, a virtual watchdog, and a global time scheduling state monitoring module. The system operation monitor, state management and switching module, and standby state management and switching module are deployed on the A core, the virtual watchdog is deployed on the M core / R core, and the global time scheduling state monitoring module is deployed in the A core system operation monitor or in the M core / R core. The system operation monitor includes a status monitoring module, a status switching self-test module, a fault handling measure module, a watchdog management module, and a global time scheduling status monitoring module; the system operation monitor is used to monitor failures of application layer software, middle layer software, and operating system, and an external virtual watchdog is designed to monitor them. The virtual watchdog cooperates with the watchdog management module to restart the system; the global time scheduling status monitoring module is used to monitor the global time scheduling business.

[0008] Furthermore, in the architecture of a single operating system (OS), the watchdog management module in the system monitoring manager is associated with a virtual watchdog deployed in an external MCU or an independent safety island (FSI). Under normal circumstances, the watchdog is fed within a specified time window. Under abnormal circumstances, the watchdog is stopped from being fed or fed incorrectly, and the SOC chip is restarted by triggering the virtual watchdog deployed in the external MCU or FSI.

[0009] Furthermore, in a multi-operating system based on a hypervisor, the watchdog management module in the system monitoring manager is associated with the software virtual watchdog in the hypervisor. Under normal circumstances, the watchdog is fed within a specified time window. Under abnormal circumstances, if the watchdog is stopped or fed incorrectly, the virtual watchdog in the hypervisor triggers the restart of the corresponding virtual machine VM. The watchdog management module deployed in the hypervisor is associated with the virtual watchdog deployed in the external MCU or the independent safety island FSI. Under normal circumstances, the watchdog is fed within the specified time window. Under abnormal circumstances, the watchdog is stopped or fed incorrectly, and the virtual watchdog deployed in the external MCU or FSI triggers the restart of the SOC chip.

[0010] Furthermore, the global time scheduling status supervision module is used to realize the supervision of the global time scheduling business, set the global time and execution duration for the corresponding task to be scheduled and pulled up, and compare it with the scheduling table set by the user. If a single task execution times out or the pull-up order is incorrect, the corresponding fault is output.

[0011] Furthermore, the status monitoring module realizes the supervision of the running status of the application layer and the middle layer software through the flexible configuration of basic supervision into global supervision. Basic supervision includes survival supervision, timeout supervision and logical supervision. The global supervision combination includes: single basic supervision combination: a single basic supervision is directly defined as a global supervision; multiple basic supervision combination: global supervision consists of two or more basic supervisions.

[0012] Furthermore, the state switching self-test module periodically sends a test signal to the state management and switching module by simulating a fault signal. The specific steps are as follows: Set the self-test mode period: set a period to trigger the self-test mechanism regularly; Simulated fault signal: In self-test mode, the system generates a simulated fault signal and sends it to the status management and switching module; Verify response time: Monitor the state management and switching module's response to the received simulated fault signal to ensure that it can correctly issue the switching instruction within the set time range. If no correct feedback is received within the set time, the state switching path is considered to be faulty.

[0013] Furthermore, the causes of the failure include a failure in the state management and switching module itself, or an IPC communication failure between the system operation monitor and the state management and switching module; When the system operation monitor detects a fault in the self-test of the status management and switching module, it switches to the backup status management and switching module; the IPC communication between the system operation monitor and the status management and switching module uses the IPC primary and backup channels to form IPC communication redundancy, and automatically switches to the backup IPC channel when the IPC primary channel fails.

[0014] Furthermore, the fault handling measures module includes implementing application layer and middle layer fault handling measures, state switching self-check failure handling measures, operating system OS fault handling measures, Hypervisor fault handling measures, and global time scheduling state supervision fault handling measures.

[0015] Furthermore, the application layer and middle layer fault handling measures include switching the corresponding software component operating status and notifying the application layer. The two measures occur in parallel without priority. The switching status is configured according to different business and functional safety goals, including switching to another backup function group to enable, or configuring the faulty software component to switch to shutdown and then pull it up for fault self-recovery.

[0016] Furthermore, the state switching self-test failure handling measures are divided into two steps. The first step is to switch to the backup state management and switching module & switch to the backup IPC mode, and perform self-test again. If the self-test passes, it will not enter the second step; if the self-test fails again, it will enter the second step, stop watching the dog, and restart the system through the external watchdog.

[0017] The beneficial effect of this invention is that, compared with the existing technology, the Autosar solution commonly used in automotive controllers currently provides a combination approach, but as long as a basic monitoring failure in the combination is detected, it will be judged as a fault. The global monitoring state combination designed by this application provides a more diverse fault determination method, facilitating flexible monitoring based on business needs and configuring diverse treatment measures. Compared to the current common solution that only monitors the software execution status of a single controller, the global time scheduling state monitoring implemented by this application can realize cross-domain software execution timing and execution logic monitoring.

[0018] The hierarchical fault monitoring and handling method ensures self-recovery in the event of a single functional software or middle-layer software failure without affecting the normal operation of other functional software. In the case of a failure of a virtual machine VM in a hypervisor-based multi-operating system architecture, the single virtual machine VM can be restarted independently without affecting other virtual machines VM. The entire system will only be restarted when the operating system and hypervisor that cannot self-recover have serious faults. Compared to the software fault monitoring solutions currently on the market, most fault handling measures are to restart the system. The multi-level monitoring and handling measures of this application can improve the availability of the product while meeting the functional safety requirements.

[0019] This application uses a configurable virtual watchdog and watchdog management module to monitor software module failures and cover hardware failures in processing units, clocks, and other hardware at different ASIL levels. In HPC, deploying a virtual watchdog outside of the SOC in an MCU or independent safety island (FSI) can meet the functional safety requirements for independence. Compared to the virtual watchdogs commonly used in current market solutions, which usually only monitor software modules, this application offers the advantage of saving hardware costs if hardware component failure coverage, especially ASIL D hardware failure coverage, is typically achieved using hardware watchdogs. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Schematic diagram of the multi-layered functional safety monitoring system for complex software systems according to the present invention; Figure 2 It is a global supervision concept map; Figure 3 It is a single operating system architecture diagram; Figure 4 This is a diagram of a multi-operating system architecture based on Hypervisor; Figure 5 This is a diagram of the entire vehicle architecture based on the central computing unit HPC; Figure 6 It is a flowchart of application layer fault handling measures; Figure 7 This is a flowchart of the measures to be taken when the state switching self-test fails; Figure 8 This is a flowchart of operating system OS troubleshooting measures. DETAILED DESCRIPTION

[0021] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and are not intended to limit the scope of protection of this application.

[0022] like Figure 1As shown, the multi-layer architecture functional safety monitoring system of the complex software system described in the present invention includes a system operation monitor, a state management and switching module, a standby state management and switching module, a virtual watchdog, and a global time scheduling state supervision module.

[0023] Used to monitor software failures in complex operating system architectures running on high-performance processor chips and execution failures of cross-domain scheduled software, including: monitoring the timing and logic errors of application layer software operation; monitoring middle layer software service failures; self-checking the status management and switching modules; monitoring operating system failures that affect the operation of application layer software; monitoring hypervisor (virtual machine monitor) failures; monitoring the execution time and logic of cross-domain scheduling software.

[0024] The system operation monitor, state management and switching module, and standby state management and switching module are deployed on the A core, while the virtual watchdog is deployed on the M / R core. The A core is the SOC, while the M / R core is the ZCU / MCU. Global time scheduling supervision is deployed in the system operation monitor in the A core and can also be deployed in the M / R core.

[0025] The system operation monitor includes a status monitoring module, a state transition self-test module, a fault handling module, a watchdog management module, and a global time scheduling status monitoring module. The system operation monitor is used to monitor faults in the A core's application layer software, middleware software, and operating system (OS). The system operation monitor is the primary monitoring module for the A core. To prevent common cause failures, it is supervised by an external virtual watchdog (alternatively, a hardware watchdog can be used).

[0026] The virtual watchdog primarily works with the watchdog management module. It remains inactive when it receives the correct "feed" signal at the correct time. However, if it fails to do so, it triggers a restart signal, effectively restarting the virtual machine (VM) or SoC. Three watchdog types are configurable: timed, windowed, and acknowledged, to address different functional safety levels. The timed watchdog meets ASIL A requirements, the windowed watchdog meets ASIL B requirements, and the acknowledged watchdog meets ASIL C / D requirements. The virtual watchdog can be deployed in an external MCU, an independent safety island (FSI), or a hypervisor.

[0027] The multi-layer architecture functional safety monitoring system for complex software systems described in the present invention can be adapted to fault monitoring of different software architectures such as a single operating system (OS) and a hypervisor-based multi-operating system (OS) in different deployment forms.

[0028] like Figure 3As shown in the figure, in a single-OS architecture, the watchdog management module in the system monitoring manager is associated with a virtual watchdog deployed in an external MCU or an independent safety island (FSI). Under normal circumstances, the watchdog is fed within a specified time window. Under abnormal circumstances, if the watchdog stops feeding or feeds incorrectly, the virtual watchdog deployed in the external MCU or FSI triggers a restart of the SoC chip.

[0029] like Figure 4 As shown in the figure, in a hypervisor-based multi-operating system (OS) architecture, the watchdog management module in the system monitoring manager is associated with the software virtual watchdog in the hypervisor. Under normal circumstances, the watchdog is fed within a specified time window. Under abnormal circumstances, the watchdog is stopped or fed incorrectly, triggering a restart of the corresponding virtual machine (VM). The watchdog management module deployed in the hypervisor is associated with a virtual watchdog deployed in an external MCU or an independent safety island (FSI). Under normal circumstances, the watchdog is fed within a specified time window. Under abnormal circumstances, the watchdog is stopped or fed incorrectly, triggering a restart of the SoC chip by the virtual watchdog deployed in the external MCU or FSI.

[0030] Figure 4 The number of ZCUs and virtual machines (VMs) shown in the figure is for reference only. Different vehicle architectures may have different numbers of ZCUs and virtual machines (VMs).

[0031] like Figure 5 The figure below illustrates a typical vehicle architecture based on a central computing unit (HPC). Key business software is deployed on the HPC's A core, but vehicle functionality may require collaboration between the HPC and the ZCU. For example, ZCU1 collects and processes data, the HPC A core processes business software algorithms, and ZCU2 controls output. The A core and M / R cores collaborate to implement vehicle functionality, and sequential logic is often required between the software running on each core.

[0032] A global time scheduling status monitoring module is designed to monitor the global time scheduling business. The global time and execution duration for the corresponding task to be scheduled and pulled up can be set and compared with the user-set schedule. If a single task execution times out or the pull-up order is incorrect, the corresponding fault will be output.

[0033] The system operation monitor includes a status supervision module, a status switching self-check module, a fault handling measure module, a watchdog management module, and a global time scheduling status supervision module.

[0034] The status monitoring module can be combined into global monitoring through flexible configuration of basic monitoring to monitor the running status of application layer and middle layer software.

[0035] like Figure 2 As shown, survival monitoring, timeout monitoring, and logical monitoring can all be used as individual basic monitoring units. These monitoring can be performed in application software and middleware to achieve monitoring purposes. The status monitoring module determines the status of individual basic monitoring units based on their checkpoint status and the running status of the process in which they reside.

[0036] When the monitored process is not activated, the basic monitoring status is "inactive", if no fault is detected, it is "normal", and if a fault is detected, it is "faulty".

[0037] Because the software architecture of an HPC controller can coexist with multiple functional groups similar to those in traditional automotive ECU architectures, the basic supervisory faults set in each functional group may correspond to different failure modes, or multiple basic supervisory faults may map to the same failure mode. Therefore, a mechanism is needed to combine different basic supervisors so that the global supervisory status can be derived from their states.

[0038] The state supervision determines the global supervision state according to the basic supervision state and the state of the process. Each functional group can have multiple global supervisions at the same time.

[0039] Global supervision combination rules: Single basis supervision combination: A single basis supervision can be directly defined as a global supervision.

[0040] Combination of multiple basic supervisions: When a global supervision consists of two or more basic supervisions, the following two logical modes can be used to determine the status of the global supervision: AND mode: A global supervisor is in the "faulty" state only when all base supervisors within it are in the "faulty" state.

[0041] Example: If global supervision includes basic supervisions A, B, and C, the global supervision will be marked as "faulty" only when A, B, and C all fail.

[0042] "OR" mode (OR): As long as any basic supervision within the global supervision is in the "faulty" state, the global supervision state is "faulty".

[0043] Example: If the global supervisor includes base supervisors A, B, and C, then as soon as any one of A, B, or C fails, the global supervisor will be marked as "failed".

[0044] Failure to switch states in time after detecting a fault is considered a multi-point fault (a fault must occur first and then fail to switch states). Therefore, a state switching self-check module is designed to cover the potential fault of state switching failure.

[0045] The state switching self-test module periodically sends test signals to the state management and switching module by simulating fault signals. This process is intended to verify whether the state management and switching module can accurately respond and issue the corresponding switching instructions within the predetermined time. The specific steps are as follows: 1. Set the self-test mode cycle: Set a cycle to trigger the self-test mechanism regularly.

[0046] 2. Simulated fault signal: In self-test mode, the system will generate a simulated fault signal and send it to the status management and switching module.

[0047] 3. Verify response time: Monitor the state management and switching module's response to the received simulated fault signal to ensure that it can correctly issue the switching instruction within the set time range. If no correct feedback is received within the set time, the state switching path is considered to be faulty.

[0048] This fault may be caused by a fault in the Status Management and Switching Module itself or by a failure in the IPC communication between the System Operation Monitor and the Status Management and Switching Module. If the System Operation Monitor detects a fault in the Status Management and Switching Module's self-test, it will switch to the backup Status Management and Switching Module. IPC communication between the System Operation Monitor and the Status Management and Switching Module uses two different communication methods to provide redundancy. The IPC communication between the System Operation Monitor and the Status Management and Switching Module is a critical safety channel, serving as the primary and backup IPC channels. If the primary channel fails, it automatically switches to the backup IPC channel.

[0049] The watchdog management module primarily functions as a "watchdog feed" module at specified times. It can be configured with a virtual watchdog to provide three watchdog feeding modes: timed watchdog, windowed watchdog, and response watchdog. The watchdog management module can be deployed in the A-core system operation monitor or hypervisor.

[0050] The operating system (OS) is the foundation for application operations. OS failures can be categorized and handled rather than directly restarting the OS to improve product availability. OS failure monitoring and categorization are performed for user-mode services, and the OS kernel restarts the service process. If a critical system process crashes or freezes, the system monitor's watchdog management module is notified, triggering a reset of the external watchdog.

[0051] In products with a hypervisor, deploy a watchdog management module that regularly sends "feeding the dog" signals to an external virtual watchdog. Bind the key processes of the hypervisor to the watchdog management module. When deadlock, resource exhaustion, or other faults occur, the module triggers a reset signal through abnormal "feeding the dog" to restart the SOC chip and hypervisor.

[0052] The fault handling module can configure different fault handling measures based on different business scenarios and functional safety objectives. Fault handling measures include switching the functional group state, notifying the application layer (which then designs safety measures such as alarms), and notifying the watchdog to reset.

[0053] like Figure 6 As shown, the fault handling measures for the application and middle layers involve switching the corresponding software component's operating state and notifying the application layer. These two measures occur in parallel, regardless of priority. The switching state can be configured to meet different business and functional safety goals, including switching to a backup functional group for activation or shutting down and then restarting the faulty software component for self-recovery.

[0054] like Figure 7 As shown in the figure, the state switching self-test failure handling measures are divided into two steps. The first step is to switch to the backup state management and switching module & switch to the backup IPC mode, and then perform self-test again. If the self-test passes, it will not enter the second step; if the self-test fails again, it will enter the second step, stop feeding the watchdog, and restart the system through the external watchdog.

[0055] like Figure 8 As shown in the figure, OS fault handling measures, OS fault classification and processing, user-mode services, and restarting service processes by the OS kernel. If a critical system process crashes or locks, the system monitor's watchdog management is notified, which triggers an external watchdog reset. Restarting the SOC through the watchdog also notifies the M and application SWC, entering functional safety minimum safety mode. This maintains correct vehicle operation control and issues an alarm signal if the HPC stops functioning.

[0056] Hypervisor fault handling measures: In products with hypervisors, deploy a watchdog management module, regularly send "dog feeding" signals to the external virtual watchdog, bind the key hypervisor processes to the watchdog management, and when faults such as deadlock and resource exhaustion occur, abnormal "dog feeding" will trigger the external virtual watchdog to trigger a reset signal, restarting the SOC chip and hypervisor.

[0057] The global time scheduling status monitors fault handling measures, notifies the status management and switching modules to switch the software component status, notifies the application layer software, and re-synchronizes the global time. If synchronization cannot be restored, the software remains in the safe state it has switched to.

[0058] The beneficial effect of this invention is that, compared with the existing technology, the Autosar solution commonly used in automotive controllers currently provides a combination approach, but as long as a basic monitoring failure in the combination is detected, it will be judged as a fault. The global monitoring state combination designed by this application provides a more diverse fault determination method, facilitating flexible monitoring based on business needs and configuring diverse treatment measures. Compared to the current common solution that only monitors the software execution status of a single controller, the global time scheduling state monitoring implemented by this application can realize cross-domain software execution timing and execution logic monitoring.

[0059] The hierarchical fault monitoring and handling method ensures self-recovery in the event of a single functional software or middle-layer software failure without affecting the normal operation of other functional software. In the case of a failure of a virtual machine VM in a hypervisor-based multi-operating system architecture, the single virtual machine VM can be restarted independently without affecting other virtual machines VM. The entire system will only be restarted when the operating system and hypervisor that cannot self-recover have serious faults. Compared to the software fault monitoring solutions currently on the market, most fault handling measures are to restart the system. The multi-level monitoring and handling measures of this application can improve the availability of the product while meeting the functional safety requirements.

[0060] This application uses a configurable virtual watchdog and watchdog management module to monitor software module failures and cover hardware failures in processing units, clocks, and other hardware at different ASIL levels. In HPC, deploying a virtual watchdog outside of the SOC in an MCU or independent safety island (FSI) can meet the functional safety requirements for independence. Compared to the virtual watchdogs commonly used in current market solutions, which usually only monitor software modules, this application offers the advantage of saving hardware costs if hardware component failure coverage, especially ASIL D hardware failure coverage, is typically achieved using hardware watchdogs.

[0061] The applicant of the present invention has made a detailed explanation and description of the implementation examples of the present invention in conjunction with the drawings in the specification. However, those skilled in the art should understand that the above implementation examples are only preferred implementation plans of the present invention, and the detailed description is only to help readers better understand the spirit of the present invention, and is not a limitation on the scope of protection of the present invention. On the contrary, any improvements or modifications based on the inventive spirit of the present invention should fall within the scope of protection of the present invention.

Claims

1. A multi-layered functional safety monitoring system for a complex software system, characterized by: It includes a system operation monitor, a state management and switching module, a standby state management and switching module, a virtual watchdog, and a global time scheduling state monitoring module. The system operation monitor, state management and switching module, and standby state management and switching module are deployed on the A core, the virtual watchdog is deployed on the M core / R core, and the global time scheduling state monitoring module is deployed in the A core system operation monitor or in the M core / R core. The system operation monitor includes a status monitoring module, a status switching self-test module, a fault handling measure module, a watchdog management module, and a global time scheduling status monitoring module; the system operation monitor is used to monitor failures of application layer software, middle layer software, and operating system, and an external virtual watchdog is designed to monitor them. The virtual watchdog cooperates with the watchdog management module to restart the system; the global time scheduling status monitoring module is used to monitor the global time scheduling business.

2. The multi-layer architecture functional safety monitoring system for complex software systems according to claim 1, characterized in that: In the architecture of a single operating system (OS), the watchdog management module in the system monitoring manager is associated with the virtual watchdog deployed in the external MCU or in the independent safety island (FSI). Under normal circumstances, the watchdog is fed within the specified time window. Under abnormal circumstances, the watchdog is stopped or fed incorrectly, and the virtual watchdog deployed in the external MCU or FSI triggers the restart of the SOC chip.

3. The multi-layer architecture functional safety monitoring system for complex software systems according to claim 1, characterized in that: In a hypervisor-based multi-operating system (OS), the watchdog management module in the system monitoring manager is associated with the software virtual watchdog in the hypervisor. Under normal circumstances, the watchdog is fed within a specified time window. Under abnormal circumstances, if the watchdog is stopped or fed incorrectly, the virtual watchdog in the hypervisor triggers a restart of the corresponding virtual machine (VM). The watchdog management module deployed in the hypervisor is associated with the virtual watchdog deployed in the external MCU or the independent safety island FSI. Under normal circumstances, the watchdog is fed within the specified time window. Under abnormal circumstances, the watchdog is stopped or fed incorrectly, and the virtual watchdog deployed in the external MCU or FSI triggers the restart of the SOC chip.

4. The multi-layer architecture functional safety monitoring system for complex software systems according to claim 1, characterized in that: The global time scheduling status supervision module is used to implement supervision of the global time scheduling business. It sets the global time and execution duration for the corresponding task to be scheduled and pulled up, and compares it with the schedule table set by the user. If a single task execution times out or the pull-up order is incorrect, the corresponding fault is output.

5. The multi-layer architecture functional safety monitoring system for complex software systems according to claim 1, characterized in that: The status monitoring module realizes the supervision of the running status of the application layer and the middle layer software through the flexible configuration of basic supervision into global supervision. Basic supervision includes survival supervision, timeout supervision and logical supervision. The global supervision combination includes: single basic supervision combination: a single basic supervision is directly defined as a global supervision; multiple basic supervision combination: global supervision consists of two or more basic supervisions.

6. The multi-layer architecture functional safety monitoring system for complex software systems according to claim 1, characterized in that: The state switching self-test module periodically sends test signals to the state management and switching module by simulating fault signals. The specific steps are as follows: Set the self-test mode period: set a period to trigger the self-test mechanism regularly; Simulated fault signal: In self-test mode, the system generates a simulated fault signal and sends it to the status management and switching module; Verify response time: Monitor the state management and switching module's response to the received simulated fault signal to ensure that it can correctly issue the switching instruction within the set time range. If no correct feedback is received within the set time, the state switching path is considered to be faulty.

7. The multi-layer architecture functional safety monitoring system for complex software systems according to claim 6, characterized in that: The causes of the fault include a fault in the status management and switching module itself, or an IPC communication fault between the system operation monitor and the status management and switching module; When the system operation monitor detects a fault in the self-test of the status management and switching module, it switches to the backup status management and switching module; the IPC communication between the system operation monitor and the status management and switching module uses the IPC primary and backup channels to form IPC communication redundancy, and automatically switches to the backup IPC channel when the IPC primary channel fails.

8. The multi-layer architecture functional safety monitoring system for complex software systems according to claim 1, characterized in that: The fault handling module includes implementation of application layer and middle layer fault handling measures, state switching self-test failure handling measures, operating system OS fault handling measures, Hypervisor fault handling measures, and global time scheduling state supervision fault handling measures.

9. The multi-layer architecture functional safety monitoring system for complex software systems according to claim 8, characterized in that: Fault handling measures for the application layer and middle layer include switching the operating status of the corresponding software components and notifying the application layer. These two measures occur in parallel without priority. The switching status is configured according to different business and functional safety goals, including switching to another backup functional group for activation, or configuring the faulty software component to switch to shutdown and then pull it up for fault self-recovery.

10. The multi-layer architecture functional safety monitoring system for complex software systems according to claim 8, characterized in that: The state switching self-test failure handling measures are divided into two steps. The first step is to switch to the backup state management and switching module & switch to the backup IPC mode, and then perform self-test again. If the self-test passes, it will not enter the second step. If the self-test fails again, it will enter the second step, stop watchingdog feeding, and restart the system through the external watchdog.

Citation Information

Patent Citations

  • Fault detection and recovery method and system for virtual machine

    CN108762886A

  • Distributed control system (DCS) data distribution calculation method and system based on multi-task central processing unit (CPU)

    CN120179388A

  • Decision unit for fail operational sensors

    US20240427303A1

  • Intelligent power distribution controller and safety monitoring system and controller thereof, and vehicle

    WO2025118501A1