Fault recovery method and vehicle

By building a virtual system with a shared microcontroller in a system-on-a-chip, the problems of wasted hardware resources and high costs in traditional vehicles are solved, enabling accurate identification and rapid recovery of fault types and improving the reliability of vehicle systems.

CN121492983APending Publication Date: 2026-02-10GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202512020932.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In traditional vehicle systems, the TBox and the main unit are separate hardware components, which leads to wasted hardware resources and high hardware costs.

Method used

Multiple virtual systems, including a host instrumentation system, a host entertainment system, and a TBox system, are built within a single system-on-a-chip (SoC). They share the same microcontroller unit and achieve accurate fault identification and recovery through a heartbeat monitoring strategy.

Benefits of technology

It reduces hardware resource waste and hardware costs, enables accurate identification and rapid recovery of fault types, and improves system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121492983A_ABST
    Figure CN121492983A_ABST
Patent Text Reader

Abstract

The invention provides a fault recovery method and a vehicle, the method is applied to the technical field of vehicle control, and the method comprises the steps that after a host instrument system, a host entertainment system and a TBox system are started, a heartbeat monitoring strategy is started; carrying out heartbeat monitoring based on the heartbeat monitoring strategy to obtain a target monitoring result; determining a fault type based on the target monitoring result; and restarting a target module based on the fault type, wherein the target module comprises any one of a host entertainment system, a TBox system and a system-on-chip. According to the method, multiple virtual systems can be constructed in one system-on-chip, the multiple virtual systems comprise a host instrument system, a host entertainment system and a TBox system, and under the condition that the multiple virtual systems share the same microcontroller unit, accurate recognition of fault types is achieved, and therefore accurate recovery operation is executed according to the fault types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle control technology, and more specifically, to a fault recovery method and a vehicle in the field of vehicle control technology. Background Technology

[0002] In traditional vehicle systems, the telematics box (TBox) and the host are two separate hardware components, each with its own system-on-chip (SOC), microcontroller unit (MCU), and controller area network (CAN) transceiver. However, this leads to a waste of hardware resources and high hardware costs.

[0003] Therefore, while reducing the waste of hardware resources and hardware costs, how to perform accurate fault recovery has become a technical problem that needs to be solved. Summary of the Invention

[0004] This application provides a fault recovery method and a vehicle. The method can accurately identify fault types when multiple virtual systems are built in a system-on-a-chip and the multiple virtual systems share the same microcontroller unit, thereby performing precise recovery operations for the fault types.

[0005] Firstly, a fault recovery method is provided, applied to a vehicle comprising multiple virtual systems built on a system-on-a-chip (SoC). These virtual systems include a main instrument cluster system, a main infotainment system, and a TBox system, all sharing a common microcontroller unit. The method includes: activating a heartbeat monitoring strategy after the main instrument cluster system, main infotainment system, and TBox system have started; performing heartbeat monitoring based on the heartbeat monitoring strategy to obtain target monitoring results, including heartbeat monitoring results between the multiple virtual systems and heartbeat monitoring results between the virtual systems and the microcontroller unit; determining the fault type based on the target monitoring results; and restarting a target module based on the fault type, the target module including any one of the main infotainment system, the TBox system, and the SoC.

[0006] The above technical solution enables the construction of multiple virtual systems within a single system-on-a-chip (SoC). These virtual systems include a host instrumentation system, a host entertainment system, and a TBox system. With multiple virtual systems sharing the same microcontroller unit, the system can accurately identify fault types based on the heartbeat monitoring results between the multiple virtual systems and between the virtual systems and the microcontroller unit, thereby enabling precise recovery operations to be performed based on the fault type.

[0007] Secondly, a fault recovery device is provided, which can be applied in a vehicle. The vehicle includes multiple virtual systems built on a system-on-a-chip (SoC). The multiple virtual systems include a host instrument cluster system, a host entertainment system, and a TBox system, and the multiple virtual systems share the same microcontroller unit. The device includes: a first startup module, used to start a heartbeat monitoring strategy after the host instrument cluster system, host entertainment system, and TBox system have started; a heartbeat monitoring module, used to perform heartbeat monitoring based on the heartbeat monitoring strategy to obtain target monitoring results, the target monitoring results including heartbeat monitoring results between the multiple virtual systems and heartbeat monitoring results between the virtual systems and the microcontroller unit; a fault determination module, used to determine the fault type based on the target monitoring results; and a restart module, used to restart a target module based on the fault type, the target module including any one of the host entertainment system, TBox system, and SoC.

[0008] Thirdly, a vehicle is provided, comprising: a memory for storing executable program code; and a processor for calling and running the executable program code from the memory, causing the vehicle to perform the methods described in the first aspect or any possible implementation thereof.

[0009] Fourthly, a computer program product is provided, comprising: computer program code, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.

[0010] Fifthly, a computer-readable storage medium is provided that stores computer program code, which, when executed on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.

[0011] The possible implementations of aspects two through five have similar effects to those of aspect one and its possible implementations, and will not be elaborated upon here. Attached Figure Description

[0012] Figure 1 This is a system architecture diagram of the TBox and host in a traditional vehicle system; Figure 2 This is a system architecture diagram of the vehicle system provided in this application embodiment after the TBox and host are integrated; Figure 3 This is a schematic flowchart of a fault recovery method provided in an embodiment of this application; Figure 4 This is a schematic diagram illustrating the startup process of each virtual system provided in the embodiments of this application; Figure 5 This is a schematic diagram illustrating heartbeat monitoring between a microcontroller unit, a host instrumentation system, a host entertainment system, and a TBox system, provided in an embodiment of this application. Figure 6 This is a schematic diagram illustrating the specific process of the fault recovery method provided in the embodiments of this application; Figure 7 This is a system architecture diagram of a vehicle system after the TBox and host are integrated in the vehicle system, as provided in the embodiments of this application, for vehicle fault recovery; Figure 8 This is a swimlane diagram of the hierarchical monitoring strategy corresponding to the fault recovery method provided in the embodiments of this application; Figure 9 This is a schematic diagram of the structure of a fault recovery device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application. Detailed Implementation

[0013] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0014] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0015] To facilitate understanding of the technical solutions in the embodiments of this application, some terms involved in the embodiments of this application will be briefly explained below.

[0016] TBox: It is the core component in the vehicle responsible for remote communication and data processing. For example, when users remotely control the vehicle's air conditioning, windows and other components through mobile applications, they need to transmit relevant control commands through the TBox.

[0017] Head unit: refers to the component in a vehicle used to realize in-vehicle entertainment and instrument display functions, which can also be called the vehicle information system (head unit system, HUT).

[0018] Currently, in traditional vehicle systems, the TBox and the host are two independent hardware components, each with its own system-on-a-chip, microcontroller unit, and CAN transceiver.

[0019] like Figure 1 As shown, the TBox includes its corresponding system-on-a-chip (SoC), microcontroller unit, and CAN transceiver. The SoC and microcontroller unit in the TBox can be connected via a serial peripheral interface (SPI), and the microcontroller unit is connected to the CAN transceiver, enabling the microcontroller unit to transmit and receive CAN messages via the CAN transceiver and CAN bus.

[0020] Correspondingly, the host also includes its corresponding system-on-a-chip, microcontroller unit and CAN transceiver. The system-on-a-chip in the host and the microcontroller unit in the host can also be connected via SPI. The microcontroller unit in the host is connected to the CAN transceiver in the host, so that the microcontroller unit in the host can realize some CAN message sending and receiving functions through the CAN transceiver in the host and the CAN bus.

[0021] Therefore, in traditional vehicle systems, the TBox and the host are two independent hardware components, resulting in good hardware fault isolation between them. However, since the TBox and the host each have their own independent system-on-a-chip, microcontroller unit, and CAN transceiver, this leads to wasted hardware resources and high hardware costs.

[0022] To reduce the waste of hardware resources and hardware costs, embodiments of this application can combine the TBox and the host into one hardware unit, so that the TBox and the host can share the same microcontroller unit and the same system-on-a-chip. Of course, the TBox and the host can also share the same CAN transceiver.

[0023] It should be noted that in this embodiment, where the TBox and the host share the same system-on-a-chip (SoC) and the same microcontroller unit, the number of SoCs and microcontroller units can be reduced, thereby reducing hardware resource waste and hardware costs. Furthermore, in this embodiment, the TBox and the host can also share the same CAN transceiver, which can further reduce the number of CAN transceivers, thereby reducing hardware resource waste and hardware costs.

[0024] Furthermore, since the TBox and the host can share the same microcontroller unit and the same system-on-a-chip, a virtual machine monitor (i.e., a hypervisor) can be used to build multiple virtual systems on this shared system-on-a-chip. These multiple virtual systems can be Linux systems.

[0025] For example, such as Figure 2 As shown, a virtual machine monitor is used to build three virtual systems on a shared system-on-a-chip. One virtual system is the host instrumentation system (i.e., the server OS system, which is the service operating system in the virtualization environment), and the other two virtual systems are guest OS systems (i.e., guest operating systems in the virtualization environment). The two guest OS systems are the host entertainment system (i.e., the guest OS Android system) and the TBox system (i.e., the guest OS Linux system).

[0026] In other words, embodiments of this application can construct multiple virtual systems within a single system-on-a-chip (SoC). These virtual systems may include a host instrument system, a host entertainment system, and a TBox system, all sharing the same microcontroller unit. Since the host instrument system, host entertainment system, and TBox system are constructed within a single SoC, they also share the same SoC. Of course, in some embodiments, the host instrument system, host entertainment system, and TBox system may also share the same CAN transceiver.

[0027] like Figure 2 As shown, the shared system-on-a-chip, shared microcontroller unit, and shared CAN transceiver are located on the hardware platform, while the host instrument system, host entertainment system, and TBox system are located on the virtual system layer.

[0028] It should be understood that a virtual machine monitor (VM monitor) is a software layer installed on physical hardware. It can virtualize the physical hardware into multiple virtual systems, meaning that a VM monitor can virtualize hardware resources, allowing multiple virtual systems to run simultaneously on a single physical hardware device. There are two types of VM monitors: Type-1 and Type-2.

[0029] Type-1 virtual machine monitors, also known as bare-metal virtual machine monitors or native virtual machine monitors, do not require a pre-loaded underlying operating system. They directly access the underlying hardware without the need for other software (such as operating systems and device drivers), and can be installed directly on the hardware, splitting the hardware into multiple virtual machines, on which virtual systems are installed. Type-2 virtual machine monitors are typically installed on top of an existing operating system; these are called managed virtual machine hypervisors. For Type-2 virtual machine monitors, the presence of the underlying operating system inevitably introduces latency.

[0030] Therefore, in this embodiment of the application, the virtual machine monitor can be a Type-1 virtual machine monitor.

[0031] It should be understood that the host instrument system and host entertainment system are virtual systems used to implement the relevant functions of the host, and the TBox system is a virtual system used to implement the relevant functions of the TBox.

[0032] like Figure 2 As shown, the host entertainment system and the host instrumentation system can communicate via virtual socket (VSOCK), the TBox system and the host instrumentation system can also communicate via VSOCK, the TBox system and the shared microcontroller unit can communicate via SPI, and the host instrumentation system and the shared microcontroller unit can also communicate via SPI.

[0033] Therefore, while building virtual systems such as the main instrument cluster, main infotainment system, and TBox system within a single system-on-a-chip (SoC), with all three systems sharing the same microcontroller unit, can reduce hardware resource waste and costs, the corresponding changes in the vehicle's system architecture also lead to changes in the fault recovery strategy within that architecture. In this case, how to perform accurate fault recovery becomes a technical problem that needs to be solved.

[0034] To address the aforementioned issues, this application provides a fault recovery method. After the host instrument system, host entertainment system, and TBox system have started, a heartbeat monitoring strategy is initiated. Heartbeat monitoring is performed based on this strategy to obtain target monitoring results, which include heartbeat monitoring results between multiple virtual systems and heartbeat monitoring results between virtual systems and the microcontroller unit. The fault type is determined based on the target monitoring results. The target module is restarted based on the fault type. The target module includes any one of the host entertainment system, TBox system, and system-on-a-chip (SoC). Therefore, this application can construct multiple virtual systems within a single SoC. These virtual systems include the host instrument system, host entertainment system, and TBox system, and all virtual systems share the same microcontroller unit. Based on the heartbeat monitoring results between the virtual systems and between the virtual systems and the microcontroller unit, accurate fault type identification is achieved, enabling precise recovery operations to be performed based on the fault type.

[0035] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments.

[0036] Figure 3 This is a flowchart illustrating a fault recovery method provided in an embodiment of this application. This fault recovery method can be applied to a vehicle, which includes multiple virtual systems built on a system-on-a-chip (SoC). These virtual systems include a main instrument cluster system, a main infotainment system, and a TBox system, and all virtual systems share the same microcontroller unit. Figure 3 As shown, the fault recovery method may specifically include the following steps: S301: After the host instrument system, host entertainment system and TBox system have finished booting, start the heartbeat monitoring policy.

[0037] The TBox system is the first of multiple virtual systems built within a system-on-a-chip (SoC), and it is the virtual system used to implement the functions related to the TBox. The host entertainment system is the second of multiple virtual systems built within this SoC, and the host instrumentation system is the third of multiple virtual systems built within this SoC. The host entertainment system and the host instrumentation system are virtual systems used to implement the functions related to the host.

[0038] After the host instrumentation system, host entertainment system, and TBox system have started up, the heartbeat monitoring policy can be activated. This policy is used to monitor heartbeats between multiple virtual systems and between virtual systems and microcontroller units.

[0039] Specifically, the heartbeat monitoring strategy is used to: monitor the heartbeat between the host entertainment system and the host instrumentation system, monitor the heartbeat between the TBox system and the host instrumentation system, monitor the heartbeat between the TBox system and the microcontroller unit, and monitor the heartbeat between the host instrumentation system and the microcontroller unit.

[0040] S302, heartbeat monitoring is performed based on the heartbeat monitoring strategy to obtain target monitoring results. The target monitoring results include heartbeat monitoring results between multiple virtual systems and heartbeat monitoring results between virtual systems and microcontroller units.

[0041] In this embodiment of the application, heartbeat monitoring can be performed between multiple virtual systems to obtain heartbeat monitoring results between multiple virtual systems, and heartbeat monitoring can also be performed between a virtual system and a microcontroller unit to obtain heartbeat monitoring results between the virtual system and the microcontroller unit.

[0042] In one possible implementation, the above-mentioned S302 "perform heartbeat monitoring based on heartbeat monitoring strategy to obtain target monitoring results" may specifically include the following steps: performing heartbeat monitoring between the host entertainment system and the host instrument system to obtain a first heartbeat monitoring result; performing heartbeat monitoring between the TBox system and the host instrument system to obtain a second heartbeat monitoring result; performing heartbeat monitoring between the TBox system and the microcontroller unit to obtain a third heartbeat monitoring result; and performing heartbeat monitoring between the host instrument system and the microcontroller unit to obtain a fourth heartbeat monitoring result.

[0043] It should be understood that the first and second heartbeat monitoring results are heartbeat monitoring results between multiple virtual systems, while the third and fourth heartbeat monitoring results are heartbeat monitoring results between the virtual system and the microcontroller unit.

[0044] S303, determine the fault type based on the target monitoring results.

[0045] In this embodiment of the application, the fault type can be determined based on the first heartbeat monitoring result, the second heartbeat monitoring result, the third heartbeat monitoring result, and the fourth heartbeat monitoring result.

[0046] The types of malfunctions include any one of the following: malfunction of the main entertainment system, malfunction of the TBox system, and malfunction of the main instrument system.

[0047] S304, restarts the target module based on the fault type. The target module includes any one of the host entertainment system, TBox system, and system-on-a-chip.

[0048] In this embodiment of the application, after determining the fault type based on the target monitoring results, different target modules can be restarted for different fault types.

[0049] In one possible implementation, the above-mentioned S304 "restarting the target module based on the fault type" may specifically include the following steps: when the fault type is that the host entertainment system has failed, the host instrument system restarts the host entertainment system; when the fault type is that the TBox system has failed, the host instrument system restarts the TBox system; when the fault type is that the host instrument system has failed, the microcontroller unit restarts the system-on-a-chip.

[0050] Specifically, the microcontroller unit restarts the system-on-a-chip (SoC) by powering down the SoC and then powering it back on.

[0051] Since the impact of failures on virtual systems and physical links differs, this application embodiment designs differentiated recovery strategies for different levels of failures. Specifically, if the host entertainment system fails, only the host entertainment system is restarted, without affecting other virtual systems; if the TBox system fails, only the TBox system is restarted, without affecting other virtual systems; and if the host instrument system fails, the system-on-a-chip is restarted to ensure the consistency of each virtual system.

[0052] This can resolve system-level risks caused by single-point failures in the main instrument system and establish a rapid detection and automatic recovery mechanism for abnormal states of the main entertainment system and TBox system.

[0053] Using the above-mentioned fault recovery strategies, the average fault recovery time for the host entertainment system or TBox system can be less than 30 seconds, and the average fault recovery time for restarting the system-on-a-chip to recover the host instrument system can be less than 60 seconds.

[0054] Thus, the embodiments of this application can ensure timely fault detection based on a multi-level monitoring mechanism, avoid service interruption caused by excessive recovery through a layered recovery strategy, and prevent the spread of single-point faults through communication isolation design (i.e., SPI communication design between the TBox system and the shared microcontroller unit, and SPI communication design between the host instrument system and the shared microcontroller unit). This enables accurate identification of fault types and execution of precise recovery operations for each fault type, thereby improving system reliability.

[0055] The startup process of each virtual system is explained in detail below.

[0056] In one possible implementation, the fault recovery method may further include the following steps: the virtual machine monitor sends a first startup command to the host instrumentation system; after the host instrumentation system completes startup based on the first startup command, it sends a first startup request and a second startup request to the virtual machine monitor; the virtual machine monitor sends a second startup command to the host entertainment system based on the first startup request, and sends a third startup command to the TBox system based on the second startup request; after the host entertainment system completes startup based on the second startup command, it sends a first notification message to the virtual machine monitor, the first notification message indicating that the host entertainment system has completed startup; after the TBox system completes startup based on the third startup command, it sends a second notification message to the virtual machine monitor, the second notification message indicating that the TBox system has completed startup; and upon receiving the first and second notification messages, the virtual machine monitor sends a third notification message to the host instrumentation system, the third notification message indicating that the host instrumentation system, the host entertainment system, and the TBox system have completed startup.

[0057] For example, Figure 4 This is a schematic diagram illustrating the startup process of various virtual systems provided in the embodiments of this application. The modules involved in the process include a virtual machine monitor, a host instrumentation system, a host entertainment system, and a TBox system. The process may specifically include the following steps S401 to S410: S401, the virtual machine monitor sends the first startup command to the host instrumentation system.

[0058] The first start command is used to trigger the start operation of the host instrument system.

[0059] S402, the main instrument system performs a startup operation based on the first startup command.

[0060] After receiving the first start command, the host instrument system can initialize based on the first start command to start the host instrument system.

[0061] S403 After the host instrumentation system has finished booting, the host instrumentation system sends a first boot request and a second boot request to the virtual machine monitor.

[0062] It should be understood that the first and second startup requests sent by the host instrumentation system to the virtual machine monitor can be sent simultaneously or at different times. Taking the host instrumentation system sending the first and second startup requests to the virtual machine monitor at different times as an example, the host instrumentation system can send the first startup request to the virtual machine monitor first and then send the second startup request, or the host instrumentation system can send the second startup request to the virtual machine monitor first and then send the first startup request.

[0063] S404, the virtual machine monitor sends a second startup command to the host entertainment system based on the first startup request.

[0064] The second startup command is used to trigger the startup operation of the host entertainment system.

[0065] S405, the virtual machine monitor sends a third boot command to the TBox system based on the second boot request.

[0066] The third startup command is used to trigger the startup operation of the TBox system.

[0067] S406, the host entertainment system performs a startup operation based on the second startup command.

[0068] After receiving the second boot command, the host entertainment system can initialize based on the second boot command to start the host entertainment system.

[0069] The S407 TBox system performs the boot operation based on the third boot instruction.

[0070] After the TBox system receives the third boot command, the TBox system can initialize based on the third boot command to start the TBox system.

[0071] S408 After the host entertainment system has finished booting, the host entertainment system sends a first notification message to the virtual machine monitor. The first notification message is used to indicate that the host entertainment system has finished booting.

[0072] S409. After the TBox system has finished booting, the TBox system sends a second notification message to the virtual machine monitor. The second notification message is used to indicate that the TBox system has finished booting.

[0073] S410, after the virtual machine monitor receives the first notification message and the second notification message, the virtual machine monitor sends a third notification message to the host instrumentation system. The third notification message is used to indicate that the host instrumentation system, the host entertainment system and the TBox system have all completed startup.

[0074] Thus, based on the steps S401 to S410 described above, the virtual machine monitor can first start the host instrumentation system. After the host instrumentation system has finished starting, the virtual machine monitor then starts the host entertainment system and the TBox system, waiting for confirmation that all virtual systems have finished starting. Afterward, the heartbeat monitoring policy execution process can be triggered.

[0075] The process of heartbeat monitoring between the microcontroller unit, host instrumentation system, host entertainment system, and TBox system is described in detail below.

[0076] In one possible implementation, the above-mentioned S301 "starting the heartbeat monitoring strategy after the host instrument system, host entertainment system and TBox system have started" may specifically include the following steps: after the host instrument system, host entertainment system and TBox system have started, the host instrument system starts a timer; when the timer reaches a first preset duration, the heartbeat monitoring strategy is started; when the timer does not reach the first preset duration, the heartbeat monitoring strategy is not started.

[0077] Thus, after the host instrumentation system receives the third notification message indicating that all virtual systems have completed startup, it can initiate a first preset stabilization period timer. This allows the application layer initialization of each virtual system to be completed within this stabilization period. During this first preset stabilization period, no heartbeat monitoring is performed. Heartbeat monitoring is then initiated uniformly across all channels after the first preset stabilization period. In other words, after the host instrumentation system receives the third notification message from the virtual machine monitor, it starts a timer. If the timer reaches the first preset duration, the heartbeat monitoring strategy is activated; otherwise, no heartbeat monitoring is performed.

[0078] This is because the state of each virtual system is uncertain during initialization. When multiple virtual systems start in parallel, the heartbeat monitoring strategy needs to be activated at the appropriate time. Activating the heartbeat monitoring strategy too early may lead to misjudgment, while activating it too late may miss early faults. Therefore, in this embodiment, the heartbeat monitoring strategy is delayed for a first preset time before activation. This avoids misjudgment during the initialization process of each virtual system, preventing accidental triggering of fault recovery operations due to temporary state fluctuations. It ensures that each virtual system is fully ready before entering strict heartbeat monitoring and reduces the possibility of missing early faults.

[0079] In some embodiments, the first preset duration can be set to a fixed value according to actual conditions. For example, the first preset duration can be 1 minute, or the first preset duration can be 30 seconds, or the first preset duration can be 2 minutes, etc.

[0080] In other embodiments, the host instrumentation system can also dynamically adjust the delay start time of the heartbeat monitoring strategy based on the startup speed of each virtual system. Specifically, the faster the startup speed of each virtual system, the shorter the delay start time of the heartbeat monitoring strategy; and the slower the startup speed of each virtual system, the longer the delay start time of the heartbeat monitoring strategy.

[0081] For example, Figure 5This is a schematic diagram illustrating heartbeat monitoring between a microcontroller unit, a host instrumentation system, a host entertainment system, and a TBox system, as provided in an embodiment of this application. The modules involved in the process include the host instrumentation system, the host entertainment system, the TBox system, and the microcontroller unit. Specifically, the process may include the following steps S501 to S514: S501, Start timer for main instrument system.

[0082] After the host instrumentation system receives the third notification message sent by the virtual machine monitor, it starts a timer, that is, it executes step S501 after step S410 above.

[0083] S502, when the timer reaches the first preset duration, the host instrument system activates the heartbeat monitoring strategy.

[0084] If the timer has not reached the first preset duration, the host instrument system will not activate the heartbeat monitoring strategy until the timer reaches the first preset duration.

[0085] S503, the host entertainment system sends a first monitoring message to the host instrument system. The first monitoring message includes a first preset cycle.

[0086] The first monitoring message is used to indicate that the host entertainment system will subsequently send the first heartbeat packet to the host instrument system periodically at a first preset period, so that the host instrument system can determine whether the first heartbeat packet it receives has timed out based on the first preset period.

[0087] S504, the TBox system sends a second monitoring message to the host instrument system. The second monitoring message includes a second preset cycle.

[0088] The second monitoring message indicates that the TBox system will subsequently send the second heartbeat packet to the host instrument system periodically at a second preset period, so that the host instrument system can determine whether the second heartbeat packet it receives has timed out based on the second preset period.

[0089] S505, the TBox system sends a third monitoring message to the microcontroller unit, the third monitoring message including a third preset cycle.

[0090] The third monitoring message is used to indicate that the TBox system will subsequently send the third heartbeat packet to the microcontroller unit periodically at a third preset period, so that the microcontroller unit can determine whether the third heartbeat packet it receives has timed out based on the third preset period.

[0091] S506, the host instrument system sends a fourth monitoring message to the microcontroller unit. The fourth monitoring message includes a fourth preset cycle.

[0092] The fourth monitoring message is used to indicate that the host instrument system will subsequently send the fourth heartbeat packet to the microcontroller unit periodically at a fourth preset period, so that the microcontroller unit can determine whether the fourth heartbeat packet it receives has timed out based on the fourth preset period.

[0093] It should be noted that the steps S503 to S506 described above do not have a fixed order. They can be executed simultaneously or sequentially. This application embodiment does not limit this.

[0094] S507, the host entertainment system sends the first heartbeat packet to the host instrument system.

[0095] S508, upon receiving the first heartbeat packet, the host instrument system returns a first heartbeat response to the host entertainment system.

[0096] It should be understood that the host entertainment system sends the first heartbeat packet to the host instrument system periodically according to a first preset period. Furthermore, each time the host instrument system receives a first heartbeat packet, it returns a first heartbeat response to the host entertainment system. The first heartbeat response indicates that the host instrument system has received the first heartbeat packet sent by the host entertainment system.

[0097] S509, the TBox system sends a second heartbeat packet to the host instrument system.

[0098] S510: When the host instrument system receives the second heartbeat packet, the host instrument system returns a second heartbeat response to the TBox system.

[0099] It should be understood that the TBox system sends the second heartbeat packet to the host instrument system periodically according to a second preset period. Furthermore, each time the host instrument system receives a second heartbeat packet, it returns a second heartbeat response to the TBox system, indicating that the host instrument system has received the second heartbeat packet sent by the TBox system.

[0100] The S511 TBox system sends a third heartbeat packet to the microcontroller unit.

[0101] S512: When the microcontroller unit receives the third heartbeat packet, the microcontroller unit returns a third heartbeat response to the TBox system.

[0102] It should be understood that the TBox system sends the third heartbeat packet to the microcontroller unit periodically according to a third preset period. Furthermore, each time the microcontroller unit receives a third heartbeat packet, it returns a third heartbeat response to the TBox system, indicating that the microcontroller unit has received the third heartbeat packet sent by the TBox system.

[0103] S513, the host instrumentation system sends the fourth heartbeat packet to the microcontroller unit.

[0104] S514: When the microcontroller unit receives the fourth heartbeat packet, the microcontroller unit returns the fourth heartbeat response to the host instrumentation system.

[0105] It should be understood that the host instrumentation system sends the fourth heartbeat packet to the microcontroller unit periodically according to a fourth preset cycle. Furthermore, each time the microcontroller unit receives a fourth heartbeat packet, it returns a fourth heartbeat response to the host instrumentation system. This fourth heartbeat response indicates that the microcontroller unit has received the fourth heartbeat packet sent by the host instrumentation system.

[0106] It should be noted that S507 to S514 are executed periodically according to the corresponding preset cycle, and there is no fixed order for the steps S507, S509, S511 and S513; they can be executed simultaneously.

[0107] Thus, through step S507 above, the host instrumentation system can monitor the heartbeat interaction between the host entertainment system and the host instrumentation system to determine whether the first heartbeat monitoring result is abnormal. Correspondingly, through step S509 above, the host instrumentation system can monitor the heartbeat interaction between the TBox system and the host instrumentation system to determine whether the second heartbeat monitoring result is abnormal. Correspondingly, through step S511 above, the microcontroller unit can monitor the heartbeat interaction between the TBox system and the microcontroller unit to determine whether the third heartbeat monitoring result is abnormal. Correspondingly, through step S513 above, the microcontroller unit can monitor the heartbeat interaction between the host instrumentation system and the microcontroller unit to determine whether the fourth heartbeat monitoring result is abnormal.

[0108] In one possible implementation, the above-mentioned S303 "determine the fault type based on the target monitoring result" may specifically include the following steps: if the first heartbeat monitoring result is abnormal, determine the fault type as a fault in the host entertainment system; if the second or third heartbeat monitoring result is abnormal, determine the fault type as a fault in the TBox system; if the fourth heartbeat monitoring result is abnormal, determine the fault type as a fault in the host instrument system.

[0109] Therefore, the types of failures include any one of the following: failure of the main entertainment system, failure of the TBox system, and failure of the main instrument system.

[0110] In one possible implementation, the fault recovery method may further include the following steps: If the host instrument system receives M consecutive first heartbeat packets that all time out, determine that the first heartbeat monitoring result is abnormal. The first heartbeat packets are periodically sent by the host entertainment system to the host instrument system using a first preset period, where M is a positive integer. If the host instrument system receives N consecutive second heartbeat packets that all time out, determine that the second heartbeat monitoring result is abnormal. The second heartbeat packets are periodically sent by the TBox system to the host instrument system using a second preset period, where N is a positive integer. If the microcontroller unit receives K consecutive third heartbeat packets that all time out, determine that the third heartbeat monitoring result is abnormal. The third heartbeat packets are periodically sent by the TBox system to the microcontroller unit using a third preset period, where K is a positive integer. If the microcontroller unit receives X consecutive fourth heartbeat packets that all time out, determine that the fourth heartbeat monitoring result is abnormal. The fourth heartbeat packets are periodically sent by the host instrument system to the microcontroller unit using a fourth preset period, where X is a positive integer.

[0111] Specifically, the host instrument system determines whether a first heartbeat packet has timed out by checking if it receives one every first preset period. If the time interval between the reception time of the first heartbeat packet received by the host instrument system this time and the reception time of the last first heartbeat packet received by the host instrument system is less than or equal to the first preset period, the host instrument system determines that the first heartbeat packet received this time has not timed out. If the time interval between the reception time of the first heartbeat packet received by the host instrument system this time and the reception time of the last first heartbeat packet received by the host instrument system is greater than the first preset period, the host instrument system determines that the first heartbeat packet received this time has timed out.

[0112] Furthermore, each time the host instrument system determines that the first heartbeat packet has timed out, it increments the count by one, thus allowing the host instrument system to track the number of timeouts of the first heartbeat packets it receives. When the host instrument system receives M instances of timeouts for the first heartbeat packets, it can determine that the first heartbeat monitoring result is abnormal.

[0113] It should be understood that if the number of times the first heartbeat packet received by the host instrument system times out has not reached M, and if the host instrument system receives the first heartbeat packet again without timeout, the host instrument system can reset the number of times the first heartbeat packet times out to zero.

[0114] Accordingly, the host instrument system determines whether a second heartbeat packet is received every second preset period to determine if the second heartbeat packet has timed out. If the time interval between the reception time of the second heartbeat packet received by the host instrument system this time and the reception time of the second heartbeat packet received by the host instrument system last time is less than or equal to the second preset period, then the host instrument system determines that the second heartbeat packet received this time has not timed out. If the time interval between the reception time of the second heartbeat packet received by the host instrument system this time and the reception time of the second heartbeat packet received by the host instrument system last time is greater than the second preset period, then the host instrument system determines that the second heartbeat packet received this time has timed out.

[0115] Furthermore, each time the host instrument system determines that the second heartbeat packet has timed out, it increments the count by one, thus allowing the host instrument system to track the number of timeouts it receives. When the host instrument system receives N timeouts for the second heartbeat packet, it can determine that the second heartbeat monitoring result is abnormal.

[0116] It should be understood that if the number of times the second heartbeat packet received by the host instrument system times out has not reached N, and the host instrument system receives another second heartbeat packet without it timing out, then the host instrument system can reset the number of times the second heartbeat packet times out to zero.

[0117] Accordingly, the microcontroller unit determines whether a third heartbeat packet is received every third preset period to determine if the third heartbeat packet has timed out. If the time interval between the reception time of the third heartbeat packet received by the microcontroller unit this time and the reception time of the third heartbeat packet previously received by the microcontroller unit is less than or equal to the third preset period, then the microcontroller unit determines that the third heartbeat packet received this time has not timed out. If the time interval between the reception time of the third heartbeat packet received by the microcontroller unit this time and the reception time of the third heartbeat packet previously received by the microcontroller unit is greater than the third preset period, then the microcontroller unit determines that the third heartbeat packet received this time has timed out.

[0118] Furthermore, each time the microcontroller unit determines that a third heartbeat packet has timed out, it increments the count by one, thus allowing the microcontroller unit to track the number of timeouts of the received third heartbeat packets. When the number of timeouts of the received third heartbeat packets reaches K, the microcontroller unit can determine that the third heartbeat monitoring result is abnormal.

[0119] It should be understood that if the number of times the third heartbeat packet received by the microcontroller unit times out has not reached K, and if the microcontroller unit receives another third heartbeat packet without it times out, the microcontroller unit can reset the number of times the third heartbeat packet times out to zero.

[0120] Accordingly, the microcontroller unit determines whether a fourth heartbeat packet is received every four preset periods to determine if the fourth heartbeat packet has timed out. If the time interval between the reception time of the fourth heartbeat packet received by the microcontroller unit this time and the reception time of the fourth heartbeat packet previously received by the microcontroller unit is less than or equal to the four preset periods, then the microcontroller unit determines that the fourth heartbeat packet received this time has not timed out. If the time interval between the reception time of the fourth heartbeat packet received by the microcontroller unit this time and the reception time of the fourth heartbeat packet previously received by the microcontroller unit is greater than the four preset periods, then the microcontroller unit determines that the fourth heartbeat packet received this time has timed out.

[0121] Furthermore, each time the microcontroller unit determines that the fourth heartbeat packet has timed out, it increments the count by one, thus allowing the microcontroller unit to track the number of timeouts of the received fourth heartbeat packets. If the number of timeouts of the fourth heartbeat packets received by the microcontroller unit reaches X, then the microcontroller unit can determine that the fourth heartbeat monitoring result is abnormal.

[0122] It should be understood that if the number of times the fourth heartbeat packet received by the microcontroller unit times out has not reached X, and the microcontroller unit receives the fourth heartbeat packet again without timeout, the microcontroller unit can reset the number of times the fourth heartbeat packet times out to zero.

[0123] For example, Figure 6 This is a schematic diagram illustrating the specific process of the fault recovery method provided in the embodiments of this application. For example... Figure 6 As shown, the fault recovery method may specifically include the following steps S601 to S618: S601: After the main instrument system, main entertainment system and TBox system have finished booting, the main instrument system starts a timer.

[0124] S602, when the timer reaches the first preset duration, the host instrument system activates the heartbeat monitoring strategy.

[0125] S603 performs heartbeat monitoring based on a heartbeat monitoring strategy.

[0126] Specifically, the heartbeat monitoring strategy is used to: perform heartbeat monitoring between the host entertainment system and the host instrument system (i.e., steps S507 and S508 above), perform heartbeat monitoring between the TBox system and the host instrument system (i.e., steps S509 and S510 above), perform heartbeat monitoring between the TBox system and the microcontroller unit (i.e., steps S511 and S512 above), and perform heartbeat monitoring between the host instrument system and the microcontroller unit (i.e., steps S513 and S514 above).

[0127] It should be noted that after performing step S603 above, steps S604, S608, S610 and S615 below are performed simultaneously.

[0128] S604, the host instrument system determines whether the number of times the first heartbeat packet has timed out has reached M times.

[0129] If the number of times the first heartbeat packet times out reaches M, execute step S605 below; if the number of times the first heartbeat packet times out does not reach M, execute step S603 above to continue heartbeat monitoring.

[0130] S605, if the number of times the first heartbeat packet times out reaches M, the host instrument system determines the fault type as a fault in the host entertainment system.

[0131] S606, when the fault type is a fault in the main entertainment system, the main instrument system restarts the main entertainment system.

[0132] S607 records the recovery log of the host entertainment system.

[0133] In one example, the host instrumentation system can record the recovery log of the host entertainment system. The recovery log of the host entertainment system refers to the relevant information when the host instrumentation system restarts the host entertainment system in the event of a failure, such as the restart time.

[0134] S608, the host instrument system determines whether the number of times the second heartbeat packet has timed out has reached N.

[0135] If the number of times the second heartbeat packet times out reaches N, execute step S609 below; if the number of times the second heartbeat packet times out does not reach N, execute step S603 above to continue heartbeat monitoring.

[0136] S609, if the number of times the second heartbeat packet times out reaches N, the host instrument system determines the fault type as a fault in the TBox system.

[0137] S610, the microcontroller unit determines whether the number of times the third heartbeat packet has timed out has reached K.

[0138] If the number of times the third heartbeat packet times out reaches K, execute step S611 below; if the number of times the third heartbeat packet times out does not reach K, execute step S603 above to continue heartbeat monitoring.

[0139] S611, if the number of times the third heartbeat packet times out reaches K, the microcontroller unit determines the fault type as a fault in the TBox system.

[0140] S612, when the microcontroller unit determines that the fault type is a fault in the TBox system, the microcontroller unit sends a restart command to the host instrumentation system. This restart command is used to instruct the host instrumentation system to restart the TBox system.

[0141] S613, the main instrument system restarts the TBox system.

[0142] Thus, in S609 above, when the host instrument system determines that the fault type is a fault in the TBox system, step S613 can be executed directly to restart the TBox system via the host instrument system. In S611 above, when the microcontroller unit determines that the fault type is a fault in the TBox system, the microcontroller unit can send a restart command to the host instrument system, causing the host instrument system to restart the TBox system based on the restart command.

[0143] S614 records the recovery log of the TBox system.

[0144] In one example, the host instrumentation system can record the recovery log of the TBox system. The recovery log of the TBox system refers to the relevant information when the host instrumentation system restarts the TBox system in the event of a failure, such as the restart time.

[0145] S615, the microcontroller unit determines whether the fourth heartbeat packet has timed out X times.

[0146] If the fourth heartbeat packet times out X times, execute step S616 below; if the fourth heartbeat packet times out X times, execute step S603 above to continue heartbeat monitoring.

[0147] S616, if the number of times the fourth heartbeat packet times out reaches X, the microcontroller unit determines the fault type as a fault in the host instrument system.

[0148] S617, the microcontroller unit restarts the system-on-a-chip.

[0149] Specifically, after the microcontroller unit determines that the fault type is a fault in the host instrument system, the microcontroller unit powers down the system-on-a-chip and then powers it back on, thereby restarting the system-on-a-chip.

[0150] S618 records the recovery log of the system-on-a-chip.

[0151] In one example, the microcontroller unit can record the system-on-a-chip (SoC) recovery log. The SoC recovery log refers to relevant information such as the restart time when the microcontroller unit restarts the SoC in the event of a failure in the host instrumentation system.

[0152] In summary, step S607 can record the recovery log of the host entertainment system, step S614 can record the recovery log of the TBox system, and step S618 can record the recovery log of the system-on-a-chip. Therefore, in one possible implementation, after the above "rebooting the target module based on the fault type", the following step may also be included: recording the recovery log of the target module, where the target module includes any one of the host entertainment system, the TBox system, and the system-on-a-chip.

[0153] Thus, since the embodiments of this application adopt a layered recovery strategy to restart one of the host entertainment system, TBox system and system-on-a-chip for different fault types, and record the recovery logs of the host entertainment system, TBox system and system-on-a-chip for subsequent analysis, it is convenient for fault location and problem investigation.

[0154] In addition, after the above-mentioned S607 "Record the recovery log of the host entertainment system", the above-mentioned S614 "Record the recovery log of the TBox system", and the above-mentioned S618 "Record the recovery log of the system-on-a-chip", the above-mentioned S603 step can be executed again to continue heartbeat monitoring.

[0155] In some embodiments, the first preset period, the second preset period, the third preset period, and the fourth preset period are all equal, and M, N, K, and X are all equal.

[0156] For example, the first preset period, the second preset period, the third preset period, and the fourth preset period can all be 5 seconds, and M, N, K, and X can all be 12, which enables the embodiments of this application to achieve a 1-minute (the standard duration for judging the timeout of a heartbeat packet is 5 seconds, and the number of timeouts used in fault judgment is 12) timeout judgment.

[0157] By setting the first, second, third, and fourth preset periods to be equal, and setting M, N, K, and X to be equal, the heartbeat monitoring strategy can be simplified and unified. Standardized monitoring parameters reduce the complexity of operation and maintenance, and facilitate system performance tuning and problem diagnosis, thereby improving the maintainability of the system.

[0158] Furthermore, by reasonably setting the specific values ​​of the first, second, third, and fourth preset cycles, as well as the specific values ​​of M, N, K, and X, the timeliness of fault detection and the suppression of false alarms can be balanced.

[0159] Of course, in other embodiments, different heartbeat cycles can be set for different virtual systems. For example, at least two of the first, second, third, and fourth preset cycles mentioned above may not be equal. Specifically, different heartbeat cycles can be set according to the importance of different virtual systems. For example, virtual systems with higher importance have shorter heartbeat cycles, and virtual systems with lower importance have longer heartbeat cycles.

[0160] In summary, combining Figures 4 to 6 This document details the startup process of each virtual system, the heartbeat monitoring process between the microcontroller unit, the host instrumentation system, the host entertainment system, and the TBox system, as well as the specific procedures for fault recovery. The following section provides a detailed explanation of these procedures in conjunction with... Figure 7 In the system architecture after the TBox and host are integrated in the vehicle system, the interaction process between various virtual systems and between virtual systems and microcontroller units is explained.

[0161] For example, Figure 7 This is a system architecture diagram provided in this application embodiment for vehicle fault recovery after the TBox and host are integrated in the vehicle system. Figure 7 As shown, the system architecture includes virtual systems such as the host instrument system, host entertainment system, and TBox system built on a system-on-a-chip, and the host instrument system, host entertainment system, and TBox system share the same microcontroller unit and the same system-on-a-chip.

[0162] like Figure 7 As shown, the shared system-on-a-chip and shared microcontroller unit are located on the hardware platform, while the host instrumentation system, host entertainment system and TBox system are located in the virtual system layer.

[0163] Specifically, the host entertainment system periodically sends a first heartbeat packet to the host instrument system at a first preset period, the TBox system periodically sends a second heartbeat packet to the host instrument system at a second preset period, the TBox system periodically sends a third heartbeat packet to the microcontroller unit at a third preset period, and the host instrument system periodically sends a fourth heartbeat packet to the microcontroller unit at a fourth preset period.

[0164] In this way, we can monitor whether the number of times the first heartbeat packet times out reaches M, the number of times the second heartbeat packet times out reaches N, the number of times the third heartbeat packet times out reaches K, and the number of times the fourth heartbeat packet times out reaches X, etc., to determine the fault type and adopt different recovery strategies for different fault types.

[0165] like Figure 7 As shown, when the fault type is a failure of the main entertainment system, the recovery strategy is to restart the main entertainment system; when the fault type is a failure of the TBox system, the recovery strategy is to restart the TBox system; and when the fault type is a failure of the main instrument system, the recovery strategy is to restart the system-on-a-chip.

[0166] Thus, this embodiment of the application establishes a heartbeat monitoring system centered on the host instrument system to uniformly manage the status of the host entertainment system and the TBox system. The host instrument system monitors the status of the host entertainment system and the TBox system through VSOCK, and the microcontroller unit monitors the hardware links of each virtual system through SPI. The comprehensiveness of status detection is ensured through dual-level monitoring (application layer + hardware layer).

[0167] For example, Figure 8 This is a swimlane diagram of the hierarchical monitoring strategy corresponding to the fault recovery method provided in the embodiments of this application. For example... Figure 8 As shown, the swimlane diagram of the hierarchical monitoring strategy can include a time control layer, a monitoring execution layer, a decision-making layer, and an execution layer.

[0168] Taking a first preset duration of 1 minute, and first, second, third, and fourth preset periods of 5 seconds each, with M, N, K, and X all being 12, as an example, in the time control layer, the heartbeat monitoring strategy is started with a 1-minute delay. That is, after the host instrument system receives the third notification message sent by the virtual machine monitor, the timer is started. The host instrument system only starts the heartbeat monitoring strategy when the timer reaches the first preset duration (e.g., 1 minute). Furthermore, in the time control layer, a 5-second heartbeat period is set, meaning that the sending period for the first, second, third, and fourth heartbeat packets is 5 seconds. In addition, in the time control layer, a 1-minute timeout judgment is set, meaning that if 12 consecutive heartbeat timeouts occur within a 5-second heartbeat period, a fault is determined to exist.

[0169] like Figure 8As shown, in the monitoring execution layer, heartbeat monitoring can be performed between the host entertainment system and the host instrument system to obtain the first heartbeat monitoring result; heartbeat monitoring can be performed between the TBox system and the host instrument system to obtain the second heartbeat monitoring result; heartbeat monitoring can be performed between the TBox system and the microcontroller unit to obtain the third heartbeat monitoring result; and heartbeat monitoring can be performed between the host instrument system and the microcontroller unit to obtain the fourth heartbeat monitoring result.

[0170] like Figure 8 As shown, in the decision-making layer, fault decisions can be made on the host entertainment system, TBox system, and host instrument system based on the first heartbeat monitoring results, the second heartbeat monitoring results, the third heartbeat monitoring results, and the fourth heartbeat monitoring results.

[0171] like Figure 8 As shown, in the execution layer, if the host entertainment system fails, the host entertainment system is restarted; if the TBox system fails, the TBox system is restarted; and if the host instrument system fails, the system-on-a-chip is restarted.

[0172] In summary, the embodiments of this application, through the design mechanism of master-slave collaborative heartbeat monitoring, hierarchical fault recovery, delayed startup monitoring, and unified heartbeat cycle, realize independent state monitoring and hierarchical fault handling of multiple virtual systems in the domain fusion architecture, ensuring system startup stability, operational reliability, and rapid recovery capability.

[0173] The above combination Figures 2 to 8 The fault recovery method provided in the embodiments of this application has been described. The apparatus for performing the above method provided in the embodiments of this application is described below.

[0174] Figure 9 This is a schematic diagram of a fault recovery device provided in an embodiment of this application. This fault recovery device can be applied to a vehicle, which includes multiple virtual systems built on a system-on-a-chip (SoC). These virtual systems include a main instrument cluster system, a main infotainment system, and a TBox system, and all virtual systems share the same microcontroller unit. Figure 9 As shown, the fault recovery device 900 may include: a first startup module 901, a heartbeat monitoring module 902, a fault determination module 903, and a restart module 904.

[0175] The system comprises the following modules: First startup module 901, which initiates a heartbeat monitoring strategy after the host instrumentation system, host entertainment system, and TBox system have started up; Heartbeat monitoring module 902, which performs heartbeat monitoring based on the strategy to obtain target monitoring results, including heartbeat monitoring results between multiple virtual systems and between virtual systems and microcontroller units; Fault determination module 903, which determines the fault type based on the target monitoring results; and Restart module 904, which restarts the target module based on the fault type, including any one of the host entertainment system, TBox system, and system-on-a-chip (SoC).

[0176] In one possible implementation, the heartbeat monitoring module 902 is specifically used to: perform heartbeat monitoring between the host entertainment system and the host instrument system to obtain a first heartbeat monitoring result; perform heartbeat monitoring between the TBox system and the host instrument system to obtain a second heartbeat monitoring result; perform heartbeat monitoring between the TBox system and the microcontroller unit to obtain a third heartbeat monitoring result; and perform heartbeat monitoring between the host instrument system and the microcontroller unit to obtain a fourth heartbeat monitoring result.

[0177] In one possible implementation, the fault determination module 903 is specifically used to: determine the fault type as a fault in the host entertainment system when the first heartbeat monitoring result is abnormal; determine the fault type as a fault in the TBox system when the second or third heartbeat monitoring result is abnormal; and determine the fault type as a fault in the host instrument system when the fourth heartbeat monitoring result is abnormal.

[0178] In one possible implementation, the fault recovery device 900 may further include an anomaly detection module, which is used to: determine that the first heartbeat monitoring result is abnormal when the host instrument system receives M consecutive first heartbeat packets that all time out, wherein the first heartbeat packets are periodically sent by the host entertainment system to the host instrument system using a first preset period, and M is a positive integer; determine that the second heartbeat monitoring result is abnormal when the host instrument system receives N consecutive second heartbeat packets that all time out, wherein the second heartbeat packets are periodically sent by the TBox system to the host instrument system using a second preset period, and N is a positive integer; determine that the third heartbeat monitoring result is abnormal when the microcontroller unit receives K consecutive third heartbeat packets that all time out, wherein the third heartbeat packets are periodically sent by the TBox system to the microcontroller unit using a third preset period, and K is a positive integer; and determine that the fourth heartbeat monitoring result is abnormal when the microcontroller unit receives X consecutive fourth heartbeat packets that all time out, wherein the fourth heartbeat packets are periodically sent by the host instrument system to the microcontroller unit using a fourth preset period, and X is a positive integer.

[0179] In one possible implementation, the first preset period, the second preset period, the third preset period, and the fourth preset period are all equal, and M, N, K, and X are all equal.

[0180] In one possible implementation, the restart module 904 is specifically used to: restart the host entertainment system when the fault type is a fault in the host entertainment system; restart the TBox system when the fault type is a fault in the TBox system; and restart the system-on-a-chip (SoC) in the microcontroller unit when the fault type is a fault in the host instrumentation system.

[0181] In one possible implementation, the first startup module 901 is specifically used to: after the host instrument system, host entertainment system and TBox system have finished starting, start a timer in the host instrument system; if the timer reaches a first preset duration, start a heartbeat monitoring strategy; if the timer does not reach the first preset duration, do not start the heartbeat monitoring strategy.

[0182] In one possible implementation, the fault recovery device 900 may further include a second startup module, which is configured to: send a first startup command to the host instrumentation system; after the host instrumentation system completes startup based on the first startup command, send a first startup request and a second startup request to the virtual machine monitor; the virtual machine monitor sends a second startup command to the host entertainment system based on the first startup request, and sends a third startup command to the TBox system based on the second startup request; after the host entertainment system completes startup based on the second startup command, send a first notification message to the virtual machine monitor, the first notification message indicating that the host entertainment system has completed startup; after the TBox system completes startup based on the third startup command, send a second notification message to the virtual machine monitor, the second notification message indicating that the TBox system has completed startup; and upon receiving the first and second notification messages, the virtual machine monitor sends a third notification message to the host instrumentation system, the third notification message indicating that the host instrumentation system, the host entertainment system, and the TBox system have completed startup.

[0183] In one possible implementation, the fault recovery device 900 may further include a recording module for recording the recovery log of the target module.

[0184] Figure 10 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application. For example, as shown... Figure 10As shown, the vehicle 1000 includes a memory 1001 and a processor 1002. The memory 1001 stores executable program code 10011, and the processor 1002 is used to call and execute the executable program code 10011 to perform a fault recovery method.

[0185] In some embodiments, the vehicle 1000 may include a controller, which includes a system-on-a-chip (SoC) and a microcontroller unit. The SoC may include a memory and a processor, and the microcontroller unit may also include a memory and a processor. The memory 1001 described above may actually include multiple memories, such as the memory in the SoC and the memory in the microcontroller unit. Similarly, the processor 1002 described above may actually include multiple processors, such as the processor in the SoC and the processor in the microcontroller unit, such that the processor in the SoC and the processor in the microcontroller unit each call their corresponding executable program code to execute the fault recovery method provided in this application embodiment.

[0186] Furthermore, embodiments of this application also protect an apparatus that may include a memory and a processor, wherein the memory stores executable program code, and the processor is used to call and execute the executable program code to perform a fault recovery method provided in embodiments of this application.

[0187] This embodiment can divide the device into functional modules based on the above method example. For example, each module can correspond to a separate function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0188] When each functional module is divided according to its corresponding function, the device may further include a first startup module, a heartbeat monitoring module, a fault determination module, a restart module, an anomaly judgment module, a second startup module, and a recording module. It should be noted that all relevant content regarding the steps involved in the above method embodiments can be referenced from the functional descriptions of the corresponding functional modules, and will not be repeated here.

[0189] It should be understood that the apparatus provided in this embodiment is used to perform the above-described fault recovery method, and therefore can achieve the same effect as the above-described implementation method.

[0190] When using an integrated unit, the device may include a processing module and a storage module. When the device is applied to a vehicle, the processing module can be used to control and manage the vehicle's movements. The storage module can be used to support the vehicle in executing relevant program code.

[0191] The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits shown in conjunction with the disclosure of this application. The processor may also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module may be a memory.

[0192] In addition, the device provided in the embodiments of this application may specifically be a chip, component or module. The chip may include a connected processor and a memory. The memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute a fault recovery method provided in the above embodiments.

[0193] This embodiment also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the aforementioned related method steps to implement a fault recovery method provided in the above embodiment.

[0194] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement a fault recovery method provided in the above embodiment.

[0195] In this embodiment, the device, computer-readable storage medium, computer program product, or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0196] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0197] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0198] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A fault recovery method, characterized in that, Applied to vehicles, the vehicle includes multiple virtual systems built on a system-on-a-chip, the multiple virtual systems including a host instrument cluster system, a host entertainment system, and a TBox system, and the multiple virtual systems share the same microcontroller unit, the method comprising: After the host instrument system, the host entertainment system, and the TBox system have finished starting up, the heartbeat monitoring strategy is activated. Based on the heartbeat monitoring strategy, heartbeat monitoring is performed to obtain target monitoring results, which include heartbeat monitoring results between the multiple virtual systems and heartbeat monitoring results between the virtual system and the microcontroller unit. The fault type is determined based on the target monitoring results; The target module is restarted based on the fault type, and the target module includes any one of the host entertainment system, the TBox system, and the system-on-a-chip.

2. The method according to claim 1, characterized in that, The process of monitoring heartbeats based on the aforementioned heartbeat monitoring strategy to obtain target monitoring results includes: Heartbeat monitoring is performed between the host entertainment system and the host instrument system to obtain the first heartbeat monitoring result; Heartbeat monitoring is performed between the TBox system and the host instrument system to obtain a second heartbeat monitoring result; Heartbeat monitoring is performed between the TBox system and the microcontroller unit to obtain a third heartbeat monitoring result; Heartbeat monitoring is performed between the host instrument system and the microcontroller unit to obtain a fourth heartbeat monitoring result.

3. The method according to claim 2, characterized in that, Determining the fault type based on the target monitoring results includes: If the first heartbeat monitoring result is abnormal, the fault type is determined to be a fault in the host entertainment system; If the second heartbeat monitoring result or the third heartbeat monitoring result is abnormal, the fault type is determined to be a fault in the TBox system; If the fourth heartbeat monitoring result is abnormal, the fault type is determined to be a fault in the host instrument system.

4. The method according to claim 3, characterized in that, The method further includes: If the host instrument system receives the first heartbeat packet M times consecutively and all of them time out, it is determined that the first heartbeat monitoring result is abnormal. The first heartbeat packet is periodically sent by the host entertainment system to the host instrument system using a first preset period, where M is a positive integer. If the host instrument system receives N consecutive second heartbeat packets that all time out, it is determined that the second heartbeat monitoring result is abnormal. The second heartbeat packet is periodically sent by the TBox system to the host instrument system using a second preset period, where N is a positive integer. If the third heartbeat packet received by the microcontroller unit times out for K consecutive times, the monitoring result of the third heartbeat is determined to be abnormal. The third heartbeat packet is periodically sent by the TBox system to the microcontroller unit using a third preset period, where K is a positive integer. If the fourth heartbeat packet received by the microcontroller unit times out X consecutive times, the monitoring result of the fourth heartbeat is determined to be abnormal. The fourth heartbeat packet is periodically sent by the host instrument system to the microcontroller unit using a fourth preset period, where X is a positive integer.

5. The method according to claim 4, characterized in that, The first preset period, the second preset period, the third preset period, and the fourth preset period are all equal, and M, N, K, and X are all equal.

6. The method according to claim 1, characterized in that, The restarting of the target module based on the fault type includes: In the event that the fault type is a fault in the host entertainment system, the host instrument system restarts the host entertainment system; In the event that the fault type is a fault in the TBox system, the host instrument system restarts the TBox system; In the event that the fault type is a fault in the host instrument system, the microcontroller unit restarts the system-on-a-chip.

7. The method according to any one of claims 1 to 6, characterized in that, After the host instrument system, the host entertainment system, and the TBox system have finished starting, the heartbeat monitoring strategy is activated, including: After the host instrument system, the host entertainment system, and the TBox system have finished starting up, the host instrument system starts a timer; When the timer reaches the first preset duration, the heartbeat monitoring strategy is activated; If the timer has not reached the first preset duration, the heartbeat monitoring strategy will not be activated.

8. The method according to any one of claims 1 to 6, characterized in that, The method further includes: The virtual machine monitor sends a first startup command to the host instrumentation system; After the host instrumentation system completes startup based on the first startup instruction, it sends a first startup request and a second startup request to the virtual machine monitor. The virtual machine monitor sends a second startup command to the host entertainment system based on the first startup request, and sends a third startup command to the TBox system based on the second startup request; After the host entertainment system completes startup based on the second startup instruction, it sends a first notification message to the virtual machine monitor. The first notification message is used to indicate that the host entertainment system has completed startup. After the TBox system completes startup based on the third startup instruction, a second notification message is sent to the virtual machine monitor. The second notification message is used to indicate that the TBox system has completed startup. Upon receiving the first notification message and the second notification message, the virtual machine monitor sends a third notification message to the host instrumentation system. The third notification message is used to indicate that the host instrumentation system, the host entertainment system, and the TBox system have completed startup.

9. The method according to claim 1, characterized in that, After restarting the target module based on the fault type, the method further includes: Record the recovery log of the target module.

10. A vehicle, characterized in that, The vehicles include: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the vehicle to perform the method as described in any one of claims 1 to 9.