System and method for protecting system security from guest virtual machine (GVM) induced global system memory management unit (SMMU) failures in automotive hosted hypervisor system

By identifying and resetting the flow identifier of the guest virtual machine in the vehicle-managed management program system, the global SMMU failure caused by GVM was resolved, ensuring system stability and security, avoiding system-level restarts, and maintaining the continuity of autonomous driving.

CN120917427APending Publication Date: 2025-11-07QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480021035.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-18
Filing Date
2024-03-19
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In a vehicle-managed management system, a Global System Memory Management Unit (SMMU) failure caused by a guest virtual machine (GVM) may lead to system safety issues, affecting the stability and safety of autonomous driving.

Method used

By using a managed hypervisor to identify the Flow Identifier (SID) associated with a global SMMU failure, and by resetting only the Guest Virtual Machine (GVM) when a failure is identified, system-wide reboots are avoided, thus maintaining system stability and functionality.

Benefits of technology

It effectively prevents system restarts caused by global SMMU failures triggered by GVM, ensuring the stable operation of the autonomous driving system under fault conditions and avoiding safety vulnerabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120917427A_ABST
    Figure CN120917427A_ABST
Patent Text Reader

Abstract

A system for global system memory management unit (SMMU) fault handling, the system comprising: a peripheral device having a guest virtual machine (GVM), the peripheral device configured to access a memory (DDR) through a system memory management unit (SMMU); a hosted hypervisor associated with the peripheral device, the GVM, the SMMU, and the memory (DDR), where, upon identification of a faulty memory transaction and a global SMMU fault issued, the hosted hypervisor is configured to identify that a flow identifier (SID) associated with the global SMMU fault is assigned to the GVM, and the hosted hypervisor is configured to reset only the GVM, and to identify that the flow identifier (SID) associated with the global SMMU fault is assigned to the GVM. And restarting of the complete system is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 492,596, filed March 28, 2023, entitled “SYSTEM AND METHOD FORPROTECTING SYSTEM SAFETY FROM GUEST VIRTUAL MACHINE (GVM) ORIGINATED GLOBALSYSTEM MEMORY MANAGEMENT UNIT (SMMU) FAULTS IN AUTOMOTIVE HOSTED HYPERVISORSYSTEMS,” the contents of which are hereby incorporated herein by reference in their entirety, as fully set forth below and used for all applicable purposes. Technical Field

[0003] This disclosure relates generally to electronic devices, and more specifically to guest virtual machine (GVM) fault tolerance in automotive systems. Background Technology

[0004] Driver assistance technologies in the automotive field are constantly expanding. For example, assisted driving and autonomous driving technologies rely on a variety of different sensors and processing functions to ensure safety. Ensuring safety is part of maintaining a high Automotive Safety Integrity Level (ASIL) score. In some automotive systems, the hypervisor and guest system can interact. The hypervisor adds a different software layer on top of the host operating system, and the guest operating system becomes a third software layer on top of the hardware. For example, the hypervisor (e.g., a Type 2 hypervisor) can interact with one or more peripheral devices, such as communication devices that can use WiFi, Bluetooth, USB, or other connectivity. In some cases, the peripheral device may have a software element that can be called a guest virtual machine (GVM). The guest virtual machine can interact with the hypervisor and memory (such as Double Data Rate (DDR) Synchronous Dynamic Random Access Memory (SDRAM)) in a system controlled by the hypervisor. The System Memory Management Unit (SMMU) can also control the GVM's access to DDR memory. However, it is expected that anomalies related to the guest virtual machine will not affect the entire system in which the SMMU operates. For example, any action of the GVM that could lead to a global SMMU failure could cause undesirable safety issues related to providing continuous autonomous driving support.

[0005] Accordingly, it is desirable to prevent guest virtual machine exceptions that can cause global SMMU faults from impacting system safety. SUMMARY

[0006] The details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following drawings can not be drawn to scale.

[0007] The details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following drawings can not be drawn to scale.

[0008] One aspect of the disclosure provides a system for global system memory management unit (SMMU) fault handling, the system comprising: a peripheral device having a guest virtual machine (GVM), the peripheral device configured to access a memory (DDR) through a system memory management unit (SMMU); a hosted hypervisor associated with the peripheral device, the GVM, the SMMU, and the memory (DDR), wherein upon identifying a fault memory transaction and a global SMMU fault being issued, the hosted hypervisor is configured to identify a stream identifier (SID) associated with the global SMMU fault being assigned to the GVM, and the hosted hypervisor is configured to reset only the GVM such that a full system reboot is avoided.

[0009] One aspect of the disclosure provides a system for global system memory management unit (SMMU) fault handling, the system comprising: a peripheral device having a guest virtual machine (GVM), the peripheral device configured to access a memory (DDR) through a system memory management unit (SMMU); a hosted hypervisor associated with the peripheral device, the GVM, the SMMU, and the memory (DDR), wherein upon identifying the GVM being terminated and a global SMMU fault being issued, the hosted hypervisor is configured to clear a level 1 SMMU translation register, and upon identifying a stream identifier (SID) associated with the global SMMU fault being assigned to the GVM that is rebooting, ignore the SMMU fault and allow the GVM to reboot.

[0010] Another aspect of the present disclosure provides a method for global system memory management unit (SMMU) fault handling, the method comprising: issuing a global system memory management unit (SMMU) fault; discovering that a stream ID (SID) associated with the global SMMU fault is assigned to a guest virtual machine (GVM); and rebooting only the GVM, thereby avoiding a complete system reboot and allowing continuous cluster functionality during the GVM reboot. BRIEF DESCRIPTION OF DRAWINGS

[0011] In the drawings, like reference numerals refer to like parts throughout the various views unless otherwise indicated. For reference numerals with following characters, the character following the reference numeral is a letter designation and is used to distinguish between two similar items for which the reference numeral is common. The letter designation both precedes the reference numeral and follows the reference numeral in the description of each drawing. For example, reference designators 102a, 102b refer to the same component having a first instance 102a and a second instance 102b.

[0012] Figure 1 is a diagram illustrating components of an automotive autonomous driving system.

[0013] Figure 2 is a block diagram illustrating a processing system.

[0014] Figure 3 is a block diagram illustrating components of a processing system of Figure 2

[0015] Figure 4 is a diagram illustrating a computer system and a related execution environment.

[0016] Figure 5 is a diagram illustrating a computer system and a related execution environment.

[0017] Figure 6 is a diagram illustrating a system memory management unit (SMMU).

[0018] Figure 7 is a diagram illustrating an example of a computing system.

[0019] Figure 8 is a call flow diagram in accordance with an example embodiment of the present disclosure.

[0020] Figure 9 is a call flow diagram in accordance with an example embodiment of the present disclosure.

[0021] Figure 10 is a flow diagram describing an example of operations for a method of global SMMU fault handling.

[0022] Figure 11 is a functional block diagram of an apparatus for global SMMU fault handling. ​

[0023] Figure 12 is a flowchart of an example of operations describing a method for global SMMU fault handling.

[0024] Figure 13 is a functional block diagram of an apparatus for global SMMU fault handling. DETAILED DESCRIPTION

[0025] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.

[0026] According to exemplary embodiments, a system and method for preventing guest virtual machine exceptions from causing system memory management unit (SMMU) global faults and system resets in a hosted hypervisor system are disclosed.

[0027] In exemplary embodiments, a vehicle such as an automobile has a cluster that can include various gauges and displays (such as, for example, displays, indicators, gauges, fault indicators, system warning indicators, etc.) that allow a driver to safely operate the vehicle. For safety reasons, it is desirable to allow the automobile cluster to maintain functionality in the event of a system fault such as, for example, a global SMMU fault.

[0028] Figure 1 is a diagram 100 illustrating components of an automobile autonomous driving system. The automobile autonomous driving system can include a processing module 110 and a drive-by-wire (DBW) system controller 136. The processing module 110 can include one or more object detection elements 112 and one or more camera perception elements 114. For example, the object detection elements 112 can receive input from one or more sensors 113; and the camera perception elements 114 can receive input from one or more cameras 117.

[0029] In exemplary embodiments, the processing module 110 can also include a localization engine 118, a map fusion and arbitration element 122, and a route planning element 124. In exemplary embodiments, the localization engine 118 can receive input from the cameras 117 and localization input 123. The localization input 123 can be, for example, global positioning system (GPS) data, inertial measurement unit (IMU) data, controller area network (CAN) data, etc. For example, the map fusion and arbitration element 122 and the route planning element 124 can receive map input from a high-definition map element 127.

[0030] In an example embodiment, processing module 110 can also include a sensor fusion and road world model (RWM) management element 130, a motion planning and control element 132, and a behavior planning and prediction element 134. In an example embodiment, sensor fusion and road world model (RWM) management element 130 can receive inputs from object detection element 112, camera perception element 114, map fusion and arbitration element 122, and route planning element 124 to develop a road world model. In an example embodiment, the road world model can be an intelligent world model for an autonomous self-driving car.

[0031] In an example embodiment, sensor fusion and road world model (RWM) management element 130 can provide outputs to motion planning and control element 132 and behavior planning and prediction element 134. Behavior planning and prediction element 134 can also provide outputs to motion planning and control element 132. The outputs of processing module 110 can be provided to a drive-by-wire (DBW) system controller 136, which can provide autonomous driving instructions to car 140.

[0032] Figure 2 is a block diagram 200 illustrating a processing system. The processing system can include a processing element 202, a system clock 204, and a voltage regulator 206.

[0033] In an example embodiment, processing element 202 can include a camera 212, an image and object recognition processor 214, a mobile display processor (MDP) 216, an application processor 218, and a co-processor 222.

[0034] In an example embodiment, processing element 202 can include a digital signal processor (DSP) 226, a modem processor 228, a memory 232, analog and custom circuitry 234, system components and resources 236, and a resource and power management (RPM) processor 238. Each of the elements of processing element 202, except for co-processor 222, can be connected to an interconnect bus 224. Co-processor 222 can be connected to application processor 218.

[0035] In an example embodiment, camera 212, image and object recognition processor 214, and mobile display processor (MDP) 216 can cooperate to provide a visual display to an operator of car 140 Figure 1 .

[0036] In an example embodiment, application processor 218 and co-processor 222 can divide processing tasks of processing element 202.

[0037] In an example embodiment, DSP 226 can perform processing on digital signals, and model processor can control communications of processing element 202. Memory 232 can be static memory, dynamic memory, and can be a combination of persistent and non-persistent memory. Although shown as a single element, memory 232 can also be distributed memory.

[0038] In an example embodiment, analog and custom circuits 234 can provide analog signal processing, system components and resources 236 can provide various signal processing and signal conditioning circuits, including, for example, voltage regulators, oscillators, phase-locked loops, peripheral memory controllers, memory controllers, system controllers, access ports, timers, and other components to support processors and software clients, and resource and power management (RPN) processor 238 controls resource power management.

[0039] Clock 204 can provide a system clock to processing element 202, and voltage regulator 206 can provide a regulated system voltage to processing element 202.

[0040] Figure 3 is a block diagram 300 showing components of a processing system of Figure 2 In an example embodiment, application program / process element 310 can include aspects of application processor 218 of Figure 2

[0041] In an example embodiment, application program / process element 310 can include software portion 312 and hardware portion 314. In an example embodiment, software portion 312 can include applications 320, application programming interfaces (APIs) 322, application binary interfaces (ABIs) 324, libraries 326, and operating system (OS) 328. In an example embodiment, APIs 322 define source interfaces and are expressed in source code. ABIs 324 define low-level binary interfaces between two or more software on a particular architecture and are expressed in compiled code, not source code.

[0042] In an example embodiment, hardware portion 314 can include peripherals 342, system memory management unit (SMMU) 344, CPU 346, CPU MMU 347 (CPU for memory management unit), and memory 348. In an example embodiment, memory 348 can include double data rate (DDR) synchronous dynamic random access memory (SDRAM), commonly referred to as DDR. In an example embodiment, memory 348 can be shared among multiple processes, such as multiple instances of CPU 346, and shared by multiple peripherals 342.

[0043] ​An industry standard architecture (ISA) 332 can define the boundary between the software portion 312 and the hardware portion 314.

[0044] Figure 4 is a diagram 400 illustrating a computer system and related execution environments. In an example embodiment, the computer system 410 can include a guest module 412 with an application / process module 413, and a runtime software module 414 that includes a portion of a virtualization module 415. A host module 416 includes an operating system 417, a portion of the virtualization module 415, and hardware 418. In an example embodiment, the virtualization module 415 allows hardware resources to be partitioned into multiple virtual machines.

[0045] In an example embodiment, a guest computer system 430 can include an application / process module 432 that is also part of the guest module 412, and a guest virtual machine (GVM) 434. In an example embodiment, the GVM is a software component of a virtual machine (VM).

[0046] Figure 5 is a diagram 500 illustrating a computer system and related execution environments. In an example embodiment, the computer system 510 can include a guest module 512 with an application / process module 513, and an operating system 517. In an example embodiment, the computer system 510 also includes a hosted hypervisor 520 embodied by a virtualization module 525. A host module 526 can include hardware 528 and can run a host operating system (OS) (not shown). In some embodiments, the hosted hypervisor 520 and the host module 526 can be combined in the same module. In an example embodiment, the hosted hypervisor 520 will run the host OS and will perform tasks associated with the hypervisor.

[0047] In an example embodiment, a guest computer system 530 can include an application / process module 532 that is also part of the guest module 512. A guest operating system module 535 can also be part of the guest module 512. A guest virtual machine (GVM) 534 can be part of the hosted hypervisor 520 and the host module 526.

[0048] Figure 6is a diagram 600 illustrating a system memory management unit (SMMU). In example embodiments, the SMMU 610 can be connected to one or more devices. For example, peripheral devices 602, 604, 606, and 608 can be connected to the SMMU 610. The peripheral devices 602, 604, 606, and 608 can be peripheral devices that can be connected to and operate on a computing system to which the SMMU 610 is also connected. Examples of peripheral devices include, but are not limited to, devices that can be connected to the SMMU 610 through a universal serial bus (USB) connection, a WiFi device, a Bluetooth device, or other devices.

[0049] In example embodiments, communication from the peripheral devices to the SMMU 610 occurs through communication streams, such as, for example, a communication stream 603 (or multiple communication streams 603a and 603b) that can correspond to the peripheral device 602, a communication stream 605 that can correspond to the peripheral device 604, a communication stream 607 that can correspond to the peripheral device 606, and a communication stream 609 that can correspond to the peripheral device 608. For example, each peripheral device includes one or more stream identifiers (stream IDs). For example, the peripheral device 602 has a stream ID 1 and a stream ID 2, the peripheral device 604 has a stream ID 3, the peripheral device 606 includes a stream ID 4, and the peripheral device 608 includes a stream ID 5. In some embodiments, a peripheral device can have multiple communication streams, where each communication stream can be assigned a unique stream ID, and each stream ID can be independently assigned to a hosted hypervisor 520 and a GVM. For example, the peripheral device 602 (or any other peripheral device) can be assigned to a hosted hypervisor and a GVM, such that the communication stream 603 can include unique communication streams 603a and 603b, for example, where the communication stream 603a can have a stream ID assigned to the hosted hypervisor 520 (e.g., stream ID 1), and the communication stream 603b can have a stream ID assigned to the GVM 534 (e.g., stream ID 2).

[0050] In example embodiments, the SMMU 610 includes a stream mapping table 611 that includes stream mapping registers (SMRs) and translation registers. For example, the SMMU 610 can include a stream mapping table 611, a level 1 translation register 612, a level 2 translation register 614, a level 1 translation register 616, a level 2 translation register 618, and an attribute translation register 622.

[0051] In example embodiments, the stream mapping table 611 can receive data streams from the peripheral devices 602, 604, 606, and 608, and can map the stream ID 1, the stream ID 2, the stream ID 3, and the stream ID 4 to the respective translation registers 612, 614, and 616, and the attribute translation register 622.

[0052] In an example embodiment, a guest virtual machine (GVM) 714 (not shown in FIG. 7) can be configured with a level 1 translation register (e.g., 712) in this example. In an example embodiment, the level 1 translation register 712 can translate a virtual address (VA) from the peripheral device 712 to a physical address (PA) in the system memory 724. Figure 6

[0053] In an example embodiment, a level 2 translation register 714 can translate an intermediate physical address (IPA) from the peripheral device 712 to a physical address (PA) in the system memory 724.

[0054] In an example embodiment, a level 1 translation register 716 can translate a virtual address (VA) from the peripheral device 712 to an intermediate physical address (IPA), and a level 2 translation register 718 can translate the intermediate physical address (IPA) from the level 1 translation register 716 to a physical address (PA) in the system memory 724.

[0055] In an example embodiment, an attribute translation register 722 can translate a physical address (PA) from the peripheral device 712 to a physical address (PA) in the system memory 724.

[0056] In an example embodiment, the outputs of the translation registers 712, 714, 718, and the attribute translation register 722 are provided to the system memory 724. In an example embodiment, the system memory 724 can include a DDR SDRAM. The system memory 726 also includes a translation table 726. The translation table 726 contains data structures for translating one address to another address, such as translating a VA to an IPA or translating an IPA to a PA. When accessing data in memory, the system looks up the physical memory address that matches the virtual address.

[0057] Figure 7 is a diagram 700 showing an example of a computing system 710. In an example embodiment, the computing system 710 can include a peripheral device 712, a GVM 714, a hosted hypervisor 716, an SMMU 718, and a physical memory 724. In an example embodiment, the physical memory 724 can be an example of a DDR SDRAM. In an example embodiment, the peripheral device 712 can be an example of the peripheral device 602, 604, 606, or 608 of Figure 6 ; the GVM 714 can be an example of the GVM 534 (or the GVM shown in FIG. 5) of Figure 5 ; the hosted hypervisor 716 can be an example of the hosted hypervisor 520 of Figure 6 ; and the SMMU 718 can be an example of the SMMU 518 of Figure 5 Figure 6 ​​of SMMU 610; and physical memory 724 can be an example of DDR SRAM 624. Figure 6

[0058] Figure 8 is a call flow diagram 800 according to example embodiments of the present disclosure. A variety of different scenarios can cause a global SMMU fault. For example, an unexpected GVM restart during an ongoing direct memory access (DMA) operation can cause a global SMMU fault. For example, when a GVM restarts, the hosting hypervisor can clear the SMMU level 1 translation registers assigned to that GVM. If a peripheral device assigned to the GVM is in the middle of initiating a DMA transaction during which the GVM is suddenly disabled, a global SMMU fault can result because the SMMU will recognize an unidentified communication stream. In some instances, such global SMMU faults can be avoided if the hypervisor powers down the peripheral device during the GVM restart, but there are instances in which the peripheral device cannot be turned off during a system restart and global SMMU fault (e.g., if the peripheral device is shared between the hypervisor and a GVM or between two GVMs as a virtual device). In some embodiments, a peripheral device can have multiple communication streams, where each stream can be assigned a unique stream ID, and each stream ID can be independently assigned to a hypervisor and a GVM, as mentioned above.

[0059] As another example, a compromised GVM can invoke a global SMMU fault to cause a system-level restart.

[0060] As another example, an improper transaction from a GVM peripheral device that is not properly configured can cause an unidentified stream fault. This unidentified stream fault can cause a global SMMU fault, where, for example, a peripheral device stream ID (SID) is not configured in any stream mapping register (SMR) in the SMMU (e.g., in stream mapping table 611 in FIG. 6B). Figure 6 If a stream ID (SID) of a peripheral device is configured in more than one SMR, this can also cause a stream match conflict. Similarly, if a SID of a peripheral device is configured in one SMR register, but the type field of that SMR register is set to a fault context, a stream match conflict can result.

[0061] Accordingly, it is desirable to have the ability to avoid a complete system restart resulting from a global SMMU fault initiated by a GVM.

[0062] ​The example embodiment described in call flow diagram 800 describes global SMMU fault handling such that a GVM reset does not automatically result in a full system level reboot. For example, a system level reboot resulting from a GVM initiated SMMU global fault can cause a significant security vulnerability because a system level reset can be a serious safety issue for an automotive autonomous driving system. An autonomous driving system must be fully operational for the duration of an autonomous driving event. Therefore, even if a global SMMU fault is detected, it is desirable to maintain the stability and functionality of the autonomous driving system.

[0063] In an example embodiment, call flow diagram 800 illustrates peripheral device 812, GVM 814, hosted hypervisor 816, SMMU 818, and physical memory 824 in operable communication. In an example embodiment, peripheral device 812 can be an example of peripheral device 712, GVM 814 can be an example of GVM 714, hosted hypervisor 816 can be an example of hosted hypervisor 716, SMMU 818 can be an example of SMMU 718, and memory 824 can be an example of physical memory 724. Figure 7 Figure 7 Figure 7 Figure 7 Figure 7

[0064] In call 822, in an example embodiment, GVM 814 incorrectly updates an associated level 1 register of SMMU 818.

[0065] In call 825, hosted hypervisor 816 emulates a write to GVM 814 to SMMU 818.

[0066] In call 826, GVM 814 requests to initiate a transaction with peripheral device 812. Initiating a transaction refers to reading or writing from DDR (physical memory) via direct memory access (DMA) mode. The transaction goes through SMMU 818 and then to physical memory 824.

[0067] In call 828, peripheral device 812 initiates a transaction to access physical memory 824 via SMMU 818.

[0068] In call 832, SMMU 818 identifies a faulted transaction. The faulted transaction can be a result of the incorrectly updated level 1 SMMU register in call 822.

[0069] In call 834, SMMU 818 initiates a global fault. For example, SMMU 818 can initiate a global fault as a result of the faulted transaction in call 832.

[0070] ​​​​​In call 836, the hosted hypervisor 816 handles the global fault. For example, in call 838, the hosted hypervisor 816 discovers that the stream ID (SID) associated with the global SMMU fault is assigned to the GVM 814.

[0071] In call 842, only the GVM is reset, resulting in continuous (e.g., lossless) cluster (system) functionality. For example, the hosted hypervisor 816 maintains a database of SIDs assigned to both the GVM 814 and the hosted hypervisor 816. During handling of the global SMMU fault, the hosted hypervisor 816 will search for the SID that caused the global fault, and then if the SID is assigned to the GVM 814, will restart the GVM, resulting in only the GVM 814 being restarted and avoiding a system-level restart.

[0072] If the global fault is caused by a SID assigned to the host or the SID is not found in the database, the hosted hypervisor 816 continues the default behavior of a system-level reset, as the global SMMU fault was not triggered by the GVM 814.

[0073] Figure 9 Call flow diagram 900 is in accordance with an example embodiment of the present disclosure.

[0074] The example embodiment described in call flow diagram 900 describes global SMMU fault handling during a direct memory access (DMA) transaction during a GVM sudden termination.

[0075] In an example embodiment, call flow diagram 900 shows a peripheral device 912, a GVM 914, a hosted hypervisor 916, an SMMU 918, and a physical memory 924 in operable communication. In an example embodiment, the peripheral device 912 can be an example of the peripheral device 712, Figure 7 the GVM 914 can be an example of the GVM 714, Figure 7 the hosted hypervisor 916 can be an example of the hosted hypervisor 716, Figure 7 the SMMU 918 can be an example of the SMMU 718, and Figure 7 the memory 924 can be an example of the physical memory 724. Figure 7

[0076] In call 922, in an example embodiment, the GVM 914 correctly updates the associated level 1 registers of the SMMU 918.

[0077] In call 925, the hosted hypervisor 916 emulates a write to the GVM 914 to the SMMU 918. ​

[0078] In call 926, GVM 914 requests to initiate a transaction with peripheral device 912. Initiating a transaction refers to reading or writing from physical memory 924 (DDR) via direct memory access (DMA) mode. The transaction goes through SMMU 918 and then to physical memory 924.

[0079] In block 928, peripheral device 912 initiates a transaction to access physical memory 924 via SMMU 918.

[0080] In call 932, SMMU 918 validates the transaction. For example, SMMU 918 identifies that the transaction initiated by peripheral device 912 is an approved transaction.

[0081] In call 934, SMMU 918 allows the transaction to pass through.

[0082] In block 936, GVM 914 is abruptly terminated. For example, GVM 914 can be killed or can have encountered a fatal failure that abruptly halted its operation.

[0083] In call 938, as part of the recovery after GVM 914 is abruptly terminated, hypervisor 916 clears the level 1 translation registers assigned to GVM 914.

[0084] In call 942, hypervisor 916 handles the SMMU failure.

[0085] In call 944, hypervisor 916 determines that the SID associated with the SMMU failure is assigned to GVM 914 that is being restarted.

[0086] In call 946, hypervisor 916 ignores the global SMMU failure and allows GVM 914 to restart.

[0087] In block 948, normal processing flow continues without a system restart. In this way, a global SMMU failure caused by an abrupt termination of GVM 914 does not result in a complete system restart.

[0088] Figure 10 FIG. 1000 is a flowchart 1000 that describes an example of the operations of a method for global SMMU failure handling. Blocks in the method 1000 can or can not be performed in the order shown, and in some embodiments, can be performed at least in part in parallel.

[0089] In block 1002, in an example embodiment, a GVM updates an associated level 1 register of an SMMU. In an example embodiment, GVM 814 can have incorrectly updated an associated level 1 translation register 602 of SMMU 610, 718, 818.

[0090] In block 1004, the hosted hypervisor emulates writes to the GVM to the SMMU. For example, the hosted hypervisor 816 emulates writes to the GVM 814 to the SMMU 818.

[0091] In block 1006, the GVM requests to initiate a transaction with a peripheral device. For example, the GVM 814 requests the peripheral device 812 to initiate a memory transaction.

[0092] In block 1008, the peripheral device initiates a transaction to access physical memory via the SMMU. For example, the peripheral device 812 initiates a transaction to read from or write to physical (DDR) memory 824 via the SMMU 818.

[0093] In block 1012, the SMMU identifies a faulty transaction. The faulty transaction can be a result of the incorrect update of the level 1 SMMU registers in the call 822. For example, the SMMU 818 identifies a faulty transaction. The faulty transaction can be a result of the incorrect update of the level 1 SMMU registers in block 1002.

[0094] In block 1014, the SMMU raises a global fault. For example, the SMMU 818 can raise a global fault due to the faulty transaction in block 1012.

[0095] In block 1016, the hosted hypervisor handles the global fault. For example, the hosted hypervisor 818 can handle the global fault raised in block 1014.

[0096] In block 1018, the hosted hypervisor discovers that the stream ID (SID) associated with the global SMMU fault is assigned to a GVM. For example, the hosted hypervisor 816 discovers that the stream ID (SID) associated with the global SMMU fault is assigned to the GVM 814.

[0097] In block 1022, the GVM is reset only, resulting in lossless cluster (system) functionality. For example, the hosted hypervisor 816 maintains a database of SIDs assigned to both the GVM 814 and the hosted hypervisor 816. During the handling of the global SMMU fault identified in block 1014, the hosted hypervisor 816 will search for the SID that caused the global SMMU fault and then, if the SID is assigned to the GVM 814, will only restart the GVM, avoiding a system level restart.

[0098] If the global fault is caused by a SID assigned to the host or the SID is not found in the database, the hosted hypervisor continues the default behavior of a system level reset, as the global fault was not triggered by a GVM.

[0099] Figure 11 is a functional block diagram of an apparatus 1100 for global SMMU fault handling. The apparatus 1100 includes means 1102 for updating a level-1 SMMU register. In certain embodiments, the means 1102 for updating a level-1 SMMU register can be configured to perform one or more of the functions described in operation block 1002 of method 1000 Figure 10 In an example embodiment, the means 1102 for updating a level-1 SMMU register can include a GVM 814 that improperly updates an associated level-1 translation register 602 of an SMMU 610, 718, 818.

[0100] The apparatus 1100 can also include means 1104 for emulating a write to a GVM to an SMMU. In certain embodiments, the means 1104 for emulating a write to a GVM to an SMMU can be configured to perform one or more of the functions described in operation block 1004 of method 1000 Figure 10 In an example embodiment, the means 1104 for emulating a write to a GVM to an SMMU can include a hypervisor 816 that emulates a write to a GVM 814 to an SMMU 818.

[0101] The apparatus 1100 can also include means 1106 for requesting a peripheral device to initiate a transaction. In certain embodiments, the means 1106 for requesting a peripheral device to initiate a transaction can be configured to perform one or more of the functions described in operation block 1006 of method 1000 Figure 10 In an example embodiment, the means 1106 for requesting a peripheral device to initiate a transaction can include a GVM 814 that requests a peripheral device 812 to initiate a memory transaction.

[0102] The apparatus 1100 can also include means 1108 for initiating a transaction to access a physical memory via a system memory management unit (SMMU). In certain embodiments, the means 1108 for initiating a transaction to access a physical memory via a system memory management unit (SMMU) can be configured to perform one or more of the functions described in operation block 1008 of method 1000 Figure 10 In an example embodiment, the means 1108 for initiating a transaction to access a physical memory via a system memory management unit (SMMU) can include a peripheral device 812 that initiates a transaction to access a physical memory 824 via an SMMU 818.

[0103] The apparatus 1100 can also include means 1112 for identifying a fault transaction. In certain embodiments, the means 1112 for identifying a fault transaction can be configured to perform one or more of the functions described in operation block 1010 of method 1000Figure 10 This refers to one or more of the functions described in operation block 1012. In an exemplary embodiment, component 1112 for identifying faulty transactions may include SMMU 818 for identifying faulty transactions. A faulty transaction may be the result of GVM 814 incorrectly updating the Level 1 SMMU registers.

[0104] The apparatus 1100 may further include a component 1114 for initiating a global failure. In some embodiments, the component 1114 for initiating a global failure may be configured to perform method 1000. Figure 10 One or more of the functions described in operation block 1014. In an exemplary embodiment, component 1114 for triggering a global failure may include SMMU 818 for triggering a global failure due to a faulty transaction.

[0105] The apparatus 1100 may further include a component 1116 for handling global faults. In some embodiments, the component 1116 for handling global faults may be configured to perform method 1000. Figure 10 One or more of the functions described in operation block 1016. In an exemplary embodiment, component 1116 for handling global faults may include a managed administration program 818 for handling global faults.

[0106] The apparatus 1100 may further include a component 1118 for detecting SIDs associated with global SMMU faults that are assigned to the GVM. In some embodiments, the component 1118 for detecting SIDs associated with global SMMU faults that are assigned to the GVM may be configured to perform method 1000. Figure 10 The functions described in operation block 1018 are one or more of the functions described in the example implementation. In an exemplary implementation, component 1118 for discovering that a SID associated with a global SMMU failure has been assigned to a GVM may include a managed hypervisor 816 for discovering that a flow ID (SID) associated with a global SMMU failure has been assigned to a GVM 814.

[0107] Apparatus 1100 may also include component 1122 for resetting the GVM without losing cluster (system) functionality. In some embodiments, component 1122 for resetting the GVM without losing cluster functionality may be configured to perform method 1000. Figure 10 One or more of the functions described in operation block 1022. In an exemplary embodiment, component 1122 for resetting the GVM without losing cluster functionality may include a managed hypervisor 816 that searches for the SID causing a global SMMU failure and then, if the SID is assigned to a GVM 814, only restarts that GVM, thereby avoiding a system-wide reboot.

[0108] Figure 10 is a flow diagram 1200 that describes an example of operations for a global SMMU fault handling method. Blocks in the method 1200 can be executed in the order shown or can be executed in a different order, and in some embodiments, can be executed in parallel, at least in part.

[0109] In block 1202, in an example embodiment, a GVM updates an associated level 1 register of an SMMU. In an example embodiment, the GVM 814 can correctly update the associated level 1 translation register 602 of the SMMU 610, 718, 918.

[0110] In block 1204, a hosted hypervisor emulates a write to the GVM to the SMMU. For example, the hosted hypervisor 916 emulates a write to the GVM 914 to the SMMU 918.

[0111] In block 1206, the GVM requests to initiate a transaction with a peripheral device. For example, the GVM 914 requests the peripheral device 912 to initiate a memory transaction.

[0112] In block 1208, the peripheral device initiates a transaction to access physical memory via the SMMU. For example, the peripheral device 912 initiates a transaction to access physical memory 924 via the SMMU 918.

[0113] In block 1212, the SMMU validates the transaction. For example, the SMMU 918 validates the transaction initiated by the peripheral device 912.

[0114] In block 1214, the SMMU allows the transaction to pass. For example, the SMMU 918 can allow the transaction initiated by the peripheral device 912 to pass.

[0115] In block 1215, the GVM is abruptly terminated. For example, the operation of the GVM 914 can be terminated or the GVM 914 can have encountered a fatal fault that halted its operation.

[0116] In block 1216, the hosted hypervisor clears level 1 translation registers in the SMMU assigned to the GVM. For example, as part of recovery after the GVM 914 is abruptly terminated, the hosted hypervisor 916 clears the level 1 SMMU translation registers 602 assigned to the GVM 914.

[0117] In block 1218, the hosted hypervisor handles a global SMMU fault. For example, the hosted hypervisor 918 can handle a global SMMU fault raised by the SMMU 918 as a result of the abrupt termination of the GVM 914.

[0118] In block 1222, the hypervisor discovers that the stream ID (SID) associated with the global SMMU fault is assigned to the GVM that is rebooting. For example, the hypervisor 916 discovers that the stream ID (SID) associated with the global SMMU fault is assigned to the GVM 914 that is currently rebooting.

[0119] In block 1224, the hypervisor ignores the global SMMU fault and allows the GVM to reboot. For example, the hypervisor 916 ignores the SMMU fault and allows the GVM 914 to reboot, allowing normal processing flow to continue without a system reboot.

[0120] Figure 12 is a functional block diagram of an apparatus 1300 for global SMMU fault handling. The apparatus 1300 includes means 1302 for updating a 1st level SMMU register. In certain embodiments, the means 1302 for updating a 1st level SMMU register can be configured to perform one or more of the functions described in operation block 1202 of the method 1200 Figure 13 of the exemplary embodiments, the means 1302 for updating a 1st level SMMU register can include the GVM 914 that properly updates the associated 1st level translation registers 602 of the SMMU 610 / 918.

[0121] The apparatus 1300 can also include means 1304 for emulating writes to the SMMU for the GVM. In certain embodiments, the means 1304 for emulating writes to the SMMU for the GVM can be configured to perform one or more of the functions described in operation block 1204 of the method 1200 Figure 12 of the exemplary embodiments, the means 1304 for emulating writes to the SMMU for the GVM can include the hypervisor 916 emulating writes to the SMMU 918 for the GVM 914.

[0122] The apparatus 1300 can also include means 1306 for requesting the peripheral device to initiate a transaction. In certain embodiments, the means 1306 for requesting the peripheral device to initiate a transaction can be configured to perform one or more of the functions described in operation block 1206 of the method 1200 Figure 12 of the exemplary embodiments, the means 1306 for requesting the peripheral device to initiate a transaction can include the GVM 914 requesting the peripheral device 912 to initiate a memory transaction.

[0123] Apparatus 1300 can also include means 1308 for initiating a transaction to access physical memory via a system memory management unit (SMMU). In certain embodiments, means 1308 for initiating a transaction to access physical memory via a system memory management unit (SMMU) can be configured to perform one or more of the functions described in operation block 1208 of method 1200 Figure 12 In an example embodiment, means 1308 for initiating a transaction to access physical memory via a system memory management unit (SMMU) can comprise peripheral device 912 initiating a transaction to access physical memory 924 via SMMU 918.

[0124] Apparatus 1300 can also include means 1312 for validating a transaction. In certain embodiments, means 1312 for validating a transaction can be configured to perform one or more of the functions described in operation block 1212 of method 1200 Figure 12 In an example embodiment, means 1312 for validating a transaction can comprise SMMU 918 validating a transaction initiated by peripheral device 912.

[0125] Apparatus 1300 can also include means 1314 for allowing a transaction to pass. In certain embodiments, means 1314 for allowing a transaction to pass can be configured to perform one or more of the functions described in operation block 1214 of method 1200 Figure 12 In an example embodiment, means 1314 for allowing a transaction to pass can comprise SMMU 918 allowing a transaction initiated by peripheral device 912 to pass.

[0126] Apparatus 1300 can also include means 1316 for clearing level 1 SMMU translation registers assigned to a GVM that was abruptly terminated. In certain embodiments, means 1316 for clearing level 1 SMMU translation registers assigned to a GVM that was abruptly terminated can be configured to perform one or more of the functions described in operation block 1216 of method 1200 Figure 12 In an example embodiment, means 1316 for clearing level 1 SMMU translation registers assigned to a GVM that was abruptly terminated can comprise hosted hypervisor 916 clearing level 1 SMMU translation registers 602 assigned to GVM 914 as part of a recovery after GVM 914 was abruptly terminated.

[0127] Apparatus 1300 can also include means 1318 for handling SMMU faults. In certain embodiments, means 1318 for handling SMMU faults can be configured to perform one or more of the functions described in operation block 1220 of method 1200 Figure 12one or more of the functions described in operation block 1218 of the method 1200. In an example embodiment, the means for handling SMMU faults 1318 can include the hypervisor 916 handling a global SMMU fault raised by the SMMU 918 due to a sudden termination of the GVM 914.

[0128] The apparatus 1300 can also include means for discovering that a stream ID (SID) associated with a global SMMU fault is assigned to a GVM that is rebooting 1322. In certain embodiments, the means for discovering that a stream ID (SID) associated with a global SMMU fault is assigned to a GVM that is rebooting 1322 can be configured to perform one or more of the functions described in operation block 1222 of the method 1200 Figure 12 In an example embodiment, the means for discovering that a stream ID (SID) associated with a global SMMU fault is assigned to a GVM that is rebooting 1322 can include the hypervisor 916 discovering that a stream ID (SID) associated with a global SMMU fault is assigned to a GVM 914 that is currently rebooting.

[0129] The apparatus 1300 can also include means for ignoring the fault and allowing the GVM to reboot 1324. In certain embodiments, the means for ignoring the fault and allowing the GVM to reboot 1324 can be configured to perform one or more of the functions described in operation block 1224 of the method 1200 Figure 12 Figure 12 Figure 12 In an example embodiment, the means for ignoring the fault and allowing the GVM to reboot 1324 can include the hypervisor 916 ignoring the fault and allowing the GVM 914 to reboot so that normal processing flow continues without a system reboot.

[0130] In the following numbered clauses are described various implementation examples:

[0131] 1. A system for global system memory management unit (SMMU) fault handling, the system comprising: a peripheral device having a guest virtual machine (GVM), the peripheral device configured to access a memory (DDR) through a system memory management unit (SMMU); a hypervisor associated with the peripheral device, the GVM, the SMMU, and the memory (DDR), wherein upon identifying a faulty memory transaction and a global SMMU fault being raised, the hypervisor is configured to identify that a stream identifier (SID) associated with the global SMMU fault is assigned to the GVM; and the hypervisor is configured to reset only the GVM such that a complete system reboot is avoided.

[0132] 2. The system of clause 1, wherein the global SMMU fault is a result of the GVM incorrectly updating a level 1 SMMU translation register.

[0133] 3. The system of any of clauses 1 or 2, wherein avoiding a full system restart allows continuous cluster functionality during the GVM reset.

[0134] 4. The system of any of clauses 1-3, wherein the memory (DDR) is shared among a plurality of peripheral devices.

[0135] 5. The system of clause 4, wherein the plurality of peripheral devices each have a respective GVM.

[0136] 6. The system of any of clauses 1-5, wherein the peripheral device includes a plurality of communication streams.

[0137] 7. The system of clause 6, wherein the plurality of communication streams include a plurality of respective unique stream identifiers (SIDs) and correspond to a GVM and a hosted hypervisor.

[0138] 8. A system for global system memory management unit (SMMU) fault handling, the system comprising: a peripheral device having a guest virtual machine (GVM), the peripheral device configured to access a memory (DDR) through a system memory management unit (SMMU); and a hosted hypervisor associated with the peripheral device, the GVM, the SMMU, and the memory (DDR), wherein upon identifying that the GVM is terminated and a global SMMU fault is issued, the hosted hypervisor is configured to clear a level 1 SMMU translation register and ignore the SMMU fault and allow the GVM to restart upon identifying that a stream identifier (SID) associated with the global SMMU fault is assigned to the GVM that is restarting.

[0139] 9. The system of clause 8, wherein the global SMMU fault is a result of the GVM abruptly terminating.

[0140] 10. The system of any of clauses 8 or 9, wherein ignoring the SMMU fault and allowing the GVM to restart avoids a full system restart and allows continuous cluster functionality during the GVM restart.

[0141] 11. The system of any of clauses 8-10, wherein the memory (DDR) is shared among a plurality of peripheral devices.

[0142] 12. The system of clause 11, wherein the plurality of peripheral devices each have a respective GVM.

[0143] 13. The system of any of clauses 8-12, wherein the peripheral device comprises a plurality of communication streams.

[0144] 14. The system of clause 13, wherein the plurality of communication streams comprise a plurality of respective unique stream identifiers (SIDs) and correspond to GVMs and hosted hypervisors.

[0145] 15. A method for global system memory management unit (SMMU) fault handling, the method comprising: issuing a global system memory management unit (SMMU) fault; discovering that a stream ID (SID) associated with the global SMMU fault is assigned to a guest virtual machine (GVM); and restarting only the GVM, thereby avoiding a complete system restart and allowing continuous cluster functionality during the GVM restart.

[0146] 16. The method of clause 15, wherein the global SMMU fault is a result of a faulted memory transaction, and a hosted hypervisor connected to the SMMU and connected to the GVM identifies that a stream identifier (SID) associated with the global SMMU fault is assigned to the GVM.

[0147] 17. The method of any of clauses 15 or 16, wherein the SMMU fault is a result of the GVM being terminated, and the method further comprises: the hosted hypervisor clearing level 1 SMMU translation registers; and upon identifying that a stream identifier (SID) associated with the global SMMU fault is assigned to the GVM that is restarting, ignoring the SMMU fault and allowing the GVM to restart.

[0148] 18. The method of any of clauses 15-17, wherein the GVM is associated with a peripheral device, and a plurality of peripheral devices share double data rate (DDR) memory.

[0149] 19. The method of clause 18, wherein the plurality of peripheral devices each have a respective GVM.

[0150] 20. The method of any of clauses 15-19, wherein the peripheral device comprises a plurality of communication streams.

[0151] The circuit architectures described herein can be implemented on one or more ICs, analog ICs, RFICs, mixed-signal ICs, ASICs, printed circuit boards (PCBs), electronic devices, etc. The circuit architectures described herein can also be fabricated using various IC process technologies, such as complementary metal-oxide-semiconductor (CMOS), N-channel MOS (NMOS), P-channel MOS (PMOS), bipolar junction transistor (BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), heterojunction bipolar transistor (HBT), high electron mobility transistor (HEMT), silicon-on-insulator (SOI), etc.

[0152] The apparatuses implementing the circuits described herein can be standalone devices, or can be part of a larger device. The devices can be (i) standalone ICs, (ii) a collection of one or more ICs that can include a memory IC for storing data and / or instructions, (iii) an RFIC such as an RF receiver (RFR) or RF transmitter / receiver (RTR), (iv) an ASIC such as a mobile station modem (MSM), (v) a module that can be embedded within other devices, (vi) a receiver, cellular telephone, wireless device, handset, or mobile unit, (vii) etc.

[0153] Although selected aspects have been illustrated and described, it will be understood that various substitutions and alterations can be made by those skilled in the art, without departing from the spirit and scope of the present application, as defined by the following claims.

Claims

1. A system for global system memory management unit (SMMU) fault handling, the system comprising: a peripheral device having a guest virtual machine (GVM), the peripheral device configured to access memory (DDR) through a system memory management unit (SMMU); a hosted hypervisor associated with the peripheral device, the GVM, the SMMU, and the memory (DDR), wherein upon identifying a fault memory transaction and a global SMMU fault being issued, the hosted hypervisor is configured to identify that a stream identifier (SID) associated with the global SMMU fault is assigned to the GVM; and the hosted hypervisor is configured to reset only the GVM such that a full system reboot is avoided.

2. The system of claim 1, wherein the global SMMU fault is a result of the GVM incorrectly updating a level 1 SMMU translation register.

3. The system of claim 1, wherein avoiding a full system reboot allows continuous cluster functionality during the GVM reset.

4. The system of claim 1, wherein the memory (DDR) is shared among a plurality of peripheral devices.

5. The system of claim 4, wherein the plurality of peripheral devices each have a respective GVM.

6. The system of claim 1, wherein the peripheral device includes a plurality of communication streams.

7. The system of claim 6, wherein the plurality of communication streams include a plurality of respective unique stream identifiers (SIDs) and correspond to a GVM and a hosted hypervisor.

8. A system for global system memory management unit (SMMU) fault handling, the system comprising: a peripheral device having a guest virtual machine (GVM), the peripheral device configured to access memory (DDR) through a system memory management unit (SMMU); and a hosted hypervisor associated with the peripheral device, the GVM, the SMMU, and the memory (DDR), wherein upon identifying that the GVM is terminated and a global SMMU fault is issued, the hosted hypervisor is configured to clear a level 1 SMMU translation register and ignore the SMMU fault and allow the GVM to reboot upon identifying that a stream identifier (SID) associated with the global SMMU fault is assigned to the GVM that is rebooting.

9. The system of claim 8, wherein the global SMMU fault is a result of the GVM abruptly terminating.

10. The system of claim 9, wherein ignoring the SMMU fault and allowing the GVM to reboot avoids a full system reboot and allows continuous cluster functionality during the GVM reboot.

11. The system of claim 8, wherein the memory (DDR) is shared among a plurality of peripheral devices.

12. The system of claim 11, wherein the plurality of peripheral devices each have a respective GVM.

13. The system of claim 8, wherein the peripheral device includes a plurality of communication streams.

14. The system of claim 13, wherein the plurality of communication streams comprise a plurality of respective unique stream identifiers (SIDs) and correspond to GVMs and hosted hypervisors.

15. A method for global system memory management unit (SMMU) fault handling, the method comprising: issuing a global system memory management unit (SMMU) fault; discovering that a stream ID (SID) associated with the global SMMU fault is assigned to a guest virtual machine (GVM); and restarting only the GVM, thereby avoiding a complete system restart and allowing continuous cluster functionality during the GVM restart.

16. The method of claim 15, wherein the global SMMU fault is a result of a fault memory transaction, and a hosted hypervisor connected to the SMMU and connected to the GVM identifies that a stream identifier (SID) associated with the global SMMU fault is assigned to the GVM.

17. The method of claim 15, wherein the SMMU fault is a result of the GVM being terminated, and the method further comprises: the hosted hypervisor clearing level 1 SMMU translation registers; and ignoring the SMMU fault and allowing the GVM to restart upon identifying that a stream identifier (SID) associated with the global SMMU fault is assigned to the GVM that is restarting.

18. The method of claim 15, wherein the GVM is associated with a peripheral device, and a plurality of peripheral devices share double data rate (DDR) memory.

19. The method of claim 18, wherein the plurality of peripheral devices each have a respective GVM.

20. The method of claim 18, wherein the peripheral device comprises a plurality of communication streams.