Secure and reliable virtualized domain-specific hardware accelerator

By introducing multi-HWA functional controllers and trusted sandboxed communication interfaces into embedded computing systems, the security and confidentiality challenges of HWA in embedded computing systems are solved, and higher system stability and reliability are achieved.

CN119939628APending Publication Date: 2025-05-06TEXAS INSTRUMENTS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510033159.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-04-08
Filing Date
2019-12-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In embedded computing systems, security, confidentiality, and virtualization challenges exist for domain-specific hardware accelerators (HWAs), especially when faced with malware and other malicious security threats.

Method used

By introducing a multi-HWA functional controller and a trusted sandboxed communication interface in the embedded computing system, communication between HWA thread users and multi-HWA functional controllers is facilitated, and the secure transmission and allocation of message requests and privileged credential information is realized.

Benefits of technology

It improves the security and confidentiality of HWA in embedded computing systems, prevents confidentiality intrusions such as fraud, and ensures the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939628A_ABST
    Figure CN119939628A_ABST
Patent Text Reader

Abstract

The invention relates to a secure and reliable virtualized domain-specific hardware accelerator. The present disclosure relates to various implementations of an embedded computing system (200). An embedded computing system (200) includes a hardware accelerator (HWA) thread user (202A, 202B, 204) and a second HWA thread user (202A, 202B, 204) that creates and issues a message request. The HWA thread user (202A, 202B, 204) and the second HWA thread user (202A, 202B, 204) are in communication with a microcontroller (MCU) subsystem (214). The embedded computing system also includes a first inter-processor communication (IPC) interface between the HWA thread user and the MCU subsystem and a second IPC interface between the second HWA thread user and the MCU subsystem, wherein the first IPC interface is isolated from the second IPC interface. The MCU subsystem is also in communication with the first domain-specific HWA (208, 210, 212) and the second domain-specific HWA (208, 210, 212).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application 201911374689.3, entitled “Secure and Reliable Virtualization Domain Specific Hardware Accelerator”, filed on December 27, 2019.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims priority to U.S. Provisional Application No. 62 / 786,616, filed on December 31, 2018, the entire contents of which are incorporated herein by reference. Background Art

[0004] Today's embedded computing systems are commonly found in a variety of applications such as consumer, medical, and automotive products. Design engineers often create embedded computing systems to perform specific tasks rather than act as general-purpose computing systems. For example, due to security and / or availability requirements, some embedded computing systems need to meet certain real-time performance constraints. To achieve real-time performance, embedded computing systems typically include a microprocessor that loads and executes software to perform various functions and specialized hardware that improves the computational operations of certain tasks. An example of specialized hardware found in an embedded system is a hardware accelerator (HWA), which improves the confidentiality and performance of the embedded computing system.

[0005] As today's products continue to increasingly utilize embedded computing devices, design engineers are continually working to improve the safety, security, and performance of these devices. For example, like any other computing system, embedded computing systems are susceptible to malware or other malicious security threats. For embedded computing systems employed in applications that directly impact safety or security applications or are critical to safety and security applications, security breaches can be problematic. For example, embedded computing systems found in advanced driver assistance systems are designed to reduce human operator error and road fatalities caused by motor vehicles. A malicious computer program that intentionally accesses and corrupts an advanced driver assistance system can cause a system failure that could result in a life-threatening or dangerous situation. Summary of the invention

[0006] A simplified overview of the disclosed subject matter is presented below to provide a basic understanding of some aspects of the subject matter disclosed herein. This overview is not an exhaustive overview of the technology disclosed herein. This overview is not intended to identify key or important elements of the present invention or to describe the scope of the present invention, and its sole purpose is to introduce some concepts in a simplified form as a preface to a more detailed description discussed later.

[0007] In one implementation, a non-transitory program storage device includes instructions stored thereon that cause one or more processors to create a trusted sandboxed communication interface to facilitate communication between a designated HWA thread user and a multi-HWA function controller, wherein the multi-HWA function controller is configured to provide a message request from the HWA thread user to a destination, domain-specific HWA. The one or more processors may filter out a first message request received from a second HWA thread user for a destination domain-specific HWA, and write a second message request and privileged credential information received from the designated HWA into a buffer of the trusted sandboxed communication interface. The one or more processors provide the second message request and privileged credential information from the buffer of the trusted sandboxed communication interface to the multi-function HWA function controller.

[0008] In another implementation, a system includes an HWA thread user, a microcontroller unit (MCU) subsystem in communication with the HWA thread user, and a domain-specific HWA in communication with the MCU subsystem, wherein the domain-specific HWA includes an HWA thread. The MCU subsystem is configured to: receive a message request and privileged credential information from the HWA thread user, assign an HWA thread of the domain-specific HWA to perform the message request, classify the message request into one of a plurality of classes based on whether the domain-specific HWA is able to verify the privileged credential information, and forward the privileged credential information to the HWA thread based on a determination that the message request belongs to a first class indicating that the HWA thread is able to process the privileged credential information.

[0009] In yet another implementation, a system includes a HWA thread user and a second HWA thread user that creates and issues a message request. The HWA thread user and the second HWA thread user communicate with an MCU subsystem. The embedded computing system also includes a first inter-processor communication (IPC) interface between the HWA thread user and the MCU subsystem and a second IPC interface between the second HWA thread user and the MCU subsystem, wherein the first IPC interface is isolated from the second IPC interface. The MCU subsystem also communicates with a first domain-specific HWA and a second domain-specific HWA. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] For a detailed description of various examples, reference will now be made to the accompanying drawings, in which:

[0011] Figure 1 is a block diagram of an embedded computing system according to various implementations.

[0012] Figure 2 is a high-level block diagram of an example embedded computing system that includes a multi-HWA function controller.

[0013] Figure 3is a block diagram of an example embedded computing system including an MCU subsystem as an example of a multi-HWA function controller and an IPC interface as an example of a trusted sandboxed communication interface.

[0014] Figure 4 is a block diagram of another example embedded computing system that includes an HWA thread but does not have a privileged generator.

[0015] Figure 5 yes Figure 3 and Figure 4 A block diagram of an example implementation of an IPC interface is shown.

[0016] Figure 6 is a flow chart of an implementation of a method for exchanging communications between an HWA thread user and a multi-HWA function controller.

[0017] Figure 7 is a flow chart of an implementation of a method for classifying message requests according to capabilities of a destination domain-specific HWA.

[0018] Although certain implementations will be described in conjunction with the illustrative implementations shown herein, the invention is not limited to those implementations. Instead, all alternatives, modifications, and equivalents are included within the spirit and scope of the invention as defined by the claims. In the drawings that are not drawn to scale, throughout the specification and in the drawings, the same reference numerals are used for components and elements having the same structure, and reference numerals with quotation marks are used for components and elements having similar functions and configurations as components and elements having the same unquoted reference numerals. DETAILED DESCRIPTION

[0019] Various example implementations of improving the security, confidentiality, and virtualization of domain-specific hardware accelerators (HWAs) within an embedded computing system are disclosed herein. In one or more implementations, the embedded computing system includes a multi-HWA function controller that facilitates communication between one or more HWA thread users and one or more domain-specific HWAs (e.g., visual HWAs). The embedded computing system creates a trusted sandboxed communication interface that independently transmits message requests from HWA thread users to the multi-HWA function controller. A "trusted" communication interface is an interface in which the source device of a communication message is confirmed to be allowed to send messages through the specific communication interface (only predefined source devices are allowed to send messages through a given communication interface). Sandboxing refers to the embedded computing system isolating each communication interface from each other. In this way, confidentiality and / or system failures that affect one HWA thread user (e.g., a host CPU) do not affect another HWA thread user (e.g., a digital signal processor (DSP)). The trusted sandboxed communication interface also transmits privileged credential information for each message request to the multi-HWA function controller to prevent confidentiality intrusions such as spoofing.

[0020] After obtaining the message request, the multi-HWA function controller schedules and allocates a hardware thread for the message request to be executed on the destination domain-specific HWA. As part of the scheduling operation, the multi-HWA function controller performs an intelligent scheduling operation that classifies the message request into a plurality of classes (referred to as hardware-assisted classes) according to the capabilities of the destination domain-specific HWA. For example, if the destination domain-specific HWA includes a privilege generator, the multi-HWA function controller classifies the message request for the destination domain-specific HWA into a class representing a domain-specific HWA with privileged credential information checking capability. For a destination domain-specific HWA without a privilege generator, the multi-HWA function controller may classify the associated message request into a different class indicating that other hardware components (e.g., an input / output (IO) memory management unit (MMU)) will assist in checking the privileged credential information. In some cases, when the embedded computing system is unable to check the associated privileged credential information, the multi-HWA function controller may classify the message request into another class. In one or more implementations, the multi-HWA function controller is also able to convert between different address space sizes (eg, from a 64-bit address space to a 32-bit address space) to additionally accommodate domain-specific HWAs with varying capabilities (eg, legacy domain-specific HWAs).

[0021] As used herein, the term "programmable accelerator" refers to a custom hardware device that can be programmed to perform a specific operation (e.g., a process, calculation, function, or task). A programmable accelerator is different from a general-purpose processor (e.g., a central processing unit (CPU)) built to perform conventional computing operations. Typically, a programmable accelerator performs specified operations faster than software running on a standard or general-purpose processor. Examples of programmable accelerators dedicated to performing specific operations include graphics processing units (GPUs), digital signal processors (DSPs), vector processors, floating-point processing units (FPUs), application-specific integrated circuits (ASICs), embedded processors (e.g., universal serial bus (USB) controllers), and domain-specific HWAs.

[0022] For the purposes of this disclosure, the term "domain-specific HWA" refers to a specific type of programmable accelerator with customized hardware units and pipelines designed to perform tasks that fall within a specific domain. Compared to other types of programmable accelerators (such as GPUs, DSPs, and vector processors), domain-specific HWAs provide relatively less computational flexibility, but when performing tasks belonging to a specific domain, domain-specific HWAs have higher efficiency in terms of power and performance efficiency. A domain-specific HWA contains one or more HWA threads, each of which represents a hardware thread that receives and executes one or more tasks associated with a given domain. As a hardware thread, an HWA thread is different from a software thread generated when a software application runs on an operating system (OS). A domain-specific HWA can execute HWA threads in serial and / or parallel manner. Examples of domains include imaging domains, video domains, visual domains, radar domains, deep learning domains, and display domains. Examples of domain-specific HWAs include visual preprocessing accelerators (VPACs), digital media preprocessing accelerators (DMPACs), video processing engines (VPEs), and image and video accelerators (IVAs) (e.g., video encoders and decoders).

[0023] Illustrative Hardware and Use Cases

[0024] Figure 1 is a simplified block diagram of an embedded computing system 100 according to various implementations. Figure 1As an example, the embedded computing system 100 is a multi-processor system on a chip (SOC) designed to support computer vision processing in a camera-based advanced driver assistance system. The embedded computing system 100 includes a general purpose processor (GPP) 102, a digital signal processor (DSP) 104, a vision processor 106, and a domain-specific HWA 112 coupled via a high-speed interconnect 122. The GPP 102 hosts a high-level operating system (HLOS) that provides control operations for one or more software applications running on the embedded computing system 100. For example, the HLOS controls the scheduling of various tasks that the software applications generate when running on the embedded computing system 100. The DSP 104 provides support for real-time computer vision processing such as object detection and classification. Although Figure 1 The embedded computing system 100 is shown to include a single GPP 102 and a single DSP 104 , but other embodiments of the embedded computing system 100 may have multiple GPPs 102 and / or multiple DSPs 104 coupled to one or more domain-specific HWAs 112 and one or more vision processors 106 .

[0025] In one or more implementations, the domain-specific HWA 112 is a VPAC that communicates with the vision processor 106. The VPAC includes one or more HWA threads configured to perform various vision pre-processing operations on incoming camera images and / or image sensor information. For example, the VPAC includes four HWA threads, an embedded hardware thread scheduler, and embedded shared memory, all of which communicate with each other when performing vision domain tasks. Each HWA thread is set up to perform a specific vision domain task, such as, for example, a lens distortion correction operation, an image scaling operation, a noise filter operation, and / or other vision-specific image processing operations. Storage blocks in the shared memory act as buffers to store data blocks processed by the HWA threads. Figure 1 In FIG. 1 , the vision processor 106 is a vector processor custom tuned for computer vision processing such as gradient calculation, direction merging, and histogram normalization by utilizing the output of the VPAC.

[0026] The embedded computing system 100 also includes a direct memory access (DMA) component 108, a camera capture component 110 coupled to a camera 124, a display management component 114, an on-chip random access memory (RAM) 116 (e.g., a non-transitory computer readable medium), and various input / output (I / O) peripherals 120, all of which are coupled to the processor and domain-specific HWA 112 via an interconnect 122. The RAM 116 can store some or all of the instructions (software, firmware) described herein for execution by the processor. In addition, the embedded computing system 100 includes a security component 118 that includes safety-related functions to enable compliance with automotive safety requirements. Such functions can include support for CRC (cyclic redundancy check) of data, clock comparators for drift detection, error signaling, windowed watchdog timers, and self-testing of the embedded computing system 100 for damage and failure.

[0027] although Figure 1 A specific implementation of the embedded computing system 100 is shown, but the present disclosure is not limited to Figure 1 For example, Figure 1 Not all components found within the embedded computing system 100 may be shown, and other components known to those of ordinary skill in the art may be included depending on the use case of the embedded computing system 100. For example, the embedded computing system 100 may also include components that are beneficial for certain use cases. Figure 1 Other programmable accelerator components not shown. Alternatively or in addition, even if Figure 1 While one or more components within embedded computing system 100 are shown as separate components, other implementations may combine the components into a single component. Figure 1 The use and discussion of are merely examples to facilitate description and explanation.

[0028] Multi-HWA function controller and trusted sandboxed communication interface

[0029] Figure 2 is a high-level block diagram of an example embedded computing system 200 including a multi-HWA function controller 214. Figure 2The multi-HWA function controller 214 is shown interfacing with one or more HWA thread users (also referred to as HWA thread user devices, and including, for example, the host CPU 202A, the host CPU 202B, and the DSP 204) and one or more domain-specific HWAs (visual domain HWA 210, video domain HWA 212). In one or more implementations, the multi-HWA function controller 214 is a microcontroller unit (MCU) subsystem that supports communication between the HWA thread users 202A, 202B, 204 and the domain-specific HWAs 208, 210, and 212. The MCU subsystem includes one or more MCU processors and embedded memory to control and manage HWA threads between one or more domain-specific HWAs. Due to scalability, design and development costs, and loss of chip area, the MCU subsystem may preferably manage communications with multiple domain-specific HWAs. For example, the MCU subsystem provides flexibility by being able to assign any HWA thread within a domain-specific HWA to any HWA thread user. The MCU subsystem may also be scalable by updating the MCU firmware with revised or new policy settings (eg, when the number of virtual machines (VMs) that the MCU subsystem needs to manage changes).

[0030] HWA threads delegate the offloading of one or more tasks to the underlying hardware resources of one or more domain-specific HWAs. Figure 2 In the example, host CPUs 202A and 202B and DSP 204 represent HWA thread users that send message requests to visual domain HWA 208, display domain HWA 210, and / or video domain HWA 212. In the example, visual domain HWA 208 is restricted to perform visual domain tasks; display domain HWA 210 is restricted to perform display domain tasks; and video domain HWA 212 is restricted to perform visual domain tasks. In other words, visual domain HWA 208, display domain HWA 210, and video domain HWA 212 are limited in processing flexibility compared to general-purpose processors (such as host CPUs 202A and 202B) and / or other types of programmable accelerators (such as DSP 204). However, visual domain HWA 208, display domain HWA 210, and video domain HWA 212 are more efficient in performing each of their respective domain tasks than host CPUs 202A and 202B and DSP 204.

[0031] In order to improve operational efficiency (e.g., power efficiency and / or performance efficiency), HWA thread users offload domain tasks to their respective domain-specific HWAs by sending message requests. Each message request typically contains a command representing a domain task that can be executed by a domain-specific HWA. For example, a virtual machine (VM) runs a software application through the host CPU 202A to generate a set of visual domain tasks. Although the host CPU 202A has the ability to execute and process visual domain tasks, the host CPU 202A offloads the visual domain task group to the visual domain HWA 208 to obtain operational efficiency. By offloading domain tasks, the amount of time and / or power consumption required for the visual domain HWA 208 to complete the execution of the visual domain task group is relatively less than the amount of time and / or power consumption that the host CPU 202A has processed the visual domain task group.

[0032] The multi-HWA function controller 214 manages and controls message requests sent between HWA thread users and domain-specific HWAs. In one or more implementations, to enhance security and confidentiality, the embedded computing system 200 creates a trusted sandboxed communication interface that securely transmits message requests from HWA thread users to the multi-HWA function controller 214. The trusted sandboxed communication interface acts as a confidentiality interface that separates and filters data from non-specified HWA thread users. In other words, the trusted sandboxed communication interface controls whether the underlying hardware resource (e.g., host CPU 202A, 202B or DSP 204) is a trusted source with permission to transmit message requests to the multi-HWA function controller 214. For example, if the trusted sandboxed communication interface is set to only identify the host DSP 204 as a trusted source, the trusted sandboxed communication interface will not transmit message requests received from the host CPU 202A and / or 202B to the multi-HWA function controller 214. Having a separate trusted sandboxed communication interface limits the impact of system failures and / or confidentiality breaches. The trusted sandboxed communication interface also provides privileged credential information for each message request to the multi-HWA function controller 214 to provide an additional layer of security to prevent malicious attacks such as spoofing.

[0033] After receiving the message request, the multi-HWA function controller 214 schedules and allocates HWA threads to execute the message request. The multi-HWA function controller 214 can schedule message requests to different domain-specific HWAs. Figure 2As an example, the host CPU 202A may generate a message request including a set of visual domain tasks, a second message request including a set of display domain tasks, and a third message request having a set of video domain tasks. The multi-HWA function controller 214 receives three different message requests through one or more trusted sandboxed communication interfaces, and then assigns each message request to an HWA thread based on the type of domain task. In other words, the multi-HWA function controller 214 does not assign HWA threads that are incompatible with domain tasks associated with other domains or that cannot process domain tasks associated with other domains. For example, the multi-HWA function controller 214 assigns at least one of the visual HWA threads 216A-216D to execute the visual domain task group, assigns at least one of the display HWA threads 218A and 218B to execute the display domain task group, and assigns at least one of the video HWA threads 220A and 220B to execute the video domain task group. When an HWA thread becomes available, the multi-HWA function controller 214 assigns a compatible HWA thread to execute the message request. In the event that a compatible HWA thread is busy, the multiple HWA function controller 214 may temporarily push the message request into one or more different queues to wait for a compatible HWA thread to become available.

[0034] In one or more implementations, as part of the scheduling operation, the multi-HWA function controller 214 performs intelligent scheduling operations to take into account the capabilities of the destination domain-specific HWAs. In one or more implementations, the multi-HWA function controller 214 categorizes each domain-specific HWA according to the capabilities of the HWA threads within the domain-specific HWA. Figure 2 As an example, after the multi-HWA function controller 214 schedules one of the visual HWA threads 216A-216D to process a message request, the multi-HWA function controller 214 determines whether the visual HWA thread 216A-216D belongs to a class of HWA threads that includes a privilege generator for dynamically processing privileged credential information. If the visual HWA thread 216A-216D includes a privilege generator, the multi-HWA function controller 214 may replay the privileged credential information obtained from the trusted sandboxed communication interface to the assigned visual HWA thread 216A-216D. The multi-HWA function controller 214 also provides privilege configuration information to the IO MMU ( Figure 2 If the assigned visual thread belongs to a class that cannot process privileged credential information but can be assisted by the IO MMU, the data output from the visual domain HWA 208 is rerouted to the IO MMU to confirm the privileged credential information.

[0035] When determining the HWA thread class, the smart scheduling operation of the multi-HWA function controller 214 also supports hardware virtualization and / or address space size conversion. In one or more implementations, the HWA thread user (e.g., the host CPU 202A) can host one or more virtualized computing systems (e.g., VMs). Due to hardware virtualization, the message request sent from the HWA thread user can include a command written to a specific virtualized destination address. In order to support hardware virtualization, the multi-HWA function controller 214 converts the virtualized destination address to a physical address. When the domain-specific HWA utilizes different address space sizes, the multi-HWA function controller 214 can also perform address space size conversion. For example, the address information received by the multi-HWA function controller 214 can utilize a 64-bit address space. However, the domain-specific HWA can utilize a 32-bit address space. As part of the smart scheduling operation, the multi-HWA function controller 214 converts the address information from the 64-bit address space to a lower-bit address space (e.g., a 32-bit address space).

[0036] MCU subsystem and IPC interface

[0037] Figure 3 3 is a block diagram of an example embedded computing system 300 that includes an MCU subsystem 328 as an example of a multi-HWA function controller and an IPC interface 320 as an example of a trusted sandboxed communication interface. The IPC interface 320 is an example of a communication interface. The example includes one IPC interface 320 for each device, such as one IPC interface 320 for the host CPU 202A, one IPC interface 320 for the host CPU 202B, and one IPC interface 320 for the DSO 204. Each IPC interface 320 communicatively couples its respective device 202A, 202B, and 204 to the MCU subsystem 328. Each IPC interface 320 provides a processor-independent application program interface (API) for communicating with a processing component. For example, the IPC interface 320 can be used for communication between processors in a multi-processor environment (e.g., between cores), communication with other hardware threads on the same processor (e.g., between processes), and communication with peripheral devices (e.g., between devices). Typically, as a software API, IPC interface 320 utilizes one or more processing resources, such as a multiprocessor heap, a multiprocessor linked list, and a message queue, to facilitate communications between processing components.

[0038] exist Figure 3, the embedded computing system 300 creates an IPC interface 320 between the MCU subsystem 328 and each virtual computing system (e.g., VM or virtual container) running on the HWA thread user. For example, the embedded computing system 300 allocates one IPC interface 320 to transmit message requests between VM 302A and the MCU subsystem 328, and allocates another IPC interface 320 to transmit message requests between VM 302B and the MCU subsystem 328. The embedded computing system 300 also creates an IPC interface 320 between the DSP 204 and the MCU subsystem 328. VMs 302A and 302B each run a separate high-level OS (HLOS) in the embedded computing system 300. For the purposes of this disclosure, an HLOS represents an embedded OS that is the same as or similar to an OS used in a non-embedded environment, such as a desktop computer and a smart phone. Reference Figure 3 For example, VMs 302A and 302B may run the same type of HLOS (e.g., both run Android TM (Android) OS) or a different type of HLOS (e.g., VM 302A running Linux TM OS, while VM 302B runs Android TM OS).

[0039] Creating separate and isolated IPC interfaces 320 for DSP 204 and each virtual computing system (e.g., VM or virtual container) running on host CPUs 202A and 202B enhances security and confidentiality by isolating faults and / or confidentiality breaches. Figure 3 In the DSP 204, a real-time operating system (RTOS) 304 is run that provides features such as threads, semaphores, and interrupts. Compared to the HLOS, the RTOS can provide a relatively fast interrupt response at a lower memory cost. In an advanced driver assistance system application, by utilizing the RTOS, the DSP 204 can manage automotive safety functions (e.g., emergency braking) by processing real-time data from one or more sensors (e.g., a camera). If other HWA thread users (e.g., the host CPU 202A) suffer a system failure or a confidentiality breach, the automotive safety features managed by the DSP 204 are not affected because the IPC interface 320 assigned to the DSP 204 is isolated and separated from the other IPC interfaces 320. This disclosure will refer to the RTOS later. Figure 5 The IPC interface 320 is discussed in more detail.

[0040] Figure 3The MCU subsystem 328 is shown to include an engine 308 that configures the MCU subsystem 328 to pair with HWA threads within the visual domain HWA 208, the display domain HWA 210, and the video domain HWA 212. By pairing with different types of HWA threads, the engine 308 can control and manage different types of HWA threads and is not limited to communicating with a specific type of HWA thread. Figure 3 As an example, after the MCU subsystem 328 receives the message request via the IPC interface 320, the engine 308 schedules and forwards the message request received from the DSP 204 and / or from the host CPUs 202A and 202B to one or more of the HWA threads in the visual domain HWA 208, the display domain HWA 210, and / or the video domain HWA 212. In one or more implementations, the engine 308 is firmware that supports policy settings (such as priority and access control for each thread) to support scheduling and forwarding the message request to one or more HWA threads.

[0041] The engine 308 can support priority-based queue services for each domain-specific HWA (e.g., the visual domain HWA 208). Figure 3 As shown, the MCU subsystem 328 includes a priority queue 306 that receives message requests from the IPC interface 320. Each priority queue 306 is configured to receive message requests from one or more of the IPC interfaces 320. Different priorities can be assigned to the priority queues 306 according to the type of HWA thread user that sends the message request. For example, due to real-time constraints, the MCU subsystem 328 can assign a higher priority to the priority queue 306 that receives the message request from the DSP 204 than the priority queues 306 assigned to the host CPUs 202A and 202B. The engine 308 can also arrange the received message requests in each priority queue according to a priority operation. As an example, the priority operation can arrange the message request in one of the priority queues 306 based on a first-in, first-out (FIFO) operation. Other examples can use other priority assignment operations to sort the message requests within a single priority queue 306. When the engine 308 extracts the message request from the priority queue 306 according to the priority, the engine 308 assigns the HWA thread identifier to the message request. The HWA thread identifier indicates which HWA thread will execute the message request. In the case that the assigned HWA thread is busy, the engine 308 pushes the pending message request to the pending queue to wait until the assigned HWA thread is available to process the message request. If the assigned HWA thread is already available or idle, the engine 308 schedules the message request for execution.

[0042] The engine 308 may also perform intelligent scheduling operations to support multiple classes of HWA threads. As previously discussed, the embedded computing system 300 may include domain-specific HWAs with different processing capabilities. Since domain-specific HWAs may have different capabilities, the engine 308 is configured to schedule message requests for different classes of HWA threads. In order to support multiple classes of HWA threads, the MCU subsystem 328 includes a privilege configuration engine 310 that sends privilege configuration information to the domain-specific HWA through a privilege generator 322 and / or a supporting device (such as an IO MMU 314). The privilege configuration information includes policy information indicating the type of privilege level used to access certain portions of the memory 318. The privilege generator 322 and / or the IO MMU 314 within the HWA thread utilizes the privilege configuration information to check the privilege credential information associated with each message request.

[0043] Different classes of HWA threads include a class of HWA threads that are capable of checking privileged credential information. For example, the IO-MMU 314 can be used to check privileged credential information. The first class of HWA threads identifies an HWA thread (e.g., the visual HWA thread 216A) that has a privilege generator 322 for dynamically processing privileged credential information. If the assigned HWA thread includes a privilege generator 322, the engine 308 replays the privileged credential information obtained from the IPC interface 320 to the assigned HWA thread. The second class of HWA threads includes HWA threads that do not have a privilege generator 322, but can be assisted by other hardware components to check privileged credential information. For example, Figure 3 The IO MMU 314 shown in FIG. 3 can assist and check the privileged credential information obtained from the IPC interface 320. The third type of HWA thread represents an HWA thread that does not have a privilege generator and cannot use other hardware components to check the privileged credential information. For the third type of HWA thread, the engine 308 may not be able to use the privileged credential information to perform additional security checks. In some implementations, the third type of HWA thread represents an HWA thread that supports hardware virtualization without checking the privileged credential information.

[0044] Figure 3The visual HWA thread 216A within the visual domain HWA 208 is depicted to also include a privilege generator 322 and a visual HWA thread 326. The privilege generator 322 supports determining whether the privileged credential information associated with the message request satisfies the privilege level to access the data and write the data into the destination storage space. The privilege generator 322 evaluates the privileged credential information, such as the VM identifier, the secure or non-secure mode identifier, the user or supervisor mode identifier, and / or the HWA thread user identifier (e.g., the host processor identifier) ​​to determine whether the visual HWA thread 326 should access the destination storage space within the memory 318. In one or more implementations, the privilege generator 322 includes an initiator confidentiality controller and a quality of service engine. The initiator confidentiality controller supports tracking and evaluating privileged credential information, such as VM identifiers and channelized firewalls, via MMR settings. When the visual HWA thread 326 executes the message request, the quality of service engine supports priority-based policies via MMR settings. The visual HWA thread 326 represents a hardware thread that executes the message request after verifying the privileged credential information of all message requests. After executing the message request, the visual HWA thread 326 outputs the data to the memory 318 .

[0045] The engine 308 may also classify HWA threads based on address space utilization. In one or more implementations, the engine 308 performs address space translation when a domain-specific HWA utilizes an address space size (e.g., a 64-bit HLOS) that is different from the address space size employed by the hardware thread user. As part of the smart scheduling operation, the engine 308 translates address information from a larger address space to a smaller address space when message requests are sent to certain HWA threads (e.g., the visual HWA thread 216A). For example, the visual domain HWA 208 includes a visual HWA thread 216A that has an address expander 324 to support a larger address space (e.g., a 64-bit HLOS). Figure 3 In the embodiment of the present invention, the address expander 324 allows the visual HWA thread 216A that utilizes a smaller address space (e.g., a 32-bit address space) to be compatible with a larger address space (e.g., a 36-bit, 40-bit, and 48-bit address space). In one or more implementations, the address expander 324 performs a regional address translation (RAT) that supports address translation from 32 bits to 36-bit, 40-bit, and / or 48-bit address spaces. The RAT supports multiple high address spaces that can be mapped to lower 32-bit address spaces via memory mapping register (MMR) settings.

[0046] After the HWA thread (e.g., visual HWA thread 216A) completes executing the message request, the HWA thread sends an interrupt completion notification back to the MCU subsystem 328. The MCU subsystem 328 includes an interrupt controller (INTC) 312 to receive and process interrupt completion notifications from one or more HWA threads. For each interrupt completion notification received by the INTC 312, the INTC 312 sends a confirmation message back to the HWA thread user to indicate that the execution of the message request is completed. The INTC 312 also notifies the engine 308 that the HWA thread that sent the interrupt completion notification is now available to process the message request. Because one or more HWA threads are asynchronous hardware threads, the INTC 312 can be beneficial.

[0047] Figure 4 is a block diagram of another example embedded computing system 400 including HWA threads without privileged generators. The embedded computing system 400 is similar to Figure 3 The embedded computing system 300 is shown, except that the visual HWA thread 216A does not include a privilege generator. Figure 4 As shown, because the visual HWA thread 216A cannot check the privileged credential information for the message request, the MCU subsystem 328 provides instructions to the visual HWA thread 216A to reroute the output data to the IO MMU 314 for processing. When the IO MMU 314 receives the output data from the visual HWA thread 216A, the IO MMU 314 checks the privileged credential information against the privileged configuration information received from the privileged configuration engine 310. If the IO MMU 314 determines that the message request is from a trusted source and has the necessary privileged credentials, the IO MMU 314 stores the output data to the destination memory address in the memory 318.

[0048] Figure 5 yes Figure 3 and Figure 4 . As previously described, the IPC interface 320 facilitates communication between the host CPU 202A and the MCU subsystem 328. Figure 5 As shown, the host CPU 202A creates and runs the VM 302A with the HLOS. When the host CPU 202A sends a message request 510 to the domain-specific HWA of the VM 302A, the firewall 502 processes the message request 510. The firewall 502 has a setting that allows hardware access to the IPC interface 320 based on a hardware resource identifier (e.g., a CPU identifier). In other words, in order to isolate the IPC interface 320 from other IPC interfaces 320 that transmit message requests from other HWA thread users, the firewall 502 blocks and filters out data from other HWA thread users (e.g., CPU 202B).

[0049] After the message request 510 passes through the firewall 502, the message request 510 encounters the first hardware agent 504, which writes the message request 510 and the privileged credential information 512 for the message request 510 into the IPC queue 506. The message request 510 may include the destination HWA thread information, one or more commands to be executed, and the destination memory address (e.g., input / output (IO) buffer address) to store the output data from the destination domain-specific HWA. The privileged credential information 512 includes sub-attributes such as an identifier for a virtual computing system (e.g., a VM or a virtual container), an indication of whether the message request is associated with a secure mode or a non-secure mode and / or a user mode or a supervisor mode, and an HWA thread user identifier (e.g., an identifier of the host CPU 202A or 202B). Subsequently, the second hardware agent 508 reads the message request 510 and the privileged credential information 512 from the IPC queue 506 and passes both the message request 510 and the privileged credential information 512 to the MCU subsystem 328. In one or more implementations, IPC queue 506 represents a FIFO buffer where second hardware agent 508 reads message requests 510 based on the order in which first hardware agent 504 wrote message requests 510 into IPC queue 506. Other implementations may implement IPC queue 506 using other types of buffers.

[0050] Figure 6 6 is a flow chart of an implementation of a method 600 for exchanging communications between an HWA thread user and a multi-HWA function controller. The method 600 may be used as follows: Figures 3 to 5 Specifically, the method 600 creates an IPC interface 320 for each virtual computing system hosted by the HWA thread user to facilitate communication between the HWA thread user and the MCU subsystem. Figure 6 The use of MCU subsystem 328 and IPC interface 320 is described, but other implementations may use other types of multi-HWA function controllers and trusted sandboxed communication interfaces. Figure 6 The blocks of method 600 are shown to be implemented as sequential operations, but method 600 is not limited to this order of operations, and other implementations of method 600 may have one or more blocks implemented as parallel operations.

[0051] The method 600 begins at block 602 by creating a trusted sandboxed IPC interface to facilitate communication between an HWA thread user and an MCU subsystem that communicates with a requested domain-specific HWA. In one or more implementations, the method 600 creates a separate IPC interface for each virtual computing system running on an HWA thread user. Creating separate and isolated IPC interfaces prevents system failures or confidentiality intrusions from affecting other HWA thread users. The method 600 then moves to block 604. At block 604, the method 600 allows the HWA thread user to access a message request and provides the message request to the created trusted sandboxed IPC interface. As an example, the method 600 may utilize a firewall to filter out message requests from other non-specified HWA thread users.

[0052] The method 600 may move to block 606 to store the message request and the privileged credential information in a buffer of the trusted sandboxed IPC interface. The method 600 then continues to block 608 and receives the message request and the privileged credential information from the trusted sandboxed IPC interface. The method 600 moves to block 610 to determine whether the HWA thread of the domain-specific HWA is available for execution. If the HWA thread is not available, the message request is pushed to a pending queue to wait for an available HWA thread. Otherwise, at block 612, when the allocated HWA thread is not available, the method 600 provides the message request and the privileged credential information to a queue within the MCU subsystem. The method 600 moves to block 614 and schedules the message request to be sent from the MCU subsystem to the domain-specific HWA when the HWA thread is available.

[0053] Figure 7 7 is a flow chart of an implementation of method 700, which classifies message requests according to the capabilities of a destination domain-specific HWA. Figures 2 to 5 214 or MCU subsystem 328 referenced in the above. Recall that as part of the scheduling operation of the multi-HWA function controller, the multi-HWA function controller organizes message requests into classes according to the capabilities of the domain-specific HWAs that are to execute the message requests. By having the method 700 divide the message requests into classes, the method 700 can schedule message requests for various domain-specific HWAs, where each domain-specific HWA includes one or more HWA threads. Similar to Figure 6 ,although Figure 7 The blocks of method 700 are shown to be implemented as sequential operations, but method 700 is not limited to this order of operations, and other implementations of method 700 may have one or more blocks implemented as parallel operations.

[0054] Method 700 begins at block 702 by determining whether the HWA thread assigned to execute the message request supports privileged credential verification. In one or more implementations, the HWA thread supports privileged credential verification as previously described with reference to FIG. Figure 3 If the method 700 determines that the assigned HWA thread supports privileged credential verification, the method 700 moves to block 704 to replay the privileged credential information captured by the trusted sandboxed IPC interface to the assigned HWA thread. Thereafter, the method 700 moves to block 716 and sends the message request to the assigned HWA thread for execution.

[0055] Returning to block 702, if the method 700 determines that the assigned HWA thread does not support privileged credential verification, the method 700 moves to block 706 and determines whether hardware assistance via the IO MMU is available. In one or more implementations, the multi-HWA function controller provides privileged configuration information to other hardware components in addition to the domain-specific HWA (e.g., IO MMU). Providing the privileged configuration information allows the IO MMU or other hardware components to check the privileged credential information associated with the message request. If the method 700 determines that hardware assistance is available, the method 700 moves to block 708 and provides instructions to cause the specific domain HWA to reroute the output to the hardware assistance component (e.g., IO MMU). Alternatively, if the method 700 determines that no hardware assistance is available, the method 700 may move to block 710 to convert the destination virtual address to a physical address. At block 710, the method 700 does not verify or check the privileged credential information for the message request.

[0056] After box 708 or 710, method 700 then moves to box 712 and determines whether the physical destination address needs to be converted to another address space size. As previously described, some HWA threads can utilize address expanders to support the address capabilities of one or more OS systems (e.g., 64-bit OS systems) that utilize larger address spaces. Because the address space used by the HWA thread is different from the address adopted by the HWA thread user, method 700 determines whether to convert to another address space size. If the physical address needs to be converted to the target address space size and the privileged credential information captured by the trusted sandboxed IPC interface is replayed to the assigned HWA thread, method 700 moves to box 714. Thereafter, method 700 moves to box 716 and sends the message request to the assigned HWA thread for execution. Alternatively, if address space conversion is not required, method 700 moves to box 716.

[0057] Although several implementations have been provided in the present disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The examples herein should be considered illustrative rather than restrictive, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.

[0058] In addition, without departing from the scope of the present disclosure, the techniques, systems, subsystems and methods described and illustrated in a discrete or separate manner in various implementations may be combined or integrated with other systems, modules, techniques or methods. Other items shown or discussed as coupled or directly coupled or communicating with each other may be indirectly coupled or communicated by electrical, mechanical or other means through some interface, device or intermediate component.

Claims

1. A system comprising: processor: a controller coupled to the processor; as well as a hardware accelerator (HWA) coupled to the controller, wherein the HWA is configured to run an HWA thread; The controller is configured as follows: receiving a message from the processor; and In response to determining that the message belongs to a first category indicating that the HWA is capable of processing privileged credential information using the HWA thread, the privileged credential information is forwarded to the HWA.

2. The system of claim 1, wherein the controller is further configured to: In response to the message, allocating the HWA thread for execution; and The message is assigned a class based on whether the HWA has verified privileged credential information.

3. The system of claim 2, wherein the privileged credential information includes a virtual machine identifier, a secure or non-secure mode identifier, a user or supervisor mode identifier, and an HWA thread user device identifier.

4. The system of claim 2, wherein the controller is configured to assign the class to the message by: assigning the message to the first category based on determining that the HWA is capable of processing privileged credential information; assigning the message to a second category based on determining that the HWA is unable to verify privileged credential information and that a hardware component is configured to assist in verifying privileged credential information; and Based on the lack of privileged credential information verification, the message is assigned to the third category.

5. The system of claim 4, further comprising an input / output (IO) memory management unit (MMU) coupled to the controller and configured to check the privileged credential information when the message is assigned to the second class. 6 . The system of claim 1 , wherein the message includes destination HWA thread information, one or more commands executed on the HWA, and a destination memory address. The system of claim 6 , wherein the processor is a general purpose processor.

8. The system of claim 6, wherein the processor is a programmable accelerator.

9. The system of claim 1, wherein the message is obtained from an inter-processor communication interface (IPC) interface.

10. The system of claim 1, wherein the controller is configured to perform address space translation when the HWA thread performs an address extender operation.

11. The system of claim 1 , wherein the processor is a first processor, the system further comprising: a first inter-processor communication interface, i.e., a first IPC interface, coupled between the first processor and the controller; A second processor; as well as A second IPC interface is coupled between the second processor and the controller.

12. The system of claim 11, wherein the first IPC interface is isolated from the second IPC interface.

13. The system of claim 11, wherein the first IPC interface comprises a firewall configured to prevent the second processor from sending a message request to the first IPC interface.

14. The system of claim 13, wherein the firewall is configured to allow hardware access to the first IPC interface based on a hardware resource identifier.

15. The system of claim 1, wherein the HWA is a first HWA, the system further comprising a second HWA coupled to the controller.

16. The system according to claim 1 further includes an inter-processor communication interface (IPC interface) coupled between the processor and the controller, the IPC interface including a first hardware agent, the first hardware agent being configured to write a message request received from the processor into a queue; and a second hardware agent, the second hardware agent being configured to read the message request from the processor.

17. The system of claim 16, wherein the queue comprises a first-in-first-out (FIFO) buffer.

18. A controller comprising: The queue configured to receive messages; and An engine is configured to assign an HWA thread to execute the message and forward the privileged credential information based on determining that the message belongs to a first category indicating that the HWA thread is capable of processing privileged credential information.

19. The controller of claim 18, wherein the engine is further configured to assign a class to the message based on whether the HWA is able to verify the privileged credential information.

20. The controller of claim 19, wherein the engine is configured to assign the class to the message by: assigning the message to the first category based on determining that the HWA is capable of processing privileged credential information; assigning the message to a second class based on determining that the HWA is unable to verify privileged credential information and that a hardware component is configured to assist in verifying privileged credential information; and Based on the lack of privileged credential information verification, the message is assigned to the third category.

21. The controller of claim 18, wherein the message includes destination HWA thread information, one or more commands executed on the HWA, and a destination memory address.

22. The controller of claim 18, wherein the engine is configured to perform address space translation when the HWA thread performs an address extender operation.

23. The controller of claim 18, wherein the queue is a first priority queue, and wherein the controller comprises a plurality of priority queues, each priority queue of the plurality of priority queues being configured to be coupled to one or more inter-processor communication interfaces (IPC interfaces), wherein, The plurality of priority queues include the first priority queue.