System memory management unit execution isolation

By implementing TCU and TBU on the SoC, combined with MMU to manage memory access, and determining priorities based on identifiers, the problem of process interference on the SoC is solved, and efficient memory isolation and resource utilization are achieved.

CN120909751APending Publication Date: 2025-11-07NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510568529.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-07
Filing Date
2025-04-30
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

On a system-on-a-chip (SoC), applications and/or tasks may interfere with each other during memory access, limiting application performance.

Method used

By implementing multiple translation control units (TCUs) and translation buffer units (TBUs) on the SoC, combined with a memory management unit (MMU), priority is determined based on process and client identifier information, memory usage is managed to prevent interference, and memory isolation is provided.

Benefits of technology

It effectively prevents low-priority processes from interfering with high-priority processes, reduces the size, weight, and power requirements of hardware, and improves resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909751A_ABST
    Figure CN120909751A_ABST
Patent Text Reader

Abstract

The invention discloses system memory management unit execution isolation. Systems and methods according to the present disclosure may prevent interference of memory operations being performed, such as through high priority or high security applications. In various examples, a memory management unit (MMU) may receive a request from a client to execute a process. The MMU may select a target memory manager of a plurality of memory managers of the MMU based on an identifier of at least one of the client or the process. The MMU may cause the target memory manager to perform a memory translation operation for the request.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] A memory management unit (MMU) can receive a virtual address and translate the virtual address to a physical address in memory. An MMU can be useful, for example, for system on a chip (SoC) hardware, for example, to facilitate management of how applications executing on the SoC are accessed and efficient utilization of memory. However, applications and / or tasks in the applications can cause interference with each other in memory accesses, which can limit performance of the applications. SUMMARY

[0002] Embodiments of the present disclosure relate to systems and methods for preventing interference between processes being implemented on a hardware architecture, such as a system on a chip (SoC). In contrast to conventional systems, systems and methods according to the present disclosure can implement MMU components, such as multiple translation control units (TCUs) and / or multiple sets of translation buffer units (TBUs), to facilitate memory usage in a manner that presents interference. For example, a TCU and / or TBU can retrieve identifier information about a process and / or a client executing the process to determine a priority (e.g., an automotive safety integrity level (ASIL)) of the process, and can manage memory usage of the process to prevent interference, such as providing memory isolation for high-priority processes.

[0003] At least one embodiment relates to one or more processors. The one or more processors can include one or more circuits. The one or more circuits can receive, using a memory management unit (MMU), a request to execute a process from a client. The one or more circuits can select, using the MMU, a target memory manager of a plurality of memory managers of the MMU corresponding to an identifier of at least one of the client or the process based at least on the identifier. The one or more circuits can cause the target memory manager to perform a memory translation operation for the request.

[0004] In some embodiments, the one or more circuits can cause the MMU to select the target memory manager in response to the identifier indicating that the process corresponds to a priority level of the target memory manager. In some embodiments, the plurality of memory managers can include a plurality of translation control unit (TCU) instances. In some embodiments, the plurality of memory managers can include a plurality of sets of translation buffer units (TBUs).

[0005] In some embodiments, the identifier can include an identifier of the client. The one or more circuits can select the target memory manager in response to the identifier of the client indicating that the process has a constant priority level and the constant priority level corresponds to the target memory manager. In some embodiments, the identifier can include an identifier of the process. The target memory manager can be associated with a first range of identifiers. The plurality of memory managers can include a second memory manager associated with a second range of identifiers. The one or more circuits can select the target memory manager in response to the identifier of the process being within the first range of identifiers.

[0006] In some embodiments, the process can include at least one of a video encoding operation or a video decoding operation. In some embodiments, the process can be a first process having a first priority level, the client can request the MMU to perform a translation for a second process having a second priority level, the second priority level being higher than the first priority level. In some embodiments, the one or more circuits can include a system on a chip (SoC) including the MMU, and the MMU can be a system memory management unit (SMMU).

[0007] In some embodiments, the one or more processors can be included in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system containing one or more virtual machines (VMs); a system implemented using edge devices; a system implemented using robots; a system for generating synthetic data; a system for performing simulation operations; a system for performing collaborative content creation of 3D assets; a system for performing conversational AI operations; a system including one or more large language models (LLMs); a system for performing digital twin operations; a system for performing optical transport simulations; a system for performing deep learning operations; a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0008] At least one embodiment relates to a system. The system can include a plurality of memory managers. The system can include a memory management unit (MMU). The MMU can include one or more circuits. The one or more circuits can receive a request to execute a process from a client. The one or more circuits can select a target memory manager of the plurality of memory managers based at least on an identifier of at least one of the client or the process. The target memory manager can correspond to the identifier. The one or more circuits can cause the target memory manager to perform a memory translation operation for the request.

[0009] In some embodiments, the one or more circuits can select the target memory manager in response to the identifier indicating that the process corresponds to a priority level of the target memory manager. In some embodiments, the plurality of memory managers can include a plurality of translation control unit (TCU) instances. In some embodiments, the plurality of memory managers can include a plurality of groups of translation buffer units (TBUs).

[0010] In some embodiments, the identifier can include an identifier of a client. The one or more circuits can select the target memory manager in response to the identifier of the client indicating that the process has a constant priority level and the constant priority level corresponds to the target memory manager. In some embodiments, the identifier can include an identifier of a process. The target memory manager can be associated with a first range of identifiers. The plurality of memory managers can include a second memory manager associated with a second range of identifiers. The one or more circuits can select the target memory manager in response to the identifier of the process being within the first range of identifiers. In some embodiments, the one or more circuits can include a system on a chip (SoC) including the MMU. The MMU can be a system memory management unit (SMMU).

[0011] At least one embodiment relates to a method. The method can include receiving, using one or more processing circuits of a memory management unit (MMU), a request to execute a process from a client. The method can include selecting, using the one or more processing circuits of the MMU and based on an identifier of at least one of the client or the process, a target memory manager of a plurality of memory managers of the MMU. The target memory manager can correspond to the identifier. The method can include causing, using the one or more processing circuits of the MMU, the target memory manager to perform a memory translation operation for the request.

[0012] In some embodiments, the method can include selecting, using the one or more processing circuits of the MMU, the target memory manager in response to the identifier indicating that the process corresponds to a priority level of the target memory manager. In some embodiments, the plurality of memory managers can include a plurality of translation control unit (TCU) instances and a plurality of groups of translation buffer units (TBUs).

[0013] The processors, systems, and / or methods described herein can be implemented by, or can be included in, at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system containing one or more virtual machines (VMs); a system implemented using edge devices; a system implemented using robots; a system for generating synthetic data; a system for performing simulation operations; a system for performing collaborative content creation of 3D assets; a system for performing conversational AI operations; a system including one or more large language models (LLMs); a system including one or more visual language models (VLMs); a system for performing digital twin operations; a system for performing optical transport simulation; a system for performing deep learning operations; a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. BRIEF DESCRIPTION OF DRAWINGS

[0014] The present systems and methods for SMMU enforcement of isolation are described in detail below with reference to the following drawings, wherein:

[0015] Figure 1 is a block diagram of an MMU / SMMU architecture in accordance with some embodiments of the present disclosure;

[0016] Figure 2 is a block diagram of an SMMU architecture implemented by a memory management unit (MMU) in accordance with some embodiments of the present disclosure;

[0017] Figure 3 is a flowchart of a process for SMMU enforcement of isolation in accordance with some embodiments of the present disclosure;

[0018] Figure 4A is an illustration of an example autonomous vehicle in accordance with some embodiments of the present disclosure;

[0019] Figure 4B is an example of a camera position and field of view of the example autonomous vehicle of Figure 4A in accordance with some embodiments of the present disclosure;

[0020] Figure 4C is an example of a camera position and field of view of the example autonomous vehicle of Figure 4A in accordance with some embodiments of the present disclosure;

[0021] Figure 4D is a system diagram of communication between a cloud-based server and the example autonomous vehicle of Figure 4A in accordance with some embodiments of the present disclosure;

[0022] Figure 5 is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0023] Systems and methods related to SMMU enforcement of isolation are disclosed, e.g., for controlling how memory operations for clients (including applications and / or accelerators) are managed based on priority information regarding the memory operations. The systems and methods of the present disclosure prevent operations of relatively low-priority clients from interfering with operations of relatively high-priority clients, including but not limited to high-ASIL (e.g., ASIL D) clients. In non-limiting embodiments, the systems and methods disclosed herein can be implemented on the same SoC, which can reduce size, weight, and / or power requirements of isolation techniques based on multiple SoCs.

[0024] Systems and methods according to the present disclosure can enable more effective prevention of interference SoCs, such as by providing full isolation of traffic having different priorities by an SMMU. In applications including but not limited to vehicle solutions, such as L3+ driving solutions (e.g., for at least semi-autonomous vehicle operation), this is useful for preventing lower-priority processes (e.g., processes with lower automotive safety integrity level (ASIL) ratings, such as QM or lower processes) from interfering with execution of higher-priority processes (e.g., processes with higher ASIL ratings (e.g., non-QM processes)). In various such architectures, in various instances of interference, interference can occur at a memory management unit (MMU), e.g., a system memory management unit (SMMU), which can be used for functions such as memory translation for device isolation or virtualization. For example, an SMMU can be shared by multiple clients (e.g., clients providing instructions for processes to be executed), each of which can be operating processes at different priorities (e.g., different safety levels (e.g., ASIL levels), priority levels, etc.), and / or one or more of which can be operating processes having different priorities during different time intervals.

[0025] Some systems address the interference problem by running different processes on different SoCs. However, using multiple SoCs can result in increased size, weight, and / or power for operating the multiple SoCs, which can be challenging in resource-limited environments (e.g., vehicle / automotive / robotic / machine environments). Moreover, reliance on multiple SoCs can result in reduced resource utilization efficiency, as more resources on the SoCs can not be utilized under a given amount of workload for executing processes.

[0026] Systems and methods according to the present disclosure can include a SoC that can include multiple translation control unit (TCU) instances for retrieving memory translations based on requests from translation buffer units (TBUs). The TCU instances can be associated with different ranges of identifiers (e.g., stream IDs), for example to assign one or more first priorities (e.g., ASIL levels) to a first range associated with a first TCU instance and one or more second priorities to a second range associated with a second TCU instance, allowing for isolation between ASIL levels. This can allow, for example, processes from the same client to be isolated to different TCU instances at different points in time at which the processes will operate at different levels. In applications in which a given client can execute processes that vary between levels, the given client can select identifiers to provide to the SMMU to correspond to the levels of the processes.

[0027] The SMMU can have multiple groups of TBUs that correspond to associated TCU instances and / or ranges of identifiers. For example, a client can be programmed to route to a particular group of TBUs (and corresponding TCU instances) in response to the client always implementing processes at a level of the group of TBUs; where the levels of the processes can vary over time, the system can use identifiers to control the group of TBUs and / or TCU instances to which the processes are directed, allowing for level isolation.

[0028] For example, where a process of a client always operates at a lower ASIL level, the SMMU can determine, based on an identifier of the client, to use a group of TBUs and / or TCU instance of the lower ASIL level for the process. Where the process always operates at a higher ASIL level, the SMMU can determine to use a group of TBUs and / or TCU instance of the higher ASIL level for the process. Where the ASIL level of the process varies, an identifier (e.g., stream ID) can be used to route the process to a corresponding group of TBUs and / or TCU instance.

[0029] Systems and methods according to the present disclosure can allow for isolation between individual processes associated with any of the individual clients (e.g., clients associated with compute engines and / or accelerators), as appropriate to prevent interference between processes of different ASIL levels. For example, a client can be used to implement an engine such as for video encoding or decoding operations. The isolation can be performed without requiring multiple SoCs, which can reduce hardware requirements for implementing the client, for example reducing size, weight, and / or power requirements.

[0030] Although the present disclosure can be described with respect to example autonomous or semi-autonomous vehicles or machines 400 (alternatively referred to herein as“vehicles 400,”“ego vehicles 400,”“machines 400,” and / or“ego machines 400”), examples of which are described with respect to Figures 4A-4D described, this is not intended to be limiting. For example, the systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled with one or more trailers, aircraft, watercraft, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Further, although the present disclosure can be described with respect to various ASIL level requested SMMU enforcement isolation, this is not intended to be limiting, and the systems and methods described herein can be used for augmented reality, virtual reality, mixed reality, robotics, safety and supervision, autonomous or semi-autonomous machine applications, and / or any other technical space in which virtual memory addresses can be translated to physical memory addresses.

[0031] Reference is made to Figure 1 , Figure 1 An overview of the MMU / SMMU architecture 100 in accordance with some embodiments of the present disclosure is provided. It should be understood that this and other arrangements described herein are set forth only as examples. Other arrangements can be utilized and not just those set forth herein (e.g., machines, interfaces, functions, order, groupings of functions, etc.) and some elements can be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. The various functions described herein as being performed by an entity can be performed by hardware, firmware and / or software. For instance, various functions can be performed by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein can be performed using components, features, and / or functionality similar to the example autonomous vehicle 400 of Figures 4A-4D and / or the example computing device 500 of Figure 5 described herein.

[0032] Architecture 100 can include at least one SoC 105. For example, SoC 105 can include various components and / or circuitry housed and / or located within a single device. SoC 105 and / or one or more components thereof can perform at least one of the various processes and / or techniques described herein. For example, SoC 105 can perform SMMU enforcement isolation described herein. SoC 105 can include at least one memory management unit (MMU) 110 and memory 135. For example, MMU 110 can receive a virtual address from a client and perform one or more translations to determine a corresponding physical address in memory 135.

[0033] MMU 110 can perform one or more memory management functions, such as memory translations (e.g., virtual to physical memory translations) and / or memory access protocols for accessing memory 135. For example, MMU 110 can receive a request (e.g., a translation request) that includes and / or identifies a virtual memory address. Continuing with this example, MMU 110 can perform one or more translation techniques (e.g., table walk, page walk, cache lookup, etc.) to translate the virtual address to a physical address in memory 135. MMU 110 can store and / or maintain one or more caches (e.g., page cache, translation cache, etc.).

[0034] MMU 110 can include at least one processing circuit 115 and at least one memory manager 120. Processing circuit 115 can include one or more processors and one or more memory devices. The one or more memory devices can store instructions, computer code, software, etc., which, when executed by the one or more processors, cause processing circuit 115 to perform at least one of the various processes and / or techniques described herein. For example, processing circuit 115 can receive one or more requests from a client and provide access to various types of information based on the requests. As another example, processing circuit 115 can receive various requests from a client and can identify one or more components of MMU 110 that correspond to the requests. Various components of SoC 105 can be communicatively coupled to one another. For example, processing circuit 115 can be in communication with memory manager 120.

[0035] Memory manager 120 can include at least one translation buffer unit (TBU) 125 and at least one translation control unit (TCU) 130. Memory manager 120 (e.g., TBU 125 and TCU 130) can receive a virtual address and translate the virtual address to a physical address in memory 135. Memory manager 120 can manage (e.g., facilitate access to, modify, adjust, change, rearrange, etc.) memory 135. For example, memory manager 120 can provide access to content and / or information stored in a given physical address of memory 135. Memory manager 120 can use a translation table to translate a virtual memory address to a physical memory address. Memory manager 120 can perform operations such as reading data stored in memory 135, writing data to memory 135, and / or memory allocation of memory 135.

[0036] TBU 125 can interface and / or interact with one or more clients. For example, TBU 125 can interface and / or interact with an engine (e.g., a software accelerator, a hardware accelerator, a client, etc.). TBU 125 can receive one or more requests (e.g., a translation request, a memory access request, etc.) from a client. For example, TBU 125 can receive a translation request from a deep learning accelerator (e.g., a client). Continuing with this example, the translation request can include and / or identify a virtual address for translation. TBU 125 can interface and / or interact with TCU 130 to receive a translation of the virtual address. For example, TCU 130 can perform one or more page walks to access one or more page tables. Continuing with this example, TCU 130 can identify a translation for the virtual address in response to the page walks.

[0037] TBU 125 can receive a translation from TCU 130. For example, TBU 125 can receive a translation for a virtual memory address from TCU 130. TBU 125 can store and / or maintain the translation in one or more caches. For example, TBU 125 can include at least one translation lookaside buffer (TLB). Continuing with this example, the TLB can cache translations between virtual addresses and physical addresses.

[0038] Memory manager 120 can use the translation to provide a memory access request to memory 135. For example, memory manager 120 can use the translated address (e.g., a virtual address translated to a physical address) to request a direct memory access (DMA) transfer with memory 135.

[0039] The MMU 110 can include and / or be integrated with one or more systems. For example, the MMU 110 can include at least one of: a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system containing one or more virtual machines (VMs), a system implemented using edge devices, a system implemented using robots, a system for generating synthetic data, a system for performing simulation operations, a system for performing collaborative content creation of 3D assets, a system for performing conversational AI operations, a system including one or more large language models (LLMs), a system including one or more visual language models (VLMs), a system for performing digital twin operations, a system for performing optical transport simulation, a system for performing deep learning operations, a system implemented at least partially in a data center, and / or a system implemented at least partially using cloud computing resources.

[0040] Figure 2 is an example of an SMMU architecture 200 implemented by the MMU 110 according to some embodiments. For example, the SMMU architecture 200 can be implemented for performing TCU isolation, including operations of TCU and TBU groups. The MMU 110 implementing the SMMU architecture 200 can refer to and / or include SMMU performing isolation and / or various examples described herein. For example, the MMU 110 can order, resolve, filter, direct, and / or otherwise select a given TBU 125 to receive and / or process various translation requests to reduce interference of lower ASIL requests on higher ASIL requests. As shown, the MMU 110 can interface with, interact with, and / or otherwise communicate with one or more clients 210 (shown as client 210a, client 210b, and client 210n). For example, the MMU 110 can receive one or more requests from the clients 210. Continuing this example, the MMU 110 can provide one or more responses to the requests. Figure 2

[0041] The MMU 110 can include at least one instance (e.g., representation, virtual software configuration, replication, reproduction, etc.) of the memory manager 120. For example, as shown, the MMU 110 includes a memory manager 120a (e.g., a first instance) and a memory manager 120n (e.g., a second instance). In this example, the MMU 110 is shown as including two instances of the memory manager 120. The MMU 110 can include any possible number of memory managers 120. A given instance of the memory manager 120 can refer to and / or include at least one SMMU. For example, the memory manager 120a can represent a first SMMU, and the memory manager 120n can represent a second SMMU. Figure 2 ​​

[0042] One or more instances of memory manager 120 can reduce interference between various memory operations requested by clients 210, including, for example, translation requests. For example, memory manager 120a can be associated with high-level ASIL requests. As another example, memory manager 120n can be associated with low-level ASIL requests. In these examples, MMU 110 can forward high-level requests to memory manager 120a, and can forward low-level requests to memory manager 120n. In other words, MMU 110 can reduce interference by direct traffic to different memory managers 120 based on the level of the request.

[0043] MMU 110 can receive one or more requests from clients 210. For example, as shown, MMU 110 can receive at least one request 0 from client 210a. MMU 110 can receive at least one request 1 and at least one request 2 from client 210b. MMU 110 can receive at least one request N from client 210n. Various requests (e.g., request 0, request 1, request 2, and / or request N) can correspond to various ASIL levels. For example, request 1 can refer to and / or include an ASIL B request. Request N can refer to and / or include an ASIL QM request. Figure 2

[0044] As described herein, a client (e.g., an accelerator, a device, an engine, client 210, etc.) can send one or more first requests with a high-level priority and one or more second requests with a low-level priority. For example, client 210b can send one or more request 1 corresponding to a high-level priority. Client 210b can send one or more request 2 corresponding to a low-level priority.

[0045] Clients 210 can store, maintain, and / or otherwise hold at least one identifier 215 (in Figure 2 ​The flow IDs 215 can be used to indicate and / or identify various priority levels (shown as flow IDs 215a, 215b, and 215n). For example, the client 210a can only handle high priority levels. Continuing with this example, the client 210a can send a flow ID 215a including a value, identifier, flag, etc. with and / or after a request to indicate that the request sent by the client 210a corresponds to a high priority level. As another example, the client 210b can handle high priority levels and low priority levels. Continuing with this example, the client 210b can send one or more first flow IDs 215b indicating high priority levels when sending high level requests, and one or more second flow IDs 215b indicating low priority levels when sending low level requests. As an example, the client 210n can only handle low priority levels. Continuing with this example, the client 210n can send a flow ID 215n including a value, identifier, flag, etc. with and / or after a request to indicate that the request sent by the client 210n corresponds to a low priority level. The identifiers 215 can refer to and / or include at least one of a tag, flag, indicator, and / or various identifiable information.

[0046] The MMU 110 can receive one or more requests to execute one or more processes. For example, the MMU 110 can receive request 0 from the client 210a. Continuing with this example, the request 0 can be a request to execute a process (e.g., a memory translation, a memory access, a data input, etc.). The MMU 110 can receive one or more requests consecutively and / or semi-consecutively. For example, the MMU 110 can receive request 1 and request 2 sequentially. The MMU 110 can receive request N at a first general purpose input / output (GPIO) port and request 1 at a second GPIO port.

[0047] The requests (e.g., request 0, request 1, request 2, and / or request N) can specify and / or indicate a given process. For example, the request 0 can correspond to a memory translation. Continuing with this example, the request 0 can include and / or indicate a virtual memory address to indicate that the request 0 is a request for a memory translation (e.g., a process). The request 1 can correspond to a data input request. The request 1 can include a string of data to indicate that the request 1 is a request for a data input (e.g., a process). The request 2 can correspond to a video encoding operation and / or a video decoding operation.

[0048] A request may include one or more identifiers 215. For example, request 0 may include stream ID 215a. As another example, request 1 may include one or more identifiers and / or indicators corresponding to one or more processes (e.g., identifying which process request 1 corresponds to). As another example, a request may include information identifying a given client 210. In this example, request N may include information indicating that the request was sent by client 210. Continuing this example, client 210n may correspond only to low-level requests, and request N includes identifiable information to provide MMU 110 with an indication that client 210n is associated only with low-level requests (e.g., client 210n only sends low-level priority requests). In other words, client 210n may be associated with a constant priority level (e.g., client 210n only sends requests with the same priority level).

[0049] MMU 110 can select one or more memory managers 120, for example, to allocate and / or isolate a given memory manager 120 to reduce interference between high-priority requests and low-priority requests (e.g., high-priority requests go to a first memory manager 120, and low-priority requests go to a second memory manager 120). For example, MMU 110 can select memory manager 120a in response to receiving a request corresponding to a high-priority request. MMU 110 can select memory manager 120a based on a given request including a stream ID 215, which corresponds to a given value within a range of stream IDs 215 indicating a high-priority request.

[0050] like Figure 2 As shown, MMU 110 can select memory manager 120a to receive requests 0 and 1, and select memory manager 120n to receive requests 2 and 2. For example, memory manager 120a may correspond to stream ID 215a included in request 0. MMU 110 can select memory manager 120a based on request 0 including stream ID 215a. MMU 110 can select memory manager 120n to receive request 2 based on request 2 including stream ID 215b corresponding to memory manager 120n.

[0051] The MMU 110 can cause the memory managers 120 to perform one or more operations (e.g., but not limited to, memory translations), for example, to perform one or more memory operations. For example, the MMU 110 can cause the memory manager 120a to perform a memory translation associated with the request 0 by sending and / or forwarding the request 0 to the memory manager 120a. The MMU 110 can send one or more control signals to the TBU 125a to cause the TBU 125a to query a cache to search for a translation of the virtual memory address included in the request 0.

[0052] The memory managers 120 (e.g., the memory manager 120a and the memory manager 120n) can be associated with one or more flow IDs 215. For example, the memory manager 120a can correspond to a first range of values of the flow IDs 215. The memory manager 120n can correspond to a second range of values of the flow IDs 215. The client 210 can include a data structure and / or a database that includes flow IDs 215 corresponding to a high level of priority and / or a low level of priority. For example, the client 210a can know that a given flow ID 215 corresponds to a high level of priority. Continuing this example, the request 0 can correspond to a high level of priority, and the client 210a can attach and / or include the flow ID 215 with a value indicating that the request 0 is a high level of priority.

[0053] The client 210 can send requests with different priority levels. For example, the client 210b can send a first request (e.g., request 1) with a first priority level. The client 210b can send a second request (e.g., request 2) with a second priority level. The priority level of the request 1 can be less than, greater than, and / or equal to the priority level of the request 2. The client 210b can include one or more given flow IDs 215b to indicate whether the request 1 and / or the request 2 is a high level of priority and / or a low level of priority.

[0054] Although some examples described herein can refer to the memory manager 120a as corresponding to a high level of ASIL requests and the memory manager 120n as corresponding to a low level of ASIL requests, these examples are merely illustrative and not limiting.

[0055] Reference is now made to Figure 3Each block of the method 300 described herein includes a computational procedure that can be performed using any combination of hardware, firmware, and / or software. For instance, various functions can be performed by a processor executing instructions stored in memory. These methods can also be embodied in computer-usable instructions stored on a computer storage medium. These methods can be provided by a standalone application, a service, or a plug-in to a hosted service, independent of or in combination with other products (to name only a few examples). Further, the method 300 is described with respect to the architecture 100 of Figure 1 However, these methods can additionally or alternatively be performed by any one system or combination of systems, including but not limited to those described herein.

[0056] Figure 3 is a flowchart illustrating a method 300 for performing SMMU execution isolation for ASIL requests in accordance with some embodiments of the present disclosure. The method 300 can be performed in response to receiving a request (e.g., from a client, an application, and / or an accelerator) to perform a memory operation. The method 300 can be performed in response to detecting memory operations and / or requests for memory operations having different priority levels, and / or at least a first of the memory operations being a high-priority operation.

[0057] The method 300 includes receiving one or more requests to perform a process at block 302. For example, the MMU can receive at least one request from a client. Continuing with this example, the request can be a request to perform a process (e.g., a memory translation, a data input, a data read, etc.). The MMU can receive one or more requests at block 302. For example, the MMU can receive a first request from a first client and a second request from a second client. Continuing with this example, the MMU can receive the first request and the second request sequentially. In other examples, the MMU can receive the second request prior to the first request.

[0058] The method 300 includes selecting a target memory manager from a plurality of memory managers at block 304. For example, the MMU can select at least one memory manager (e.g., a target memory manager) based on the request received at block 302. As another example, the request received at block 302 can include at least one identifier (e.g., a stream ID) indicating a priority level of the request. In this example, the request can include a stream ID indicating that the request is a high-level priority. Continuing with this example, the MMU can select a given memory manager (e.g., a target memory manager) corresponding to the high-level priority request. As another example, the MMU can select a second given memory manager based at least on the request including a stream ID corresponding to a low-level priority.

[0059] A memory manager can refer to and / or include one or more instances of a TCU and / or TBU. For example, a first memory manager can include a first TCU and / or one or more TBU groups. Continuing this example, the first TCU can refer to a first instance of a TCU. In this example, the first TCU can maintain and / or store a first cache related to translations of a given priority level. As another example, a second memory manager can include a second TCU and / or one or more second TBU groups. In this example, the second TCU can maintain and / or store a second cache related to translations of a second given priority level.

[0060] The above examples regarding the first TCU and the second TCU illustrate examples of the SMMU performing isolation in that the MMU can select the first TCU in response to a request having a given priority level, and the MMU can select the second TCU in response to a request having a second given priority level (e.g., the MMU directs the request to the corresponding TCU instance). As another example, the MMU can reduce and / or eliminate interference by acting as a gateway between clients and the memory managers (e.g., both low priority level requests and high priority level requests are forwarded to the same TCU instance). In other words, the MMU can receive a request, and subsequently forward the request to a corresponding TCU instance based on an identifier included in the request.

[0061] The method 300 includes, at block 306, causing the memory manager to perform a memory translation. For example, the MMU can forward and / or send the request received at block 302 to the memory manager selected at block 304. Continuing this example, the sending and / or subsequent receiving of the request can cause the target memory manager to perform a memory translation of the virtual memory address included in the request.

[0062] The systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled with one or more trailers, aircraft, watercraft, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Further, the systems and methods described herein can be used for various purposes, such as, for example and without limitation, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twin, autonomous or semi-autonomous machine applications, deep learning, environment simulation, object or actor simulation and / or digital twin, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0063] Embodiments of the present disclosure can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, marine systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in data centers, systems for performing conversational AI operations, systems implementing at least one large language model (LLM), systems implementing at least one visual language model (VLM), systems for hosting real-time streaming applications, systems for presenting one or more of virtual reality content, augmented reality content, or mixed reality content, systems for performing light transport simulation, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0064] Figure 4AFIG. 1 is a diagram of an example autonomous or semi-autonomous vehicle or machine 400 in accordance with some embodiments of the present disclosure. Autonomous vehicle 400 (alternatively referred to herein as “vehicle 400”) can include, but is not limited to, a passenger vehicle such as a car, truck, bus, ambulance, shuttle, electric or motorized bicycle, motorcycle, fire truck, police car, ambulance, boat, construction vehicle, underwater vehicle, robotic vehicle, drone, airplane, vehicle coupled to a trailer (e.g., a semi-truck for transporting goods), and / or other types of vehicles (e.g., driverless and / or capable of accommodating one or more passengers). Autonomous vehicles are often described in terms of levels of automation as defined by a division of the United States Department of Transportation, the National Highway Traffic Safety Administration (NHTSA), and the Society of Automotive Engineers (SAE) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (Standard No. J3016-201506 published June 15, 2018, Standard No. J3016-201609 published September 30, 2016, and prior and future versions of this standard). Vehicle 400 is capable of implementing functionality that complies with one or more of Levels 3-5 of autonomous driving. Vehicle 400 is capable of implementing functionality that complies with one or more of Levels 1-5 of autonomous driving. For example, depending on the embodiment, vehicle 400 is capable of implementing the ability for driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). As used herein, the term “autonomous” can include any and / or all types of autonomy of 400 or other machines, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assistance autonomous, semi-autonomous, primarily autonomous, or other designations.

[0065] Vehicle 400 can include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. Vehicle 400 can include a propulsion system 450 such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. Propulsion system 450 can be connected to a drivetrain of vehicle 400 that can include a transmission in order to effectuate propulsion of vehicle 400. Propulsion system 450 can be controlled in response to receiving a signal from a throttle / accelerator 452.

[0066] A steering system 454, which can include a steering wheel, can be used to steer the vehicle 400 (e.g., along a desired path or route) while the propulsion system 450 is operating (e.g., while the vehicle is in motion). The steering system 454 can receive signals from a steering actuator 456. For full automation (level 5) functionality, the steering wheel can be optional.

[0067] A braking sensor system 446 can be used to operate the vehicle brakes in response to receiving signals from a braking actuator 448 and / or a braking sensor.

[0068] One or more controllers 436, which can include one or more system on a chip (SoC) 404 Figure 4C The one or more controllers 436 can provide signals (e.g., signals representing commands) to one or more components and / or systems of the vehicle 400. For example, the one or more controllers can send signals to operate the vehicle brakes via one or more braking actuators 448, to operate the steering system 454 via one or more steering actuators 456, to operate the propulsion system 450 via one or more throttle / accelerators 452. The one or more controllers 436 can include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 400. The one or more controllers 436 can include a first controller 436 for autonomous driving functionality, a second controller 436 for functional safety functionality, a third controller 436 for artificial intelligence functionality (e.g., computer vision), a fourth controller 436 for infotainment functionality, a fifth controller 436 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 436 can handle two or more of the above functionalities, two or more controllers 436 can handle a single functionality, and / or any combination thereof.

[0069] One or more controllers 436 can provide signals for controlling one or more components and / or systems of vehicle 400 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data can be received from, for example and without limitation, a global navigation satellite system (“GNSS”) sensor 458 (e.g., a global positioning system sensor), a RADAR sensor 460, an ultrasonic sensor 462, a LIDAR sensor 464, an inertial measurement unit (IMU) sensor 466 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 796, a stereo camera 468, a wide-angle camera 470 (e.g., a fisheye camera), an infrared camera 472, a surround camera 474 (e.g., a 360-degree camera), a long-range and / or mid-range camera 498, a speed sensor 444 (e.g., for measuring the speed of vehicle 400), a vibration sensor 442, a steering sensor 440, a brake sensor (e.g., as part of a brake sensor system), and / or other sensor types.

[0070] One or more of controllers 436 can receive inputs (e.g., represented by input data) from an instrument cluster 432 of vehicle 400 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 434, an audible annunciator, a speaker, and / or via other components of vehicle 400. These outputs can include information such as vehicle speed, velocity, time, map data (e.g., a high-definition (“HD”) map 422 of Figure 4C

[0071] ​The vehicle 400 further includes a network interface 424 that can communicate over one or more networks using one or more wireless antennas 426 and / or modems. For example, the network interface 424 can be capable of communicating over Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”), and / or the like. The one or more wireless antennas 426 can also enable communication between objects (e.g., vehicles, mobile devices, and / or the like) in the implementation environment using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, and / or the like and / or one or more low power wide area networks (LPWANs) such as LoRaWAN, SigFox, and / or the like.

[0072] Figure 4B For example autonomous vehicle 400 for Figure 4A An example camera position and field of view of the example autonomous vehicle 400 according to some embodiments of the present disclosure. The cameras and respective fields of view are one example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras can be included and / or these cameras can be located at different positions on the vehicle 400.

[0073] The camera type for the cameras can include, but is not limited to, a digital camera that can be suitable for use with components and / or systems of the vehicle 400. The cameras can operate at Automotive Safety Integrity Level (ASIL) B and / or at another ASIL. The camera type can have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, and / or the like, depending on the embodiment. The cameras can be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array can include a Red-White-White-White (RCCC) color filter array, a Red-White-White-Blue (RCCB) color filter array, a Red-Blue-Green-White (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a clear pixel camera such as a camera with a RCCC, RCCB, and / or RBGC color filter array can be used in efforts to improve light sensitivity.

[0074] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more (e.g., all) of the cameras can simultaneously record and provide image data (e.g., video).

[0075] One or more of the cameras can be mounted in mounting assemblies such as custom designed (three-dimensional ("3D") printed) assemblies in order to cut off stray light and reflections from within the car (such as reflections from the dashboard reflected in the windshield mirror) that can interfere with the image data capture capabilities of the cameras. With respect to wing mirror mounting assemblies, the wing mirror assemblies can be custom 3D printed such that the camera mounting plates match the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side view cameras, one or more cameras can also be integrated into the four pillars of each corner of the cab.

[0076] Cameras with fields of view that include the portion of the environment in front of the vehicle 400 (e.g., front-facing cameras) can be used for surround view to help identify the forward path and obstacles, and to assist in providing information critical to generating an occupancy grid and / or determining a preferred vehicle path with the help of one or more controllers 436 and / or control SoCs. Front-facing cameras can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. Front-facing cameras can also be used for ADAS functions and systems, including lane departure warning (LDW), adaptive cruise control (ACC), and / or other functions such as traffic sign recognition.

[0077] A wide variety of cameras can be used in the front-facing configuration, including, for example, monocular camera platforms including complementary metal-oxide-semiconductor ("CMOS") color imagers. Another example can be a wide-angle camera 470, which can be used to perceive objects (e.g., pedestrians, intersection traffic, or bicycles) entering the field of view from the periphery. Although Figure 4B Although only one wide-angle camera is illustrated in FIG. 4, there can be any number (including zero) of wide-angle cameras 470 on the vehicle 400. In addition, any number of long-range cameras 498 (e.g., long-view stereo camera pairs) can be used for depth-based object detection, especially for objects for which a neural network has not been trained. Long-range cameras 498 can also be used for object detection and classification and basic object tracking.

[0078] Any number of stereo cameras 468 can also be included in the front-facing configuration. In at least one embodiment, one or more of the stereo cameras 468 can include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor with integrated controller area network ("CAN") or Ethernet interfaces and a field programmable gate array ("FPGA") on a single chip. Such a unit can be used to generate a 3D map of the vehicle's environment, including distance estimates for all points in the image. Alternative stereo cameras 468 can include compact stereo vision sensors that can include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 468 can be used in addition to or instead of those described herein.

[0079] Cameras with fields of view that include portions of the environment to the side of the vehicle 400 (e.g., side-view cameras) can be used for surround view, providing information used to create and update the occupancy grid and to generate side-crash collision warnings. For example, surround cameras 474 (e.g., four surround cameras 474 as shown in FIG. 15) can be placed on the vehicle 400. The surround cameras 474 can include wide-view cameras 470, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be placed on the front, back, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 474 (e.g., left, right, and back), and can utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround view camera. Figure 4B

[0080] Cameras with fields of view that include portions of the environment to the rear of the vehicle 400 (e.g., rear-view cameras) can be used for assist parking, surround view, rear collision warnings, and to create and update the occupancy grid. A wide variety of cameras can be used, including but not limited to cameras that are also suitable as front-facing cameras (e.g., long- and / or mid-range cameras 498, stereo cameras 468, infrared cameras 472, etc.) as described herein.

[0081] Figure 4C For use in accordance with some embodiments of the present disclosure Figure 4A ​FIG. 1 is a block diagram of an example system architecture of an example autonomous vehicle 400. It should be understood that this arrangement and other arrangements described herein are set forth merely as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) can be used in addition to or instead of those shown, and some elements can be wholly omitted. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components, and in any suitable combinations and locations. Various functions described herein as being performed by an entity can be implemented in hardware, firmware, and / or software. For instance, various functions can be implemented by a processor executing instructions stored in a memory.

[0082] Figure 4C Each of the components, features, and systems of vehicle 400 are illustrated as being connected via bus 402. Bus 402 can include a controller area network (CAN) data interface (alternatively referred to herein as a "CAN bus"). The CAN can be a network within vehicle 400 that is used to assist in controlling various features and functions of vehicle 400, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have tens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, revolutions per minute (RPM) of the engine, button positions, and / or other vehicle status indicators. The CAN bus can be ASIL B compliant.

[0083] Although bus 402 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet can be used in addition to or instead of a CAN bus. Further, although bus 402 is represented with a single line, this is not intended to be limiting. For example, there can be any number of buses 402, which can include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses that use different protocols. In some examples, two or more buses 402 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 402 can be used for collision avoidance functions, and a second bus 402 can be used for drive control. In any example, each bus 402 can communicate with any component of vehicle 400, and two or more buses 402 can communicate with the same components. In some examples, each SoC 404, each controller 436, and / or each computer within the vehicle can have access to the same input data (e.g., inputs from sensors of vehicle 400), and can be connected to a common bus, such as a CAN bus.

[0084] The vehicle 400 can include one or more controllers 436, such as those described herein with respect to Figure 4A The controllers 436 can be used for a wide variety of functions. The controllers 436 can be coupled to any other distinct components and systems of the vehicle 400 and can be used for control of the vehicle 400, artificial intelligence of the vehicle 400, infotainment for the vehicle 400, and / or the like.

[0085] The vehicle 400 can include one or more system on a chip (SoC) 404. The SoC 404 can include a CPU 406, a GPU 408, a processor 410, a cache 412, an accelerator 414, a data store 416, and / or other components and features not illustrated. The SoC 404 can be used to control the vehicle 400 in a wide variety of platforms and systems. For example, one or more SoCs 404 can be used in a system (such as a system of the vehicle 400) in conjunction with an HD map 422 that can obtain map refreshes and / or updates from one or more servers (such as the one or more servers 478) via a network interface 424. Figure 4D

[0086] The CPU 406 can include a CPU cluster or CPU complex (alternatively referred to herein as a “CCPLEX”). The CPU 406 can include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 406 can include eight cores in a coherent multi-processor configuration. In some embodiments, the CPU 406 can include four dual-core clusters with each cluster having a dedicated L2 cache (such as a 2 MB L2 cache). The CPU 406 (e.g., the CCPLEX) can be configured to support simultaneous cluster operation such that any combination of clusters of the CPU 406 can be active at any given time.

[0087] The CPU 406 can implement power management capabilities including one or more of the following features: individual hardware blocks can be automatically clock-gated when idle to save dynamic power; each core clock can be gated when the core is not actively executing instructions due to execution of WFI / WFE instructions; each core can be independently power-gated; each core cluster can be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster can be independently power-gated when all cores are power-gated. The CPU 406 can further implement an enhanced algorithm for managing power states in which the allowed power states and the desired wake-up time are specified, and the hardware / microcode determines the best power state for the cores, clusters, and CCPLEX to enter. The processing cores can support a simplified power state entry sequence in software, with the work offloaded to microcode.​

[0088] GPU 408 can include an integrated GPU (alternatively referred to herein as an “iGPU”). GPU 408 can be programmable and efficient for parallel workloads. In some examples, GPU 408 can use an enhanced tensor instruction set. GPU 408 can include one or more streaming microprocessors, where each streaming microprocessor can include an LI cache (e.g., an LI cache having at least 96 KB of storage capacity), and two or more of the streaming microprocessors can share an L2 cache (e.g., an L2 cache having 512 KB of storage capacity). In some embodiments, GPU 408 can include at least eight streaming microprocessors. GPU 408 can use a compute application programming interface (API). In addition, GPU 408 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA).

[0089] In the case of automotive and embedded uses, GPU 408 can be power-optimized for best performance. For example, GPU 408 can be fabricated on a fin field effect transistor (FinFET). However, this is not intended to be limiting, and GPU 408 can be fabricated using other semiconductor fabrication processes. Each streaming microprocessor can incorporate several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a thread warp scheduler, a dispatch unit, and / or a 64 KB register file. In addition, the streaming microprocessor can include independent parallel integer and floating point data paths to provide efficient execution of workloads with a mix of compute and address compute. The streaming microprocessor can include independent thread scheduling capabilities to allow for more fine-grained synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined LI data cache and shared memory unit to improve performance while simplifying programming.

[0090] GPU 408 can include a high bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem that provides approximately 900 GB / s of peak memory bandwidth in some examples. In some examples, in addition to or alternatively from HBM memory, a synchronous graphics random access memory (SGRAM) can be used, such as a fifth generation graphics double data rate synchronous random access memory (GDDR5).

[0091] GPU 408 can include a unified memory technology that includes access counters to allow memory pages to be migrated more precisely to the processors that access them most frequently, improving efficiency of memory ranges shared between processors. In some examples, address translation services (ATS) support can be used to allow GPU 408 to directly access CPU 406 page tables. In such examples, when a GPU 408 memory management unit (MMU) experiences a miss, an address translation request can be sent to CPU 406. In response, CPU 406 can look up the virtual-to-physical mapping for the address in its page tables and send the translation back to GPU 408. In this way, the unified memory technology can allow a single unified virtual address space for memory of both CPU 406 and GPU 408, simplifying GPU 408 programming and porting applications to GPU 408.

[0092] Further, GPU 408 can include access counters that can track how frequently GPU 408 accesses other processors' memory. The access counters can help ensure that memory pages are migrated to the physical memory of the processor that accesses these pages most frequently.

[0093] SoC 404 can include any number of caches 412, including those described herein. For example, caches 412 can include an L3 cache available to both CPU 406 and GPU 408 (e.g., connected to both CPU 406 and GPU 408). Caches 412 can include a write-back cache that can track the state of a line, for example, by using a cache coherency protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache can include 4 MB or more, although smaller cache sizes can also be used.

[0094] SoC 404 can include one or more arithmetic logic units (ALUs) that can be used to perform processing with respect to any of a variety of tasks or operations of vehicle 400, such as processing a DNN. Further, SoC 404 can include a floating point unit (FPU) or other mathematical co-processor or digital co-processor type for performing mathematical operations within the system. For example, SoC 404 can include one or more FPUs integrated as execution units within CPU 406 and / or GPU 408.

[0095] The SoC 404 can include one or more accelerators 414 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 404 can include a hardware acceleration cluster that can include optimized hardware accelerators and / or a large on-chip memory. This large on-chip memory (e.g., 4 MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement the GPU 408 and offload some of the tasks of the GPU 408 (e.g., freeing up more cycles of the GPU 408 for performing other tasks). As one example, the accelerators 414 can be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs), etc.) that are stable enough to be easily controlled for acceleration. As used herein, the term “CNN” can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0096] The accelerators 414 (e.g., hardware acceleration cluster) can include a deep learning accelerator (DLA). The DLA can include one or more tensor processing units (TPUs) that can be configured to provide an additional 100 billion operations per second for deep learning applications and inferencing. The TPU can be an accelerator that is configured to perform and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). The DLA can be further optimized for a specific set of neural network types and floating point operations and inferencing. The design of the DLA can provide higher performance per mm than a general purpose GPU and far exceeds the performance of a CPU. The TPU can perform several functions, including single instance convolution functions, support for INT8, INT16, and FP16 data types for both features and weights, for example, and post-processor functions.

[0097] The DLA can perform neural networks, especially CNNs, on processed or unprocessed data for any of a wide variety of functions, such as and not limited to: CNNs for object recognition and detection using data from a camera sensor; CNNs for distance estimation using data from a camera sensor; CNNs for emergency vehicle detection and identification and detection using data from a microphone; CNNs for face recognition and vehicle owner identification using data from a camera sensor; and / or CNNs for safety and / or safety related events.

[0098] The DLA can perform any of the functions of the GPU 408, and by using an inferencing accelerator, the designer can target the DLA or the GPU 408 for any function. For example, the designer can focus the processing and floating point operations of the CNNs on the DLA and leave other functions to the GPU 408 and / or other accelerators 414.

[0099] Accelerator 414 (e.g., hardware acceleration cluster) can include a programmable vision accelerator (PVA), which can be alternatively referred to herein as a computer vision accelerator. The PVA can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA can include, for example and without limitation, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0100] The RISC cores can interact with image sensors (e.g., image sensors of any of the cameras described herein), image signal processors, and / or the like. Each of the RISC cores can include any number of memories. Depending on the embodiment, the RISC cores can use any of several protocols. In some examples, the RISC cores can execute a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include an instruction cache and / or a tightly coupled RAM.

[0101] The DMA can enable components of the PVA to access system memory independently of the CPU 406. The DMA can support any number of features to provide optimizations to the PVA, including, but not limited to, supporting multi-dimensional addressing and / or circular addressing. In some examples, the DMA can support addressing up to six or more dimensions, which can include block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.

[0102] The vector processors can be programmable processors that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystems can operate as the main processing engines of the PVA and can include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., VMEM). The VPU core can include a digital signal processor, such as, for example, a single instruction multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and rate.

[0103] Each of the vector processors can include an instruction cache and can be coupled to a dedicated memory. As a result, in some examples, each of the vector processors can be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA can be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA can execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA can execute different computer vision algorithms on the same image simultaneously, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs can be included in the hardware acceleration cluster, and any number of vector processors can be included in each of the PVAs. Moreover, the PVAs can include additional error-correcting code (ECC) memory to enhance overall system security.

[0104] The accelerator 414 (e.g., hardware acceleration cluster) can include an on-chip computer vision network and SRAM to provide high bandwidth, low latency SRAM for the accelerator 414. In some examples, the on-chip memory can include at least 4 MB of SRAM composed of, for example and without limitation, eight field-programmable memory blocks, which can be accessed by both the PVA and the DLA. Each pair of memory blocks can include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVA and the DLA can access the memory via a backbone that provides high-speed memory access to the PVA and the DLA. The backbone can include an on-chip computer vision network that interconnects the PVA and the DLA to the memory (e.g., using APB).

[0105] The on-chip computer vision network can include an interface that determines that both the PVA and the DLA provide ready and valid signals before sending any control signals / addresses / data. Such an interface can provide separate phases and separate channels for sending control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can comply with ISO 26262 or IEC 61508 standards, but other standards and protocols can also be used.

[0106] In some examples, the SoC 404 can include a real-time ray tracing hardware accelerator, such as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine locations and extents of objects (e.g., within a world model) in order to generate real-time visualizations simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison with LIDAR data for purposes of localization and / or other functions, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) can be used to perform one or more ray tracing related operations.

[0107] The accelerator 414 (e.g., hardware accelerator cluster) has a wide range of autonomous driving uses. The PVA can be a programmable vision accelerator that can be used for key processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithm domains that require predictable processing, low power, and low latency. In other words, the PVA performs well on semi-dense or dense regular computations, and even on small data sets that require predictable runtimes with low latency and low power. Thus, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are effective at object detection and integer math operations.

[0108] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. In some examples, a semi-global matching based algorithm can be used, although this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require instant motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA can perform computer stereo vision functions on input from two monocular cameras.

[0109] In some examples, the PVA can be used to perform dense optical flow. According to processing raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used for time-of-flight depth processing, such as by processing raw time-of-flight data to provide processed time-of-flight data.

[0110] The DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence metric for each object detection. Such a confidence value can be interpreted as a probability, or as providing a relative "weight" for each detection compared to other detections. The confidence value enables the system to make further decisions about which detections should be considered true positive detections and not false positive detections. For example, the system can set a threshold for confidence, and only consider detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections would cause the vehicle to automatically perform an emergency brake, which is obviously undesirable. Thus, only the most confident detections should be considered a trigger for AEB. The DLA can run a neural network for regression of a confidence value. The neural network can take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), inertial measurement unit (IMU) sensor 466 outputs related to vehicle 400 orientation, distance, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LIDAR sensor 464 or RADAR sensor 460), etc.

[0111] SoC 404 can include one or more data stores 416 (e.g., memory). Data stores 416 can be on-chip memory of SoC 404, which can store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and safety, data stores 416 can be large enough in capacity to store multiple instances of a neural network. Data stores 412 can include L2 or L3 cache 412. References to data stores 416 can include references to memory associated with PVAs, DLAs, and / or other accelerators 414 as described herein.

[0112] SoC 404 can include one or more processors 410 (e.g., embedded processors). The processors 410 can include a boot and power management processor, which can be a specialized processor and subsystem for handling boot power and management functions and related security implementations. The boot and power management processor can be part of the SoC 404 boot sequence and can provide run-time power management services. The boot power and management processor can provide clock and voltage programming, auxiliary system low power state transitions, SoC 404 thermal and temperature sensor management, and / or SoC 404 power state management. Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 404 can use the ring oscillator to detect the temperature of the CPU 406, GPU 408, and / or accelerator 414. If it is determined that the temperature exceeds a threshold, the boot and power management processor can enter a temperature fault routine and place the SoC 404 in a lower power state and / or place the vehicle 400 in a driver safe park mode (e.g., safely park the vehicle 400).

[0113] The processors 410 can further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows for full hardware support for multi-channel audio over multiple interfaces and a range of widely flexible audio I / O interfaces. In some examples, the audio processing engine is a specialized processor core with a digital signal processor with dedicated RAM.

[0114] The processors 410 can further include an always-on processor engine, which can provide the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine can include a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0115] The processors 410 can further include a safety cluster engine, which includes a specialized processor subsystem that handles safety management for automotive applications. The safety cluster engine can include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In safety mode, the two or more cores can operate in lockstep mode and act as a single core with comparison logic that detects any differences between their operations.

[0116] The processors 410 can further include a real-time camera engine, which can include a specialized processor subsystem for handling real-time camera management.

[0117] The processor 410 can further include a high dynamic range signal processor, which can include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0118] The processor 410 can include a video image compositor, which can be a processing block (e.g., implemented on a microprocessor), that implements video post-processing functions needed by the video playback application to produce the final image for the player window. The video image compositor can perform lens distortion correction on the wide-angle camera 470, surround camera 474, and / or on the cab-in monitor camera sensors. The cab-in monitor camera sensors are preferably monitored by a neural network running on another instance of the advanced SoC, configured to recognize cab-in events and respond accordingly. The cab-in system can perform lip reading to activate mobile phone services and place a call, dictate an email, change the vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode, and are disabled otherwise.

[0119] The video image compositor can include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, where motion is present in the video, the noise reduction appropriately weights the spatial information, reducing the weight of information provided by neighboring frames. Where the image or portions of the image do not include motion, the temporal noise reduction performed by the video image compositor can use information from previous images to reduce noise in the current image.

[0120] The video image compositor can also be configured to perform stereo correction on input stereo lens frames. The video image compositor can further be used for user interface composition when the operating system desktop is in use and the GPU 408 does not need to continuously render new surfaces. Even when the GPU 408 is powered on and active, doing 3D rendering, the video image compositor can be used to offload the GPU 408 to improve performance and responsiveness.

[0121] The SoC 404 can further include a Mobile Industry Processor Interface (MIPI) camera serial interface for receiving video and input from the cameras, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions. The SoC 404 can further include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a particular role.

[0122] SoC 404 can further include a wide range of peripheral device interfaces to enable communication with peripherals, audio codecs, power management, and / or other devices. SoC 404 can be used to process data from cameras (connected over Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., LIDAR sensor 464, RADAR sensor 460, etc. that can be connected over Ethernet), data from bus 402 (e.g., speed of vehicle 400, steering wheel position, etc.), data from GNSS sensor 458 (connected over Ethernet or CAN bus). SoC 404 can further include a dedicated high-performance mass storage controller, which can include their own DMA engine, and which can be used to free up CPU 406 from routine data management tasks.

[0123] SoC 404 can be an end-to-end platform with a flexible architecture that spans automation levels 3-5, providing an integrated functional safety architecture for a platform that leverages and efficiently uses computer vision and ADAS technology to achieve diversity and redundancy, along with deep learning tools. SoC 404 can be faster, more reliable, and even more energy and space efficient than conventional systems. For example, accelerators 414, when combined with CPU 406, GPU 408, and data storage 416, can provide a fast and efficient platform for level 3-5 autonomous vehicles.

[0124] The technology thus provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs that can be configured using high-level programming languages such as the C programming language to perform a wide variety of processing algorithms across a wide variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to, for example, execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for on-board ADAS applications and for practical level 3-5 autonomous vehicles.

[0125] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the technology described herein allows multiple neural networks to be executed simultaneously and / or sequentially, and the results to be combined together to achieve level 3-5 autonomous driving functionality. For example, a CNN executed on a DLA or dGPU (e.g., GPU 420) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs for which a neural network has not been specifically trained. The DLA can further include a neural network that is able to recognize, interpret, and provide a semantic understanding of the sign, and pass that semantic understanding to a path planning module running on the CPU complex.

[0126] As another example, multiple neural networks can be run simultaneously as required for level 3, 4, or 5 driving. For example, a warning sign consisting of the words "Flashing lights indicate icy conditions" along with a light can be interpreted by several neural networks independently or collectively. The sign itself can be recognized by a first deployed neural network (e.g., a trained neural network) as a traffic sign, the text "Flashing lights indicate icy conditions" can be interpreted by a second deployed neural network that informs the vehicle's path planning software (preferably executing on the CPU complex) that icy conditions exist when flashing lights are detected. The flashing lights can be recognized by operating a third deployed neural network over multiple frames that informs the vehicle's path planning software of the presence (or absence) of flashing lights. All three neural networks can be run simultaneously, for example, within the DLA and / or on the GPU 408.

[0127] In some examples, a CNN for face recognition and owner recognition can use data from the camera sensors to recognize the presence of an authorized driver and / or owner of the vehicle 400. A processing engine always on the sensor can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in a safe mode, disable the vehicle when the owner leaves the vehicle. In this way, the SoC 404 provides security against theft and / or carjacking.

[0128] In another example, a CNN for emergency vehicle detection and recognition can use data from the microphones 796 to detect and recognize emergency vehicle sirens. In contrast to conventional systems that use a general classifier to detect sirens and manually extract features, the SoC 404 uses a CNN to classify ambient and urban sounds as well as to classify visual data. In a preferred embodiment, a CNN running on the DLA is trained to recognize the relative closing speed of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to recognize emergency vehicles specific to the local area in which the vehicle is operating as recognized by the GNSS sensor 458. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when in the United States, the CNN will seek to recognize sirens that are only North American. Once an emergency vehicle is detected, a control program can be used to execute an emergency vehicle safety routine that slows the vehicle, pulls over to the side of the road, stops the vehicle, and / or idles the vehicle until the emergency vehicle passes, with the assistance of the ultrasonic sensors 462.

[0129] The vehicle can include a CPU 418 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 404 via a high-speed interconnect (e.g., PCIe). The CPU 418 can include, for example, an X86 processor. The CPU 418 can be used to perform any of a wide variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and the SoC 404, and / or monitoring the status and health of the controller 436 and / or infotainment SoC 430.

[0130] The vehicle 400 can include a GPU 420 (e.g., a discrete GPU or dGPU) that can be coupled to the SoC 404 via a high-speed interconnect (e.g., NVIDIA’s NVLINK). The GPU 420 can provide additional artificial intelligence functionality, for example, by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on input (e.g., sensor data) from sensors of the vehicle 400.

[0131] The vehicle 400 can further include a network interface 424 that can include one or more wireless antennas 426 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 424 can be used to enable wireless connections through the Internet with a cloud (e.g., with the server 478 and / or other network devices), with other vehicles, and / or with computing devices (e.g., client devices of passengers). For communication with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across a network and through the Internet). The direct link can be provided using a car-to-car communication link. The car-to-car communication link can provide the vehicle 400 with information about vehicles that are approaching the vehicle 400 (e.g., vehicles in front of, to the side of, and / or behind the vehicle 400). This functionality can be part of a cooperative adaptive cruise control functionality of the vehicle 400.

[0132] The network interface 424 can include a SoC that provides modulation and demodulation functionality and enables the controller 436 to communicate over a wireless network. The network interface 424 can include a radio frequency front end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. The frequency conversion can be performed through well-known processes, and / or can be performed using a super-heterodyne process. In some examples, the radio frequency front end functionality can be provided by a separate chip. The network interface can include wireless functionality for communication over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0133] The vehicle 400 can further include a data store 428, which can include off-chip (e.g., off-SoC 404) storage. The data store 428 can include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disks, and / or other components and / or devices that can store data for at least one bit.

[0134] The vehicle 400 can further include a GNSS sensor 458. The GNSS sensor 458 (e.g., GPS, assisted GPS sensor, differential GPD (DGPS) sensor, etc.) is used to assist in mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 458 can be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.

[0135] The vehicle 400 can further include a RADAR sensor 460. The RADAR sensor 460 can be used by the vehicle 400 for long-range vehicle detection, even in darkness and / or adverse weather conditions. The RADAR functional safety level can be ASIL B. The RADAR sensor 460 can use the CAN and / or bus 402 (e.g., to send data generated by the RADAR sensor 460) for control as well as access to object tracking data, in some examples, Ethernet for access to raw data. A wide variety of RADAR sensor types can be used. For example and without limitation, the RADAR sensor 460 can be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor is used.

[0136] The RADAR sensor 460 can include different configurations, such as long-range with narrow field of view, short-range with wide field of view, short-range side coverage, and so on. In some examples, long-range RADAR can be used for adaptive cruise control functionality. Long-range RADAR systems can provide a wide field of view (e.g., 250 m range) implemented through two or more independent scans. The RADAR sensor 460 can help distinguish between static and moving objects, and can be used by the ADAS system for emergency brake assist and forward collision warning. The long-range RADAR sensor can include a single-station multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas, as well as high-speed CAN and FlexRay interfaces. In examples with six antennas, the central four antennas can create focused beam patterns designed to record the surroundings of the vehicle 400 at higher speed with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, making it possible to quickly detect vehicles entering or leaving the lane of the vehicle 400.

[0137] As one example, a mid-range RADAR system can include a range of up to 460 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 450 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor the rear and the blind spot next to the vehicle.

[0138] A short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assist.

[0139] The vehicle 400 can further include ultrasonic sensors 462. The ultrasonic sensors 462, which can be placed on the front, rear, and / or sides of the vehicle 400, can be used for parking assist and / or to create and update an occupancy grid. A wide variety of ultrasonic sensors 462 can be used, and different ultrasonic sensors 462 can be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensors 462 can operate at an ASIL B functional safety level.

[0140] The vehicle 400 can include LIDAR sensors 464. The LIDAR sensors 464 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensors 464 can be at an ASIL B functional safety level. In some examples, the vehicle 400 can include multiple LIDAR sensors 464 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0141] In some examples, the LIDAR sensors 464 can be capable of providing a list of objects and their distances for a 360-degree field of view. A commercially available LIDAR sensor 464 can have, for example, an advertised range of approximately 400 m, a precision of 2 cm - 3 cm, and support for a 400 Mbps Ethernet connection. In some examples, one or more flush-mounted LIDAR sensors 464 can be used. In such examples, the LIDAR sensors 464 can be implemented as small devices that can be embedded into the front, rear, sides, and / or corners of the vehicle 400. In such examples, the LIDAR sensors 464 can provide a field of view of up to 120 degrees horizontal and 35 degrees vertical, with a range of 200 m, even for low reflectivity objects. Front-mounted LIDAR sensors 464 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0142] In some examples, LIDAR technology such as 3D Flash LIDAR can also be used. 3D Flash LIDAR uses a flash of laser light as a source of emission to illuminate the vehicle’s surroundings up to about 200 m. The flash LIDAR unit includes a receptor that records laser pulse send times and reflected light on each pixel, which in turn corresponds to a range from the vehicle to the object. Flash LIDAR can allow for generation of highly accurate and distortion-free images of the surroundings with each laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the vehicle 400. Available 3D flash LIDAR systems include solid-state 3D staring array LIDAR cameras (e.g., non-scanning LIDAR devices) that have no moving parts other than fans. The flash LIDAR device can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of 3D range point clouds and co-registered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 464 can be less susceptible to motion blur, vibration, and / or jostling.

[0143] The vehicle can further include an IMU sensor 466. In some examples, the IMU sensor 466 can be located at the center of the rear axle of the vehicle 400. The IMU sensor 466 can include, for example and without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, the IMU sensor 466 can include an accelerometer and a gyroscope, for example in a six-axis application, and an accelerometer, a gyroscope, and a magnetometer in a nine-axis application.

[0144] In some embodiments, the IMU sensor 466 can be implemented as a micro-electro-mechanical systems (MEMS) based high-performance GPS-aided inertial navigation system (GPS / INS) that combines MEMS inertial sensors, high-sensitivity GPS receivers, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude. As such, in some examples, the IMU sensor 466 can enable the vehicle 400 to estimate heading without input from a magnetic sensor by directly observing changes in velocity from GPS to the IMU sensor 466 and correlating them. In some examples, the IMU sensor 466 and the GNSS sensor 458 can be combined into a single integrated unit.

[0145] The vehicle can include a microphone 496 placed in and / or around the vehicle 400. The microphone 496 can be used for emergency vehicle detection and identification, among other things.

[0146] The vehicle can further include any number of camera types, including stereo cameras 468, wide-view cameras 470, infrared cameras 472, surround-view cameras 474, long and / or mid-range cameras 498, and / or other camera types. These cameras can be used to capture image data around the entire periphery of the vehicle 400. The types of cameras used depend on the embodiment and requirements of the vehicle 400, and any combination of camera types can be used to provide the necessary coverage around the vehicle 400. Further, the number of cameras can vary depending on the embodiment. For example, the vehicle can include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As one example and without limitation, the cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described in more detail herein with respect to Figure 4A and Figure 4B are described in more detail.

[0147] The vehicle 400 can further include vibration sensors 442. The vibration sensors 442 can measure vibrations of components of the vehicle, such as axles. For example, changes in vibration can indicate changes in the road surface. In another example, when two or more vibration sensors 442 are used, differences between the vibrations can be used to determine the friction or slip of the road surface (e.g., when there is a difference in vibration between a power driven axle and a free spinning axle).

[0148] The vehicle 400 can include an ADAS system 438. In some examples, the ADAS system 438 can include a SoC. The ADAS system 438 can include adaptive / automatic / autonomous cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functionality.

[0149] The ACC system can use RADAR sensors 460, LIDAR sensors 464, and / or cameras. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately ahead of the vehicle 400 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and, if necessary, suggests a lane change for the vehicle 400. Lateral ACC is related to other ADAS applications such as LCA and CWS.

[0150] CACC uses information from other vehicles, which can be received from other vehicles via a wireless link via the network interface 424 and / or wireless antenna 426 or indirectly through a network connection, such as through the Internet. Direct links can be provided by vehicle-to-vehicle (V2V) communication links, while indirect links can be infrastructure-to-vehicle (I2V) communication links. Generally, V2V communication concepts provide information about the immediately preceding vehicles, such as vehicles immediately ahead of and in the same lane as the vehicle 400, while I2V communication concepts provide information about traffic further ahead. A CACC system can include either or both of I2V and V2V information sources. Given information about vehicles ahead of the vehicle 400, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.

[0151] FCW systems are designed to alert the driver to a hazard so that the driver can take corrective action. FCW systems use a front-facing camera and / or RADAR sensor 460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. FCW systems can provide warnings in the form of, for example, sound, visual warnings, vibrations, and / or quick brake pulses.

[0152] AEB systems detect an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. AEB systems can use a front-facing camera and / or RADAR sensor 460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When an AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the effects of a predicted collision. AEB systems can include technologies such as dynamic brake support and / or crash imminent braking.

[0153] LDW systems provide visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle 400 is crossing lane markings. The LDW system is not activated when the driver indicates an intentional lane departure by activating a turn signal. LDW systems can use a front-side facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0154] An LKA system is a variation of the LDW system. If the vehicle 400 begins to leave the lane, the LKA system provides a steering input or brake to correct the vehicle 400.

[0155] A BSW system detects and warns the driver of vehicles in the car's blind spot. The BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses a turn signal. The BSW system can use one or more rear-side facing cameras and / or one or more RADAR sensors 460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0156] A RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear-facing camera while the vehicle 400 is backing up. Some RCTW systems include AEB to ensure application of the vehicle brakes to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0157] Conventional ADAS systems can be prone to false positive results, which can annoy and distract the driver, but typically are not catastrophic because the ADAS system alerts the driver and allows the driver to decide whether the safety condition is truly present and act accordingly. However, in an autonomous vehicle 400, in the case of conflicting results, the vehicle 400 itself must decide whether to heed the results from the primary computer or the secondary computer (e.g., the first controller 436 or the second controller 436). For example, in some embodiments, the ADAS system 438 can be a secondary and / or auxiliary computer for providing perception information to a backup computer plausibility module. The backup computer plausibility monitor can run redundant diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 438 can be provided to a supervisory MCU. If the outputs from the primary and secondary computers conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0158] In some examples, the host computer can be configured to provide a confidence score to the supervisory MCU indicating the host computer's confidence in the selected result. If the confidence score exceeds a threshold, then the supervisory MCU can follow the host computer's direction, regardless of whether the secondary computer provides conflicting or inconsistent results. In the event that the confidence score does not satisfy the threshold and in the event that the host computer and the secondary computer indicate different results (e.g., a conflict), the supervisory MCU can arbitrate between the computers to determine the appropriate result.

[0159] The supervisory MCU can be configured to run a neural network that is trained and configured to determine conditions under which the secondary computer provides false alarms based on the output from the host computer and the secondary computer. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying metal objects that are not in fact dangerous, such as drain grates or manhole covers that trigger false alarms. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to disregard the LDW when a cyclist or pedestrian is present and lane departure is in fact the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU can include at least one of a DLA or a GPU suitable for running a neural network with associated memory. In preferred embodiments, the supervisory MCU can include and / or be included as a component of the SoC 404.

[0160] In other examples, the ADAS system 438 can include secondary computers that perform ADAS functions using traditional computer vision rules. As such, the secondary computers can use classic computer vision rules (if-then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, the diverse implementation and intentional non-identity make the overall system more fault-tolerant, especially with respect to faults caused by software (or software-hardware interface) functions. For example, if there is a software bug or error in the software running on the host computer and the non-identical software code running on the secondary computer provides the same overall result, then the supervisory MCU can be more confident that the overall result is correct and that the bug in the software or hardware on the host computer did not cause a substantial error.

[0161] In some examples, the output of the ADAS system 438 can be fed to a perception block of the host computer and / or a dynamic driving task block of the host computer. For example, if the ADAS system 438 indicates a forward collision warning due to an object immediately ahead, the perception block can use this information in identifying the object. In other examples, the secondary computer can have its own neural network that is trained and thus reduces the risk of false positives as described herein.

[0162] The vehicle 400 can further include an infotainment SoC 430 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as a SoC, the infotainment system can not be a SoC and can include two or more discrete components. The infotainment SoC 430 can include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephony (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, park assist, radio data system, vehicle related information such as fuel level, total distance covered, brake fuel level, oil level, doors open / closed, air filter information, etc.) to the vehicle 400. For example, the infotainment SoC 430 can include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, an in-car computer, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, a heads-up display (HUD), the HMI display 434, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 430 can further be used to provide information (e.g., visual and / or audible) to a user of the vehicle, such as information from the ADAS system 438, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0163] The infotainment SoC 430 can include GPU functionality. The infotainment SoC 430 can communicate with other devices, systems, and / or components of the vehicle 400 over the bus 402 (e.g., a CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 430 can be coupled to a supervisory MCU such that, in the event of a failure of the host controller 436 (e.g., a primary and / or backup computer of the vehicle 400), the GPU of the infotainment system can perform some autonomous driving functions. In such examples, the infotainment SoC 430 can place the vehicle 400 in a driver safe park mode as described herein.

[0164] Vehicle 400 may further include an instrument cluster 432 (e.g., a digital instrument panel, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 432 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 432 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 430 and the instrument cluster 432. In other words, the instrument cluster 432 may be included as part of the infotainment SoC 430, or vice versa.

[0165] Figure 4D For cloud-based servers and according to some embodiments of this disclosure Figure 4A This is a system diagram illustrating communication between example autonomous vehicles 400. System 476 may include server 478, network 490, and vehicles including vehicle 400. Server 478 may include multiple GPUs 484(A)-484(H) (collectively referred to herein as GPU 484), PCIe switches 482(A)-482(H) (collectively referred to herein as PCIe switch 482), and / or CPUs 480(A)-480(B) (collectively referred to herein as CPU 480). GPUs 484, CPUs 480, and PCIe switches may be interconnected with high-speed interconnects and / or PCIe connections 486, such as, but not limited to, NVLink interface 488 developed by NVIDIA. In some examples, GPUs 484 are connected via NVLink and / or NVSwitch SoCs, and GPUs 484 and PCIe switches 482 are connected via PCIe interconnects. Although eight GPUs 484, two CPUs 480, and two PCIe switches are shown in the diagram, this is not intended to be limiting. Depending on the embodiment, each of the servers 478 may include any number of GPUs 484, CPUs 480, and / or PCIe switches. For example, each of the servers 478 may include eight, sixteen, thirty-two, and / or more GPUs 484.

[0166] The server 478 can receive image data from vehicles over the network 490 and representing images showing unexpected or changing road conditions such as a recently started road work. The server 478 can send neural networks 492, updated neural networks 492, and / or map information 494, including information about traffic and road conditions, to vehicles over the network 490. Updates to the map information 494 can include updates to the HD map 422, e.g., information about construction sites, potholes, curves, flooding, or other obstacles. In some examples, the neural networks 492, updated neural networks 492, and / or map information 494 can have been generated from experience using training performed at a data center (e.g., using the server 478 and / or other servers) and / or from data received from any number of vehicles in the environment.

[0167] The server 478 can be used to train machine learning models (e.g., neural networks) based on training data. The training data can be generated by vehicles and / or can be generated in simulations (e.g., using game engines). In some examples, the training data is labeled (e.g., in the case of supervised learning benefiting the neural network) and / or undergoes other pre-processing, while in other examples, the training data is not labeled and / or pre-processed (e.g., in the case of unsupervised learning not requiring the neural network). The training can be performed according to any class or more classes of machine learning techniques, including but not limited to the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including spare dictionary learning), rule-based machine learning, anomaly detection, and any variants or combinations thereof. Once the machine learning models are trained, the machine learning models can be used by vehicles (e.g., sent to vehicles over the network 490) and / or the machine learning models can be used by the server 478 to remotely monitor vehicles.

[0168] In some examples, the server 478 can receive data from vehicles and apply the data to the latest real-time neural networks for real-time intelligent inference. The server 478 can include a deep learning supercomputer and / or a specialized AI computer powered by GPUs 484, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, the server 478 can include deep learning infrastructure of a data center powered by CPUs only.

[0169] The deep learning infrastructure of the server 478 can be capable of fast real-time inference, and can use this capability to assess and validate the health of the processors, software, and / or associated hardware in the vehicle 400. For example, the deep learning infrastructure can receive periodic updates from the vehicle 400, such as a sequence of images and / or objects located in the sequence of images that the vehicle 400 has located (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify the objects and compare them to the objects identified by the vehicle 400, and if the results do not match and the infrastructure concludes that the AI in the vehicle 400 is malfunctioning, the server 478 can send a signal to the vehicle 400 instructing the fail-safe computer of the vehicle 400 to take control, notify the passengers, and complete a safe parking operation.

[0170] For inference, the server 478 can include GPUs 484 and one or more programmable inference accelerators (such as NVIDIA’s TensorRT 3). The combination of GPU-powered servers and inference-accelerated can make real-time responses possible. In other examples, such as where performance is less important, CPU-, FPGA-, and other processor-powered servers can be used for inference.

[0171] Figure 5 A block diagram of an example computing device 500 suitable for implementing some embodiments of the present disclosure is shown. The computing device 500 can include an interconnection system 502 coupling the following devices: a memory 504, one or more central processing units (CPUs) 506, one or more graphics processing units (GPUs) 508, a communication interface 510, input / output (I / O) ports 512, I / O components 514, a power supply 516, one or more presentation components 518 (e.g., a display), and one or more logic units 520. In at least one embodiment, the computing device 500 can include one or more virtual machines (VMs), and / or any component thereof can include virtual components (e.g., virtual hardware components). For non-limiting examples, the one or more GPUs 508 can include one or more vGPUs, the one or more CPUs 506 can include one or more vCPUs, and / or the one or more logic units 520 can include one or more virtual logic units. Thus, the computing device 500 can include discrete components (e.g., a full GPU dedicated to the computing device 500), virtual components (e.g., a portion of a GPU dedicated to the computing device 500), or a combination thereof.

[0172] Although Figure 5various blocks are shown as being connected via the interconnection system 502 having a line, but this is intended to be a simplified representation of a more complex connection that can be present. For example, in some embodiments, a rendering component 518, such as a display device, can be considered an I / O component 514 (e.g., if the display is a touchscreen). As another example, the CPU 506 and / or GPU 508 can include memory (e.g., the memory 504 can represent a storage device in addition to the memory of the GPU 508, CPU 506, and / or other components). In other words, Figure 5 The computing device of FIG. 1 is merely illustrative. Distinctions are not made between “workstation,” “server,” “laptop,” “desktop,” “tablet,” “client device,” “mobile device,” “hand-held device,” “game console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types, as all are contemplated Figure 5 within the scope of the computing device of FIG. 1.

[0173] The interconnection system 502 can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnection system 502 can include one or more link or bus types, such as an industry standard architecture (ISA) bus, an extended industry standard architecture (EISA) bus, a video electronics standards board (VESA) bus, a peripheral component interconnect (PCI) bus, a peripheral component interconnect express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 506 can be directly connected to the memory 504. Also, the CPU 506 can be directly connected to the GPU 508. Where there are direct or point-to-point connections between components, the interconnection system 502 can include a PCIe link to perform the connection. In these examples, a PCI bus need not be included in the computing device 500.

[0174] The memory 504 can include any of a wide variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 500. Computer-readable media can include volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, computer-readable media can comprise computer storage media and communication media.

[0175] Computer storage media can include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and / or other data types. For example, memory 504 can store computer readable instructions (e.g., which represent program and / or program elements, such as an operating system). Computer storage media can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by computing device 500. When information is here said to be "stored on" a computer storage medium, such as memory 504, it is meant that the information is stored in memory 504, or in another computer storage medium accessible by computing device 500.

[0176] Computer storage media can include computer readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, computer storage media can include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of the any of the above should also be included within the scope of computer readable media.

[0177] CPUs 506 can be configured to execute at least some of the computer readable instructions in order to control one or more components of computing device 500 to perform one or more of the methods and / or processes described herein. Each of CPUs 506 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a large number of software threads simultaneously. CPUs 506 can include any type of processors and can include different types of processors depending on the type of computing device 500 being implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 500, the processors can be Advanced RISC Machines (ARM) processors implemented using reduced instruction set computing (RISC) or x86 processors implemented using complex instruction set computing (CISC). Computing device 500 can include one or more CPUs 506 in addition to one or more microprocessors or supplemental co-processors such as math co-processors.

[0178] In addition or alternatively to CPU 506, GPU 508 can be configured to execute at least some computer-readable instructions to control one or more components of computing device 500 to perform one or more methods and / or processes described herein. GPU(s) 508 can be integrated GPUs (e.g., with CPU(s) 506) and / or GPU(s) 508 can be discrete GPUs. In embodiments, GPU(s) 508 can be co-processors to CPU(s) 506. Computing device 500 can use GPU(s) 508 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPU(s) 508 can be used for general-purpose computing on GPUs (GPGPU). GPU(s) 508 can include hundreds or thousands of cores capable of processing hundreds or thousands of software threads concurrently. GPU(s) 508 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from CPU(s) 506 received via a host interface). GPU(s) 508 can include graphics memory, such as display memory, for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory can be included as part of memory 504. GPU(s) 508 can include two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or through a switch (e.g., using NVSwitch). When combined together, each GPU 508 can generate pixel data or GPGPU data for different portions of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or can share memory with other GPUs.

[0179] In addition or alternatively to CPU 506 and / or GPU 508, logic unit(s) 520 can be configured to execute at least some computer-readable instructions to control one or more components of computing device 500 to perform one or more methods and / or processes described herein. In embodiments, CPU(s) 506, GPU(s) 508, and / or logic unit(s) 520 can execute any combination of methods, processes, and / or portions thereof discretely or jointly. Logic unit(s) 520 can be part of and / or integrated with CPU(s) 506 and / or GPU(s) 508 and / or logic unit(s) 520 can be discrete components of or otherwise external to CPU(s) 506 and / or GPU(s) 508. In embodiments, logic unit(s) 520 can be processors of CPU(s) 506 and / or GPU(s) 508.

[0180] Examples of logic units 520 include one or more processing cores and / or components thereof, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel visual cores (PVCs), visual processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multi-processors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application specific integrated circuits (ASICs), floating point units (FPUs), input / output (I / O) elements, peripheral component interconnects (PCI) or peripheral component interconnect express (PCIe) elements, and the like.

[0181] Communication interface 510 can include one or more receivers, transmitters, and / or transceivers that enable computing device 500 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communication. Communication interface 510 can include components and functionality to enable communication over any of a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, and the like), wired networks (e.g., communication over Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, and the like), and / or the Internet. In one or more embodiments, logic units 520 and / or communication interface 510 can include one or more data processing units (DPUs) to send data received over a network and / or over interconnect system 502 directly to one or more GPUs 508 (e.g., memory in GPUs 508).

[0182] I / O port 512 enables computing device 500 to be logically coupled to other devices, including I / O component 514, presentation component 518, and / or other components, some of which may be built into (e.g., integrated into) computing device 500. Illustrative I / O component 514 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, browsers, printers, wireless devices, and so on. I / O component 514 can provide a Natural User Interface (NUI) for processing user-generated air gestures, voice, or other physiological input. In some instances, input may be sent to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of computing device 500 (described in more detail below). Computing device 500 may include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof for gesture detection and recognition. In addition, computing device 500 may include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by computing device 500 to render immersive augmented reality or virtual reality.

[0183] Power supply 516 may include a hard-wired power supply, a battery power supply, or a combination thereof. Power supply 516 may supply power to computing device 500 so that the components of computing device 500 can operate.

[0184] The presentation component 518 may include a display (such as a monitor, touch screen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 518 may receive data from other components (such as GPU 508, CPU 506, DPU, etc.) and output that data (e.g., as images, videos, sounds, etc.).

[0185] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may... Figure 5 The implementation may be carried out on one or more instances of computing device 500—for example, each device may include similar components, features, and / or functions of computing device 500. Furthermore, in the case of implementing back-end devices (e.g., servers, NAS, etc.), the back-end devices may be included as part of a data center.

[0186] Components of a network environment can communicate with each other over a network, which can be wired, wireless, or both. The network can include multiple networks, or a network of multiple networks. By way of example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the public switched telephone network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (among other components) can provide wireless connectivity.

[0187] Compatible network environments can include one or more peer-to-peer network environments (in which case servers can not be included in the network environment), as well as one or more client-server network environments (in which case one or more servers can be included in the network environment). In a peer-to-peer network environment, functionality described herein with respect to servers can be implemented on any number of client devices.

[0188] In at least one embodiment, a network environment can include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which can include one or more core network servers and / or edge servers. The framework layer can include a framework for supporting one or more software layers and / or application programs of an application layer. The software or application programs can include network-based service software or applications, respectively. In embodiments, one or more client devices can use the network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer can be, without limitation, a type of free and open-source software web application framework, such as can be used for large-scale data processing (e.g., “big data”) using the distributed file system.

[0189] A cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functionality described herein (or one or more portions thereof). Any of these various functionalities can be distributed across multiple locations from central or core servers (e.g., one or more data centers that can be distributed across states, regions, countries, globally, and the like). If a connection to a user (e.g., a client device) is relatively close to an edge server, the core server can designate at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0190] A client device can include at least some components, features, and functionality of the example computing device 500 described herein with respect to Figure 5 As examples and not by way of limitation, a client device can embody a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a video camera, a surveillance device or system, a vehicle, a watercraft, an aircraft, a virtual machine, a drone, a robot, a hand-held communication device, a hospital device, a gaming device or system, an entertainment system, an in-vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of these described devices, or any other suitable device.

[0191] The present disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, being executed by a computer or other machine, such as a personal digital assistant or other handheld devices. Generally, program modules include routines, programs, objects, components, data structures, and the like, that perform particular tasks or implement particular abstract data types. The present disclosure can be practiced in a variety of system configurations, including hand-held devices, consumer electronics, general- purpose computers, more specialty computing devices, and the like. The present disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network.

[0192] As used herein, the term “and / or,” when used in a list of two or more elements, means that any one of the elements can be present, or combinations of two or more of the elements can be present. For example, “element A, element B, and / or element C” can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. Also, “at least one of element A or element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, “at least one of element A and element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0193] The subject matter of the present disclosure is described with specificity herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have contemplated that the claimed subject matter might also be embodied in other ways, to include different steps or combinations of steps similar to the ones described in this document, in conjunction with other present or future technologies. Moreover, although the terms "step" and / or "block" might be used herein to connote different elements of methods employed, the terms should not be interpreted as implying any particular order among or between various steps herein disclosed unless and except when the order of individual steps is explicitly described.

Claims

1. One or more processors comprising: one or more circuits to: receive, using a memory management unit (MMU), a request from a client to execute a process; select, using the MMU, a target memory manager of a plurality of memory managers of the MMU corresponding to an identifier of at least one of the client or the process based on at least the identifier; and cause the target memory manager to perform a memory translation operation for the request.

2. The one or more processors of claim 1, wherein the one or more circuits are to cause the MMU to select the target memory manager in response to the identifier indicating that the process corresponds to a priority level associated with the target memory manager.

3. The one or more processors of claim 1, wherein the plurality of memory managers comprises a plurality of translation control unit (TCU) instances.

4. The one or more processors of claim 1, wherein the plurality of memory managers comprises a plurality of groups of translation buffer units (TBUs).

5. The one or more processors of claim 1, wherein: the identifier comprises an identifier of the client; and the one or more circuits are to select the target memory manager in response to the identifier of the client indicating that the process has a constant priority level and the constant priority level corresponds to the target memory manager.

6. The one or more processors of claim 1, wherein: the identifier comprises an identifier of the process; the target memory manager is associated with a first range of identifiers; the plurality of memory managers comprises a second memory manager associated with a second range of identifiers; and the one or more circuits are to select the target memory manager in response to the identifier of the process being within the first range of identifiers.

7. The one or more processors of claim 1, wherein the process comprises at least one of a video encoding operation or a video decoding operation.

8. The one or more processors of claim 1, wherein the process is a first process having a first priority level, and the client requests the MMU to perform a translation for a second process having a second priority level that is higher than the first priority level.

9. The one or more processors of claim 1, wherein the one or more circuits comprise a system on chip (SoC), the SoC comprising the MMU, and wherein the MMU is a system memory management unit (SMMU).

10. The one or more processors of claim 1, wherein the one or more processors are comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system containing one or more virtual machines (VMs); a system implemented using an edge device; a system implemented using a robot; a system for generating synthetic data; a system for performing simulation operations; a system for collaborative content creation of 3D assets; ​ ​ ​ Systems for performing conversational AI operations Systems including one or more large language models (LLMs) Systems including one or more visual language models (VLMs) Systems for performing digital twin operations Systems for performing optical transport simulation Systems for performing deep learning operations Systems implemented at least in part in a data center; or Systems implemented at least in part using cloud computing resources.

11. A system comprising: a plurality of memory managers; and a memory management unit (MMU) including one or more circuits to: receive, from a client, a request to execute a process; select, based at least on an identifier of at least one of the client or the process, a target memory manager of the plurality of memory managers, the target memory manager corresponding to the identifier; and cause the target memory manager to perform a memory translation operation associated with the request.

12. The system of claim 11, wherein the one or more circuits select the target memory manager in response to the identifier indicating that the process corresponds to a priority level associated with the target memory manager.

13. The system of claim 11, wherein the plurality of memory managers comprises a plurality of translation control unit (TCU) instances.

14. The system of claim 11, wherein the plurality of memory managers comprises a plurality of groups of translation buffer units (TBUs).

15. The system of claim 11, wherein: the identifier comprises an identifier of the client; and the one or more circuits select the target memory manager in response to the identifier of the client indicating that the process has a constant priority level and the constant priority level corresponds to the target memory manager.

16. The system of claim 11, wherein: the identifier comprises an identifier of the process; the target memory manager is associated with a first range of identifiers; the plurality of memory managers comprises a second memory manager associated with a second range of identifiers; and the one or more circuits select the target memory manager in response to the identifier of the process being within the first range of identifiers.

17. The system of claim 11, wherein the one or more circuits comprise a system on a chip (SoC), the SoC including the MMU, and wherein the MMU is a system memory management unit (SMMU).

18. A method comprising: receiving, using one or more processing circuits of a memory management unit (MMU), a request to execute a process from a client; selecting, using the one or more processing circuits of the MMU and based at least on an identifier of at least one of the client or the process, a target memory manager of a plurality of memory managers of the MMU, the target memory manager corresponding to the identifier; and causing, using the one or more processing circuits of the MMU, the target memory manager to perform a memory translation operation associated with the request. ​ 19. The method of claim 18, comprising: selecting, using the one or more processing circuits of the MMU, the target memory manager in response to the identifier indicating that the process corresponds to a priority level corresponding to the target memory manager.

20. The method of claim 18, wherein the plurality of memory managers comprises: a plurality of translation control unit (TCU) instances; and a plurality of groups of translation buffer units (TBUs). ​

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2