DATA PROCESSING USING ACCELERATORS WITH MULTI-FRAME SUPPORT

The DMA system with frame formats and independent memory processing addresses inefficiencies in conventional accelerators, improving computational efficiency and reducing memory latency in robotic systems.

DE102025128317A1Pending Publication Date: 2026-02-05NVIDIA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102025128317
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-31
Filing Date
2025-07-17
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional processing accelerators, such as those used in robotic systems, face inefficiencies due to system latencies in memory operations and inefficient configuration techniques, limiting their performance in applications like automated vehicle operations and computer vision.

Method used

Implementing a direct memory access (DMA) system that performs DMA transfers between source and target memories in a sequence, with frame formats that allow for multiple frame types, and configuring accelerators to process data independently from a common memory source, adjusting addressable bit width, and batch processing DMA transfers.

Benefits of technology

Reduces memory latency effects and improves the computational efficiency by mitigating the ratio of computational to memory operations, enhancing the performance of processing accelerators in robotic systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Various examples disclose systems and methods relating to data processing using accelerators in a system-on-a-chip. For example, a Direct Access (DMA) system can be programmed to perform one or more DMA transfers between source and destination memory in a sequence. The DMA system can signal to an accelerator that the DMA transfers are complete and the data is available in the destination memory. In some examples, the DMA system can be configured to perform the one or more data transfers according to the frame formats associated with one of several frame types.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUNDProcessing accelerators, including vector processing units (VPUs), may be used to perform single instruction, multiple data (SIMD) operations in parallel during operation of robotic systems, for example, in automated vehicle operation (e.g., semi- or fully automated). These operations may be implemented to enable computer vision-based applications such as image processing, signal processing, and / or the like. Conventional accelerators are limited by system latencies, such as those associated with reading and writing to memory between and / or during performance of SIMD operations. Moreover, conventional techniques for configuring the operation of these accelerators as well as the devices supporting these accelerators may be inefficient.SUMMARYThe invention is defined by the claims. To illustrate the invention, aspects and embodiments are described herein which may or may not fall within the scope of the claims.Disclosed are systems and methods related to data processing using accelerators in a system on a chip. For example, a direct memory access (DMA) system may be programmed to perform one or more DMA transfers between source memory and target memory in a sequence. The DMA system may signal an accelerator that the DMA transfers are complete and the data is available in the target memory. In some examples, the DMA system may be configured to perform the one or more data transfers according to frame formats associated with one of multiple frame types.Some embodiments of the present disclosure relate to systems and methods for data processing using accelerators in one or more system-on-a-chip (SoCs). In some examples, systems and methods are disclosed that include implementing a pixel processing engine (PPE) for data processing using a two-dimensional (2D) array of processing engines. In contrast to conventional systems such as those described above, the systems and methods and the techniques implemented herein provide the ability to process data independently from a common memory source and adjust the scale of the addressable bit width in at least one dimension. This may reduce the effects of latencies associated with reading and writing to memory before, during, and after data processing. Additionally, ratios of computational operations to memory operations that would be constrained in conventional systems may be mitigated (or even improved) as described herein.Some embodiments of the present disclosure relate to systems and methods for coordinating performance of direct memory access (DMA) transfers using multi-frame support accelerators, for example, in one or more SoCs. In some examples, systems and methods are disclosed that include configuration of frame formats that can cause multiple types of DMA transfers. For example, a DMA system may receive data associated with a frame format representing a set of DMA transfers. The DMA system may then determine a frame type of the frame format and configure, using the DMA system, the amount of the set of DMA transfers to be performed (e.g., by a hardware accelerator such as a vector processing unit or a pixel processing engine) based at least on the frame format. In these embodiments, the DMA system may be configured to process frame formats associated with multiple frame types. In contrast to conventional systems, the systems and methods described herein, as well as the techniques implemented, enable batch processing of multiple DMA transfers in conjunction with a single frame format and configuration of a single DMA system to cause the set of DMA transfers to be performed. By configuring frame formats for certain types of DMA transfers, the complexity of the frame format can be reduced. This, in turn, may reduce the computational effort for configuring DMA transfers between a source memory and a target memory.At least one aspect relates to one or more processors. At least one aspect relates to a system. The system may include one or more processors to obtain, via a direct memory access (DMA) hardware sequencer, data representing a frame format including a set of DMA transfers to be performed in a sequence. In some implementations, the frame format includes a set of descriptor identifiers corresponding to the descriptors. The sequence may be based on at least one frame type of the frame format. In some implementations, the one or more processors are to determine, by the DMA system, the frame type of the frame format from a set of frame types based at least on the frame format. In some implementations, the one or more processors are to obtain, by the DMA system, data associated with the descriptors based at least on the descriptor identifiers of the frame format and the frame type of the frame format. In some implementations, the one or more processors are to cause the DMA system to perform the set of DMA transfers according to the sequence between a source memory and a destination memory based at least on the frame format and descriptors. The DMA system may be configured to process the frame formats associated with each frame type from the group of frame types.In some implementations, the frame format is associated with a frame addressing frame format. The system may include one or more processors operable to cause the DMA system to perform the set of DMA transfers based at least on the frame addressing frame format. The frame addressing frame format may indicate that the set of DMA transfers is to be performed according to a raster scan sequence. In some implementations, the frame format is associated with a descriptor that addresses the frame format. The system may include one or more processors operable to cause the DMA system to perform the set of DMA transfers based at least on the descriptor addressing frame format. The descriptor addressing frame format may indicate that the set of DMA transfers is to be performed based at least on a configuration of the DMA transfers by an accelerator. In some implementations, the frame format is associated with a frame format for random region addressing. The system may include one or more processors that cause, using the DMA system, the set of DMA transfers to be performed based at least on the set of descriptors corresponding to the regions of interest identified by the frame format within a frame.In some implementations, the system may include one or more processors that are to determine the frame type based at least on the frame format. The frame type may indicate that one or more byte fields of the frame format are reserved byte fields. The system may include one or more processors to obtain the data associated with the descriptors based at least on the frame format type. In some implementations, the frame type includes a descriptor addressing frame type associated with one or more updated descriptors generated by an accelerator, or a random region frame type associated with one or more descriptors indicating at least one offset, and at least one descriptor associated with a tile bounding a region of interest within a frame, the frame indicated by the frame format.In some implementations, the system may include one or more processors to cause the set of DMA transfers to be performed in a single channel based at least on the frame format type and descriptors. The frame format may include one or more reserved byte fields. The system may include one or more processors to obtain the data associated with the descriptors, the descriptors comprising one or more descriptor byte fields corresponding to one or more of the reserved byte fields of the frame format.In some implementations, the system may include one or more processors to determine one or more aspects of the set of DMA transfers based at least on the frame type. The system may include one or more processors to determine the sequence of the set of DMA transfers based at least on one or more aspects of the set of DMA transfers.In some implementations, the one or more processors include at least one of: an autonomous or semi-autonomous machine control system; an autonomous or semi-autonomous machine perception system; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system for hosting one or more real-time streaming applications; a system for implementing large language models (LLMs); a system for implementing vision language models (VLMs); a system implemented using an edge device; a system implemented using a robot; a system for performing one or more conversational AI operations; a system for performing one or more generative AI operations; a system for generating synthetic data; a system including one or more virtual machines (VMs); a system at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources.At least one aspect relates to one or more processors. The one or more processors may include one or more circuits to: obtain data representing a frame format comprising a set of DMA transfers to be performed in a sequence. The frame format may include a set of descriptor identifiers corresponding to the descriptors, and wherein the sequence is based at least on a frame type of the frame format. The one or more circuits may determine the frame type of the frame format from a set of frame types based at least on the frame format. In some implementations, the one or more circuits may obtain the data associated with the descriptors based at least on the descriptor identifiers of the frame format and the frame type of the frame format. The one or more circuits may cause the set of DMA transfers to be performed according to the sequence between a source memory and a destination memory based at least on the frame format and descriptors. The DMA system may be configured to process frame formats associated with each frame type of the set of frame types.In some implementations, the frame format is associated with a frame addressing frame format. The one or more circuits may cause the set of DMA transfers to be performed based at least on the frame addressing frame format indicating that the set of DMA transfers is to be performed according to a raster scan sequence. The raster scan sequence may be associated with at least one pass order of a plurality of pass orders. In some implementations, the frame format is associated with a descriptor that addresses the frame format. The one or more circuits may cause the set of DMA transfers to be performed based at least on the descriptor addressing frame format indicating that the set of DMA transfers is to be performed based at least on a configuration of the DMA transfers by an accelerator. In some implementations, the frame format is associated with a frame format for random region addressing. The one or more circuits may cause the set of DMA transfers to be performed based at least on the set of descriptors corresponding to the regions of interest identified by the frame format within a frame.In some implementations, the one or more circuits are to determine the frame type based at least on the frame format, the frame type indicating that one or more byte fields of the frame format are reserved byte fields. The one or more circuits may obtain the data associated with the descriptors based at least on the frame format type. In some implementations, the frame type includes a descriptor addressing frame type associated with one or more updated descriptors generated by an accelerator, or a random region frame type associated with one or more descriptors indicating at least one offset, and at least one descriptor associated with a tile bounding a region of interest within a frame, the frame indicated by the frame format.In some implementations, the one or more circuits are to: cause the set of DMA transfers to be performed in a single channel based at least on the frame format type and descriptors. The frame format may include one or more reserved byte fields. The one or more circuits may receive the data associated with the descriptors, the descriptors comprising one or more descriptor byte fields corresponding to one or more of the reserved byte fields of the frame format.In some implementations, the one or more processors include at least one of: an autonomous or semi-autonomous machine control system; an autonomous or semi-autonomous machine perception system; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system for hosting one or more real-time streaming applications; a system for implementing large language models (LLMs); a system for implementing vision language models (VLMs); a system implemented using an edge device; a system implemented using a robot; a system for performing one or more conversational AI operations; a system for performing one or more generative AI operations; a system for generating synthetic data; a system including one or more virtual machines (VMs); a system at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources.At least one aspect relates to a method. The method may include obtaining data representing a frame format including a set of DMA transfers to be performed in a sequence. The frame format may include a set of descriptor identifiers corresponding to the descriptors, and wherein the sequence is based at least on a frame type of the frame format. In some implementations, the method includes determining the frame type of the frame format from a set of frame types based at least on the frame format. The method may include obtaining data associated with the descriptors based at least on the descriptor identifiers of the frame format and the frame type of the frame format. In some implementations, the method includes causing the set of DMA transfers to be performed according to the sequence between a source memory and a destination memory based at least on the frame format and descriptors.The disclosure extends to all novel aspects or features described and / or illustrated herein.Further features of the disclosure are characterized by the independent and dependent claims.Each feature in one aspect of the disclosure may be applied to other aspects of the disclosure in any suitable combination. In particular, method aspects may be applied to device or system aspects and vice versa.Furthermore, features implemented in hardware may also be implemented in software and vice versa. Accordingly, any reference to software and hardware features in the present description is to be construed as such.Any system or device feature as described herein may also be provided as a method feature, and vice versa. Functionally described system and / or device aspects (including means and features) may alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.It should also be appreciated that certain combinations of the various features described and defined in any aspects of the disclosure may be implemented and / or provided and / or used independently of one another.The disclosure also provides computer programs and computer program products comprising software code adapted, when executed on a data processing device, to perform any of the methods described herein and / or to embody any of the device and system features described herein, including any or all of the component steps of a method.The disclosure also provides a computer or computing system (including networked or distributed systems) having an operating system that supports a computer program for performing any of the methods described herein and / or for embodying any of the apparatus or system features described herein.The disclosure also provides a computer readable medium having stored thereon one or more of the aforementioned computer programs.The disclosure also provides a signal carrying one or more of the aforementioned computer programs.The disclosure extends to methods and / or apparatuses and / or systems as described herein with reference to the accompanying drawings.Aspects and embodiments of the disclosure will now be described, by way of example only, with reference to the accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGSThe present systems and methods for processing data using one or more systems on a chip (Systems-on-a-chip, SoCs) and other processing hardware are described above with reference to the appended drawing figures. The following are shown: FIG. 1A illustrates an example computing environment in which one or more devices operate to process data using a SoC, in accordance with some embodiments of the present disclosure; FIG. 1B is an example diagram of a pixel processing engine (PPE) in accordance with some embodiments of the present disclosure; FIG. 1C is an example diagram of a processing element (PE) of a PPE, in accordance with some embodiments of the present disclosure; FIG. 2 illustrates an example PPE configuration, in accordance with some embodiments of the present disclosure; FIG. 3 is a flow diagram of an example method for processing data using accelerators in a system on a chip, in accordance with some embodiments of the present disclosure; FIGS. 4A-C illustrate example frame formats, in accordance with some embodiments of the present disclosure; FIG. 4D illustrates an example set of tile sequence orders, in accordance with some embodiments of the present disclosure; FIG. 5 is a flow diagram of an example method for processing data based at least on associating frame types, in accordance with some embodiments of the present disclosure; FIG. 6 is an exemplary illustration of a frame in accordance with some embodiments of the present disclosure; FIG. 7 is a flow diagram of an example method 700 for processing data based on random regions in a frame, in accordance with some embodiments of the present disclosure; FIGS. 8A-8F illustrate example representations of data transfers between accelerators, in accordance with some embodiments of the present disclosure; FIG. 9 is an exemplary illustration of a data layout across registers in PEs of a two-dimensional accelerator, in accordance with some embodiments of the present disclosure; FIG. 10A is a flow diagram of an example method for performing transmissions between speeds, in accordance with some embodiments of the present disclosure; FIG. 10B is a flow diagram of an example implementation of the method of claim 10A, in accordance with some embodiments of the present disclosure; FIGS. 11A-11C illustrate example sequences of frame transmissions using accelerators, in accordance with some embodiments of the present disclosure; FIG. 12 is a flow diagram of an example method for sequencing frame transmissions using accelerators, in accordance with some embodiments of the present disclosure; FIG. 13 is a diagram illustrating an implementation of a process for generating an example accelerator instruction, in accordance with some embodiments of the present disclosure; FIG. 14 is a flow diagram of an example method for generating accelerator instructions, in accordance with some embodiments of the present disclosure; FIG. 15A is an illustration of an example autonomous vehicle, in accordance with some embodiments of the present disclosure; FIG. 15B illustrates an example of camera locations and fields of view for the example autonomous vehicle of FIG. 15A, in accordance with some embodiments of the present disclosure; FIG. 15C is a block diagram of an example system architecture for the example autonomous vehicle of FIG. 15A, in accordance with some embodiments of the present disclosure; FIG. 15D is a system diagram for communication between one or more cloud-based servers and the example autonomous vehicle of FIG. 15A, in accordance with some embodiments of the present disclosure; FIG. 16 is a block diagram of an example computing device suitable for use in implementing some embodiments of the present disclosure; and FIG. 17 is a block diagram of an example data center suitable for use in implementing some embodiments of the present disclosure.DETAILED DESCRIPTIONSystems and methods are disclosed with respect to various components of one or more SoCs, and techniques using one or more components of the one or more SoCs. Some embodiments described herein include a pixel processing engine (PPE) and / or a direct memory access (DMA) system (e.g., including a DMA hardware sequencer) and may be described with respect to an example autonomous or semi-autonomous vehicle or machine 1500 (alternatively referred to herein as "vehicle 1500", "ego vehicle 1500", "machine 1500", or "ego machine 1500", an example of which is described with reference to FIGS. 15A-15D ), which is not intended to be limiting. The systems and methods described herein may be used without limitation on non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS)), autonomous vehicles or machines, steered and unrouted robots or robot platforms, storage vehicles, off-road vehicles, vehicles coupled to one or more trailers, foil boats, boats, shuttles, emergency vehicles, motor cycles, electric or motorized bicycles, aircraft, construction vehicles, subsea vehicles, drones, and / or other types of vehicles. While the present disclosure is described in terms of computer vision, machine learning, artificial intelligence, image processing, and / or the like, this is not intended to be limiting, and the systems and methods described herein may additionally be used in augmented reality, virtual reality, mixed reality, robotics, security and monitoring, autonomous or semi-autonomous machine applications, and / or any other field of technology, where a vector processing unit (VPU), a DMA system, an instruction set architecture (ISA), a programmable vision accelerator (PVA), a decoupled accelerator, a decoupled look-up table (DLUT) accelerator, a hardware sequencer, a single instruction, multiple data (SIMD) architecture, and / or multiple other components of one or more SoCs may be used. Although the components and associated processes described herein may be described with respect to one or more SoCs, this is not intended to be limiting in nature and these components may be implemented as stand-alone components, as individual components of a system, and / or as integrated components of a device. In some embodiments, systems, components, features, functionality, and / or methods of the present disclosure may be incorporated into the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.HARDWARE ACCELERATOR IMPLEMENTED BY SOCSingle command, multiple data (SIMD) processors may be incorporated into hardware accelerators, such as a programmable vision accelerator (PVA) from NVIDIA, to enable image and video processing pipelines involved in real-time operation of robotic systems. In particular, accelerators may implement compute-intensive portions of image and video processing pipelines and enable functions such as image filtering, feature extraction, object recognition, image segmentation, etc. This hardware architecture may meet the increasing demands on faster and more efficient hardware and provide complex processing pipelines developed for image processing based industries, such as the automated vehicle and robot industry.As the algorithms implemented by these video processing pipelines become more and more complex, certain bottleneckes in conventional hardware implementations may limit the efficiency of such processing pipelines. In one example, the bit width of an accelerator may be limited by the bit width of a local memory (e.g., a local memory of the SIMD). This limitation on memory throughput may result in constraints on the ratio of computational operations to memory operations associated with other systems within the accelerator. A ratio of computational operations to memory operations measures the relative bandwidth of the arithmetic logic unit (ALU) of a particular system as compared to the memory access bandwidth of the system. For example, an accelerator may have two vector units (each with a processing bit width of 384 bits or 784 bits) and three memory units (each supporting 512 bit read / write functions). The 1.5 times ratio of computational operations to memory operations ( 784 / 512) may assist in expanding the dynamic range in generating data associated with intermediate values, however, the bit width of vector processing may not be increased without correspondingly increasing the memory bit width of the memory units. Conventional approaches to cancelling bandwidth bottlenecking generally consist in redesigning the hardware architecture to achieve larger bit widths, often resulting in increased power consumption, which may be extremely burdensome for numerous applications such as vehicle automation (especially in electric vehicle automation).The systems and methods described herein relate to system architectures and control architectures that enable the scaling of fixed arrays of processing elements (PEs) involved in the processing of images and video (e.g., in a pixel processing engine) and address different size inputs without scaling the bit width of the local data store. More particularly, the present disclosure describes a system including a plurality of PEs in a PPE operatively coupled together. The system also includes a control system for determining a processing machine configuration representing connections between the plurality of PEs, determining a magnitude of an input to the system, and dividing the input into a plurality of sub-inputs to be processed by the array of PEs. The split inputs may then be loaded into (e.g., via) the PEs to cause the PEs to perform operations included in the above-mentioned processing pipelines.In implementation, the disclosed system architecture and control techniques enable the inputs to the architecture to be scaled as needed to support more sophisticated algorithms implemented by these image and video processing pipelines. In particular, by sharing the data to be processed among multiple PEs, the data may be read from or written to the PEs based at least on the size of the input data (e.g., corresponding to the size of the input vector defined for a particular algorithm) according to the fixed bit width associated with each row of PEs, and portions of the input data may be processed simultaneously. This can ease the restrictions of the PVA associated with the bit width of the PVA's local memory. The PPE configuration and techniques disclosed herein also reduce or eliminate the need for intermediate read and / or write operations from memory external to the accelerator, and thereby reduce and eliminate the effects of bottleneck in loading / storing data before, during, and after processing the data and the corresponding power consumption in managing and moving the data, as is usually required.FIG. 1A is an example computing environment (referred to as environment 100) in which one or more devices operate to process data using a SoC, in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein may be performed using similar components, features, and / or functionality to that of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.Environment 100 may include processor 102, memory 104, instruction switches 106, memory 108 (sometimes referred to as dynamic random access memory (DRAM)), and function blocks 110 a, 110 b(individually referred to as function block 110 and collectively referred to as function blocks 110, unless otherwise indicated). In some embodiments, the processor 102, memory 104, instruction switch 106, memory 108, and function blocks 110 may interconnect via wired or wireless connections (e.g., connect to communicate and / or the like). In some embodiments, the components of environment 100 may be included in a system on a chip (SoC). For example, the components of the environment 100 may be included in one or more SoCs that form integrated circuits by combining some or all of the components of the environment 100.The processor 102 may include one or more processors, such as one or more central processing units (CPUs), graphics processing units (GPUs), microprocessors, microcontrollers, and / or the like. In some embodiments, processor 102 may include a controller (referred to as a PVA controller), where environment 100 corresponds to the PVA. Processor 102 may be coupled to an instruction cache (not explicitly shown) that stores instructions that processor 102 is to execute. In some embodiments, processor 102 may be configured to output data associated with the configuration and / or control of one or more of the devices of FIG. 1A. For example, processor 102 may be configured to output data associated with the configuration of a DMA system 114 aand / or DMA system 114 b(sometimes referred to as a DMA hardware sequencer) to control DMA transfers to and / or from vector memory (VMEM) 112 aand / or VMEM 112 bof function block 110 aand function block 110 b, respectively.Memory 104 (sometimes also referred to as an L2 buffer or an L2 cache) may include a storage device coupled to DMA system 114 aand / or DMA system 114 bof function blocks 110. In some embodiments, memory 104 may be configured to receive and store data from DMA system 114 aand / or DMA system 114 bof function blocks 110, as described herein. In some embodiments, memory 104 may include one or more (e.g., 2) banks that enable simultaneous read or write requests. For example, memory 104 may include a first bank associated with DMA system 114 aand a second bank associated with DMA system 114 b. In some embodiments, memory 104 may enable cross communication between DMA system 114 aand DMA system 114 bby granting each of the DMA systems access to both banks.The instruction switch 106 may include one or more processors configured to scan the memory 108, receive data from the memory 108, cause data stored in the memory 108 and / or local memory of the instruction switch 106 to be loaded into the VMEM 112, and / or the like. For example, the instruction switch 106 may be coupled to the memory 108 and / or may include internal memory on which instructions required to operate one or more devices of the respective function blocks 110 are stored. In an example, the instruction switch 106 may be configured to receive and provide instruction-associated data to perform one or more DMA transfers as described herein. In another example, the instruction switch 106 may be configured to receive and provide instructions associated data for performing one or more operations specific to one or more devices of the function blocks 110. In an example, the instruction switch 106 may be configured to receive and provide instructions associated data for performing one or more filtering operations (e.g., finite impulse response (FIR) filtering, min / max filtering, 3x3 filtering, 5x5 filtering, 7x7 filtering, and / or the like), and the instruction switch 106 may transmit the data to the caches 120 of the respective function blocks 110. In this example, the respective caches 120 may be configured to transfer (e.g., load) the data associated with the instructions to the VPU 116 or PPE 118 to cause (e.g., execute) the respective device to perform the one or more filtering operations. In some embodiments, the instruction switch 106 may be configured to receive data from memory 104 as well as memory 108 (e.g., system memory). By obtaining the data from memory 104 and memory 108, instruction switch 106 may reduce a penalty for an instruction cache error (e.g., by reducing the time period associated with an error from 100 cycles to 10 cycles).The memory 108 may include a storage device connected to the DMA system 114 aand / or the DMA system 114 bof the function blocks 110. In some embodiments, memory 108 may receive and store sensor data generated by one or more sensors of a robot, such as example autonomous vehicle 1500 of FIGS. 15A-15D. For example, during operation of the robot, the memory 108 may be configured to receive data based at least on a direct connection to the one or more sensors or an indirect connection to the one or more sensors (e.g., via communication via a CAN bus and / or the like). In these examples, the sensor data may include image data associated with one or more images generated or obtained using the one or more cameras, LiDAR data associated with one or more point clouds generated by one or more LiDAR sensors, radar data associated with one or more radar images generated by one or more radar sensors, and / or the like. In some embodiments, the memory 108 may be configured to provide (e.g., transmit) the sensor data stored therein to one or more components of the function blocks 110. For example, during processing of the one or more images generated by the one or more cameras of the robot and / or another machine, the DMA system 114 aand / or the DMA system 114 bmay obtain the image data from the memory 108 and cause the image data to be stored in the VMEM 112 aand / or VMEM 112 b, respectively. In some embodiments, memory 108 may receive and store data from DMA system 114 aand / or DMA system 114 bof function blocks 110. For example, the DMA system 114 aand / or the DMA system 114 bmay provide image data updated based at least on the processing of the image data to the memory 108, and the memory 108 may store the updated image data in the memory 108.Function blocks 110 may include VMEMs 112 a, 112 b; DMA systems 114 a, 114 b; vector processing units (VPUs) 116 a, 116 b(alternatively referred to as vision processing units); pixel processing engines (PPEs) 118 a, 118 b; caches 120 a, 120 b, 120 c, 120 d; and decoupled look-up tables (DLUTs) 122 a, 122 b(and / or other decoupled accelerators). For clarity, the individual components are individually referred to as VMEM 112, DMA system 114, VPU 116, PPE 118, cache 120, and DLUT 122, and collectively referred to as VMEMs 112, DMA systems 114, VPUs 116, PPEs 118, caches 120, and DLUTs 122, unless otherwise indicated. Although particular connections are illustrated, it should be understood that the illustrated connections are for convenience and that one or more devices of the function blocks 110 may be connected to one or more other devices of the function blocks 110, unless expressly stated otherwise.The VMEMs 112 may include a memory device coupled to the processor 102 and the respective DMA systems 114, VPUs 116, PPEs 118, and caches 120 of the function blocks 110. In some embodiments, the VMEMs 112 may receive and store the sensor data obtained from the memory 108. For example, the VMEMs 112 may receive and store the sensor data obtained from the memory 108 by the DMA systems 114. Additionally or alternatively, VMEMs 112 may receive and store the sensor data obtained from memory 108 based on instructions of instruction switch 106. In some embodiments, the VMEMs 112 may be connected to the PPEs 118 via Decoupled Load / Store Units (DLSUs) 124. As described herein, the DLSUs 124 may be configured to buffer data conveyed between the VMEMs 112 and the PPEs 118 to manage latencies associated with communication between the VMEMs 112 and the PPEs 118 such that any latencies do not result in a reduction in processing speed or a blockage of the PPEs.The DMA systems 114 may include one or more processors that control the execution of one or more instructions. For example, the DMA systems 114 may receive instructions from the processor 102, the respective VPUs 116 or PPEs 118, and / or a storage device (e.g., a device associated with the DMA systems 114 such as internal or external memory; not explicitly shown), and the DMA systems 114 may coordinate with the respective VPUs 116 and / or the PPEs 118 to perform one or more operations during execution of the instructions. In an example, the DMA systems 114 may receive instructions that cause the DMA systems 114 to obtain data (e.g., sensor data and / or the like) from the memory 108 and store the data in the respective VMEMs 112. In some embodiments, the DMA systems 114 may perform one or more operations based at least on the data obtained from the memory 108. For example, the DMA systems 114 may pad frames (e.g., image frames), manipulate addresses, manage overlapping data, manage different traversal orders, take into account different frame sizes, and / or the like. In some embodiments, the DMA systems 114 may receive signals (e.g., from the VPUs 116 or PPEs 118) indicating that one or more operations have been performed on the data stored in the VMEMs 112, update one or more descriptors based at least on the updates of the data, and re-perform operations on the data.The VPUs 116 may include one or more processors executing one or more instructions. For example, the VPUs 116 may receive instructions from the processor 102, and the respective VPUs 116 may coordinate with the DMA systems 114 and / or PPEs 118 to perform one or more operations during execution of the instructions. In an example, the VPUs 116 may receive instructions from the processor 102 that cause the VPUs 116 to trigger respective DMA systems 114 to obtain sensor data from the memory 108 and store the sensor data in the respective VMEMs 112. In examples, the VPUs 116 may process the data stored in the respective VMEMs 112 and write data back to the VMEMs 112. In these examples, the data written by the VPUs 116 to the respective VMEMs 112 may include updated sensor data and / or data generated based at least on an analysis performed by the VPUs 116 on the sensor data including object or feature positions within a frame, a classification indicating a type of object or agent, and / or the like. In some embodiments, the VPUs 116 may send a signal to the respective DMA systems 114 to cause the DMA systems 114 to update one or more descriptors (described herein). For example, the VPUs 116 may send a signal to the respective DMA systems 114 to cause the DMA systems 114 to update one or more descriptors based at least on the data written by the VPUs 116 to the respective VMEMs 112.The PPEs 118 may include one or more processors that execute one or more instructions. For example, the PPEs 118 may receive instructions from the processor 102, and the respective PPEs 118 may coordinate with the DMA systems 114 and / or VPUs 116 to perform one or more operations during execution of the instructions. In an example, the PPEs 118 may receive instructions from the processor 102 that cause the PPEs 118 to trigger respective DMA systems 114 to obtain sensor data from the memory 108 and store the sensor data in the respective VMEMs 112. In some examples, the PPEs 118 may process the data stored in the respective VMEMs 112 and write data back to the VMEMs 112. In these examples, the data written by the PPEs 118 to the respective VMEMs 112 may include updated sensor data and / or data generated based at least on an analysis performed by the PPEs 118 on the sensor data including object or feature positions within a frame, a classification indicating a type of object or agent, and / or the like. In some embodiments, the PPEs 118 may send a signal to the respective DMA systems 114 to cause the DMA systems 114 to update one or more descriptors (described herein). For example, the PPEs 118 may send a signal to the respective DMA systems 114 to cause the DMA systems 114 to update one or more descriptors based at least on the data written by the PPEs 118 to the respective VMEMs 112. In some embodiments, the PPEs 118 may be the same or similar to the PPE 140 of FIG. 1B.The caches 120 may include a storage device connected to the VMEMs 112 and / or the instruction switch 106. As mentioned above, the caches 120 may receive instructions associated data from the instruction switch 106 and load the instructions into one or more devices of the function blocks 110 to cause the one or more devices to operate according to the instructions.The DLUTs 122 may include a processor and / or memory configured to store one or more lookup tables. In some embodiments, the DLUTs 122 may be configured to enable communication between the processor 102 and one or more components of the function blocks 110. For example, the DLUTs 122 may be configured to communicate with the processor 102 and / or one or more storage devices of FIG. 1A (e.g., memory 108 and / or memory (104). The DLUT 122 may then manage the data storage and retrieval process between the processor 102 and the one or more storage devices of FIG. 1A. Further details on a DLUT are contained in U.S. Patent Application No. 17 / 391,491, filed April 2, 2021, the contents of which are hereby incorporated by reference in their entirety.The DLSUs 124 may include a storage device connected to the VMEMs 112 and PPEs 118 of a particular function block 110. For example, the DLSUs 124 may receive and store the sensor data obtained from the VMEMs 112 from the memory 108. Additionally or alternatively, the DLSUs 124 may receive and store the data provided as output by the PPEs 118.FIG. 1B is an example PPE 140 according to some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein may be performed using similar components, features, and / or functionality to that of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.The PPE 140 may be the same as or similar to the PPEs 118 of FIG. 1A. In some embodiments, the PPE 140 may include an array of processing elements (PEs). For example, the PPE 140 may include the PEs 152 a- 170 h. As shown in FIG. 1B, the PPE 140 includes the PEs 152 a- 170 h, with each PE 152 a- 170 hassociated with a particular row and column. In some embodiments, each PE 152 a- 170 hmay be associated with a row and a column. For example, the PE 152 amay be associated with a first row and a first column, the PE 152 bmay be associated with the first row and a second column, the PE 152 cmay be associated with the first row and a third column, and so forth. In some examples, the PE 166 amay be associated with an eighth row and the first column, the PE 168 amay be associated with a ninth row and the first column, and the PE 170 amay be associated with a tenth row and the first column. In this manner, the PEs 152 a- 170 hmay be arranged in an 8x10 array. It should be appreciated that the array of PEs 152 a- 170 hformed by the PPE 140 of FIG. 1B is a non-limiting example and that different arrays may be formed by different arrangements of PEs 152 a- 170 hin a PPE 140. For example, the PPE 140 may be updated to include a different number of PEs in each column and / or row.In some embodiments, each PE 152 a- 170 hmay include one or more devices that enable each PE 152 a- 170 hto perform one or more operations. For example, each PE 152 a- 170 hmay include one or more arithmetic logic units (ALUs), special function units (SFUs), load / store units (LSUs), registers, controllers, and / or the like. In some embodiments, the PEs 152 a- 170 hmay be the same as or similar to the PE 170 of FIG. 1C.In some embodiments, the PEs in the first PE row (PEs 152 a- 152 h) may be connected to a VMEM 112. For example, each PE of the first row of PEs 152 a- 152 hmay be connected to the VMEM 112 via respective connections 142 a- 142 h. In an example, each PE of the first row of PEs 152 a- 152 hmay be interconnected via the corresponding interconnects 142 a- 142 hto enable each PE of the first row of PEs 152 a- 152 hto establish read streams with the VMEM 112. The read streams may be associated with the transfer of data from the VMEM 112 to the corresponding PEs of the first row of PEs 152 a- 152 h.In some embodiments, the PEs in the first PE column (PEs 152 a- 170 a) may be connected to a VMEM 112 (such connection is not explicitly shown). For example, each PE of the first column of PEs 152 a- 170 amay be connected to the VMEM 112 via corresponding connections. In an example, each PE of the first column of PEs 152 a- 170 amay be interconnected via the corresponding connections to enable each PE of the first column of PEs 152 a- 170 ato establish communication connections with the VMEM 112. In some embodiments, the PEs in the last column of PEs (PEs 152 h- 170 h) may be connected to the VMEM 112 (such connection is not explicitly shown) similar to the first column of PEs 152 a- 170 a. For example, each PE of the last column of PEs 152 h- 170 hmay be connected to VMEM 112 via corresponding connections. In an example, each PE of the last column of PEs 152 h- 170 hmay be interconnected via the corresponding connections to enable each PE of the last column of PEs 152 h- 170 hto establish communication connections with VMEM 112. In these examples, where the first column of PEs 152 h- 170 hand the last column of PEs 152 h- 170 hare connected to the VMEM 112 to establish communication links, these communication links may be used by the respective PEs to allow the PEs to request and receive data. As described herein, in an example where each PE corresponds to one or more pixels of an image, the PEs of the first PE column may communicate with the VMEM 112 to obtain data associated with adjacent pixels (that were not originally loaded into the PPE 140) to perform one or more operations (e.g., filtering and / or the like) based at least on the data associated with the adjacent pixels.In some embodiments, the first row of PEs 152 a- 152 hmay be connected to one or more other PEs 152 a- 170 hin the PPE 140. For example, each PE of the PEs 152 a- 170 hmay be connected to one or more other PEs 152 a- 170 haccording to predefined connection sets. Each connection set may preset the relative position of the one or more other PEs 152 a- 170 hwith which a particular PE of the PEs 152 a- 170 hconnects when data is transmitted or received. In an example, the PE 152 amay be connected to the PE 154 a(not explicitly shown), the PE 152 b, the PE 170 a, and the PE 152 h. In this example, PE 152a connects to four separate PEs 162a (located above or "north" with respect to PE 152a), PE 152b (located to the right or "east" with respect to PE 152a), PE 170a (located below or "south" with respect to PE 152a), and PE 152h (located to the left or "west" with respect to PE 152a) to establish communication links with the PEs. In this particular example, the PEs located south and west of the PE 152a are associated with connections that wrap around the PPE 140.In some embodiments, the PEs in the last (as shown, tenth) row of PEs (PEs 170 a- 170 h) in the PPE 140 may be interconnected with the VMEM 112. For example, each PE of the last row of PEs 152 a- 152 hmay be connected to the VMEM 112 via corresponding connections 144 a- 144 h. In an example, each PE of the last row of PEs 170 a- 170 hmay be interconnected via the corresponding interconnects 144 a- 144 hto enable each PE of the last row of PEs 170 a- 170 hto establish write currents with the VMEM 112. The write currents may be associated with the transfer of data from the respective PEs 170 a- 170 hto the VMEM 112.In some embodiments, one or more PEs may be connected to one or more other PEs of PEs 152 a- 170 hvia one or more wrap-around connections. For example, each PE in the first column of PEs (e.g., PEs 152 a, 154 a, 156 a, 158 a, 160 a, 162 a, 164 a, 166 a, 168 a, 170 a, collectively referred to as PEs 152 a- 170 a) may be connected to corresponding PEs in the last column of PEs (e.g., PEs 152 h, 154 h, 156 h, 158 h, 160 h, 162 h, 164 h, 166 h, 168 h, 170 h, collectively referred to as PEs 152 h- 170 h). In another example, each PE in the first row of PEs (e.g., PEs 152 a, 152 b, 152 c, 152 d, 152 e, 152 f, 152 g, 152 h, collectively referred to as PEs 152 a- 152 h) may be connected via a wrap-around connection to corresponding PEs in the last row of PEs (e.g., PEs 170 a, 170 b, 170 c, 170 d, 170 e, 170 f, 170 g, 170 h, collectively referred to as PEs 170 a- 170 h).In some embodiments, one or more of the PEs 152 a- 170 hmay be connected to a PE controller (not explicitly shown). For example, one or more PEs may be connected to a PE controller to enable communication of instructions between the PEs 152 a- 170 h. In some embodiments, the one or more PEs 152 a- 170 hmay be directly interconnected via dedicated connections between the PE controller and the one or more PEs 152 a- 170 h. For example, if the PE controller is connected to each PE of the one or more PEs 152 a- 170 h, the PE controller may establish a one-to-all connection set with the PEs 152 a- 170 h. In this example, the PE controller may transmit instructions to load the PPE 140, as described herein, to cause each PE 152 a- 170 hto first read the data from the read streams 142 a- 142 hand via the PPE 140. In some embodiments, once data from the read streams is loaded into the respective PEs of the plurality of PEs 152 a- 170 h, the PE controller may transmit instructions to each PE of the plurality of PEs 152 a- 170 hto perform one or more data transfers between PEs. In one example, the PE controller may send a shift north ("shift to north") instruction that causes each PE of the one or more PEs 152 a- 170 hto shift data stored in at least one register of each PE to a PE that is north (e.g., above) that PE. For example, the Shift North instruction may cause the PE 152 ato shift data in a first register from PE 152 to North to PE 154 a.The PEs 152 a- 170 hmay include rows of PEs each having a predetermined bit width. In one example, each PE may have a width of 48 bits and may support one track, two tracks (each with 24 bits), etc. The bit width of each PE of the PEs 152 a- 170 hmay be scaled consistently across the PEs 152 a- 170 h, as appropriate for a particular implementation. In some embodiments, PEs 152 a- 170 hmay also include one or more vector instruction slots to enable multiple vector operation instructions to be performed in a particular set of time steps. In the example shown in FIG. 1B, the PEs 152 a- 170 hform a PPE 140 that is 8 PEs wide and 10 PEs high, with the rows corresponding to the width of the PPE 140 and the columns corresponding to the height of the PPE 140. In this example, each of the PEs 152 a- 170 hmay have a processing width of 48 bits and a total width dimension of 384 bits, which is comparable to a data storage bit width of 512 bits. Due to the two-dimensional structure of the PPE 140, the bit width is then multiplied by the height (10 PEs in this example), providing a processing width of 3 840 bits in totalFIG. 1C is an example processing element (PE) 170 in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein may be performed using similar components, features, and / or functionality to that of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.The PE 170 may be the same as or similar to the PEs 152 a- 170 hin FIG. 1B. As shown, the PE 170 includes transfer logic 172, register storage 174 (sometimes referred to as vector register files), and arithmetic logic unit (ALU) 176. In some embodiments, the PE 170 may be connected to one or more other PEs. For example, the PE 170 may be connected to one or more other PEs located north, south, east and west of the PE 170 as a portion of a PPE (e.g., a PPE that is identical or similar to the PPE 140 of FIG. 1B ).The transmission logic 172 may include one or more circuits that receive and / or transmit data as described herein. For example, transmission logic 172 may include one or more circuits configured to receive data from one or more adjacent PEs over channels 170 a- 170 d. In some embodiments, a north channel 170 amay be configured to communicate data sent from a PE that is northly located within a PPE with respect to the PE 170; a south channel 170 bmay be configured to communicate data sent from a PE that is southly located within a PPE with respect to the PE 170; an east channel 170 cmay be configured to communicate data sent from a PE that is northly located within a PPE with respect to the PE 170; and a west channel 170 dmay be configured to communicate data sent from a PE that is westly located within a PPE with respect to the PE 170. In some embodiments, the one or more circuits of the transfer logic 172 may determine that data is received over the respective channels 170 a- 170 d, and cause the received data to be stored in corresponding registers in the register memory 174.Register memory 174 may include one or more register files. In some embodiments, register memory 174 may be configured to be coupled to transfer logic 172 to receive data via input channel 172a. In some embodiments, register memory 174 may be configured to be coupled to one or more other PEs and / or transfer logic 172 to transfer data over an output channel 174 b. In embodiments, the register memory 174 may be configured to be connected to a DLSU (e.g., a DLSU identical or similar to the DLSUs 124 of FIG. 1A ) to receive and / or transmit data from and / or to the DLSU. For example, if the PE 170 is configured to receive data via a read stream (e.g., a read stream identical or similar to the read streams 142 a- 142 hin FIG. 1B ) or transmit data via a write stream (e.g., a write stream identical or similar to the write streams 144 a- 144 hin FIG. 1B ), the PE 170 may receive and / or transmit the data from and to the DLSU via a charge / storage channel 174 c. In some embodiments, register memory 174 may transfer data to and receive data from ALU 176 via output channel 174 aand input channel 178 a. For example, the ALU 176 may receive an instruction to perform one or more operations based at least on the data stored in one or more registers of the register memory 174, and the ALU 176 may receive (e.g., read) the data stored in the one or more registers via the output channel 174 a. In some examples, ALU 176 may provide (e.g., write) data (e.g., after performing one or more operations) to one or more registers of register memory 174 via input channel 178 a.In some embodiments, ALU 176 may include one or more circuits that receive, process, and / or provide data as described herein. For example, the ALU 176 may be connected to a PE controller via a broadcast channel 170e. In this example, the PE controller may transmit instructions to the ALU 176. The instructions may be configured to cause the ALU 176 to perform one or more operations. For example, the instructions may be configured to cause the ALU 176 to perform one or more operations based at least on the data stored in one or more registers of the register memory 174. In one example, the ALU 176 may receive an instruction from the PE controller via the broadcast channel 170 eto perform one or more filtering operations. In this example, the ALU 176 may receive data from the register memory 174 via the output channel 174 athat corresponds to one or more registers of the register memory 174, and the ALU 176 may determine a pixel value based at least on the instructions and the data stored in the one or more registers of the register memory 174. The ALU 176 may then provide the pixel value to a register of the register memory 174 via the input channel 178a.FIG. 2 is an example PPE configuration 200 in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein may be performed using similar components, features, and / or functionality to that of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.The PPE configuration 200 includes a two-dimensional array of PEs 202-232. In some embodiments, each PE of the first PPE configuration 200 amay include a PE identical to or similar to the PE 170 of FIG. 1C. In some embodiments, each PE of the PEs 202- 232 may be operatively coupled to one or more other PEs of the PEs 202- 232. As shown, PPE configuration 200 includes four horizontal arrays (or rows) of PEs (row 1: PEs 202-208; row 2: PEs 210-216; row 3: PEs 218-224; row 4: PEs 226-232), and four vertical arrays (or columns) of PEs (column 1: PEs 202, 210, 218, 226; column 2: PEs 204, 212, 220, 228; column 3: 206, 214, 222, 230; column 4: 208, 216, 224, 232). In these examples, each PE of PEs 202- 232 may be associated with a particular row or column. It should be appreciated that the dimensions of the PPE configuration 200 are merely an example and that other configurations may have other dimensions.In some embodiments, the PPE configuration 200 may include PEs 202- 232 connected according to one or more connection sets. As used herein, the term "connection set" refers to connections between a particular PE and other PEs of the PPE configuration 200. In some embodiments, a set of connections may represent connections between the PEs 202- 232 based at least on the relative position of each PE to the other PEs 202- 232 and / or one or more portions of memory (e.g., internal registers of each PE). In an example, a link set may be based on the position of one or more PEs 202- 232 relative to a particular PE, where the connected PEs 202- 232 are positioned north (e.g., top), south (e.g., bottom), east (e.g., right), and / or west (e.g., left) (sometimes referred to as a 4-neighborhood link set). For example, with respect to PE 212, a connector set based on the location of one or more PEs 202- 232 located north / south / east / west of PE 212 may include interconnects (e.g., wires, printed traces disposed on a printed circuit board (PCB), and / or the like) to PE 220 (North), PE 204 (South), PE 214 (East), and PE 210 (West). In another example, a link set may be based on links where the connected PEs 202- 232 are positioned in the north, south, east, and / or west (e.g., left) (as discussed above) as well as above (in a register associated with an upper portion of a frame and / or tile) and below (in a register associated with a lower frame and / or tile). As an example, referring again to PE 212, PE 212 may include north / south / east / west connections to respective PEs 220, 204, 214, and 210 as well as connections between an upper register of PE 212 and a lower register of PE 212 (also referred to as toroidal topology). In each of these examples, each PE and / or the corresponding registers of each PE 202- 232 may establish a connection to enable communication (e.g., transfers) of data therebetween.In some embodiments, the PEs 202- 232 may receive data to be processed. For example, a first row of PEs (e.g., PEs 202, 204, 206, 208) may be connected to an input interface to receive data to be processed. In an example, the input interface may establish a connection between the first row of PEs and a VMEM (e.g., a VMEM identical or similar to the VMEM 112 of FIGS. 1A and 1B ). In another example, the input interface may establish a connection between the first row of PEs and a DLSU (e.g., a DLSU identical to or similar to the DLSU 124 of FIG. 1 ). In these examples, the VMEM and / or the DLSU may store data to be input to the PEs 202- 232 via the first row of PEs.In some embodiments, the connection between the first row of PEs and the DLSU may be associated with one or more read streams. For example, the DLSU may include data generated by one or more sensors (e.g., cameras, LiDAR sensors, RADAR sensors, and / or the like). The data may then be provided to the PEs 202- 232 for processing. In this example, the data may be divided into a plurality of inputs based on the size of the data and / or on which portions of the data are presented. In one example, when processing an image, the image may be divided into a plurality of sub-inputs (e.g., values associated with the respective pixels) based on the size of the image. In particular, the image may be divided such that each partial input corresponds to a PE of the PEs 202- 232. In some embodiments, the PEs 202- 232 may be configured to receive and transmit each of the sub-inputs. For example, the data may be loaded into the PEs 202- 232 based on how (e.g., after) the data has been subdivided. In this example, the data may be loaded sequentially into the PEs 202- 232 and transferred between the PEs 202- 232 until the sub-inputs are loaded into the registers of the corresponding PEs. In the above example, where the sub-inputs represent data associated with at least a portion of an image, the sub-inputs associated with the top pixel row of the image may be loaded into the first row of PEs in a first time step. In a second time step, the sub-inputs associated with the upper pixel row may be transmitted from the PEs of the first PE row to the PEs of a second PE row (PEs 210, 212, 214, 216), and data associated with a second pixel row may be loaded into the first PE row. This process may be iteratively repeated until the sub-inputs associated with the top pixel row are sequentially transferred to a fourth (or top) row of PEs (PEs 226, 228, 230, 232).When data (e.g., partial inputs) is transferred to one or more of the PEs 202- 232, the data may be stored in the register memory (e.g., a register memory similar to or similar to the register memory 174 in FIG. 1C ) associated with the respective PE. For example, if the PE 202 receives a partial input in response to a partial input via a read stream from the VMEM or DLSU, it may be stored in a register associated with the register memory. In response to the transfer of the sub-inputs between the PEs 202-232 according to the link sets of each PE, multiple sub-inputs may be stored in corresponding registers of the register memory. In another example, in response to a second partial input received from the PE 202 in a second time step, the PE 202 may store the second partial input to another register of the register memory of the PE 202. In this way, each PE of the PEs 202- 232 may store multiple sub-inputs transmitted to the PE.The PEs 202- 232 may each be connected (either directly or indirectly) to a control system (not explicitly shown). In some embodiments, the control system (referred to herein as a "PE controller") may be configured to transmit instructions to each of the PEs 202- 232. For example, the PE controller may determine a configuration for the PPE configuration based at least on the plurality of PEs 202- 232. In some examples, the PE controller may determine the configuration based at least on the connections between the PEs 202- 232. In some embodiments, once data is loaded into the PEs 202-232, the PE controller may determine one or more instructions to send to each of the PEs. For example, the PE controller may determine one or more instructions implementing single instruction and multiple data (SIMD) parallel processing to cause the one or more instructions to be simultaneously performed by each PE of the PEs 202- 232. The PE controller may then determine that the data corresponding to the instruction is loaded into the PE array and provide (e.g., transmit) the instruction to cause the PEs 202- 232 to perform the SIMD parallel processing.In some embodiments, the PEs 202- 232 may perform one or more operations based at least on data (e.g., sub-inputs) stored in the registers of the PEs 202- 232. For example, the PEs 202- 232 may receive one or more sub-inputs that are loaded into the PEs 202- 232 via a read current. In some embodiments, the PEs 202- 232 may receive the one or more instructions from the PE controller. In an example, the PE controller may generate an instruction associated with an addition operation, and the PE controller may provide the instruction to each of the PEs 202- 232. In some embodiments, the PEs 202- 232 may each update a value (e.g., a first value) associated with a partial input stored in a register of the PE based on the one or more instructions from the PE controller and store an updated partial input (e.g., associated with a second value) in that register or another register of the PE. In this example, the PEs 202- 232 may perform additional operations based on the updated sub-input or transmit the sub-input to be provided by the PEs 202- 232 to the VMEM or the DLSU via a write stream.In some embodiments, the PEs 202- 232 may transmit one or more sub-inputs based at least on performing the one or more operations. For example, each of the PEs 202- 232 may receive a partial input (e.g., via a write current and / or via one or more other PEs 202- 232), and each of the PEs 202- 232 may perform one or more operations based at least on the partial inputs. In this example, each of the PEs 202- 232 may then transmit the received partial input and / or the updated partial input (updated based at least on the operation performed by the PE) to one or more other PEs 202- 232. The one or more other PEs 202- 232 may then perform one or more operations based at least on the transmitted sub-input. This process may be repeated by each of the PEs 202- 232 according to the instruction provided until the operations associated with the instruction are complete. In some embodiments, after completion of the operations, the PEs 202- 232 may transmit one or more of the sub-inputs to be provided by the PEs 202- 232 to the VMEM or the DLSU via the write stream.In some embodiments, the transfers of partial inputs from a read stream to one or more PEs 202-232 and to a write stream may be referred to as a data path. For example, PEs 202- 232 may receive partial inputs via a write current as well as an instruction from the PE controller to perform a set of operations. In one example, the set of operations may be associated with a filter instruction, wherein multiple sub-inputs are obtained by a first PE through a predetermined set of transmissions (sometimes referred to as transmissions between PEs) from a plurality of PEs involved in the filter instruction. The plurality of PEs involved may include any number of PEs that store sub-inputs representing pixels involved in the filter instruction. For example, in response to the operation according to a 3x3 filter instruction, the PE 212 may receive partial inputs from the PEs 202-206, 210, 214, and 218-220 through a set of transmissions between PEs. In one example, the partial input associated with the PE 218 may be transmitted to either the PE 220 or the PE 210 and then again to the PE 212. In some embodiments, these transmissions between PEs may be performed based at least on the connection sets corresponding to each of the PEs 202-232. Once the sub-inputs are obtained by transmissions between PEs according to the data path, the PE 212 may perform one or more operations to determine an updated value for the pixel that was originally associated with the PE 212. Each of the PEs 202- 232 may perform similar transmissions and operations and determine corresponding updated values for the pixels originally associated with the PEs 202- 232. The PEs 202- 232 may then transmit the sub-inputs such that the sub-inputs are provided from the PEs 202- 232 to the VMEM or the DLSU via the write current.As shown in FIG. 3, each block of the method 300 described herein includes a computing process that may be performed using any combination of hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. The methods may also be embodied as computer usable instructions stored on computer storage media. The methods may be provided by a stand-alone application, service or hosted service (stand-alone or in combination with another hosted service), or plug-in for another product, to name just a few. Moreover, the method 300 is described by way of example with respect to the apparatus of the example computing environment of FIG. 1A, PPE 140 of FIG. 1B, and / or PE 170 of FIG. 1C. However, these methods may additionally or alternatively be performed by any system or combination of systems, including, but not limited to, the systems described herein.FIG. 3 is a flow diagram illustrating a method 300 of processing data using accelerators in a system on a chip, in accordance with some embodiments of the present disclosure. The method 300 includes, at block 302, determining a processing machine configuration based at least on a plurality of PEs. For example, a PE controller (e.g., a PE controller that is the same or similar PE controller described with respect to FIG. 2 ) may determine a processing engine configuration based at least on a plurality of PEs. In some examples described herein, the PEs may be the same or similar to the PEs of the PPE 140 of FIG. 1B and / or the PE 170 of FIG. 1C, and the processing machine described herein may be the same or similar to the PPEs 118 of FIG. 1A, PPE 140 of FIG. 1B, and / or the PEs of the PPE configuration 200 of FIG. 2. In some embodiments, the configuration of the processing engine may represent one or more array sizes that may be processed by the PEs of the PPE. For example, if the PEs of the PPE 140 form a processing machine configuration having dimensions of 8 in width and 10 in height, the PE controller may determine that the processing machine configuration may process 8 x 10 INT32 integers (e.g., 4-byte integers), 16 x 10 INT32 integers (e.g., 2-byte integers), 16 x 10 INT32 integers (e.g., in a double vector), or 32 x 10 INT32 integers (e.g., in a double vector).In some embodiments, the PE controller may determine the configuration of the processing engine, wherein the configuration of the processing engine includes a set of vertical arrays (columns) and a set of horizontal arrays (rows). For example, the PE controller may determine the configuration of the processing engine based at least on the connections between each of the PEs of the plurality of PEs. In some embodiments, the configuration of the processing engine may indicate a relative position of one or more PEs with respect to one or more other PEs. For example, the configuration of the processing machine may indicate the relative position of one or more PEs with respect to one or more other PEs in a processing machine, the relative position based at least on connections between the PEs. In another example, the configuration of the processing engine may indicate connections between the PEs that enable a data flow (e.g., a set of data transfers) between PEs. In some embodiments, the configuration of the processing engine may indicate which row and column corresponds to each PE of the plurality of PEs. Additionally or alternatively, the configuration of the processing engine may indicate which PEs are associated with a particular PE based at least on the set of connections associated with the PEs of the PE configuration. For example, if the configuration of the processing machine represents the PPE configuration 200, the configuration of the processing machine may at least indicate that the PE 202 is connected to a north PE (PE 210), an east PE (PE 204), a south PE (PE 226 based on a wrap-around connection), and a west PE (PE 208 based on a wrap-around connection).In some embodiments, the PE controller may determine the configuration of the processing engine, wherein each PE of the configuration of the processing engine is configured to receive at least one partial input. For example, and again with respect to the PPE configuration 200, a 4 x 4 pixel image may be obtained and stored in memory (e.g., a VMEM identical or similar to VMEM 112 and / or a DLSU identical or similar to DLSU 124 of FIG. 1A ). The image may then be split into a plurality of sub-inputs (described below) and the sub-inputs provided to (e.g., transmitted to) PEs 202- 232.In method 300, at block 304, the PE controller may determine the magnitude of an input for the processing engine. For example, the PE controller may determine the size of an input for the processing engine, the input representing at least a portion of an image. In one example, the image may include a 4x4 pixel set and the PE controller may determine the magnitude of the input to the processing engine. In this example, the input may include four sets of sub-inputs, each set of sub-inputs including four sub-inputs corresponding to pixels of a particular image line. In another example, the image may include a 4x8 pixel set and the PE controller may determine the magnitude of the input for the processing engine. In this example, the input may include eight sets of sub-inputs to be provided to the processing engine, each PE receiving two sub-inputs to be stored in an upper and a lower register of the given PE.In method 300, at block 306, the PE controller may cause a first set of sub-inputs from a plurality of sub-inputs to be provided to one or more first PEs of the plurality of PEs. For example, the PE controller may cause the one or more first PEs of the plurality of PEs to be provided with the first set of sub-inputs from the plurality of sub-inputs based at least on the configuration of the processing engine and the size of the input. In an example where the image includes a 4x4 set of pixels divided into four sets of sub-inputs, each set of sub-inputs may be provided simultaneously to the corresponding PEs of a first row of PEs. For example, with respect to the PPE configuration 200 of FIG. 2, a first set of four sub-inputs may be provided to the PEs 202- 208 at a first time step, respectively. After the first time step, the first set of sub-inputs may be transmitted from the PEs 202- 208 to the PEs 210- 216, respectively, and a second set of four sub-inputs may be provided to the PEs 202- 208. This process may be performed iteratively according to a sequence (also referred to as a data path) until the first set of four sub-inputs are transmitted to the PEs 226- 232, the second set of four sub-inputs are transmitted to the PEs 218- 224, a third set of four sub-inputs are transmitted to the PEs 210- 216, and a fourth set of sub-inputs are provided to the first row of PEs 202- 208.In some embodiments, the PE controller may cause the one or more PEs to perform one or more operations. For example, the PE controller may provide data associated with an instruction to each of the PEs to cause each of the PEs to perform one or more operations according to the instruction. For example, the PE controller may provide data associated with each of the PEs of the instruction (also referred to as a SIMD instruction) to cause each of the PEs to perform one or more operations according to the instruction. In this example, the PEs may be caused to perform the one or more operations in parallel and / or in coordination with the one or more other PEs of the processing machine. In some embodiments, the PEs may perform the one or more operations based at least on a value associated with a partial input corresponding to the PEs. In some examples where the instruction includes one or more transmissions of sub-inputs between PEs (e.g., according to a filter instruction and / or the like), the one or more PEs may transmit the sub-inputs to corresponding PEs involved in performing the filter instruction. In this example, each PE may then perform one or more operations based on the values representing the sub-inputs.TECHNIQUES FOR PROGRAMMING DMA SYSTEMSDMA transfers involve devices reading from and writing to memory without being coordinated by the main processors of a system (e.g., central processing units (CPUs) and / or the like). The use of DMA transfers can free up powerful system resources reserved for performing complex operations, and can be particularly useful in a system involved in real-time applications, such as automated operation of a robot, such as an automated vehicle (e.g., a car, trucks, boats, shuttles, storage vehicles, a drone, and / or the like), simulated operation of a robot (such as within a simulation environment that can be hosted using a 3D content collaboration platform, such as OMNIVERSSE from NVIDIA, another platform, or a system that can use universal scene descriptor (USD) data, such as open USD, and / or a platform or system supporting easy transport simulation operations, such as ray tracing and / or path tracing). DMA transfers can be implemented by configuring a DMA system to receive one or more descriptors (e.g., from memory associated with the DMA system in which the descriptors are stored, sometimes referred to as descriptor RAM). Each descriptor may include headers with one or more fields. Each field may include information used to configure one or more operations to be performed to cause a frame or a frame tile to be loaded into a VPU or PPE. In some examples, the fields may identify an address in memory at which to begin reading frames or tiles from memory (sometimes referred to as vector memory or VMEM), define a number of pixels to be added to padding a frame or tile, define a number of frames or tiles over which to iterate, and so forth.Conventional descriptors are configured per descriptor to enable different DMA transfer types. While the use of conventional descriptors may improve the performance of a system performing DMA transfers, these conventional descriptors are generally configured in groups to enable different types of DMA transfers, such as transfers with streaming frames generated by a sensor, such as that of a vehicle (e.g., a camera, a LiDAR sensor, a RADAR sensor, an ultrasonic sensor, and / or the like). For example, one or more descriptors may be configured to transmit data during streaming of frame tiles. Because these conventional descriptors are configured into groups and independent of each other, the conventional descriptors are queued (e.g., linked) and processed sequentially. This sequential processing of different descriptors results in large scale inefficiencies. In particular, sequential processing includes separate configurations to enable independent read / write operations. In some cases, this results in an increased number of transmission gaps between descriptors, which in turn may result in unused "bubbles" (e.g., time periods) between the frames or tiles to be transmitted. This may result in increased power consumption and memory consumption.This disclosure relates to the linking of frame types as opposed to the linking of descriptors. In some implementations, schedulers involved in data transfers, such as DMA systems, are configured to receive frame formats of different frame types, so that the DMA system may be configured to more quickly initiate DMA transfers. This can result in a reduction in the amount of time and resources that would otherwise be required for individual configuration of each DMA transfer, thereby saving power and memory consumption. Moreover, the systems and methods described herein may simplify the control code of the processing elements associated with the configuration and operation of such processing elements, and thus similarly conserve the processing resources consumed in performing DMA transfers. The techniques disclosed herein also increase bandwidth utilization by allowing for faster processing of descriptors in a single channel, rather than having to process them in parallel across multiple channels to achieve the same processing speed. And by implementing the techniques disclosed herein, the complexity of the Keleel code may be reduced.FIGS. 4A-C are example frame formats 400 a, 400 b, 400 c, in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the system may be included in or may include components, features, and / or functionality similar to those of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.In some embodiments, the frame formats 400 a, 400 b, 400 cmay represent frame formats obtained and / or stored by a DMA system (e.g., a DMA system identical or similar to the DMA system 114 of FIG. 1A ), a VMEM (e.g., a VMEM identical or similar to the VMEMs 112 of FIG. 1A ), a cache memory (e.g., a cache identical or similar to the caches 120 of FIG. 1A ), and / or a system memory (e.g., a memory identical or similar to the memory 104 and / or the memory 108 of FIG. 1A ). As described herein, each frame format 400 a, 400 b, 400 cmay be associated with one or more frame types involved in one or more different DMA transfers. For example, frame format 400 amay be associated with a frame addressing frame type to perform DMA transfers for an entire frame according to a sequence (e.g., streaming sequential tiles of a frame with or without padding as described herein), frame format 400 bmay be associated with a descriptor addressing frame type to configure the DMA system by an accelerator to perform one or more subsequent DMA transfers, and frame format 400 cmay be associated with a random region addressing frame type to perform one or more DMA transfers according to instructions associated with applications performed by an accelerator as described herein.In embodiments, the data associated with the one or more frame formats may be involved in one or more DMA transfers (e.g., may cause operations involved in one or more DMA transfers). For example, the data associated with the one or more frame formats may cause one or more DMA transfers between a source memory (e.g., memory 108 in FIG. 1A ) and a destination memory (e.g., VMEMs 112 and / or DLSUs 124 in FIG. 1A ). In some embodiments, the data associated with the one or more frame formats may be provided to a DMA system to cause the DMA system to obtain data (e.g., frame data associated with at least a portion of a frame) from the source memory for one or more operations to be performed by an accelerator (e.g., an accelerator identical or similar to the VPUs 116 and / or PPEs 118 of FIG. 1A ). For example, the DMA system may receive the data associated with the one or more frame formats, obtain data (e.g., frame data) indicated by the one or more frame formats from the source memory, and provide (e.g., transmit) the frame data indicated by the one or more frame formats from the source memory to a destination memory, such as VMEM and / or DLSU. In this way, the DMA system may pre-load data involved in one or more operations performed by one or more accelerators into the target memory, thereby reducing the number of operations that the accelerator would otherwise need to perform to obtain the frame data from the source memory.Referring to FIG. 4A, the example frame format 400 arepresents a frame format associated with a frame addressing frame type. In some embodiments, frame format 400 amay include one or more byte fields corresponding to descriptors involved in the transfer of tiles of a frame (e.g., portions of a frame) from a source memory to a destination memory when the tiles are streamed as a portion of a DMA transfer, as described herein. As described herein, the frame-addressing frame-type byte fields may be included in one or more different frame formats, such as a descriptor-addressing frame-type (represented by frame-type 400b) and / or a random-region-addressing frame-type (represented by frame-type 400c). In examples, the descriptors may be associated with transmissions in a raster scan sequence (e.g., transmitting tiles of a frame according to a traversal order). For example, descriptors of a frame format with a frame addressing frame type may be associated with transmissions in a raster scan sequence in which data associated with tiles of a frame is transmitted sequentially from source memory to destination memory, from top to bottom and from left to right of the frame. In another example, the descriptors may be associated with transmissions in a raster scan sequence in which the data associated with the tiles of a frame is sequentially transmitted from left to right and top to bottom into the target memory.In some embodiments, as shown in FIG. 4D, frame formats may be associated with a tile sequence order. For example, the frame format 400 aillustrated in FIG. 4A may implement a tile sequence order from a set of tile sequence orders 400 d(sometimes referred to as traversal orders) illustrated in FIG. 4D. As shown in FIG. 4D, the set of tile sequence orders 400 dmay include a raster left-up sequence 410, a raster right-up sequence 412, a raster left-down sequence 414, a raster right-down 416, a vertical left-up sequence 418, a vertical right-up sequence 420, a vertical left-down sequence 422, and a vertical right-down sequence 424. As shown, each of the tile sequence orders 400 dmay include a different traversal order when processing a particular frame. As shown, the raster left-up sequence 410, the raster left-down sequence 414, the vertical left-up sequence 418, and the vertical left-down sequence 422 may include a positive tile offset (e.g., when tiles within a frame are shifted down or to the right), and the raster right-up sequence 412, the raster right-down sequence 416, a vertical right-up sequence 420, and the vertical right-down sequence 424 may include a negative tile offset (e.g., when tiles within a frame are shifted up or to the left). In some embodiments, a raster left-up sequence 410, a raster right-up sequence 412, a vertical left-up sequence 418, a vertical right-up sequence 420, and the raster left-down sequence 414, a raster right-down sequence 416, a vertical left-down sequence 422, and a vertical right-down sequence 424 may include a positive line offset.By aggregating multiple sequential descriptors in a single frame format, a DMA system may be configured to stream batches of tiles associated with a particular frame or set of frames, thereby optimizing bandwidth and access to the memory involved in the transfer (e.g., source memory, DMA system memory, and target memory) by allocating one or more DMA buffers to a single channel and pipelining the transfers associated with that channel. This can reduce latencies (referred to as bubbles) that occur when multiple frame types corresponding to multiple descriptors must be loaded during the configuration of the DMA system.In some embodiments, frame format 400 aincludes frame header portion 402, first descriptor set 404, and Nth descriptor set 406. It should be appreciated that the number of descriptor sets may be any number of descriptor sets, and that the present disclosure is not limited to frame formats 400 ahaving a particular number of descriptor sets. As described herein, each row in frame format 400 aincludes a field description and one or more byte fields.In some embodiments, the example frame format 400 aincludes a frame header portion 402. The frame header section 402 may include a field description section identifying the first frame header ("frame header 1"), a second frame header ("frame header 2"), and a third frame header ("frame header 3"). For example, the example frame format 400 amay include a first frame header corresponding to a set of four byte fields (each byte field includes a length of eight bits). In an example, the first frame header may include a first byte field indicating a number of descriptor sets represented by the frame format 400 a, a second byte field indicating a frame repetition factor, a third byte field indicating a second frame identifier ("FID 1") identifying a second frame, and a first frame identifier ("FID 0") identifying a first frame. In some embodiments, the first frame identifier and the second frame identifier may indicate the frame type of the frame format (e.g., the frame type is associated with a frame addressing frame type). In this example, the second frame header may correspond to a first byte field representing a frame offset and a second byte field representing a tile offset (each byte field includes a length of sixteen bits). The third frame header may include a first byte field ("pad B") indicating a padding value (e.g., corresponding to a pixel count) including a pixel count for padding a frame along a bottom portion of the frame; a second byte field ("pad L") indicating a padding value including a pixel count for padding the frame along a right portion of the frame; a third byte field ("pad T") indicating a padding value including a pixel count for padding the frame along a top portion of the frame; and a fourth byte field ("Pad R") indicating a padding value including a pixel number for padding the frame along a right portion of the frame (each byte field includes a length of eight bits).In some embodiments, the example frame format 400 aincludes a first descriptor set 404. The first descriptor set 404 may include a first column and row header ("column 1 / row 1 header") and one or more descriptor headers ("descriptor 1 and descriptor 2" through "descriptor N"). In some embodiments, the first column and row header may correspond to a first byte field ("column 1 / row 1 header") that indicates a column and row offset and a pixel row and pitch (thereby indicating a beginning point of a frame or at least a portion of a frame (also referred to as a tile or patch) and a distance between pixels), the first byte field including a length of sixteen bits. The first column and row header may correspond to a second byte field ("column 1 / row 1 repetition factor") that indicates how often the data associated with the frame is to be transferred from the source memory to the destination memory during a DMA transfer, the second byte field including a length of eight bits. In some embodiments, the first column and row headers may correspond to a third byte field ("descriptor entry number") that indicates a number of descriptors included in the first descriptor set 404, where the third byte field includes a length of eight bits.In some embodiments, the one or more descriptor headers of the first descriptor set 404 may correspond to a plurality of descriptor IDs and respective repetition factors. For example, a first descriptor header ("Descriptor 1 and Descriptor 2") may correspond to four byte fields, which in turn correspond to two descriptor identifiers, each byte field containing a length of eight bits. In examples, a descriptor header (e.g., "descriptor N") may correspond to a descriptor ID ("Nth descriptor ID").In some embodiments, the descriptor identifiers may include values corresponding to descriptors stored in memory of the DMA system. For example, the descriptor identifiers may correspond to predetermined descriptors stored in memory of the DMA system that the DMA system may access in response to receiving data associated with a frame format (e.g., from a processor, a VPU, a PPE, and / or the like). In examples, the descriptor identifiers may correspond to descriptors being updated and stored in memory (e.g., the DMA system, the VMEM, DLSU, caches, and / or the like). For example, during object tracking, a VPU may determine one or more updated positions (e.g., with respect to a subsequent image) corresponding to positions of a tile for an image (e.g., a current image). In this example, the VPU may update the descriptor involved in the DMA transfer stored in the target memory (VMEM), and the DMA system may obtain the updated descriptor. The DMA system may then cause one or more additional DMA transfers based on the updated descriptor.Referring to FIG. 4B, frame format 400 brepresents a frame format associated with a descriptor addressing frame type. In some embodiments, frame format 400 bmay be similar to frame format 400 aof FIG. 4A. However, certain portions of the frame format 400 bmay be different from portions of the frame format 400 a. For example, the first frame header portion 402 bmay include a first frame header corresponding to the bit fields indicating a second frame identifier ("FID 1") identifying a second frame and a first frame identifier ("FID 0") identifying a first frame. In this embodiment, the first frame identifier and the second frame identifier may indicate the frame type of the frame format (e.g., the frame type is associated with a descriptor addressing frame type). In some embodiments, one or more byte fields of the frame format 400 bmay be reserved (e.g., not used, referred to as "RSVD" in the figures) compared to the frame format 400 a. For example, the byte fields of the first frame header and the byte fields of the second frame header may be reserved. In examples, the first byte field of each of the descriptor sets (e.g., the first descriptor set 404 b, one or more other descriptor sets (not explicitly shown), and the Nth descriptor set 406 b) may be reserved. By reserving one or more byte fields instead of re-structuring portions of the frame format 400 b, the frame format 400 bmay be provided to a DMA system that may process different frame format types without the DMA system having to be separately configured to process different frame types. This can reduce overall complexity in configuring the DMA system and improve compatibility between applications that configure DMA transfers with the same DMA system architecture.In some embodiments, frame format 400b may include one or more byte fields corresponding to the descriptor identifiers involved in one or more DMA transfers. The descriptor identifiers may be associated with descriptors stored in memory of the DMA system that include byte fields corresponding to one or more of the reserved byte fields of the frame format 400b. For example, the descriptor identifiers of the frame format 400 bmay indicate a descriptor including similar byte fields configured to store data associated with a frame offset and row pitch, a tile offset and row pitch, one or more padding values (e.g., padding values corresponding to padding of the bottom, left, top, and / or right sides of a frame and / or tile indicated by the descriptor), and / or a column / row offset and row pitch. By reserving these fields of frame format 400 band including one or more fields in the descriptor, a single frame format 400 bmay be used to bundle multiple descriptors corresponding to multiple DMA transfers. This may bundle multiple DMA transfers based at least on a common frame, thereby reducing the number of configurations to perform (e.g., by configuring a DMA system involved in the DMA transfers). In examples where an accelerator (e.g., the VPU) updates one or more frame formats (e.g., by updating one or more descriptors of the frame format) to configure subsequent (e.g., future) DMA transfers to be performed by the DMA system, the accelerator may generate a single frame format with multiple descriptors, thus also reducing the number of configurations involved in the configuration of the DMA system. These descriptors may be dynamically updated based on one or more operations performed by the VPU in connection with one or more applications. For example, when performing operations involved in tracking position movement of an object relative to a sensor (e.g., camera, radar sensor, LiDAR sensor, and / or the like) from frame to frame, the VPU may update one or more descriptors of tiles corresponding to the object as the object moves within the field of view of the camera, thereby performing DMA transfers by the DMA system that affect tiles of the subsequent frames corresponding to the position of the object over time.Referring to FIG. 4C, frame format 400 crepresents a frame format associated with an addressing frame type for random regions. In some embodiments, frame format 400 cmay be similar to frame format 400 aof FIG. 4A. However, certain portions of the frame format 400 cmay be different from portions of the frame format 400 a. For example, the first frame header portion 402 cmay include a first frame header corresponding to the bit fields indicating a second frame identifier ("FID 1") identifying a second frame and a first frame identifier ("FID 0") identifying a first frame. In this embodiment, the first frame identifier and the second frame identifier may indicate that the frame type of the frame format is associated with an addressing frame type for random regions. In some embodiments, one or more byte fields of frame format 400 cmay be reserved compared to frame format 400 aand / or frame format 400 b. For example, the byte fields of the second frame header and the byte fields of the third frame header may be reserved. In examples, the first three byte fields of each descriptor set (e.g., the first descriptor set 404 c, one or more other descriptor sets (not explicitly shown), and the Nth descriptor set 406") may be reserved.In some embodiments, frame format 400 cmay include descriptor sets 404 c, 406 c, each including a column header. For example, frame format 400 cmay include a descriptor set 404 chaving a first column header ("column header 1"), wherein the first column header corresponds to a column / row offset byte field. In this example, the data stored in the column / row offset byte field may include 32 bits. In some embodiments, the data stored in the column / row offset byte field may indicate a point along a frame or tile that is offset with respect to the frame or tile (identified by the frame header portion 402 c). In an example where a frame or tile is referenced in X and Y coordinates, the bottom left point of the frame or tile may represent the origin (0,0). The column / row offset may represent a number of pixels offset from the origin along the X axis and the Y axis.In some embodiments, frame format 400 cmay include a first descriptor set 404 cincluding four descriptor fields, similar to first frame format 400 aand second frame format 400 b. In this example, the first three byte fields of the first descriptor set 404 cmay be reserved and the fourth byte field may correspond to a descriptor identifier involved in a DMA transfer. In some embodiments, the descriptor identifier may be associated with descriptors stored in memory of the DMA system that include byte fields corresponding to one or more of the reserved byte fields of the frame format 400c. For example, the descriptor identifiers of the frame format 400 cmay indicate a descriptor identifying a frame offset and row pitch, a tile offset and row pitch, one or more padding values (e.g., padding values corresponding to padding of the bottom, left, top, and / or right sides of a frame and / or tile indicated by the descriptor), and / or a column / row offset and row pitch. In some embodiments, a DMA system configured to cause one or more DMA transfers to occur based at least on the frame format 400 cmay cause the transfer of the one or more frames or tiles specified by the data included in the frame header portion 400 cand the respective descriptor sets 404 c, 406 cin accordance with one or more parameters of the specified descriptor.In some embodiments, for processing one or more frames and / or tiles, the DMA system may receive data associated with a frame format (e.g., from a processor, the VMEM, and / or the VPU). For example, the DMA system may first receive the data associated with the frame format from the processor. In this example, the DMA system may obtain one or more descriptors indicated by the frame format and initiate one or more corresponding DMA transfers based at least on one or more descriptors. In some embodiments, when data associated with frames and / or tiles indicated by the one or more descriptors of the frame format is identified (e.g., in source memory), the DMA system may obtain the data from the source memory and provide the data associated with the frames to the target memory (e.g., the VMEM and / or the DLSU). In some embodiments, the DMA system may send a notification to one or more accelerators (e.g., the VPU and / or the PPE) that the data associated with the frames is stored in the target memory. Once the one or more accelerators have completed one or more operations based at least on the data associated with the frames stored in the target memory, the one or more accelerators may generate data in a different (e.g., updated) frame format and provide it to the target memory and / or directly to the DMA system. In these embodiments, the DMA system may cause one or more different DMA transfers to be performed based at least on the different frame format.FIG. 5 is a flow diagram of an example method 500 for processing data based at least on associating frame types, in accordance with some embodiments of the present disclosure. In some embodiments, aspects of method 500 may be performed by one or more devices identical or similar to one or more of the devices of FIGS. 1A-1C, such as DMA systems 114, VPUs 116, PPEs 118, and / or processor 102. In embodiments, one or more other devices of FIG. 1A may perform one or more aspects of methods 500. In some embodiments, one or more of the frame formats described herein may be identical or similar to the one or more frame formats of FIGS. 4A-4C.The method 500 includes, at block 502, obtaining data associated with a frame format representing a set of DMA transfers. For example, a DMA system may receive the data associated with the frame format representing the set of DMA transfers. In some examples, the DMA system may receive the data associated with the frame format from a processor. Additionally or alternatively, the DMA system may receive the data associated with the frame format from an accelerator of a functional block of an SoC (e.g., a functional block identical or similar to functional blocks 110 of FIG. 1A ). For example, the DMA system may obtain the frame format associated data from VMEM (e.g., VMEM identical or similar to the VMEMs 112 of FIG. 1A ) based at least on an accelerator (e.g., the VPU and / or the PPE) that writes the frame format associated data to the VMEM. In an example, the VPU and / or PPE may perform one or more operations based on at least one or more completed DMA transfers, and the VPU and / or PPE may determine one or more regions of interest. In this example, the VPU and / or PPE may generate a frame format associated with a descriptor addressing frame type, wherein one or more of the descriptors indicated by the frame format correspond to one or more regions of interest, wherein the regions of interest correspond to features (e.g., objects, agents, such as vehicles and / or pedestrians, and / or the like) that move with respect to a field of view of a sensor involved in generating the frames.In some embodiments, the frame format may include a set of descriptor identifiers. For example, the frame format may include one or more descriptor identifiers that form a set of descriptor identifiers. In some embodiments, the one or more descriptor identifiers may be associated with (e.g., correspond to) one or more descriptors stored in memory. For example, the one or more descriptor identifiers may be associated with one or more descriptors stored in a memory of the DMA system. In some examples, the one or more descriptor identifiers may be associated with one or more descriptors stored in VMEM.In some embodiments, the descriptors may be associated with one or more aspects related to one or more DMA transfers. For example, descriptors may indicate one or more aspects related to moving a frame (or a tile of a frame) from a source memory (e.g., a memory such as memory 108 in FIG. 1A ) to a destination memory (e.g., a VMEM such as VMEM 112 in FIG. 1A ). In examples, descriptors may accurately indicate one or more of a frame offset (e.g., with respect to a set of frames), a tile offset (e.g., a column and row offset with respect to a point, such as an origin of a frame), one or more padding values (e.g., for padding a frame or tile along a bottom portion, left portion, top portion, and / or right portion), and / or the like.In some embodiments, the DMA system may be configured to process frame formats associated with one or more different frame types. For example, a DMA system may be configured to process frame formats associated with a frame addressing frame type, a descriptor addressing frame type, and / or a random region addressing frame type, as described herein. In this way, the DMA system may be configured to perform DMA transfers according to various predetermined frame types, thereby reducing complexity in configuring DMA transfers. In some embodiments, the DMA system may process frame formats according to frame types based at least on a set of channels, each frame format corresponding to a single channel, as described herein. In this way, the DMA system may be configured to stack and process similar DMA transfers without sharing the performance of the DMA transfers among multiple channels, thus consolidating the resources and complexity involved in configuring the DMA system between DMA transfers.The method 500 includes, at block 504, determining a frame type of the frame format. For example, the DMA system may determine the frame type of the frame format based on at least one or more byte fields of the frame format. In an example (as shown in FIGS. 4A-4C ), frame formats may include two byte fields in a first frame header. In this example, the two byte fields may include values corresponding to the frame type in combination. In some embodiments, the DMA system may obtain the values of the byte fields corresponding to the frame type and determine the frame type of a particular frame format based at least on the values stored in the byte fields.In some embodiments, the DMA system may determine that one or more byte fields of the frame format are reserved byte fields. For example, the DMA system (as shown in FIGS. 4B and 4C ) may determine that one or more byte fields associated with a frame offset, a tile offset, padding values for a frame, column and row offsets, and / or one or more descriptors are reserved. These byte fields may be reserved based at least on the frame type associated with descriptors including byte fields corresponding to at least some of the reserved byte fields. Additionally or alternatively, these byte fields may be reserved because they are not used by the DMA system to configure one or more DMA transfers associated with the frame type.In some embodiments, the DMA system may determine that the frame type of the frame format is a frame addressing frame type, a descriptor addressing frame type, or a random region addressing frame type. For example, the DMA system may determine that the frame type is a frame addressing frame type that instructs the DMA system to perform DMA transfers by sequentially traversing a frame and transferring each tile of the frame from the source memory to the destination memory. In this example, the frame format may be configured using multiple descriptors grouped in batches such that the DMA system is configured once and the tiles of the frame corresponding to the multiple descriptors may be sequentially transferred from the source memory to the destination memory in a single channel. This can maximize the bandwidth of the source memory and / or target memory and reduce latencies in reconfiguring the DMA system between DMA transfers of tiles (such latencies are sometimes referred to as "bubbles").In one example, the DMA system may determine that the frame type includes a descriptor addressing frame type that includes configuration by an accelerator (e.g., the VPU or PPE) of the DMA system for subsequent DMA transfers. In this example, the accelerator may dynamically configure the frame format based at least on one or more operations performed by the accelerator (e.g., to track features represented in one or more frames). In some embodiments, frame formats associated with the descriptor addressing frame type may be used to transfer configuration data of a frame format for a particular descriptor stored in memory of the DMA system and / or configuration data of vector processing instructions stored in the instruction cache (e.g., caches identical or similar to caches 120 of FIG. 1A ). The DMA system may then obtain the data of the frame format and cause one or more DMA transfers to be performed based at least on descriptors included in the frame format.In another example, the DMA system may determine that the frame type includes a random region addressing frame type in which tiles corresponding to the regions of interest within a frame are moved from the source memory to the destination memory. In some embodiments, frame formats associated with an addressing frame type for random regions may include descriptors corresponding to one or more 2D and / or 3D regions of interest. The 2D and / or 3D regions of interest may correspond to tiles in a frame that bound the region of interest in the frame and are to be transmitted from the source memory to the VMEM before one or more instructions are performed using an accelerator. The DMA system may determine the offset of each region of interest that needs to be transferred with respect to the frame (e.g., an address indicating a point at which a frame begins in the source memory). In some embodiments, the accelerator (e.g., the VPU) may update the memory of the DMA system with a batch of regions of interest to be transferred (e.g., up to 32 regions of interest per batch) and trigger the corresponding DMA transfers. The random region addressing frame type may cause the DMA system to fetch a pipelined batch of tiles corresponding to the relevant regions in a frame, thereby maximizing bandwidth from the source memory and target memory at reduced latency by assigning all buffers to a single channel and pipelined the DMA transfers to reduce latency (e.g., bubbles) between 2D and / or 3D patches.The method 500 includes, at block 506, obtaining data associated with one or more descriptors. For example, the DMA system may obtain the data associated with the one or more descriptors from a memory of the DMA system. In some embodiments, the DMA system may obtain the data associated with the one or more descriptors based at least on one or more descriptor identifiers of a frame format corresponding to descriptors stored in memory of the DMA system. Additionally or alternatively, the DMA system may obtain the data associated with the one or more descriptors from the memory of the DMA system based at least on the frame type of the frame format. In some embodiments where the frame format includes one or more reserved byte fields, the one or more descriptors retrieved by the DMA system may receive descriptors including byte fields representing the data corresponding to the one or more byte fields.In some embodiments, the DMA system may determine a sequence of DMA transfers. For example, the DMA system may determine a sequencer of DMA transfers for a set of DMA transfers. In examples, the set of DMA transfers may be represented by a single frame format received from the DMA system. In an example where the DMA system receives a frame format associated with a frame addressing frame type, the DMA system may determine a sequence of DMA transfers that arrange the tiles of an entire frame (or portions thereof) identified by the frame format to be transferred from the source memory to the destination memory. In an example where the DMA system receives a frame format associated with a descriptor addressing frame type, the DMA system may determine a sequence of DMA transfers that arrange the tiles of the frame identified by the frame format to be transferred from the source memory to the destination memory following a previously performed or queued sequence of DMA transfers. In another example, where the DMA system receives a frame format associated with an addressing frame type for random regions, the DMA system may determine a sequence of DMA transfers that arrange the tiles of the regions of interest in the sequence to be transferred from the source memory to the destination memory.The method includes, at block 508, causing the set of DMA transfers to be performed between a source memory and a destination memory. For example, the DMA system may cause the set of DMA transfers to be performed between the source memory and the destination memory. In an example where the DMA system receives a single frame format corresponding to one or more DMA transfers that form the set of DMA transfers, the DMA system may cause the one or more DMA transfers to be performed based at least on the frame format and descriptors. In some embodiments, the DMA system may cause one or more of the DMA transfers indicated by a frame format to be performed based on at least a single channel associated with the frame format and / or the frame type (e.g., one or more aspects indicated by the frame type). Additionally or alternatively, the DMA system may cause one or more of the DMA transfers to be performed according to a sequence determined by the DMA system.FIG. 1B is an example PPE 140 according to some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein may be performed using similar components, features, and / or functionality to that of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.DMA transfers that include random regions corresponding to regions of interestFIG. 6 includes an example illustration of a frame 600 in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the frame 600 may be included in or may include components, features, and / or functionality similar to those of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.In some embodiments, the frame 600 may represent an image (e.g., a camera image and / or the like). For example, a camera may generate data associated with an image, the image representing an environment in which the camera is operated within the field of view of the camera. The data associated with the image may include one or more values representing a color and / or intensity at one or more pixels of the image. In some embodiments, the values corresponding to the pixels of the image may be stored in memory (e.g., system memory identical to or similar to memory 108 of FIG. 1A ). In embodiments, the values corresponding to pixels of the image may be stored in or transmitted to one or more other memory devices, such as a VMEM (e.g., a VMEM identical to or similar to VMEM 112 of FIG. 1A ), the DLSU (e.g., a DLSU identical to or similar to DLSUs 124), and / or a buffer (e.g., a buffer identical to or similar to memory 104 of FIG. 1A ).As described herein, a DMA system (e.g., a DMA system identical or similar to DMA system 114 of FIG. 1A ) may cause one or more DMA data transfers to be performed between source memory and target memory. For example, the DMA systems may receive instructions from an accelerator (e.g., an accelerator identical or similar to the VPUs 116 and / or PPEs 118 of FIG. 1A ) and / or from a processor (e.g., a processor identical or similar to the processor 102 of FIG. 1A ) that cause the DMA system to perform the one or more DMA transfers. In some embodiments, in response to the instructions of the accelerator or processor to cause the one or more DMA transfers to be performed, the DMA system may identify data corresponding to a frame stored in the source memory and cause a DMA transfer to be performed to move the data to the target memory indicated by the DMA transfer.The frame 600 may include an origin that is at a fixed point relative to the frame 600. In some embodiments, the origin of the frame 600 may be located in the upper left corner of the frame. It should be appreciated that any other point may be associated with the origin, such as the bottom left corner of the frame 600, a point along one of the sides of the frame 600, a point outside the frame 600, a point inside the frame 600, and / or the like. As described herein, by setting an origin at a particular point for one or more frames, the origin may be used to identify the position of one or more points along one or more regions (also referred to as random regions) of the frame 600.In some embodiments, one or more regions of the frame may be involved in performing one or more applications. For example, one or more regions of the frame may be involved in execution of applications by an accelerator, such as a VPU and / or PPE. In this example, the applications may perform operations during execution of the application, the result of the operations being based at least on the values of the pixels in one or more regions of the frame. In these examples, the accelerator may generate data that causes the DMA system to initiate one or more DMA transfers to transfer data associated with the one or more regions from the source memory to the target memory. For clarity, the description of the movement of the data associated with the one or more regions with respect to tiles 604 a- 604 dis described. Although the regions described herein are two-dimensional, the present disclosure is not limited to only two-dimensional regions, and one of ordinary skill in the art will understand that the techniques described herein can be applied to regions that are one-dimensional (1D) and three-dimensional (3D).In some embodiments, tiles 604 a- 604 dmay be associated with portions of frame 600. For example, a tile 604 amay be associated with (e.g., correspond to) a discrete portion of the frame 600. In examples, tiles 604 b- 604 dmay be associated with portions of frame 600 and portions that are outside of frame 600. For example, tile 2 604 balong a left portion (also referred to as a west portion) may be associated with a subset of pixels along a left portion of frame 600. Tile 2 604 bmay also be associated with a portion that is outside of frame 600 (e.g., beyond frame 600). Similarly, tile 3 604 cmay be associated with a right portion and a bottom portion (also referred to as a south east portion) of frame 600 as well as a portion that is outside the south east portion of frame 600. Tile 4 604 dmay also be associated with an upper portion (also referred to as a north portion) of frame 600 and a portion that is outside the north portion of frame 600.In some embodiments, the position of the tiles 604 a- 604 dmay be described as an offset (e.g., represented as a value indicating a positive or a negative offset) from the origin of the frame 600 with respect to a point (e.g., an leftmost point along each of the frames 604 a- 604 d). For example, the position of the tiles 604 a- 604 dmay be described as an offset from the origin of the frame 600 along an X-axis (running from left to right) and from the origin of the frame 600 along a Y-axis (running from top to bottom). As shown in FIG. 6, tile 604 amay be described as being located at a point offset by a distance X 1, Y 1, where X 1 and Y 1 correspond to a distance measured in pixels. Similarly, tile 604 bmay be described as being offset by X 2, Y 2; tile 604 cmay be described as being offset by X 3, Y 3; and tile 604 dmay be described as being offset by X 4, Y 4. In this particular example, the offset of tile 1 604 aand tile 3 604 cmay correspond to points that are within frame 600, and the offset of tile 2 604 band tile 4 604 dmay correspond to points that are not within frame 600.In some embodiments, the data associated with tiles 604 a- 604 dmay be transmitted from a source memory (not explicitly shown) and a destination memory, such as a VMEM 602. For example, an application performed using an accelerator may generate data configured to cause the DMA system to cause one or more DMA transfers such that the data associated with tiles 604 a- 604 din the source memory is transferred to the target memory (the data is also referred to as descriptor addressing data, which may include a descriptor addressing frame type). In this example, the accelerator may generate the data, where the data includes multiple descriptors corresponding to the frame 600 and / or tiles 604 a- 604 d. While aspects of data transmission based on frame 600 are described with respect to VMEM 602, it should be appreciated that the aspects may be applied to communications between system memory and VMEM when the applications described herein are executed by the VPU. However, it should be appreciated that the VMEM may be the source memory and the transmissions described herein may include transmissions from the VMEM to another memory, such as the DLSU.FIG. 7 is a flow diagram of an example method 700 for processing data based at least on random regions in a frame, in accordance with some embodiments of the present disclosure. In some embodiments, aspects of method 700 may be performed by one or more devices identical or similar to one or more of the devices of FIGS. 1A-1C, such as DMA systems 114, VPUs 116, PPEs 118, and / or processor 102. In embodiments, one or more other devices of FIG. 1A may perform one or more aspects of methods 700. In some embodiments, one or more of the frame formats described herein may be identical or similar to the one or more frame formats of FIGS. 4A-4C. In some embodiments, one or more of the frames and / or tiles of the frames may be the same or similar to the frame 600 and / or tiles 604 a- 604 d.The method 700 includes, at block 702, determining one or more regions of interest. For example, a VPU may determine one or more regions of interest within a frame. The regions of interest may include portions of the frame that correspond to objects (which may include physical objects such as traffic cones, traffic lights, and agents that may move in the environment, such as pedestrians, vehicles, and / or the like) in an image generated by a sensor. The sensor may include a camera installed on a robotic system, such as an automated vehicle, a warehouse vehicle, and / or the like, while generating sensor data during operation in an environment, the sensor data including data associated with the frame.In some embodiments, the VPU may determine the one or more regions of interest within a frame based at least on one or more operations performed by an application executed by the VPU. For example, the VPU may execute an application, such as an object tracking application, in which the relative motion of the objects over multiple frames is tracked over time. During execution of the object tracking application, the VPU may perform one or more operations that result in determinations about the positions of objects, movement of objects (e.g., from times t=-1 to time t=0), and / or predicted positions of objects (e.g., at time t=1). Examples of operations involved in object tracking may include object detection to determine one or more objects present in one or more frames, assignment of an identifier (ID) to correlate the position of an object over one or more frames over time, tracking movement of the object in the frames and / or environment based at least on the correlated positions of the object over the one or more frames, and predicting possible and / or likely future positions of the object in future frames and / or environment.In some embodiments, the operations in performing one or more operations may include results corresponding to requests for data associated with tiles of future frames, the tiles corresponding to a region of interest. For example, in the context of object tracking, an object may be determined to be at a position (e.g., within a region) of a particular frame (e.g., in a current frame at time t=0). In this example, the one or more operations may include determining a future region of interest, for example, an expected region in which the object is or could be located, and generating a request for data associated with one or more tiles associated with the future region of interest. The determination of the future region of interest may be based at least on the movement of the object relative to the robot system, the movement of the object relative to the environment in which the robot system is operating, a size (e.g., length and width represented in either a inferred length and width of the object or pixels bounding the object in the frame) bounding the region of interest containing the object at a current time and / or as expected at a future time, a change in size bounding the region of interest over time (e.g., at times before and / or including a current time), a change in size bounding the region of interest expected at future times, and / or the like. While the tiles represented by frame 600 in FIG. 6 have a uniform size, it should be appreciated that the operations may indicate changes in the size of one or more tiles such that the region of interest may be dynamically adjusted to correspond to the representation of the object in the one or more frames.The method 700 includes, at block 704, generating data associated with at least one descriptor based at least on a region of interest. For example, the VPU may generate data associated with the at least one descriptor based at least on the at least one region of interest. In some embodiments, the at least one descriptor may correspond to data stored in the source memory of an existing or future frame. For example, the VPU may generate the at least one descriptor to include an offset (e.g., along an X axis and a Y axis) with respect to an origin common to the frames and a size (e.g., of the offset along the X axis and the Y axis) of the region of interest bounding the object within the region of interest.In some embodiments, the VPU may generate data associated with a first descriptor and one or more second descriptors. For example, the VPU may generate the data associated with the first descriptor such that the first descriptor corresponds to the entire frame or a region of the frame that includes the region of interest within the frame. In this example, the first descriptor may include one or more second descriptors. The one or more second descriptors may correspond to any region of interest bounding each object within the frame involved in the one or more operations performed by the VPU. In this way, the VPU may be configured to process descriptors corresponding to multiple tiles associated with the region of interest in batches such that the corresponding DMA transfers are performed sequentially without requiring the DMA system to be reconfigured to perform the respective DMA transfers for each tile. In some embodiments, each of the one or more second descriptors may also be associated with an offset and / or a size that indicates a position of a point along the tile relative to a point along (or proximate to) the frame.In some embodiments, the VPU may determine one or more updates to be performed on the data associated with the one or more tiles. For example, the VPU may determine one or more updates to perform based at least on the position of the tiles relative to the frame. In some examples, the one or more updates may be associated with an overlap between the tiles and the frame. In these examples, the overlap may include an overlap of a tile with an edge of the frame such that a portion of the tile is enclosed by the frame and a portion of the tile is not enclosed by the frame. As shown in FIG. 6, examples of overlaps by tile 2 604 b, tile 3 604 c, and tile 4 604 dare illustrated. In some embodiments, the VPU may determine the one or more updates to be performed, the one or more updates including updates of values involved in an overlap between a tile and a frame, the values associated with pixels that go beyond (e.g., are not encompassed by) the frame. For example, the VPU may determine one or more padding values corresponding to pixels of a tile that project beyond a frame. The values may include a default value (e.g., a predetermined intensity and / or color value), a value corresponding to one or more pixels of the tile that are adjacent to the pixels that are not encompassed by the frame, and / or the like. In some embodiments, the VPU may generate the data associated with the one or more second descriptors that includes an overlap between a tile and a frame such that the DMA system updates the values during the DMA transfer as described herein.The method 700 includes, at block 706, providing data associated with the at least one descriptor to cause one or more DMA transfers to be performed. For example, the VPU may provide the data associated with the at least one descriptor to a DMA system by transmitting the data to the VMEM and sending a signal to the DMA system to indicate that the data associated with the at least one descriptor has been transmitted to the VMEM. In this example, the data associated with the at least one descriptor may configure the DMA system and cause one or more corresponding DMA transfers to be performed. In this way, the VPU may cause the DMA system to manage DMA transfers involved in operations performed or to be performed by the VPU to reserve processing and storage resources for the operations performed by the VPU.COMMUNICATIONS BETWEEN ACCELERATORSA PPE that includes a two-dimensional (2D) array of interconnected PEs (e.g., in a toroidal topology) can address inefficiencies associated with implementing spatially dependent algorithms using accelerators such as a VPU. In some embodiments, the PPE may read data associated with an image into the PPE, and each PE may communicate with other local PEs to perform certain operations (e.g., filtering and / or the like) in coordination with each other and with greater efficiency than the VPU or similar accelerators. Since each PE can communicate with local PEs, it is less necessary to request additional information from the working memory. However, the width of a particular PE array may affect the overall efficiency of the PPE. For example, when a 3x3 pixel filter is applied to a 32-bit image loaded into a PPE having the size of 8x10 pixels, the output of the PPE is 6x8 pixels, resulting in an efficiency of 60%. This is calculated by multiplying the length and width of the usable output of the PPE (e.g., 6 x 8) and dividing this output by the total size of the PPE arrays (e.g., (8 x 10)). This efficiency can be determined at least by making the outer rows and columns of the PEs inaccessible to the values of the adjacent pixels and therefore the 3x3 filter cannot be applied to pixels in these rows and columns.Efficiency may be increased by enabling communication between accelerators and, in some implementations, processing lower bit number images. For example, the efficiency of a PPE of size 8 x 10 may be calculated as (14 x 8) / (16 x 10) = 70%, with the size of the PPE now being 16 x 10, since two 16-bit pixels may be stored per PE instead of a single 32-bit pixel. In this particular example, the efficiency of the PPE may further decrease as the PPE implements increasing size filters (e.g., 5 x 5, 7 x 7, etc.), resulting in more unusable columns. By allowing for the transfer and storage of data within the registers of the PEs of a PPE, memory calls in processing the pixels in the PPE can be minimized. For example, with continued reference to the 3x3 filter applied to an 8x10 sample of an image loaded into a PPE, a set of data displacements is performed involving that PE and sets of other directly or indirectly interconnected PEs to obtain and store the required pixel values to apply the 3x3 filter to that PE. These include pixels at the top and left which would not otherwise be accessible. And in another example, where two adjacent tiles of an image (referred to as blocks) are loaded into the PPE, the bottom row of PEs having access to a bottom row of pixels in a first block may communicate with the top row of PEs having access to a top row of the next consecutive block. The same is possible if more blocks are loaded into the PEs in each direction (north, south, east and west), whereby the PEs can store and access data for which read and / or write operations in the working memory would otherwise be required.In implementation, the systems and methods described herein enable the use of PEs in a PPE that can perform spatially dependent algorithms in parallel on a pixel-by-pixel basis. By allowing the PEs to access data within the PPE or memory over longer distances and by loading multiple blocks simultaneously, the efficiency of a given spatial algorithm can be improved because more rows and / or columns of data can be accessed than would otherwise be the case if PEs could receive the data only at the individual PEs, thereby reducing the otherwise required memory calls. For example, if multiple blocks corresponding to adjacent portions of an image are loaded into the PEs, the PEs may reduce the number of total calls that would otherwise be required to the memory. This, in turn, allows the PPE to perform operations in fewer cycles and to eliminate (or at least minimize) the number of unusable pixels output after the operation.FIGS. 8A-8F are example representations of data transfers between accelerators, in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the example communications between accelerators may be included in and / or may include components, features, and / or functionality similar to those of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.In some embodiments, communications between accelerators may include (e.g., be implemented by) components of one or more accelerators (e.g., one or more accelerators, such as PPEs 118 of FIG. 1A ). For example, communications between accelerators may be implemented by a PE 802 and a PE 804. In these examples, the PE 802 and the PE 804 may be included in a PPE that is identical or similar to the PPEs 118 of FIG. 1, and the PE 802 and the PE 804 may be identical or similar to the PE 170 of FIG. 1C. In some embodiments, the PE 802 and the PE 804 may be adjacent to each other in the PPE. For example, the PE 802 may be physically located west of the PE 804, as shown in FIG. 8A. In another example, the PE 802 may be physically located eastward of the PE 804, as shown in FIG. 8B. In this example, the PE 802 and the PE 804 may be interconnected such that the PE 802 is configured to receive data from the PE 804 over a wrap-around connection. For clarity, in the examples described herein, data transfers between accelerators are discussed with respect to west transfers (e.g., from a first PE that transfers data to a second PE that is either physically west relative to the first PE or logically west positioned via a wrap-around connection relative to the first PE). It is understood that other transmissions and transmission sequences are also possible. For example, a PE may transmit data to another PE via a north transmission (e.g., from a first PE that transmits data to a second PE that is either physically positioned in the north relative to the first PE or logically positioned in the north via a wrap-around connection relative to the first PE), a south transmission (e.g., from a first PE that transmits data to a second PE that is either physically positioned in the south relative to the first PE or logically positioned in the south via a wrap-around connection relative to the first PE), or an east transmission (e.g., from a first PE that transmits data to a second PE, which is positioned either physically in the east with respect to the first PE or logically in the east via a wrap-around connection with respect to the first PE).Referring to FIG. 8A, the example data transfer between accelerators 800 aillustrates west transfer between the PE 802 and the PE 804, where the PE 802 and the PE 804 are both adjacent in an array of a PPE configuration (e.g., a PPE configuration that is the same as or similar to the PPE configuration 200 of FIG. 2 ). In some embodiments, PEs 802 and 804 may each include three registers (e.g., registers identical or similar to registers of register memory 174 of PE 170). For example, PEs 802 and 804 may include a first register "register 1", a second register "register 2", and a third register "register 3". In this example, the first register and the second register may include source registers (e.g., registers that store data transferred to another register of the same PE or another PE, as described herein). The third register may include a destination register (e.g., a register configured to receive and store data transferred from a source register). While the PEs 802 and 804 of FIGS. 8A and 8B are described with respect to three registers, it should be appreciated that the embodiments contemplated may include a different number of source and destination registers.In some embodiments, the source and destination registers of PEs 802 and 804 may be configured to store data associated with a pixel. For example, the source and destination registers of PEs 802 and 804 may be configured to store data represented by 32 bits (referred to as a "word" data type). In examples, the source and destination registers of PEs 802 and 804 may be configured to store data represented by 48 bits (referred to as "enhanced precision" word data type). In the examples described herein, the word data type and the half word data type (described with respect to FIGS. 8C-8F ) may represent portions of an image (e.g., pixels of an image). While image pixels are referred to in the description of at least FIGS. 8A-8F, it should be appreciated that the data that the source and destination registers of the PEs 802 and 804 may represent any suitable form of data, including LiDAR data associated with a point cloud, RADAR data associated with a RADAR image, and / or the like.In some embodiments, PEs 802 and 804 may receive data associated with a first block (also referred to as a tile) and / or a second block. For example, PEs 802 and 804 may receive the data associated with the first block and / or the second block, each block representing a portion of an image. In one example, an image from a processor (not explicitly shown) may be divided into multiple blocks and stored in system memory. In this example, the data associated with one or more blocks of the image may be transferred to a VMEM via at least one DMA transfer and then to a DLSU before being transferred to the PPE. It should be appreciated that the system memory, DMA system, VMEM, and DLSU may be the same or similar as memory 108, DMA systems 114, VMEMs 112, and DLSUs 124 of FIG. 1A ). The data associated with the first block and / or the second block may be transferred and stored by the first register and the second register of the PEs 802 and 804. In this manner, PEs 802 and 804 may store the data associated with the first block in respective first registers and the data associated with the second block in respective second registers. By storing (e.g., stacking) data associated with multiple blocks in corresponding registers of PEs 802 and 804 as described, the inter-accelerator data transfers described herein (also referred to as shifts or shifts between accelerators) may enable operations to be performed on images that are wider and / or higher than the PE configuration could otherwise support.In some embodiments, PEs 802 and 804 may each receive an instruction (e.g., a SIMD instruction) from a PE controller (not explicitly shown) connected to PEs 802 and 804 to perform a transmit-to-west operation (also referred to as transmit-to-west). For example, PEs 802 and 804 may receive an instruction to perform a transmit-to-west operation based at least on the data associated with a first block (represented as variables "x 1" and "x 0" that may correspond to values representing the corresponding pixels) stored in the first register of PEs 802 and 804. In this example, where the instruction causes the PEs 802 and 804 to perform a transmit-to-west operation based at least on the data associated with the first block, the PE 804 may transmit the data associated with the first block ("x1") stored in the first register of the PE 804 to the PE 802, and the PE 802 may store this data in the third register of the PE 802. In some embodiments, the instruction may cause the PE 802 to perform one or more additional operations. For example, the instruction may cause the PE 802 to perform one or more arithmetic operations that may include adding, subtracting, multiplying, or dividing the value stored in the third register ("x1") of the PE 802 to / from / with, or by the value stored in the first register ("x0") of the PE 802. In examples, the instruction may cause the PE 802 to perform one or more additional transmissions. For example, the instruction may cause the PE 802 to transfer the value stored in the third register ("x1") to one or more other PEs of the PPE configuration. It should be appreciated that in some embodiments, the instructions may include sequences of shifts and arithmetic operations to be performed such that one or more functions are performed by the PEs of the PE configuration. These functions may be associated with, for example, filter functions (e.g., implementation of 3x3 filters, 5x5 filters, 7x7 filters, and / or the like), bandpass filter functions, matrix multiplication functions, image processing functions (e.g., implementation of color or intensity adjustments), and / or the like. In some embodiments, the PEs may then transmit the data back (via one or more write currents) to the DLSU.Referring to FIG. 8B, the example data transfer between accelerators 800 billustrates west transfer between the PE 802 and the PE 804, where the PE 802 and the PE 804 are not adjacent in an array of the PPE configuration. As described above with respect to FIG. 8A, the PEs 802 and 804 may receive data associated with a first block and / or a second block. In some embodiments, PEs 802 and 804 may each receive an instruction from the PE controller connected to PEs 802 and 804 to perform a transmit-to-west operation. For example, PEs 802 and 804 may receive an instruction to perform a transmit-to-west operation based at least on the data associated with a first block (represented as variables "x1" and "x0") and the data associated with the second block (represented as variables "y1" and "y0"). In this example, where the instruction causes the PEs 802 and 804 to perform a transmit-to-west operation based at least on the data associated with the first and second blocks, the PE 804 may transmit the data associated with the second block ("y1") stored in the second register of the PE 804 to the PE 802, and the PE 802 may store this data in the third register of the PE 802. Similar to described above, the instruction may cause the PE 802 to perform one or more additional operations, such as subsequent shifts and / or arithmetic operations. In this manner, PEs 802 and 804 may transfer data amongst each other to perform operations on adjacent, contiguous blocks of an image without requiring additional read or write operations to or from the PEs of the PE configuration. In some embodiments, the PEs may then transmit the data back (via one or more write currents) to the DLSU.Referring to FIGS. 8C and 8D, the example data transfer between accelerators 800 cillustrates a transfer between PE 806 and PE 808, where PE 802 and PE 804 are both adjacent in an array of a PPE configuration (e.g., a PPE configuration that is the same as or similar to PPE configuration 200 of FIG. 2 ). In some embodiments, PEs 806 and 808 may each include three registers (e.g., registers that are identical or similar to the registers of register memory 174 of PE 170). For example, PEs 806 and 808 may include a first register "register 1", a second register "register 2", and a third register "register 3". In this example, the first register and the second register may include source registers (e.g., registers that store data transferred to another register of the same PE or another PE, as described herein). The third register may include a destination register (e.g., a register configured to receive and store data transferred from a source register). While the PEs 806 and 808 of FIGS. 8C-8F are described with respect to three registers, it should be appreciated that the embodiments contemplated may include a different number of source and destination registers.In some embodiments, the source and destination registers of PEs 806 and 808 may be configured to store data associated with one or more pixels. For example, the source and destination registers of PEs 806 and 808 may be configured to store data represented by 16 bits (referred to as a "halfword" data type). In examples, the source and destination registers of PEs 806 and 808 may be configured to store data represented by 24 bits (referred to as half-word data type "enhanced accuracy"). In some embodiments, the registers may be configured to store data associated with multiple pixels. For example, if the register size of the PEs 806 and 808 is 32 bits, each register may be configured to store data associated with a first pixel and / or a second pixel, each pixel represented by 16 bits. In examples where the register size of the PEs 806 and 808 is 48 bits (enhanced accuracy), each register may be configured to store data associated with a first pixel and / or a second pixel, each pixel represented by 24 bits. As described herein, when the data associated with two pixels is stored in a register of the PEs 806 and 808, the bits corresponding to each pixel may be referred to as "upper bits" or "lower bits", or may be transmitted according to a "first track" (corresponding to the lower bits) and a "second track" (corresponding to the upper bits). It will be appreciated that the data associated with the first pixel and the second pixel may be stored in a particular register or across multiple registers in a little-endian convention such that the bits are ordered such that the least significant bit (LSB) is stored at the lowest memory address and the most significant bit (MSB) is stored at the highest memory address of a particular register or group of registers.Referring to FIG. 8C, the PEs 806 and 808 may receive data associated with a first block and / or a second block, as described herein. For example, the PEs 806 and 808 may receive the data associated with the first block and / or the second block, each block representing portions of an image. In this example, the portions associated with the first block and / or the second block received from the PEs 806 and 808 may each represent multiple adjacent portions (e.g., adjacent pixels) of the image. As shown, the data associated with the first block received from the PE 806 may include a set of lower bits corresponding to a value ("x 0") representing a first pixel of the first block and a set of higher bits corresponding to a value ("x 1") representing a second pixel of the first block that is adjacent to the first pixel of the first block. Similarly, the data associated with the second block received from the PE 806 may include a set of lower bits corresponding to a value ("y0") representing a first pixel of the second block and a set of higher bits corresponding to a value ("y 1") representing a second pixel of the second block that is adjacent to the first pixel of the second block. The PE 808 is shown to similarly receive data associated with the first block and the second block, where the data associated with the first block (upper bits: "x3"; lower bits: "x2") and the second block (upper bits: "y3"; lower bits "y2") are stored in the first and second registers of the PE 808, respectively.In some embodiments, the PEs 806 and 808 may each receive an instruction (e.g., a SIMD instruction) from a PE controller connected to the PEs 806 and 808 to perform a transmit-to-west operation. For example, the PEs 806 and 808 may receive an instruction to perform a transmit-to-west operation based at least on the data associated with a first block stored in the first register of the PEs 806 and 808. In some embodiments, the instruction causes the PEs 806 and 808 to perform a transmit-to-west operation associated with the first block, where the PE 808 may transmit at least a portion of the data associated with the first block ("x 2") stored in the first register of the PE 808 to the PE 806, and the PE 806 may store that data in the third register (e.g., in the portion corresponding to the upper bits of the third register) of the PE 806. The PE 806 may also transfer at least a portion of the data associated with the first block ("x1") in the first register of the PE 806 to the third register (e.g., the portion corresponding to the lower bits of the third register) of the PE 806. In this manner, the PEs 806 and 808 may shift portions of the data stored in each of the registers involved in a transfer operation between registers of each PE and within registers of each individual PE to cause a transfer-to-west operation to be performed.In some embodiments, the instruction may cause the PE 806 to perform one or more additional operations. For example, the instruction may cause the PE 806 to perform one or more arithmetic operations, which may include adding, subtracting, multiplying, or dividing the value stored in at least a portion of the third register of the PE 806 to / from / with one or more or by one or more values stored in at least a portion of the first register of the PE 806. In examples, the instruction may cause the PE 802 to perform one or more additional transmissions. It should be appreciated that the instructions, in some embodiments, may include sequences of shifts and arithmetic operations to be performed such that one or more functions are performed by the PEs of the PE configuration, as described above. In some embodiments, the PEs may then transmit the data back (via one or more write currents) to the DLSU.In FIG. 8D, the transfer operations including PEs 806 and 808 are illustrated with respect to transfer operations along lanes. In some embodiments, the PEs 806 and 808 may receive data associated with a first block and / or a second block and store the data in the registers of each PEs 806 and 808, similar to described with respect to FIG. 8C. As shown, the data associated with the first block received from the PE 806 may be stored in conjunction with a first track (e.g., at least a portion of a register associated with a path including one or more displacements within or between the PEs 806 and 808) corresponding to a value ("x 0") representing a first pixel of the first block and a second track corresponding to a value ("x 1") representing a second pixel of the first block adjacent to the first pixel of the first block. Similarly, the data associated with the second block received from the PE 806 may be stored in association with a first track corresponding to a value ("y0") representing a first pixel of the second block and a second track corresponding to a value ("y1") representing a second pixel of the second block adjacent to the first pixel of the second block. The PE 808 is shown to similarly receive data associated with the first block and the second block, where the data associated with the first block (first track: "x2"; second track: "x3") and the second block (first track: "y2"; second track "y3") are stored in the first register and the second register of the PE 808, respectively.In some embodiments, the PEs 806 and 808 may each receive an instruction from the PE controller connected to the PEs 806 and 808 to perform a transmit-to-west operation based at least on the data associated with a first block stored in the first register of the PEs 806 and 808. In some embodiments, the instruction causes the PEs 806 and 808 to perform a transmit-to-west operation associated with the first block, where the PE 808 may transmit at least a portion of the data associated with the first block ("x 2") stored in the first track of the PE 808 to the PE 806, and the PE 806 may store that data in the third register (e.g., in the portion corresponding to the third track of the third register) of the PE 806. The PE 806 may also transfer at least a portion of the data associated with the first block ("x1") in the second track of the PE 806 to the third register (e.g., the portion corresponding to the first track of the third register) of the PE 806. In this manner, the PEs 806 and 808 may shift portions of the data stored in each of the registers across lanes involved in a transfer operation between registers of each PE and within registers of each individual PE to perform a transfer-to-west operation.Referring to FIG. 8E, the PEs 806 and 808 may receive data associated with a first block and / or a second block and store the data in the registers of the PEs 806 and 808. As shown, the data associated with the first block received and stored by the PE 806 may include a set of lower bits corresponding to a value ("x 0") representing a first pixel of the first block and a set of higher bits corresponding to a value ("x 1") representing a second pixel of the first block adjacent to the first pixel of the first block. Similarly, the data associated with the second block received from the PE 806 may include a set of lower bits corresponding to a value ("y0") representing a first pixel of the second block and a set of higher bits corresponding to a value ("y1]") representing a second pixel of the second block adjacent to the first pixel of the second block. The PE 808 is shown to similarly receive data associated with the first block and the second block, where the data associated with the first block (upper bits: "x3"; lower bits: "x2") and the second block (upper bits: "y3"; lower bits "y2") are stored in the first and second registers of the PE 808, respectively. As shown in FIG. 8E, in the PPE configuration, the PEs 806 and 808 are logically adjacent to each other.In some embodiments, the PEs 806 and 808 may each receive an instruction from a PE controller connected to the PEs 806 and 808 to perform a transmit-to-west operation based at least on the data associated with a first block and the second block stored in the first register and the second register of the PEs 806 and 808, respectively. In some embodiments, the instruction causes the PEs 806 and 808 to perform a transmit-to-west operation associated with the first block during which the PE 808 may transmit to the PE 806 at least a portion of the data associated with the second block (lower bits: "y2") stored in the second register of the PE 808, and the PE 806 may store that data in the third register (e.g., in the portion corresponding to the upper bits of the third register) of the PE 806. The PE 806 may also transfer at least a portion of the data associated with the first block (upper bits: "x1") in the first register of the PE 806 to the third register (e.g., the portion corresponding to the lower bits of the third register) of the PE 806. In this way, the PEs 806 and 808 may shift portions of the data stored in each of the registers involved in a transfer operation between registers of each PE (involving a wrap-around connection) and within registers of each individual PE to perform a transfer-to-west operation.In FIG. 8F, the transfer operations with PEs 806 and 808 are illustrated with respect to transfer operations along lanes. In some embodiments, the PEs 806 and 808 may receive data associated with a first block and / or a second block as described herein and store the data in the registers of each PEs 806 and 808, similar to described with respect to FIG. 8E. As shown, the data associated with the first block received from the PE 806 may be stored in conjunction with a first track (e.g., at least a portion of a register associated with a path including one or more displacements within or between the PEs 806 and 808) corresponding to a value ("x0") representing a first pixel of the first block and a second track corresponding to a value ("x1") representing a second pixel of the first block adjacent to the first pixel of the first block. Similarly, the data associated with the second block received from the PE 806 may be stored in conjunction with a first track corresponding to a value ("y0") representing a first pixel of the second block and a second track corresponding to a value ("y1") representing a second pixel of the second block adjacent to the first pixel of the second block. The PE 808 is shown to similarly receive data associated with the first block and the second block, where the data associated with the first block (first track: "x2"; second track: "x3") and the second block (first track: "y2"; second track "y3") is stored in the first register and the second register of the PE 808, respectively. As shown in FIG. 8F, the PEs 806 and 808 are logically adjacent to each other in the PPE configuration.In some embodiments, the PEs 806 and 808 may each receive an instruction from the PE controller connected to the PEs 806 and 808 to perform a transmit-to-west operation based at least on the data associated with a first block stored in the first register of the PEs 806 and 808. In some embodiments, the instruction causes the PEs 806 and 808 to perform a transmit-to-west operation associated with the first block, where the PE 808 may transmit at least a portion of the data associated with the first block ("x 2") stored in the first track of the PE 808 to the PE 806, and the PE 806 may store that data in the third register (e.g., in the portion corresponding to the third track of the third register) of the PE 806. The PE 806 may also transfer at least a portion of the data associated with the first block ("x1") in the second track of the PE 806 to the third register (e.g., the portion corresponding to the first track of the third register) of the PE 806. In this manner, the PEs 806 and 808 may shift portions of the data stored in each of the registers across lanes involved in a transfer operation between registers of each PE and within registers of each individual PE to perform a transfer-to-west operation.While aspects of the present disclosure are discussed in relation to a single transmit-to-west operation, it should be appreciated that transmit sequences may result in different transmission directions. For example, with respect to the two-dimensional PPE discussed in FIG. 9, the PEs may be commanded to perform multiple transmit-to-west operations such that the data is rotated. An example transfer sequence may include: transferring data in the respective registers of the PEs, as shown in FIG. 9 as follows: transferring data in PE registers storing block 00 with data in PE registers storing block 01; transferring data in PE registers storing block 01 with data in PE registers storing block 00; transferring data in PE registers storing block 10 with data in PE registers storing block 11; transferring data in PE registers storing block 11 with data in PE registers storing block 10; transferring data in PE registers storing block 20 with data in PE registers storing block 21; and transferring data in PE registers storing block 21 with data in PE registers storing block 20. In this way, the PEs can be commanded to perform data moves to allow rotation of the data within the PPE. This also allows for complex functions to be performed, such as matrix multiplication for an equation C=A*B, such that each row i of A may meet with each column j of B to contribute to C[i][j]. Additionally, while aspects of the present disclosure are discussed with respect to operations that may be performed according to local PEs (e.g., filtering operations), the PEs of the PPE may be commanded to process two-dimensional data that does not correspond to the size of the PPE. For example, as compared to a 2x3 or 3x3 block, the PPE may be configured to receive data associated with longer one- or two-dimensional shapes (e.g., 1x10 or 1x100). The PEs may then be provided with instructions to perform operations based at least on data stored in registers of PEs that are physically or logically eastern and westjacent of the PE, and in some cases, no instructions may be provided to perform operations based at least on values of PEs that are physically or logically north or south of the PE.In an example with respect to an 8x10 PPE (FIG. 1B ), for data analyzed such that the data blocks have a greater difference in width to height than other blocks (e.g., 848 x 4 (wide and short) or 8 x 1024 (thin and high)), one or more DMA transfers may include mapping the data into blocks represented as 848 x 10 and 8 x 1024 as 32 x 1030. In examples with an alternative one-dimensional organization, the PEs may receive and operate 320 x 1 blocks. This can result in images where 848 x 4 can be mapped as 2240 x 4 and 8 x 1024 can be mapped as 8 x 1280 (for the problem size "thin and high", the PPE supports transposed vector charges, where the DMA transfer required in loading the PPE involves swapping the rows and columns of data loaded into the PPE to operate at a block size of 1 x 320). This segmentation and mapping can cause the load on the PE array to be improved to 91% and 80%, respectively.FIG. 9 is an exemplary illustration of a data layout across registers in PEs of a two-dimensional accelerator, in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the example communications between accelerators may be included in and / or may include components, features, and / or functionality similar to those of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.In some embodiments, the example representation of a data layout across registers may include (e.g., be implemented by) components of one or more accelerators (e.g., one or more accelerators, such as PPEs 118 of FIG. 1A ). For example, communications between accelerators may be implemented by PEs 902 and 904 physically adjacent to each other within a PPE configuration 900. It should be appreciated that the PPE configuration 900 may include additional PEs or other PE configurations than those illustrated in FIG. 9. In some embodiments, PEs 902 and 904 may be the same or similar to PE 170 of FIG. 1C.In some embodiments, PEs 902 and 904 may receive data associated with an image. For example, PEs 902 and 904 may receive data associated with an image, where the image is segmented into multiple blocks (or tiles). In the illustrated example, the image may be segmented along two columns and three rows. For example, the image may be segmented (e.g., during one or more DMA transfers) such that corresponding portions are provided to the accelerator such that the respective bits of a particular block are loaded into the corresponding PEs 902 and 904.In some embodiments, bits of a first block (e.g., block 00) may be loaded into PEs 902 and 904. In this example, the bits of the first block may be stored in the first register of the respective PEs 902 and 904. This process may be repeated for the remaining blocks in any order. For example, blocks 01, 10, 11, 20, and 21 may be loaded into PEs 902 and 904 one after another. In another example, blocks 10, 20, 01, 11, and 21 may be loaded into PEs 902 and 904 one after another. The PEs 902 and 904 may then be commanded to perform one or more operations (e.g., shifts and arithmetic operations as described with respect to FIGS. 8A-8E ). Once the operations are complete, the PEs 902 and 904 may transmit the data associated with the blocks (which may be updated based at least on the operations performed) via a write stream from the accelerator.FIG. 10A is a flowchart of an example method 1000 for performing transmissions between accelerations, in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the frame 600 may be included in or may include components, features, and / or functionality similar to those of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.The method 1000 includes, at block 1002, receiving data associated with a first pixel and data associated with one or more instructions. For example, one or more PEs (which may be the same or similar to the PE 170 of FIG. 1C and / or the PEs 802- 808) may be interconnected to form a PPE (e.g., a PPE that may be the same or similar to the PPEs 118 of FIG. 1A ). In this example, each PE may be configured to receive the data associated with a first pixel obtained from the PPE via a read stream. The data associated with each pixel may be divided into a number of rows corresponding to a width of the PPE (e.g., measured from the number of PEs in each row of the PPE), and corresponding portions of the data (e.g., representing one or more pixels of an image) may be provided to the respective PEs in a first row of PEs within the PPE. The data may then be transmitted sequentially via the PEs of the PPE (e.g., via transmit-to-north operations) until the data associated with the pixels for a particular block or set of blocks (e.g., first block, second block, and / or the like) is received and stored in corresponding registers of the PEs (this process is also referred to as "loading" the PPE). For example, as shown in FIG. 1B, the PEs 152 a- 152 hmay each receive data associated with one or more pixels via one or more read streams corresponding to a block of an image being loaded into the PPE 140. The data may then be transmitted sequentially (e.g., from PE 152 ato PE 154 a, and so forth, until PE 170 ais reached) to one or more other PEs (e.g., via transmit-to-north operations) until the data associated with each pixel has been received and stored in the corresponding registers of the PEs. In this way, the PPE may be loaded such that data associated with multiple blocks of an image is stored in corresponding registers of the PEs. It should be understood that while the discussion regarding data transmitted between PEs includes data associated with pixels of an image, the techniques described herein are not limited to image data and may be applied to any form of data suitable for processing via a two-dimensional accelerator, such as the PPEs discussed herein.In some embodiments, the PEs may be configured to receive data associated with an instruction (e.g., a SIMD instruction). For example, the PEs may each be connected to a PE controller configured to transmit the instructions to the PEs of the PPE. In this example, the instructions may represent one or more sequences of data transfers between registers of the PEs or within registers of a single PE and / or one or more arithmetic operations to be performed based at least on data stored in the registers of the PEs. In some embodiments, the instructions may be associated with one or more DMA transfers as described herein.In some embodiments, the PEs may perform one or more data transmissions. For example, the PEs may transfer the data associated with the first pixel (corresponding to the first block) into a register of another PE in the PPE. In this example, the other PE may be physically or logically located in the north, south, east, or west relative to the PE that is transmitting the data. In another example, the PEs may transfer the data associated with the first pixel to another register within the PE. For example, if the registers are configured to contain upper and lower bits (e.g., corresponding to halfword data types), the PE may transfer data internally from a source register to a destination register. For clarity, registers containing data that is later transferred may be referred to as source registers, and registers receiving the data from another register may be referred to as destination registers.The method 1000 includes, at block 1004, determining an updated first pixel based at least on the first pixel and the one or more instructions. For example, one or more PEs of the PPE may determine an updated first pixel based on at least data associated with the first pixel and the one or more instructions by adding, subtracting, or multiplying a value representing the first pixel to determine an updated value corresponding to the updated first pixel. This process may be repeated, for example, to cause the PEs to perform uniform operations on each individual pixel loaded into the PPE.In some embodiments, one or more of the PEs may receive data associated with at least a second pixel. As described above, an image may be divided into multiple blocks, and each block may be further divided based at least on the size of the block and / or the size of a source register into which the data in the PPE is loaded. The data can then be loaded into the corresponding registers of the PEs of the PPE. In some embodiments, an instruction received from the PEs may cause the data associated with the first pixel to be transmitted to one or more different PEs by one or more data transmissions. For example, as shown in FIG. 8A, an instruction may cause the data stored in register 1 of PE 804 to be transferred to the third register of PE 802. In this example, the PE 802 may then determine an updated first pixel based at least on the data associated with the first pixel loaded into the first register of the PE 802 and the data associated with the first pixel loaded first into the first register of the PE 804 and subsequently transferred into the third register of the PE 802. In some embodiments, an instruction received from the PEs may cause the data associated with the first pixel stored in a register of a PE to be transferred to one or more other registers of the PE by one or more data transfers. For example, as shown in FIG. 8C, an instruction may cause the data stored in register 1 of PE 806 (shown as "x1") to be transferred to the third register of PE 806. The instructions may also cause data stored in register 1 of PE 808 to be transferred to the third register of PE 806. In this example, the PE 806 may then determine an updated first pixel based at least on the data associated with one or more of the pixels loaded into the first register of the PE 806 and / or the data associated with one or more pixels loaded into the third register of the PE 806.In some embodiments, the instructions provided to the PEs of the PPE by the PE controller may represent one or more sequences of data transfers between registers of the PEs or within registers of a single PE and / or one or more arithmetic operations to be performed based at least on data stored in the registers of the PEs. For example, the one or more sequences may include combinations of operations including adding, subtracting, or multiplying a value representing the first pixel and operations including transferring data associated with pixels between registers (the same PE or between PEs). In this way, the instructions may cause the PEs to perform operations that perform higher order functions, such as filter functions, bandpass filter functions, matrix multiplication functions, image processing functions, and / or the like.The method 1000 includes, at block 1006, providing as output data associated with the updated first pixel. For example, each PE of the PPE may be configured to provide (e.g., transmit) the data associated with the updated first pixel to one or more other PEs and ultimately output it via a write stream to memory (e.g., a VMEM or a DLSU). In some embodiments, the PEs may provide the data associated with the updated first pixel based on the complete performance of the operations included in the instruction.FIG. 10B is a flow diagram of an example implementation 1050 of the method of claim 10A. The implementation may be associated with implementation 1050 of a 3x3 filter. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the frame 600 may be included in or may include components, features, and / or functionality similar to those of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17.The implementation 1050 includes loading a PPE at block 1052. For example, one or more PEs (PEs configured to process data associated with halfword data types, such as PEs 806 and 808 of FIGS. 8C-8F and PEs 902 and 904 of FIG. 9 ) may receive data associated with a first pixel and a second pixel.The implementation 1050 includes, at block 1054, performing one or more transmit-to-west operations. For example, a set of interconnected PEs of a PPE may be configured to perform three sequential transmit-to-west operations. In this example, the PEs may store the data in respective registers such that once the three transmit-to-west operations have been performed, each PE contains one or more other values that are transmitted into the PE.The implementation 1050 includes performing horizontal filtering in block 1056. For example, the set of interconnected PEs of the PPE may multiply the values stored in the registers of each PE by a coefficient. In this example, the set of interconnected PEs of the PPE may perform a vector multiplication operation and one or more vector addition operations to determine values for a particular pixel. In block 1058, the PEs may round one or more values stored in the registers of each PE. The implementation 1050 includes, at block 1060, performing a transmit-to-north operation. For example, the group of PEs may perform transmit to north operations.The implementation 1050 includes performing vertical filtering in block 1062. In this example, the set of interconnected PEs of the PPE may perform a vector multiplication operation and one or more vector addition operations to determine values for a particular pixel. In block 1058, the PEs may round one or more values stored in the registers of each PE.DMA FRAME LINK SUPPORTUsing descriptors to coordinate DMA transfers can improve the performance of systems, but since these descriptors are often implemented using software, their efficient implementation may be difficult. For example, descriptors indicating criteria for DMA transfers may be configured such that DMA transfers are performed a predetermined number of times. In the case of feature tracking (determining the presence and location of objects across a set of frames), these descriptors may indicate the number of frames to be obtained and processed by accelerators such as VPUs or PPEs. The VPUs or PPEs may then implement the descriptors to obtain and process the frames as they perform operations that track the object across the frames. However, if objects remain for more frames than are indicated by the descriptors, additional descriptors may be obtained (e.g., generated) from the VPU or PPE to reconfigure the VPU or PPE as the objects are tracked further. Alternatively, objects may leave the field of view of the sensor that generates the frames, and the VPUs or PPEs may continue to perform operations according to the descriptors until the specified DMA transfers are complete. This may be inefficient because the VPUs or PPEs may be unnecessarily reconfigured (or other devices may reconfigure), thereby wasting processing resources during the reconfiguration process. Additionally or alternatively, the VPUs or PPEs may continue to perform operations according to the descriptors even though the object is no longer present in the frames. This may also result in waste of processing resources and a delay in performing subsequent operations.Systems and methods are disclosed that include configuring accelerators, such as VPUs or PPEs (alone or in coordination with a DMA system), to obtain data associated with frames from source memory (SRAM), and perform one or more operations based on the frames. In particular, in embodiments with a first mode (referred to as a "fixed frame number link"), a VPU or PPE may be configured to receive data identified by a first descriptor and a set of second descriptors in coordination with a DMA system. In examples, the VPU or PPE may obtain data associated with the frame based on the one or more descriptors (e.g., in coordination with a DMA system) and perform operations based on the data obtained in connection with the descriptors (e.g., based on the frames or portions thereof).In second mode embodiments (referred to as "continuous frame number link"), the VPU or PPE may be configured to obtain data associated with frames as represented by a first descriptor (and, in examples, one or more second descriptors) that causes the VPU or PPE to obtain the data in coordination with the DMA system. The DMA system may then be configured to iteratively perform operations using the data obtained in conjunction with the one or more descriptors (referred to as a loop) until the VPU or PPE generates and transmits to the DMA system a signal indicating that a particular loop is an end loop. This signal may be sent to the DMA system by changing a value in a frame sequence count register to indicate that the loop is no longer to continue (e.g., be interrupted).Further, descriptors may be loaded and performed in a ping-pong manier due to the configuration of VPU and PPE, such that a first descriptor may be loaded and a second descriptor may be loaded and queued for a second frame during performing DMA transfers according to the first descriptor in a first frame. This reduces down times that may occur with processors or PEs of the VPU, PPE or DMA system in connection with data acquisition according to certain descriptors.By implementing at least some of the described techniques, VPUs, PPEs, and / or DMA systems may be configured to operate independently or in coordination with each other to reduce or avoid waste (e.g., unused resources) due to "bubbles.". These bubbles may correspond to transfer gaps between performing DMA transfers corresponding to the descriptors. The techniques described herein may also save time and resources that would otherwise have to be used for individual configuration of each DMA transfer. Moreover, the systems and methods disclosed herein may reduce the complexity of control code for system construction and sequencing of frames as described herein.FIGS. 11A-11C are example sequences of frame transmissions using accelerators, in accordance with some embodiments of the present disclosure. In particular, frames 11A-11C represent performing DMA transfers to move data associated with frames generated by sensors. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the example sequences of frame transmissions using accelerators may be included in and / or may include components, features, and / or functionality similar to those of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17. In some embodiments, the example sequences of frame transfers using accelerators may include (e.g., be implemented by) components of one or more accelerators (e.g., one or more accelerators such as VPUs 116 and / or PPEs 118 of FIG. 1A ) in coordination with one or more DMA systems (e.g., one or more DMA systems such as DMA systems 114 of FIG. 1A ).Referring to FIG. 11A, an example sequence of frame transmissions using accelerators according to the first mode (the "fixed frame number link") is illustrated. In this example sequence, a VPU 1102 may be configured to obtain data associated with one or more DMA transfers (represented by one or more descriptors) when streaming tiles from one or more frames. The VPU 1102 may receive the data associated with the one or more DMA transfers from a processor (e.g., a processor identical or similar to the processor 102 of FIG. 1A ) or from a DMA system 1104. While the present disclosure is discussed with respect to a VPU 1102, it should be appreciated that other accelerators, such as a PPE, may implement some or all of the functions described herein with respect to the VPU 1102.In some embodiments, the VPU 1102 receives the data associated with the one or more DMA transfers, the data specifying a fixed number of DMA transfers to be performed. For example, the VPU 1102 may receive the data associated with the one or more DMA transfers that are performed sequentially to support one or more operations performed by the VPU 1102. In this example, the data associated with the one or more DMA transfers may be associated with (e.g., represented by) one or more frame formats, as described with reference to FIGS. 4A-4C.In the example illustrated in FIG. 11A, the VPU 1102 may receive data associated with three DMA transfers to be performed by the VPU 1102 or another device, such as a DMA system 1104. The three DMA transfers may be performed at three different times (e.g., t=0, t=1 and t=2) and / or in a sequence. In this example, the three DMA transfers may be associated with (e.g., correspond to) operations to be performed by the VPU 1102. For example, the VPU 1102 may receive an instruction to perform one or more operations on data associated with a frame represented at different resolutions (e.g., a first resolution of 2 megapixels, a second resolution of 1 megapixel, and a third resolution of 0.5 megapixels). In this example, the VPU 1102 may receive the instructions to perform the one or more operations and the data associated with the three DMA transfers from a processor (e.g., a processor identical or similar to the processor 102 of FIG. 1A ) or one or more other devices that configure the operation of the VPU 1102.In some embodiments, the VPU 1102 may cooperate with the DMA system 1104 to obtain the data specified by the DMA transfers. For example, at times t=0, t=1 and t=2, the VPU 1102 may provide data associated with discrete DMA transfers (represented by individual descriptors) to the DMA system 1104 to cause the DMA system 1104 to transfer the data associated with the individual frames to the VMEM based at least on one or more operations to be performed by the VPU 1102 (not explicitly shown in FIGS. 11A-11C ). In another example, the VPU 1102 may provide the data associated with the discrete DMA transfers to the DMA system 1104, and during performance of operations by the VPU 1102, the VPU 1102 may provide signals to the DMA system 1104 to cause (e.g., trigger) certain DMA transfers. In this example, the VPU 1102 may provide data associated with a frame format (e.g., at or before time t=0) that specify each of the DMA transfers to the DMA system 1104 at or before a particular time (e.g., at or before time t=0) and configure the DMA system 1104 to perform the DMA transfers in response to trigger signals provided by the VPU 1102.With continued reference to FIG. 11A, the VPU 1102 may provide (e.g., send) a first trigger signal to the DMA system 1104 as the VPU 1102 performs or prepares for performance the one or more operations. The first trigger signal may cause the DMA system 1104 to perform at least one DMA transfer (e.g., obtain data associated with a frame from the source memory such that the frame is scanned to generate a 2 megapixel image before being stored in the target memory). When the DMA transfer is complete (e.g., the data associated with the frame is stored in the target memory), the DMA system 1104 may send a signal to the VPU 1102 indicating that the transfer is complete. In some embodiments, the VPU 1102 may send a second trigger signal to cause the DMA system 1104 to perform at least a second DMA transfer. During the at least one second DMA transfer, at least a portion of the data associated with the frame involved in the first DMA transfer may be transferred to the target memory such that the frame is scanned based at least on the operations performed by the DMA system to form a 1 megapixel image. When the DMA transfer is complete, the DMA system 1104 may send a signal to the VPU 1102 indicating that the transfer is complete. In some embodiments, the VPU 1102 may send a third trigger signal to cause the DMA system 1104 to perform at least a third DMA transfer. During the at least one third DMA transfer, the data associated with the frame involved in the first DMA transfer may be transferred to the target memory such that the frame is scanned based at least on the operations performed by the DMA system to form a 0.5 megapixel image. When the DMA transfer is complete (e.g., the data associated with the frame is stored in the target memory), the DMA system 1104 may send a signal to the VPU 1102 indicating that the transfer is complete. In this manner, the VPU 1102 and the DMA system 1104 may cooperate to perform a fixed number of DMA transfers that include (e.g., are associated with) a common frame or a particular sequence of operations performed by the VPU 1102.Referring to FIG. 11B, an example of a continuous frame link including a configuration frame and streaming frames plus padding is shown. As shown, the VPU 1102 may receive data associated with a continuous number of DMA transfers to be performed (represented by descriptors corresponding to a sequence of DMA transfers that are not specified). For example, the VPU 1102 may receive instructions to continuously perform one or more operations. In one example, the operations may be associated with (e.g., involved in) performing object tracking across multiple frames until the one or more objects in one or more of the frames are no longer detected. In this example, the VPU 1102 may provide data associated with at least one frame format that indicates one or more regions (e.g., up to 32 regions and / or the like) within the frame that are involved in corresponding operations performed by the VPU 1102 to track the one or more objects. In some embodiments, the VPU 1102 may also specify that the one or more DMA transfers should be repeated until the VPU 1102 provides a subsequent signal indicating that the DMA transfers are complete. For example, the VPU 1102 may provide a signal to indicate that the DMA transfers are complete, the signal causing a value in a register of the DMA system indicating that the DMA transfers are complete.With continued reference to FIG. 11B, the VPU 1102 may first send data associated with the sequence of DMA transfers to the DMA system 1104. The data associated with the sequence of DMA transfers may be associated with a frame format that causes the DMA system to perform (e.g., configure) one or more DMA transfers. Once the DMA system 1104 is configured based at least on the frame format (shown as a "configure" block in FIG. 11B ), the DMA system 1104 may cause one or more DMA transfers to be performed according to a first frame ("frame 1") until a last frame ("frame n") is reached. While the DMA system 1104 is shown performing DMA transfers for frames 1- n, it should be appreciated that each frame may represent a portion of a particular frame. In these examples, the frame format may indicate an offset, a length, and a width corresponding to a region within the given frame, for example as shown in FIG. 4B.In some embodiments, once DMA system 1104 completes the DMA transfers, DMA system 1104 may send a signal to VPU 1102 indicating that the transfer sequence is completed. In this example, the DMA system 1104 may check to determine whether or not a signal is received from the VPU 1102 (e.g., a particular value is set in a register). The signal may indicate that the DMA system 1104 is to refrain from one or more DMA transfers (thereby interrupting the loop shown in FIG. 11B ). For example, the VPU 1102 may perform operations that result in the determination that one or more objects in one or more frames (or regions) are no longer detected, and the VPU 1102 may send a signal to the DMA system 1104 indicating that the DMA transfers should no longer be performed. In examples where the DMA system 1104 does not receive a signal from the VPU 1102, the DMA system 1104 may iteratively repeat or pause (e.g., jam) the DMA transfers indicated by the VPU 1102 until a signal, such as a trigger signal, is received to cause one or more subsequent DMA transfers to be performed.In some embodiments, the VPU 1102 may determine one or more updates for one or more of the DMA transfers that are continuously performed by the DMA system 1104. For example, the VPU 1102 may determine that one or more operations performed by the VPU 1102 indicate that an object associated with a particular DMA transfer has been moved from a first region within the frame to a second region within the frame. In this example, the VPU 1102 may determine an update of the portion of the frame format corresponding to the motion of the object within the frame and provide the update of the portion of the frame format to the DMA system 1104. In this example, the DMA system may continue to perform the specified DMA transfers according to the original configuration and update. In this way, the DMA system can be iteratively updated without having to reconfigure the entire sequence of DMA transfers at each iteration of the sequence. This, in turn, may allow the DMA system 1104 to perform the DMA transfers more quickly, as some (or in some cases all) data may be reused in the registers in which the instructions involved in the DMA transfers are stored without the VPU 1102 or other processors (e.g., a function block 110 of FIG. 1A ) involved.Referring to FIG. 11C, an example of a continuous frame link including configuration frames and random region access frames plus padding is shown. As described herein, the example of FIG. 11C may be implemented in the implementation of a feature tracker. As shown, the VPU 1102 may receive data associated with multiple sets of DMA transfers, and may configure the DMA system 1104 to perform the sets of DMA transfers while the DMA system 1104 performs one or more other sets of DMA transfers. For example, at a first time (t=0), the VPU 1102 may trigger the DMA system 1104 by providing the DMA system with data associated with a channel (e.g., an independent virtual path) along which DMA data transfers are performed. The DMA system 1104 may also perform one or more DMA transfers according to data associated with a frame format received at an earlier time (e.g., a time prior to time t=0).In this example, at a second time (t=1), the VPU 1102 may re-trigger the DMA system 1104 by providing the DMA system 1104 with data associated with a second frame format associated with the same channel. The DMA system 1104 may also perform the one or more DMA transfers according to data associated with a frame format received at an earlier time (time t=0). This process may be iteratively repeated (e.g., at times t=2, t=3, etc.) such that VPU 1102 configures DMA system 1104 to perform DMA transfers while DMA system 1104 concurrently performs previously configured DMA transfers. In this way, DMA transfers that would otherwise be reserved for separate channels may be configured to be performed on the same channel, thereby reducing the need for additional channels and / or enabling channels to perform additional DMA transfers. Using the example illustrated in FIG. 11C, by associating four frame formats (e.g., frame formats associated with descriptor addressing frame types (such as shown in FIG. 4B ) paired with four corresponding frame formats associated with random region addressing frame types (such as shown in FIG. 4C )), the DMA system 1104 may be configured to perform sequences of DMA transfers. And, in cases where a DMA system is configured to process 32 DMA transfers (corresponding to up to 32 objects in an object) in blocks of four frames, the VPU 1102 may configure the DMA system 125 times when up to 4000 objects are covered, as opposed to up to 500 times that would be required on four separate channels.FIG. 12 is a flow diagram of an example method 1200 for sequencing frame transmissions using speeds up, in accordance with some embodiments of the present disclosure. In some embodiments, aspects of method 1200 may be performed by one or more devices identical to or similar to one or more of the devices of FIGS. 1A-1C, such as DMA systems 114, VPUs 116, PPEs 118, and / or processor 102. In embodiments, one or more other devices of FIG. 1A may perform one or more aspects of methods 500. In some embodiments, one or more of the frame formats described herein may be identical or similar to the one or more frame formats of FIGS. 4A-4C.The method 1200 includes, at block 1202, determining a first DMA transfer. For example, a device (e.g., a processor, a VPU, and / or a PPE) may determine the first DMA transfer. For clarity, the non-limiting examples described herein are described with respect to operations performed by a VPU; however, it should be understood that the operations described herein may also be performed by one or more other devices, such as a processor, a PPE, a DMA system, or any other suitable device described herein, alone or in coordination.In some embodiments, the VPU may determine the first DMA transfer based at least on generation of frame data associated with a frame by a sensor. For example, during operation of a robotic system, such as an automated vehicle, a sensor, such as a camera, a LiDAR sensor, a RADAR sensor, and / or the like, may generate data corresponding to the frames (e.g., images, point clouds, and / or the like) generated by the sensor. In this example, the VPU may determine the first DMA transfer based at least on the sensor (also referred to as frame data) generating the data and one or more operations the VPU is to perform. The operations may include, without limitation, operations associated with one or more image processing operations (including processing frames or portions of the frames), one or more prediction operations (including identifying objects represented by one or more frames), one or more object tracking operations (including tracking objects as they move past an environment represented in successive frames), one or more trajectory prediction operations (including predicting future positions of objects as they move through an environment represented in the successive frames), and / or any other suitable operations.In some embodiments, the first DMA transfer may include the transfer of data from a source memory (e.g., a system memory that is the same or similar to memory 108 of FIG. 1A ) to a destination memory (e.g., a VMEM that is the same or similar to VMEMs 112 of FIG. 1A ). For example, the first DMA transfer may include transferring data from the source memory to the target memory to enable the VPU to perform one or more operations based at least on the data. In some embodiments, the first DMA transfer may include multiple independent DMA transfers. For example, the first DMA transfer may include a sequence of DMA transfers associated with one or more operations the VPU is configured to perform. In some embodiments, the sequence of DMA transfers may be performed independently of the VPU, a DMA system, and / or the like. For example, the VPU may configure the DMA system such that the one or more DMA transfers are performed once and a signal is returned indicating whether the transfers are complete (e.g., successful) or not complete (e.g., in progress or unsuccessful).In some embodiments, the first DMA transfer may include communication between the VPU and the DMA system during the DMA transfer. For example, if the first DMA transfer is associated with a sequence of DMA transfers, the VPU may configure the DMA system to perform one or more of the DMA transfers based at least on (e.g., in response to) the DMA system receiving signals to initiate one or more of the DMA transfers. These signals (also referred to as triggers) may be transmitted from the VPU to the DMA system based at least on (e.g., in response to) execution of one or more corresponding operations by the VPU.The method 1200 includes, at block 1204, determining at least one second DMA transfer. For example, the VPU may determine the at least one second DMA transfer. In some embodiments, the at least one second DMA transfer may be the same as or similar to the first DMA transfer. For example, the VPU may determine the at least one second DMA transfer of data from the source memory to the target memory. The at least one second DMA transfer may be based on at least one sequence of DMA transfers involved in operations performed by the VPU. In this example, the sequence of DMA transfers may correspond to operations configured to be executed by the VPU according to the frame. In an example, in the context of object tracking, the one or more second DMA transfers may correspond to the transfer of data associated with (e.g., representing) regions of the frame indicated by the first DMA transfer. In this example, the VPU may be configured to perform operations to track positions of objects relative to the frame and / or relative to regions of the frame specified by the one or more second DMA transfers.In some embodiments, the VPU may configure the DMA system to perform at least the first DMA transfer and the one or more second DMA transfers. For example, the VPU may generate and provide data associated with at least one descriptor (described below) to cause the DMA system to perform the first DMA transfer and the one or more second DMA transfers. In some examples, the VPU may configure the DMA system such that the first DMA transfer and the one or more second DMA transfers are performed without coordination with the VPU. In other examples, the VPU may configure the DMA system to perform the first DMA transfer and the one or more second DMA transfers based at least on communication with the VPU. In some of these examples, the VPU may configure the DMA system to perform the first DMA transfer and the one or more second DMA transfers based at least on indications of the VPU sent to the DMA system to initiate one or more of the first and at least one second DMA transfer. In examples, the VPU may configure the DMA system to continuously perform the first DMA transfer and the one or more second DMA transfers. In these examples, the DMA system may perform the first DMA transfer and the one or more second DMA transfers until the DMA system receives an instruction from the VPU to not perform one or more of the first DMA transfers and / or the one or more second DMA transfers. The VPU may provide the indication by changing a value in a register of the DMA system that is checked by the DMA system at each iteration.In some embodiments, the VPU may determine updates for one or more of the first DMA transfers and at least one second DMA transfer. For example, the VPU may determine updates for one or more of the first DMA transfers and the at least one second DMA transfer based on at least one or more operations performed by the VPU. In an example where the VPU performs one or more object tracking operations, the VPU may provide data to the DMA system that updates the one or more descriptors corresponding to the first DMA transfer and the one or more second DMA transfers. The updates may represent updates of an offset and the width and / or height of an object tracked across frames.The method 1200 includes, at 1206, generating data associated with at least one descriptor based at least on the first DMA transfer and the at least one second DMA transfer. For example, the VPU may generate the data associated with at least one descriptor. The VPU may generate the data associated with the at least one descriptor, the data configured to cause one or more DMA transfers to be performed based at least on the operations performed by the DMA system from the source memory to the destination memory according to the at least one descriptor. In this example, the at least one descriptor may represent the first DMA transfer and the at least one second DMA transfer. In this way, the descriptor may represent instructions that enable multiple DMA transfers. By consolidating (e.g., associating) the instructions corresponding to multiple DMA transfers (and corresponding DMA transfer types) into a single descriptor, the operations required to configure a device to perform the DMA transfers can be reduced. This may improve upon techniques in which individual descriptors are configured for individual DMA transfers. As described herein, the data associated with the at least one descriptor may be configured to cause one or more discrete sets of DMA transfers, one or more continuous DMA transfers, and / or the like.The method 1200 includes, at 1208, providing the data associated with the at least one descriptor to at least one device to cause the at least one device to obtain data according to the at least one descriptor. For example, the VPU may provide the data associated with the at least one descriptor to at least one device of an accelerator, such as the DMA system, to cause the DMA system to obtain at least data according to the at least one descriptor. In this example, the data obtained according to the at least one descriptor may correspond to the frame data associated with at least a portion of a frame stored in the source memory.In some embodiments, the VPU may provide the data associated with the at least one descriptor to the DMA system to cause the DMA system to perform the DMA transfers, wherein the at least one descriptor specifies a discrete (e.g., fixed) number of DMA transfers. For example, the data associated with the at least one descriptor may specify a discrete number of DMA transfers for a particular frame (e.g., by storing the data associated with the frame at different resolutions in target memory). In this example, the DMA system may complete the DMA transfers according to the sequence. For example, the DMA system may complete the DMA transfers according to the sequence without turning on the VPU. In examples, the DMA system may complete the DMA transfers according to the sequence with VPU turned on. For example, the DMA system may perform one or more of the DMA transfers based at least on the DMA system receiving a signal from the VPU that triggers the DMA system to perform the DMA transfers.In some embodiments, the VPU may provide the data associated with the at least one descriptor to the DMA system to cause the DMA system to perform the DMA transfers, wherein the at least one descriptor dispenses with specifying a discrete number of DMA transfers. For example, the data associated with the at least one descriptor may specify a fixed number of DMA transfers for a particular frame. In this example, the DMA system may complete the DMA transfers according to the sequence and iteratively repeat the sequence. For example, the DMA system may complete the DMA transfers according to the sequence and send a signal to the VPU that the sequence is complete. In this example, the DMA system may then repeat the DMA transfers until a signal is received from the VPU to no longer perform the DMA transfers. In this way, the DMA system may be configured to perform one or more DMA transfers without reconfiguration, thereby saving resources that would otherwise be required for reconfiguration of the DMA system.In some embodiments, the VPU may provide the DMA system with the data associated with the at least one descriptor while the DMA system is performing one or more DMA transfers according to previously generated descriptors. For example, the VPU may provide the data associated with the at least one descriptor to the DMA system to configure the DMA system to perform one or more DMA transfers based at least on one or more different DMA transfers being performed by (e.g., after performing) the DMA system. In this example, the data associated with the at least one descriptor may be associated with a frame generated at a future time. In this way, the VPU can coordinate the configuration and performance of a DMA system according to the ping-pong principle, wherein the DMA system continuously receives data from the source memory and transfers it to the destination memory. This can reduce resource downtime that would otherwise be associated with configuring and reconfiguring the DMA system to perform successive DMA transfers.PROGRAMMING MULTI-DIMENSIONAL SIMD PROCESSORSMulti-dimensional SIMD processors, such as the PPEs described herein, may significantly improve the computing power of systems implementing parallel processing algorithms. For example, multi-dimensional SIMD processors may load data (e.g., into the respective PEs of a PPE) and execute one or more SIMD instructions without additional shared memory calls, thereby saving the time spent reading and writing data associated with intermediate results in that memory. While performing such operations in a multi-dimensional SIMD processor may improve computing power, it may be difficult to configure systems operating according to higher level instructions (e.g., programmed in languages such as C / C++) such that the multi-dimensional SIMD processor is efficiently configured to perform SIMD instructions.Embodiments disclosed herein include implementing techniques for mapping higher level instructions to the SIMD instructions. Compilers are also disclosed that are capable of processing the data types associated with particular accelerators (e.g., VPUs), as well as multi-dimensional SIMD processors, such as PPEs. By mapping higher level instructions represented using programming languages such as C / C++ to SIMD instructions executable by the multi-dimensional SIMD processors disclosed herein, the present disclosure reduces complexity in programming such processors. This may also reduce the overall time required to configure higher order systems (e.g., software stacks for automated or semi-automated vehicles, image processing systems, machine learning based systems, and / or the like) to operate in accordance with the systems disclosed herein and improve interoperability.FIG. 13 is a diagram illustrating an implementation of a process 1300 for generating an example accelerator instruction, in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. In some embodiments, the example sequences of frame transmissions using accelerators may be included in and / or may include components, features, and / or functionality similar to those of the example autonomous vehicle 1500 of FIGS. 15A-15D, the example computing device 1600 of FIG. 16, and / or the example data center 1700 of FIG. 17. In some embodiments, the example sequences of frame transfers using accelerators may include (e.g., be implemented by) components of one or more accelerators (e.g., one or more accelerators such as VPUs 116 and / or PPEs 118 of FIG. 1A ) in coordination with one or more DMA systems (e.g., one or more DMA systems such as DMA systems 114 of FIG. 1A ).As shown in FIG. 13, a processor 1302 receives an instruction to be performed by an accelerator at 1320. In the examples described herein, processor 1302 may be the same or similar to processor 102 of FIG. 1A and function block 1304 may be the same or similar to function blocks 110 of FIG. 1A. In some embodiments, the accelerator may include a PPE 1306 that is identical or similar to the PPEs 118 of FIG. 1A. As described herein, the PPE 1306 may be a multi-dimensional SIMD processor that includes a plurality of PEs logically arranged in a 2D array. Although examples are described herein with reference to a PPE, it should be understood that the present disclosure is not limited to the PPE and that any other suitable multi-dimensional SIMD processor may be capable of performing one or more of the operations described herein.In some embodiments, the instruction to be performed by the accelerator may be represented in a first programming language. For example, the instruction to be performed may be represented in a programming language such as C, C++, Python, and / or other high level programming languages associated with hardware abstractions. In some embodiments, the instructions may specify one or more aspects of a data path. For example, the instructions may specify one or more memory locations, VMEMs, a DLSU, or one or more registers of one or more PEs in a PPE corresponding to one or more portions of data to be processed by the PPE, and / or may specify (e.g., indicate) one or more data transfers to be performed.In 1322, the processor 1302 generates an accelerator instruction. In examples, the processor 1302 may perform one or more operations to generate the accelerator instruction. For example, at 1324, processor 1302 may provide the instructions (obtained at 1320) to a system for mapping 2D SIMD primitives (also referred to as mapping system 1302 a). The imaging system 1302 amay include logic, a look-up table, combinations thereof, and / or the like that receives the instructions and determines a correspondence to one or more operations to be performed (e.g., in a sequence) by an accelerator. In some embodiments, the one or more operations to be performed may be represented using one or more lower level languages, such as assembler code. In 1326, the one or more operations to be performed by the accelerator may then be output by the imaging system 1302 aas accelerator instructions.As described above, the accelerator instructions may be represented using lower level languages, such as assembler code. For example, the imaging system 1302 amay be associated with a compiler that converts the instructions represented in C++ at 1324 into accelerator instructions 1326 represented in an assembly language. In some embodiments, the compiler may be configured to process data associated with one or more types of data associated with one or more instructions (also referred to as primitives or PPE primitives) that, in combination, form the accelerator instructions. In one example, the following types of data may be associated with the instructions received at 1320: class intx property(48 bit signed); class shortx property(24 bit signed); class v2d_intx property(vector intx[MATX*MATIY]); class v2d_shortx property(vector shortx[2*MATX*MATY]); class dv2d_intx property(vector intx[2*MATX*MATY]); class dv2d_shortx property(vector shortx[4*MATX*MATY]. In this example, the classes may include object-oriented C++ programmed building blocks representing integers, shorts, two-dimensional vectors (v2d), duplicate two-dimensional vectors (dv2d), etc. The mapping system 1302 amay receive data associated with these data types and map the data to primitives configured to cause a multi-dimensional SIMD processor (e.g., the PPE 1306) to perform one or more operations. For example, an instruction may be represented at 1320 as follows: and during programming, the processor may receive inputs represented as follows: The processor may then provide relevant portions of the instruction to the imaging system 1302 ato cause the imaging system 1302 ato issue a vadd (vector addition) instruction as an accelerator instruction in 1326 in response to the term in the instruction 1324 ("c=atb").In some embodiments, in addition to the arithmetic operations described above, processor 1302 may generate accelerator instructions that cause a PPE 1306 to transfer data between adjacent PEs, as shown by PEs 170 of PPE 140 in FIG. 1B. For example, the following one or more data transfers between PEs may be represented as: VXfer<direction><type> Vsrc1, Vsrc2 / Rsrc2, Vdst<direction>={North, South, East, West}<type>={H (for Halfword or 24-bit), W (for Word or 48-bit)}In this example, the accelerator instructions in 1306 may cause one or more PEs to perform one or more vector transfers (VXfer) and move data across registers of a single PE or adjacent PEs in the direction indicated for a row of PEs (when the direction is north or south) or a column (when the direction is east or west), as shown in FIGS. 8A-8F, for example. A vector source (VsrcI) may represent a primary (vector) input register, and a vector source Vsrc2or Rsrc2may provide the backup (or padding) input, which may be a vector or a scalar register.In some embodiments, the PE array (e.g., a PPE 170) may have a defined (e.g., limited) capacitance. For example, if data stored in registers of PEs in a PE array is shifted one row to north, all rows except the lowermost row may receive input from Vsc1 of a PE that is south of the PE that receives the data. In some embodiments, the PEs of the lower row may receive data from a register associated with a Vsrc2 register of an PE of an upper row, or may communicate data from Rsrc2. Similar transmissions may be made in the other directions, south, east and west, as described with respect to Figures 8A-8F. In these examples, the transmissions may be based at least on the capacity of the respective registers. For example, the transmissions may include the transmission of portions (e.g., half) of the bits stored in a particular register or all bits in a particular register according to the accelerator instructions.In some embodiments, the functionality for these instructions may also be mapped at the application layer to C-language features, for example: v2d_intx vxfer_west(v2d_intx,v2d_intx)=v2ds48 VXferWestW(v2ds48 v2ds48); v2d_intx in_00_w1= vxfer_west(in_00 in_01); v2d_intx in_01_w1= vxfer_west(in_01, 0). As shown, the instructions obtained at 1320 may specify one or more specific transmissions (e.g., a West vector transmission or "vxfer west") and one or more data types of data to be moved between registers according to the transmission. In this manner, the instructions may be configured such that the mapping system 1302 agenerates instructions corresponding to particular transmissions (e.g., transmissions of particular data between particular registers of one or more PEs). The imaging system 1302 amay then generate accelerator instructions according to transmissions indicated by the instruction received at 1324.In some embodiments, processor 1302 may provide to imaging system 1302 ain order to generate accelerator instructions 1326 that include use of a DLSU (e.g., a DLSU identical or similar to DLSUs 124 of FIG. 1A ). For example, processor 1302 may provide to imaging system 1302 aexecute 1324 that causes imaging system 1302 ato generate the accelerator instruction. In some embodiments, the accelerator instruction may include one or more instructions to transmit data from a DLSU to a multi-dimensional SIMD processor, such as the PPE 1306, as well as one or more instructions to perform one or more operations by the PPE 1306. An example set of instructions may include: class lstrm property (DLSU_AGEN_REG_SIZE bit non-consecutive); class sstrm property (DLSU_AGEN_REG_SIZE bit non-consecutive); these streams may then be coupled to stream start and load / store operations: void vload_start (agen, lsrm&)=void vload_start (aword, lsrm_t&); void vstore_start(agen, sstrm&)=void vstore_start(aword, sstrm_t&); v2d int vload_w(lstrm& a) void vstore(v2d_int s, sstrm&a)The boot and load / store operations can then be implemented, for example, in a C application: vload_start(agen1, lsrm1); (transfer age and activate load stream) vstore_start(agen2, sstrm1); (transfer age and activate store stream) for (i=0; i<(NUMBLKS_W+1)*NUMBLKS_H; i++) next_in_blk0_iorf = vload_w(lstrm1);← load from stream DSLU next_in_blkl_iorf = vload_w(lstrm1); filt_h_compute(...); vstore_i((v2d_int)filt_out_blk0, sstrml); ← store to strem DLSU vstore_1i((v2d_int)filt_out_blk1, sstrm1); }In some embodiments, processor 1302 may cause mapping system 1302 ato generate accelerator instructions at 1326, which schedule the start and load and store operations between a DLSU and PPE 1306 by providing data associated with the accelerator instructions at 1328, to a respective function block 1304. In this example, function block 1304 may then receive the accelerator instructions and, at 1330, cause a DMA system (e.g., a DMA system identical or similar to DMA systems 114 of FIG. 1A ) and PPE 1306 to perform one or more operations according to the instructions such that device stalls or wait times due to DMA transfer latencies are minimized or eliminated. For example, with respect to the PPE 1306, the function block 1304 may provide the accelerator instruction 1332 to the PPE 1306 to cause the PPE 1306 to perform one or more operations at 1334. In this example, the one or more operations may be associated with south transfer between individual PEs of the PPE 1306.FIG. 14 is a flowchart of an example method for generating accelerator instructions, in accordance with some embodiments of the present disclosure. In some embodiments, aspects of method 1400 may be performed by one or more devices identical or similar to one or more of the devices of FIGS. 1A-1C, such as DMA systems 114, VPUs 116, PPEs 118, and / or processor 102. In embodiments, one or more other devices of FIG. 1A may perform one or more aspects of methods 1400.The method 1400 includes, at block 1402, receiving an instruction to be performed by an accelerator. For example, a processor may receive the instruction to be performed by the accelerator. In some embodiments, the instruction may be presented in a first programming language. For example, the instruction may be presented in a high level programming language, such as one or more object oriented programming languages.In some embodiments, the instruction may represent one or more operations to be performed by an accelerator, such as a PPE. For example, the instruction may represent one or more operations corresponding to one or more operations (e.g., SIMD operations) to be performed by PEs of a PPE. In this example, the PEs may be logically arranged in a 2D array and configured to communicate with one or more other PEs within the PPE. In some embodiments, communication between the PEs of the PPE may be according to one or more connection sets as described herein.The method 1400 includes, at block 1404, determining one or more operations to be performed by the accelerator. For example, the processor may determine the one or more operations to be performed by the accelerator based at least on the instruction. In examples, the processor may determine the one or more operations to be performed by the accelerator based at least on the instruction and a data path. In these examples, the data path may be associated with the accelerator (e.g., indicate compatible transmissions and operations that may be performed by it) and represent one or more data transmissions within the accelerator via one or more components of the accelerator. For example, the data path may represent one or more transfers between registers of one or more PEs within the accelerator during performance of operations that cause data to be transferred between registers of a single PE, between registers of multiple PEs of a PPE, and / or combinations thereof.In some embodiments, the processor may determine the one or more operations based on the accelerator provided for executing the operations. For example, the processor may determine the one or more operations based on compatible operations for which the accelerator may be configured to execute. In an example where the accelerator is a PPE that processes data associated with images, the processor may determine the one or more operations based at least on operations associated with processing the images. In some embodiments, the processor may determine a correspondence between the instruction to be performed by the accelerator and a set of accelerator instructions. For example, if the accelerator is a PPE, the processor may determine a correspondence between the instruction to be performed by the PPE (e.g., shown in a high-level programming language such as C / C++) and the operations to be performed by the PPE according to the instruction. In this example, the processor may determine one or more instructions to be performed by the accelerator, the instructions being represented in an assembly language. As described herein, the set of operations to be performed by the accelerator may be referred to as accelerator instructions.The method 1400 includes, at block 1406, generating a set of accelerator instructions. For example, the processor may generate the set of accelerator instructions. In some embodiments, the set of accelerator instructions may be based at least on operations to be performed by the accelerator. The accelerator instructions may correspond to instructions received from the processor.In some embodiments, the instruction may correspond to operations performed by an accelerator that include moving data between registers. For example, the processor may generate the set of accelerator instructions, the accelerator instructions causing data moves between a first register of a first component (e.g., a PE of a PPE) of the accelerator and a second register of the component of the accelerator. In another example, the processor may generate the set of accelerator instructions, the accelerator instructions causing data shifts between a first register of a first component of the accelerator and a first register of another component of the accelerator. In this example, where the accelerator is a PPE, the first register of the first component may correspond to a first PE and the first register of the other component may correspond to a first register of a second PE, where the first PE and the second PE are configured to be interconnected via a link set that enables data transfer therebetween.In some embodiments, the processor may generate the set of accelerator instructions, wherein the accelerator instructions correspond to data displacements between registers and one or more arithmetic operations. For example, the processor may generate the set of accelerator instructions, the accelerator instructions corresponding to data displacements between registers within a single PE or across multiple PEs interconnected according to one or more connection sets. In this example, the accelerator instructions may also correspond to one or more of addition operations, subtraction operations, multiplication operations, or division operations (commonly referred to as arithmetic operations). In some embodiments, the accelerator instructions may include a row of sequential data displacements between registers and arithmetic operations. For example, according to an instruction to perform a 3x3 filter operation, the accelerator instructions may include a set of shift and multiply operations such that multiple sets of data (representing the values of adjacent pixels with respect to a particular PE of a PPE) are obtained and stored in registers of a single component of an accelerator, and one or more multiply operations are performed based at least on the values stored in the registers. While the present example is discussed with respect to an instruction corresponding to accelerator instructions performed to implement a 3x3 filter, it should be understood that the present disclosure is not limited to such instructions and that any other suitable instructions that may be mapped to one or more data shifts and arithmetic operations are contemplated.The method 1400 includes, at block 1408, providing data associated with the set of accelerator instructions to a system to cause the system to coordinate operation of the accelerator according to the set of accelerator instructions. For example, the processor may provide the data associated with the set of accelerator instructions to a function block. In this example, the functional block may cause one or more components of the functional block (e.g., a PPE, a VPU, a DLSU, a DMA system, and / or the like) to perform corresponding instructions included in the set of accelerator instructions. In some embodiments, the one or more components may execute the instructions individually (e.g., without waiting for one or more instructions to be performed by one or more other components of the functional block). In other embodiments, the one or more components may execute the instructions in coordination with one or more other components of the functional block.In an example, the accelerator instructions may cause a first component (e.g., a PPE) and a second component (e.g., a DLSU) to operate in coordination with each other. For example, the instructions may cause the DLSU to obtain (e.g., buffer) data associated with an image (or portions thereof, sometimes referred to as blocks or tiles). The instructions may then cause the PPE to obtain the data buffered by the DLSU. In an example where the data associated with the image is obtained from the PPE, data associated with at least a portion of an image may be read into the PPE via a read stream. In some embodiments, the PPE may perform one or more operations according to the accelerator instruction, for example, one or more data shifts between registers and one or more arithmetic operations. Once one or more operations are complete, the PPE may pass the resulting data to the DLSU via a write stream. While the principles of the present disclosure are described in relation to the operations performed by the PPE, it should be appreciated that any suitable accelerator instructions may cause each component of a functional block to operate according to any suitable instruction.Although the present disclosure may be described with respect to an example autonomous vehicle 1500 (alternatively referred to herein as "vehicle 1500" or "ego vehicle 1500", an example of which is described with reference to FIGS. 15A-15D ), this is not intended to be limiting. The systems and methods described herein may be used without limitation on non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more Adaptive Driver Assistance Systems (ADAS)), autonomous vehicles or machines, steered and unrouted robots or robot platforms, storage vehicles, off-road vehicles, vehicles coupled to one or more trailers, wing boats, shuttles, emergency vehicles, motor bicycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other types of vehicles. Moreover, although the present disclosure may be described with respect to certain implementations including processing data during operation of the automated vehicle, this is not to be understood as limiting, and the systems and methods described herein may be used in augmented reality, virtual reality, mixed reality, robotics, safety and monitoring, autonomous or semi-autonomous machine applications, and / or any other fields of technology in which accelerators may be used to process data used during operation of a robot.EXAMPLE AUTONOMOUS VEHICLEFIG. 15A illustrates an example autonomous vehicle 1500, in accordance with some embodiments of the present disclosure. In some embodiments, the example autonomous vehicle 1500 may include one or more components (e.g., SoCs and / or the like) that are the same as or similar to the function blocks 110 of FIG. 1A and / or other components as described herein. The autonomous vehicle 1500 (alternatively referred to herein as "vehicle 1500") may include, without limitation, a passenger vehicle, such as a car, truck, bus, emergency service vehicle, shuttle, electric or motorized bicycle, motorcycle, fire truck, police vehicle, lift truck, boat, construction vehicle, underwater vehicle, robotic vehicle, drone, plane, vehicle coupled to a trailer (e.g., semi-trailer used for transporting cargo), and / or another type of vehicle (e.g., unmanned and / or receiving one or more passengers). Autonomous vehicles are generally described in terms of automation levels promulgated by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Ministry of Traffic, and the Society of Automotive Engineers (SAE) Society of Automotive Engineers (SAE) Standard No. J3016-201806 published June 15, 2018, Standard No. J3016-20169 published September 30, 2016, as well as earlier and future versions of this standard). The vehicle 1500 may include functionality according to one or more of the autonomous driving level 3 through level 5. The vehicle 1500 may include functionality according to one or more of the autonomous driving level 1 through level 5. For example, depending on the embodiment, the vehicle 1500 may be capable of providing driver assistance (level 1), partial automation (level 2), conditional automation (level 3), high automation (level 4) and / or full automation (level 5). The term "autonomous" as used herein may include any and / or all types of endoscopy for the vehicle 1500 or other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, supporting endoscopy, semi-autonomous, primarily autonomous, or other designation. In some embodiments, during operation of the autonomous vehicle 1500, the autonomous vehicle 1500 may implement at least some of the systems, methods, and techniques described herein. For example, the autonomous vehicle 1500 may implement at least some of the components described with respect to the example computing environment of FIG. 1A, PPE of FIG. 1B, and / or PEs of FIG. 1C in obtaining and processing data generated by sensors of the autonomous vehicle as described herein.The vehicle 1500 may include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. The vehicle 1500 may include a propulsion system 1550, such as an internal combustion engine, a hybrid electric power plant, a pure electric motor, and / or another type of propulsion. The propulsion system 1550 may be connected to a powertrain of the vehicle 1500, which may include a transmission to enable propulsion of the vehicle 1500. The propulsion system 1550 may be controlled in response to receiving signals from the accelerator 1552.A steering system 1554 that may include a steering wheel may be used to steer the vehicle 1500 (e.g., along a desired path or route) when the propulsion system 1550 is operating (e.g., when the vehicle is in motion). The steering system 1554 may receive signals from a steering actuator 1556. The steering wheel can optionally be for full automation (stage 5).The brake sensor system 1546 may be used to actuate the vehicle brakes in response to receiving signals from the brake actuators 1548 and / or the brake sensors.The one or more controllers 1536 which may include one or more system-on-chips (SoCs) 1504 (FIG. 15C ) and / or GPUs may provide signals (e.g., representative of instructions) to one or more components and / or systems of the vehicle 1500. For example, the one or more controllers may send signals to actuate vehicle brakes via one or more brake actuators 1548, actuate the steering system 1554 via one or more steering actuators 1556, actuate the propulsion system 1550 via one or more throttle / accelerator devices 1552. The one or more controllers 1536 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and issue operating commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 1500. The one or more controllers 1536 may include a first autonomous driving function controller 1536, a second safety function controller 1536, a third artificial intelligence (e.g., computer vision) function controller 1536, a fourth infotainment function controller 1536, a fifth emergency redundancy controller 1536, and / or other controllers. In some examples, a single controller 1536 may take over two or more of the above-mentioned functionalities, two or more controllers 1536 may take over a single functionality, and / or any combination thereof.The one or more controllers 1536 may provide the signals to control one or more components and / or systems of the vehicle 1500 in response to sensor data received from one or more sensors (e.g., sensor inputs). The sensor data may be received, for example and without limitation, from one or more of the following: one or more global navigation satellite systems ("GNSS") sensors 1558 (e.g., one or more global positioning system sensors), one or more RADAR sensors 1560, one or more ultrasonic sensors 1562, one or more LiDAR sensors 1564, one or more inertial measurement unit (IMU) sensors 1566 (e.g., accelerometer(s), gyroscope(s), magnetic compass(s), magnetometer(s), etc.), one or more microphones 1596, one or more stereo cameras 1568, one or more wide-angle cameras 1570 (e.g., fish-eye cameras), one or more infrared cameras 1572, one or more environmental cameras 1574 (e.g., 360-degree cameras), one or more long-range and / or mid-range cameras 1598, one or more speed sensors 1544 (e.g., for measuring the speed of the vehicle 1500), one or more vibration sensors 1542, one or more steering sensors 1540, one or more brake sensors (e.g., as part of the brake sensor system 1546), and / or other types of sensors.One or more of the controllers 1536 may receive input (e.g., in the form of input data) from an instrument cluster 1532 of the vehicle 1500 and provide output (e.g., in the form of output data, display data, etc.) via a human-machine interface (HMI) display 1534, an acoustic detector, a speaker, and / or via other components of the vehicle 1500. The outputs may include information such as vehicle speed, speed, time, map data (e.g., High Definition ("HD") map 1522 of FIG. 15C ), location data (e.g., the location of vehicle 1500, e.g., on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and status of objects as perceived by the one or more controllers 1536, etc. For example, information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.), and / or information about driving maneuvers may be displayed on the HMI display 1534, which the vehicle has performed, is performing, or is being performed (e.g., now changing lanes, take exit 34B in two miles, etc.).The vehicle 1500 further includes a network interface 1524 that may use one or more wireless antennas 1526 and / or modems to communicate over one or more networks. The network interface 1524 may be suitable for long-term evolution ("LTE"), wideband code division multiple access ("WCDMA"), universal mobile telecommunications system ("UMTS"), global system for mobile communication ("GSM"), imt-cdma multi-carrier ("CDMA2000"), etc., communication, for example. The one or more wireless antennas 1526 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc., and / or low power wide area networks ("LPWANs") such as LoRaWAN, SigF, etc.FIG. 15B is an example of camera locations and fields of view for the example autonomous vehicle 1500 of FIG. 15A, in accordance with some embodiments of the present disclosure. The cameras and respective fields of view are an example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different locations of the vehicle 1500.The camera types for the cameras may include, but are not limited to, digital cameras that may be configured for use with the components and / or systems of the vehicle 1500. The one or more cameras may operate with the Automotive Safety Integrity Level (ASIL) B and / or with another ASIL. The camera types may be capable of any image capture rate, such as 60 images per second (fps), 120 fps, 240 fps, etc. The cameras may use shutters, global shutters, other type of shutter, or a combination thereof, depending on the embodiment. In some examples, the color filter array may include a red clear clear clear color filter array (RCCC), a red clear blue color filter array (RCCB), a red blue green clear color filter array (RBGC), a foveon X3 color filter array, a bayer sensor color filter array (RGGB), a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras with an RCCC, an RCCB, and / or an RBGC color filter array, may be used to increase light sensitivity.In some examples, one or more of the cameras may be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocamera may be installed to provide functions including lane departure warning, road sign assist, and smart headlight control. One or more of the cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).One or more of the cameras may be mounted in a mount, e.g., in a specially designed (three-dimensionally ("3D") printed) mount, to eliminate stray light and reflections from the vehicle interior (e.g., mirroring of the dashboard that mirrors in the windshield) that could interfere with the image data acquisition of the camera. Referring to the mounting of exterior mirrors, the exterior mirrors may be printed individually in 3D such that the camera mounting plate conforms to the shape of the exterior mirror. In some examples, the one or more cameras may be integrated into the exterior mirror. In side cameras, the one or more cameras may also be integrated into the four columns at each corner of the cabin.Cameras having a field of view that includes portions of the environment in front of the vehicle 1500 (e.g., forward facing cameras) may be used for the environment view to help identify forward facing paths and obstacles, and to provide information relevant to the creation of an occupancy grid and / or determination of preferred vehicle paths using one or more controllers 1536 and / or control SoCs. Forward facing cameras may be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian sensing, and crash avoidance. Forward facing cameras may also be used for ADAS functions and systems that include lane departure warnings ("LDW"), autonomous cruise control ("ACC"), and / or other functions such as road sign detection.A plurality of cameras may be used in a forward-facing configuration, including, for example, a monocular camera platform that includes a complementary metal oxide semiconductor ("CMOS") color imager. Another example is the wide-angle cameras 1570 that can be used to capture objects that come into view from the periphery (e.g., pedestrians, crossing vehicles, or bicycles). Although only one wide-angle camera is illustrated in FIG. 15B, any number (including zero) of wide-angle cameras 1570 may be present on the vehicle 1500. In addition, any number of long-range cameras 1598 (e.g., a long-range stereo camera pair) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The one or more long-range cameras 1598 may also be used for object detection and classification, as well as for basic object tracking.Any number of stereo cameras 1568 may also be included in a forward-facing configuration. In at least one embodiment, one or more of stereo cameras 1568 may include an integrated control unit that includes a scalable processing unit that may provide programmable logic ("FPGA") and a multi-core microprocessor with a controller area network ("CAN") or Ethernet interface on a single chip. Such a unit may be used to generate a 3D map of the environment of the vehicle that includes a range estimate for all points in the image. One or more alternative stereo cameras 1568 may include a compact stereo vision sensor that may include two camera lenses (one left and right each) and an image processing chip that may measure the distance between the vehicle and the target object and use the generated information (e.g., metadata) to enable the autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1568 may be used in addition to or alternatively from those described herein.Cameras having a field of view that includes portions of the environment lateral to the vehicle 1500 (e.g., side cameras) may be used for the environment view and provide information used to create and update the occupancy grid as well as to generate crash warnings in the event of a side impact. For example, the one or more environmental cameras 1574 (e.g., four environmental cameras 1574 as illustrated in FIG. 15B ) may be positioned on the vehicle 1500. The one or more surround cameras 1574 may include one or more wide-angle cameras 1570, one or more fish-eye cameras, one or more 360-degree cameras, and / or the like. For example, four fish eye cameras may be mounted on the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 1574 (e.g., left, right, and rear) and use one or more other cameras (e.g., a front facing camera) as the fourth surround camera.Cameras having a field of view that includes portions of the environment behind the vehicle 1500 (e.g., backup cameras) may be used for parking assist, environmental view, rear impact alerts, and occupancy grid creation and update. A variety of cameras may be used, including, but not limited to, cameras also suitable as one or more forward facing cameras (e.g., one or more long-range and / or mid-range cameras 1598, one or more stereo cameras 1568, one or more infrared cameras 1572, etc.), as described herein.FIG. 15C is a block diagram of an example system architecture for the example autonomous vehicle 1500 of FIG. 15A, in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, arrangements, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. Various functions described herein that are performed by entities may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory.Each of the vehicle 1500 components, features, and systems in FIG. 15C is illustrated as being connected via bus 1502. Bus 1502 may include a controller area network (CAN) data interface (alternatively referred to herein as a "CAN bus"). A CAN may be a network within the vehicle 1500 that serves to assist in controlling various features and functions of the vehicle 1500, such as actuation of brakes, acceleration, brakes, steering, windshield wipers, etc. A CAN bus may be configured to have dozens or even hundreds of nodes, each having its own unique identifier (e.g., a CAN ID). The CAN bus may be read to determine the steering wheel angle, the vehicle speed, the engine speed (U / min), the button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.Although bus 1502 is described herein as a CAN bus, this is not to be understood as limiting. For example, FlexRay and / or Ethernet may be used in addition or alternatively to the CAN bus. In addition, while a single line is used to represent bus 1502, this is not intended to be limiting. For example, there may be any number of buses 1502, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using a different protocol. In some examples, two or more buses 1502 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 1502 may be used for the collision avoidance functionality and a second bus 1502 may be used for the actuation control. In each example, each bus 1502 may communicate with one of the components of the vehicle 1500, and two or more buses 1502 may communicate with the same components. In some examples, each SoC 1504, controller 1536, and / or computer within the vehicle may have access to the same input data (e.g., inputs from sensors of the vehicle 1500) and may be connected to a common bus, such as the CAN bus.The vehicle 1500 may include one or more controllers 1536 as described herein with reference to FIG. 15A. The one or more controllers 1536 may be used for a variety of functions. The one or more controllers 1536 may be coupled to one of the various other components and systems of the vehicle 1500 and may be used to control the vehicle 1500, artificial intelligence of the vehicle 1500, infotainment for the vehicle 1500, and / or the like.The vehicle 1500 may include one or more systems on a chip (SoC) 1504. The SoC 1504 may include one or more CPUs 1506, one or more GPUs 1508, one or more processors 1510, one or more caches 1512, one or more accelerators 1514, one or more data stores 1516, and / or other components and features not illustrated. The one or more SoCs 1504 may be used to control the vehicle 1500 in a variety of platforms and systems. For example, the one or more SoCs 1504 may be combined in a system (e.g., the vehicle 1500 system) with a HD card 1522 that receives card refresh and / or updates from one or more servers (e.g., the one or more servers 1578 of FIG. 15D ) via a network interface 1524.The one or more CPUs 1506 may include a CPU cluster or complex (alternatively referred to herein as a "CCPLEX"). The one or more CPUs 1506 may include multiple cores and / or L2 caches. In some embodiments, for example, the one or more CPUs 1506 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the one or more CPUs 1506 may include four dual core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). The one or more CPUs 1506 (e.g., the CCPLEX) may be configured to support simultaneous operation of clusters such that any combination of the clusters of the one or more CPUs 1506 may be active at a particular time.The one or more CPUs 1506 may implement power management functions that include one or more of the following features: individual hardware blocks may be automatically clock controlled at idle to conserve dynamic power; each core clock may be controlled when the core is not actively executing instructions due to execution of WFI / WFE instructions; each core may be independently power controlled; each core cluster may be independently clock controlled when all cores are clock controlled or power controlled; and / or each core cluster may be independently power controlled when all cores are power controlled. The one or more CPUs 1506 may also implement an enhanced power state management algorithm in which allowed power states and expected wake-up times are set and the hardware / microcode determines the best power state to enter for the core, cluster, and CCPLEX. The processing cores may support simplified sequences for inputting the power state to the software, offloading work to microcode.The one or more GPUs 1508 may include an integrated GPU (alternatively referred to herein as "1G"). The one or more GPUs 1508 may be programmable and may be efficient for parallel workloads. The one or more GPUs 1508 may use an extended tensor instruction set, in some examples. The one or more GPUs 1508 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96 KB storage capacity), and two or more of the streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512 KB storage capacity). In some embodiments, the one or more GPUs 1508 may include at least eight streaming microprocessors. The one or more GPUs 1508 may use one or more application programming interface(s), API(s)) for computations. Moreover, the one or more GPUs 1508 may use one or more parallel computing platforms and / or programming models (e.g., CUDA from NVIDIA).The one or more GPUs 1508 may be energy optimized for best performance in automotive and embedded use cases. The one or more GPUs 1508 may be fabricated on a fin field effect transistor (FinFET), for example. However, this is not to be understood as limiting, and the one or more GPUs 1508 may also be fabricated with other semiconductor fabrication processes. Each streaming microprocessor may include a series of mixed precision processing cores divided into multiple blocks. For example and without limitation, 64 PF32 cores and 32 PF64 cores may be divided into four processing blocks. In such an example, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed precision NVIDIA TENSOR COREs for deep leaming matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file. In addition, the streaming microprocessors may include independent parallel integer and floating point data paths to enable efficient execution of workloads with a mixture of computations and addressing computations. The streaming microprocessors may include an independent thread scheduling function to enable fine grain synchronization and cooperation between parallel threads. The streaming microprocessors may include a combined L1 data cache and a shared memory unit to improve performance while simplifying programming.The one or more GPUs 1508 may include a high bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide a peak memory bandwidth of about 900 GB / second, in some examples. In some examples, in addition to or as an alternative to HBM memory, synchronous graphics random access memory (SGRAM) may be used, e.g., double data rate synchronous graphics random access memory type five (GDDR5).The one or more GPUs 1508 may include unified memory technology including access counters to enable more accurate migration of memory pages to the processor that most frequently accesses them, thereby improving efficiency for processor shared memory areas. In some examples, Address Translation Services (ATS) support may be used to allow the one or more GPUs 1508 to directly access the page tables of the one or more CPUs 1506. In such examples, if the memory management unit (MMU) of the one or more GPUs 1508 fails, an address translation request may be sent to the one or more CPUs 1506. In response, the one or more CPUs 1506 may search in their page tables for the virtual-physical mapping for the address and send the translation back to the one or more GPUs 1508. Thus, the unified memory technology may enable a single unified virtual address space for memory of both the one or more CPUs 1506 and the one or more GPUs 1508, thereby facilitating programming of the one or more GPUs 1508 and porting applications to the one or more GPUs 1508.Additionally, the one or more GPUs 1508 may include an access counter that may track the frequency of access of the one or more GPUs 1508 to memory of other processors. The access counter may help to move memory pages into physical memory of the processor that most frequently accesses the pages.The one or more SoCs 1504 may include any number of caches 1512, including those described herein. For example, the one or more caches 1512 may include an L3 cache available to both the one or more CPUs 1506 and the one or more GPUs 1508 (e.g., connected to both the one or more CPUs 1506 and the one or more GPUs 1508). The one or more caches 1512 may include a write back cache that may track the states of the lines, e.g., by using a cache coherency protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may contain 4 MB or more, depending on embodiment, although smaller cache sizes may also be used.The one or more SoCs 1504 may include the one or more arithmetic logic units (ALUs) that may be used in performing processing with respect to any of the various tasks or operations of the vehicle 1500, such as processing DNNs. Additionally, the one or more SoCs 1504 may include one or more floating point units (FPUs) or other mathematical co-processors or numerical co-processors to perform mathematical operations within the system. For example, the one or more SoCs 1504 may include one or more FPUs integrated as execution units into one or more CPUs 1506 and / or one or more GPUs 1508.The one or more SoCs 1504 may include one or more accelerators 1514 (e.g., hardware accelerators, software accelerators, or a combination thereof). The one or more SoCs 1504 may include, for example, a hardware accelerator cluster that may include optimized hardware accelerators and / or large memory on-chip. The large memory on-chip (e.g., 4MB SRAM) may allow the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster may be used as a supplement to the one or more GPUs 1508 and offload some of the tasks of the one or more GPUs 1508 (e.g., to enable more cycles of the one or more GPUs 1508 to perform other tasks). The one or more accelerators 1514 may be used, for example, for targeted workloads (e.g., perception, convolutional neural networks (CNNs), etc.) that are stable enough to be suitable for acceleration. As used herein, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).The one or more accelerators 1514 (e.g., the hardware acceleration cluster) may include a deep learning accelerator (DLA). The one or more DLAs may include one or more tensor processing units (TPUs) configured to provide additional ten billion operations per second for deep learning applications and inferencing. The TPUs may be accelerators configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). The one or more DLAs may also be optimized for a particular set of neural network types and floating point operations, as well as inferencing. The design of the one or more DLAs can provide more performance per millimeter than a general purpose GPU and far outbreaks performance of a CPU. The one or more TPUs may perform multiple functions including a single instance convolution function that supports, for example, INT8, INT16, and FP16 data types for both features and weights, and postprocessor functions.The one or more DLAs may quickly and efficiently execute neural networks, particularly CNNs, on processed or unprocessed data for a variety of functions including, for example and without limitation: a CNN for identifying and detecting objects using camera sensor data; a CNN for distance estimation using camera sensor data; a CNN for detecting and identifying emergency vehicles using microphone data; a CNN for face detection and identifying vehicle owners using camera sensor data; and / or a CNN for security and / or protection relevant events.The one or more DLAs may perform each function of the one or more GPUs 1508, and by using an inference accelerator, for example, a developer may provide either the one or more DLAs or the one or more GPUs 1508 to each function. For example, the developer may focus processing of CNNs and floating point operations on the one or more DLAs and leave other functions of the one or more GPUs 1508 and / or other accelerators 1514.The one or more accelerators 1514 (e.g., the hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may also be referred to herein as a computer vision accelerator. The one or more PVAs may be developed and configured to speed computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The one or more PVAs may provide a balance between performance and flexibility. Each PVA may include, for example and without limitation, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.The RISC cores may interact with image sensors (e.g., the image sensors of any of the cameras described herein), image signal processors, and / or the like. Each of the RISC cores may include any amount of memory. The RISC cores may use any number of protocols, depending on the embodiment. In some examples, the RISC cores may execute a real-time operating system (RTOS). The RISC cores may be implemented with one or more integrated circuits, application specific integrated circuits (ASICs), and / or memory devices. The RISC cores may include, for example, an instruction cache and / or a close coupled RAM.The DMA may allow components of the PVA(s) to access the memory of the system independent of the one or more CPUs 1506. The DMA may support any number of features that serve to optimize the PVA, including, but not limited to, support multi-dimensional addressing and / or circular addressing. In some examples, DMA may support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block locking, vertical block locking, and / or depth locking.The vector processors may be programmable processors that may be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing functions. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripheral devices. The vector processing subsystem may operate as a primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or a memory (e.g., VMEM). A VPU core may include a digital signal processor, such as a single instruction, multiple data (SIMD) and very long instruction words (VLIW). The combination of SIMD and VLIW may increase throughput and speed.Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. Thus, in some examples, each of the vector processors may be configured to operate independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to use data parallelism. For example, in some embodiments, the multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may execute different computer vision algorithms on the same image simultaneously or even execute different algorithms on successive images or portions of an image. Among other things, any number of PVAs may be included in the hardware acceleration cluster and any number of vector processors may be included in each of the PVAs. Moreover, the one or more PVAs may include additional memory for an error correcting code (ECC) to increase the overall security of the system.The one or more accelerators 1514 (e.g., the hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide a high bandwidth, low latency SRAM to the one or more accelerators 1514. In some examples, memory on the chip may include at least 4 MB SRAM, which may be, for example and without limitation, eight field configurable memory blocks accessible by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuits, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that enables the PVA and the DLA to access the memory at high speed. The backbone may include an on-chip computer vision network that connects the PVA and the DLA to the memory (e.g., using the APB).The computer vision network on the chip may include an interface that determines that both the PVA and the DLA provide ready and valid signals prior to transmission of control signals / addresses / data. Such an interface may provide separate phases and separate channels for the transmission of control signals / addresses / data as well as bursty communication for continuous data transmission. This type of interface may comply with the ISO 26262 or IEC 61508 standards, although other standards and protocols may also be used.In some examples, the one or more SoCs 1504 may include a real-time ray tracing hardware accelerator as described in U.S. patent application Ser. No. 16 / 101,232 filed on Aug. 10, 2018. The real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the positions and extents of objects (e.g., within a world model), to generate real-time visualization simulations, for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulation, for comparison with LiDAR data for purposes of localization and / or for other functions and / or for other purposes. In some embodiments, one or more Tree Traversal Units (TTUs) may be used for performing one or more operations associated with ray tracing.The one or more accelerators 1514 (e.g., the hardware accelerator cluster) have a wide range of uses for autonomous driving. The PVA may be a programmable vision accelerator that may be used for important processing steps in ADAS and autonomous vehicles. The capabilities of the PVA fit well to algorithmic areas that require predictable processing with low power consumption and low latency. In other words, the PVA is well suited for semi-dense or dense regular computations, even for small data sets requiring predictable low latency, low power consumption runtimes. Thus, in the context of autonomous vehicle platforms, the PVAs are designed to execute classical computer vision algorithms because they are efficient in object detection and operate with integer mathematics.According to one embodiment of the technology, the PVA is used, for example, to perform computer stereo vision. In some examples, a semi-global matching based algorithm may be used, although this is not intended to be limiting. Many autonomous driving level 3-5 applications require spontaneous motion estimation (e.g., structure of motion, pedestrian detection, lane detection, etc.). The PVA may perform a computer stereo vision function on inputs from two monocular cameras.In some examples, the PVA may be used to perform a dense optical flow. According to processing raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used for time-of-flight depth processing, e.g., by processing raw time-of-flight data to provide processed time-of-flight data.The DLA may be used to operate any type of network to improve control and driving safety; for example, this includes a neural network that outputs a confidence measure for each object detection. Such a confidence value may be interpreted as a probability or as providing a relative "weight" of each detection as compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positive detections and not false positive detections. For example, the system may set a threshold for confidence and consider only the detections exceeding the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections would result in the vehicle automatically performing emergency braking, which is, of course, undesirable. Therefore, only the most safe detections should be considered as triggers for AEB. The DLA may employ a neural network to regression the confidence value. The neural network may use as input at least a subset of parameters, such as, but not limited to, the dimensions of the bounding box, the ground plane estimate obtained (e.g., from another subsystem), output of the inertial measurement unit (IMU) sensor 1566 that correlates to the orientation of the vehicle 1500, distance, 3D position estimates of the object obtained from the neural network and / or other sensors (e.g., one or more LiDAR sensors 1564 or one or more RADAR sensors 1560).The one or more SoCs 1504 may include the one or more data stores 1516 (e.g., memory). The one or more data stores 1516 may be on-chip memory on the one or more SoCs 1504, in which neural networks to execute on the GPU and / or the DLA may be stored. In some examples, the one or more data stores 1516 may be large enough to store multiple instances of neural networks for redundancy and security. The one or more data stores 1512 may include one or more L2 or L3 caches 1512. The reference to the one or more of the data stores 1516 may include a reference to memory associated with the PVA, the DLA, and / or one or more other accelerators 1514, as described herein.The one or more SoCs 1504 may include one or more processors 1510 (e.g., embedded processors). The one or more processors 1510 may include a boot and power management processor, which may be a dedicated processor and subsystem to handle the boot power and management functions and associated security enforcement. The boot and power management processor may be part of the boot sequence of the one or more SoCs 1504 and may provide runtime power management services. The boot and power management processor may provide clock and voltage programming, support for transitioning the system to a low power state, management of the thermals and temperature sensors of the one or more SoCs 1504, and / or management of the one or more SoCs 1504. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the one or more SoCs 1504 may use the ring oscillators to sense the temperatures of the one or more CPUs 1506, the one or more GPUs 1508, and / or the one or more accelerators 1514. When it is determined that the temperatures exceed a threshold, the boot and power management processor may enter a temperature fault routine and place the one or more SoCs 1504 in a lower power state and / or place the vehicle 1500 in a chauffeur-to-safe stop mode (e.g., place the vehicle 1500 in a safe stop).The one or more processors 1510 may also include a number of embedded processors that may serve as an audio processing engine. The audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio over multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a dedicated RAM digital signal processor.The one or more processors 1510 may also include an Ways-On processor engine that may provide the necessary hardware functions to support sensor management with low power consumption and waking up use cases. The Ways-On-Processor engine may include a processor core, close-coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.The one or more processors 1510 may also include a safety cluster machine that includes a dedicated processor subsystem for safety management of automotive applications. The safety cluster machine may include two or more processor cores, close coupled RAM, supporting peripherals (e.g., timers, an interrupt controller, etc.), and / or routing logic. In a safety mode, the two or more cores may operate in a lockstep mode and function as a single core with comparison logic that detects all differences between their operations.The one or more processors 1510 may also include a real-time camera engine, which may include a dedicated processor subsystem for managing the real-time camera.The one or more processors 1510 may further include a high dynamic range signal processor, which may include an image signal processor that is a hardware machine that is part of the camera processing pipeline.The one or more processors 1510 may include a video image combiner, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions needed by a video playback application to generate the final image for the player window. The video image combiner may make a lens distortion correction at the one or more wide angle cameras 1570, the one or more surround cameras 1574, and / or at the monitoring camera sensors in the cabin. The sensor of the surveillance camera in the cabin is preferably monitored by a neural network running on another instance of the enhanced SoC and configured to detect events in the cabin and react accordingly. A system in the cabin may perform lipread to activate the cellular service and place a call, dictate emails, change destination, activate or change infotainment system and vehicle settings, or enable voice-controlled browsing on the Internet. Certain functions are available to the driver only when the vehicle is operating in an autonomous mode and are otherwise disabled.The video image combiner may include improved temporal noise suppression for both spatial and temporal noise suppression. For example, if motion occurs in a video, the noise suppression weights the spatial information accordingly and reduces the weight of the information provided by adjacent images. If an image or portion of an image does not contain motion, the temporal noise suppression performed by the video image combiner may use information from the previous image to reduce the noise in the current image.The video image combiner may also be configured to perform stereo equalization of the input stereo objective images. The video image combiner may also be used for user interface design when the operating system desktop is in use and the one or more GPUs 1508 do not need to continually render new surfaces. Even when the one or more GPUs 1508 are turned on and actively drive 3D rendering, the video image combiner may be used to offload the one or more GPUs 1508 to improve performance and responsiveness.The one or more SoCs 1504 may also include a Mobile Industry Processor Interface (MIPI) serial camera interface for receiving video and camera inputs, a high speed interface, and / or a video input block that may be used for camera and related pixel input functions. The one or more SoCs 1504 may also include one or more input / output controllers, which may be controlled by software and used to receive I / O signals that are not associated with any particular role.The one or more SoCs 1504 may also include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management and / or other devices. The one or more SoCs 1504 may be used to process data from cameras (e.g., via gigabit multimedia serial link and Ethernet), sensors (e.g., one or more LiDAR sensors 1564, one or more RADAR sensors 1560, etc., which may be connected via Ethernet), data from bus 1502 (e.g., speed of vehicle 1500, steering wheel position, etc.), data from one or more GNSS sensors 1558 (e.g., connected via Ethernet or CAN bus). The one or more SoCs 1504 may further include dedicated high performance mass storage controllers, which may include their own DMA engines and which may be used to offload the one or more CPUs 1506 from routine data management tasks.The one or more SoCs 1504 may be an end-to-end platform with a flexible architecture that extends across automation levels 3- 5 thereby providing a comprehensive functional safety architecture that supports and efficiently uses computer vision and ADAS techniques for diversity and redundancy and provides a platform for a flexible, reliable driving software stack along with deep learning tools. The one or more SoCs 1504 may be faster, more reliable, and even more energy and space saving than conventional systems. For example, the one or more accelerators 1514, in combination with the one or more CPUs 1506, the one or more GPUs 1508, and the one or more data stores 1516, may form a fast, efficient Level 3-5 autonomous vehicle platform.The technology thus provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs that can be configured using a high-level programming language, such as the C programming language, to execute a variety of processing algorithms for a variety of visual data. However, CPUs are often unable to meet the performance requirements of many computer vision applications, such as execution time and power consumption requirements. In particular, many CPUs are unable to execute complex object detection algorithms in real-time, which is a prerequisite for in-vehicle ADAS applications and a prerequisite for practical level 3-5 autonomous vehicles.In contrast to conventional systems, the technology described herein, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, enables multiple neural networks to be executed simultaneously and / or sequentially and the results combined to enable the level 3-5 autonomous driving functionality. A CNN executing on the DLA or the dG (e.g., the one or more GPUs 1520) may include, for example, text and word detection that enables the supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA may further include a neural network capable of identifying, interpreting, and providing a semantic understanding and passing this semantic understanding to the path planning modules running on the CPU complex.Another example is that multiple neural networks may run simultaneously, as required for level 3, 4 or 5 driving. For example, a warning sign labeled "Caution: flashing lights indicate Smoothis" may be interpreted independently or jointly together with an electrical light from multiple neural networks. The sign itself may be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), the text "flashing lights indicate smoothis" may be interpreted by a second deployed neural network informing the path planning software of the vehicle (preferably executed on the CPU complex) that when flashing lights are detected smoothis is present. The turn signal light may be identified across multiple images by a third neural network informing the path planning software of the vehicle of the presence (or absence) of turn signals. All three neural networks may run simultaneously, e.g., within the DLA and / or on the one or more GPUs 1508.In some examples, a CNN for face detection and vehicle owner identification may use data from camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 1500. The Ways-On-Sensor processing engine may be used to unlock the vehicle when the owner approaches the driver door and turn on the lights, and to disable the vehicle in the security mode when the owner leaves the vehicle. In this way, the one or more SoCs 1504 provide security against theft and / or carjacking.In another example, a emergency vehicle detection and identification CNN may use data from microphones 1596 to detect and identify emergency vehicle sirens. Unlike conventional systems that use general classifiers to detect sirens and manually extract features, the one or more SoCs 1504 use the CNN to classify environmental and urban sounds as well as to classify visual data. In a preferred embodiment, the CNN running on the DLA is trained to detect the relative approach speed of the emergency vehicle (e.g., using the Doppler effect). The CNN may also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by one or more GNSS sensors 1558. For example, the CNN will attempt to detect European sirens when operating in Europe, and when operating in the United States, the CNN will attempt to identify only North American sirens. Once an emergency vehicle is detected, a controller may be used to execute an emergency vehicle safety routine, slow the vehicle, drive to the roadside, park the vehicle, and / or idle the vehicle using ultrasonic sensors 1562 until the one or more emergency vehicles pass.The vehicle may include one or more CPUs 1518 (e.g., one or more discrete CPUs or one or more dCPUs), which may be coupled to the one or more SoCs 1504 via a high speed link (e.g., PCIe). The CPUs 1518 may include, for example, an X86 processor. For example, the CPUs 1518 may be used to perform a variety of functions including arbitration for potentially inconsistent results between ADAS sensors and the one or more SoCs 1504, and / or monitoring the status and state of the one or more controllers 1536 and / or the infotainment SoC 1530.The vehicle 1500 may include one or more GPUs 1520 (e.g., one or more discrete GPUs or one or more dGs), which may be coupled to the one or more SoCs 1504 via a high-speed link (e.g., NVLINK from NVIDIA). The one or more GPUs 1520 may provide additional artificial intelligence functions, e.g., through execution of redundant and / or different neural networks, and may be used to train and / or update neural networks based on inputs (e.g., sensor data) from vehicle 1500 sensors.The vehicle 1500 may further include the network interface 1524, which may include one or more wireless antennas 1526 (e.g., one or more wireless antennas for various communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 1524 may be used to enable wireless connection over the Internet to the cloud (e.g., to the one or more servers 1578 and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). To communicate with other vehicles, a direct connection between the two vehicles and / or an indirect connection may be established (e.g., via networks and the Internet). Direct connections may be established via vehicle-to-vehicle communication. The vehicle-to-vehicle communication may provide information to the vehicle 1500 about vehicles in the vicinity of the vehicle 1500 (e.g., vehicles in front of, beside, and / or behind the vehicle 1500). This functionality may be part of a cooperative adaptive cruise control function of the vehicle 1500.The network interface 1524 may include an SoC that provides modulation and demodulation functions and allows the one or more controllers 1536 to communicate over wireless networks. The network interface 1524 may include a radio frequency front end for up-converting from baseband to radio frequency and down-converting from radio frequency to baseband. The frequency conversions can be carried out using known methods and / or using super-heterodyne methods. In some examples, the radio frequency front end functionality may be provided by a separate chip. The network interface may include wireless functionality for communicating over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWA, and / or other wireless protocols.The vehicle 1500 may further include one or more data stores 1528, which may be off-chip (e.g., off-chip from the SoCs 1504). The one or more data stores 1528 may include one or more memory elements including RAM, SRAM, DRAM, VRAM, flash, hard drives, and / or other components and / or devices capable of storing at least one bit of data.The vehicle 1500 may further include one or more GNSS sensors 1558. The one or more GNSS sensors 1558 (e.g., GPS, assisted GPS sensors, differential GPS (DGPS) sensors, etc.) assist in mapping, perception, occupancy grid creation, and / or path planning. Any number of GNSS sensors 1558 may be used, including, for example and without limitation, a GPS using a USB port with an Ethernet-to-Serial (RS-232) bridge.The vehicle 1500 may further include one or more RADAR sensors 1560. The one or more RADAR sensors 1560 may be used by the vehicle 1500 to detect long-range vehicles, even in darkness and / or bad weather conditions. The functional safety plane of the RADAR may be ASIL B. The one or more RADAR sensors 1560 may utilize the CAN and / or bus 1502 (e.g., to transmit the data generated by the one or more RADAR sensors 1560) to control and access object tracking data, in some examples, the raw data being accessed via Ethernet. A variety of RADAR sensor types may be used. The one or more RADAR sensors 1560 may be suitable for, for example, front, rear, and side RADARs without limitation. In some examples, one or more pulse Doppler RADAR sensors are used.The one or more RADAR sensors 1560 may include various configurations, such as a long range narrow field of view, short range wide field of view, lateral coverage short range, etc. In some examples, long range RADAR may be used for the adaptive cruise control function. The long range RADAR systems may provide a wide field of view realized by two or more independent scans, such as 250 m range. The one or more RADAR sensors 1560 may aid in distinguishing between static and moving objects and may be used by the ADAS systems for emergency braking assist and frontal crash warning. Long range RADAR sensors may include a monostatic multimode RADAR with multiple (e.g., six or more) fixed RADAR antennas and a high speed CAN and FlexRay interface. In a six antenna example, the middle four antennas may generate a focused beam pattern that is to detect the surroundings of the vehicle 1500 at higher speeds with minimal interference from traffic in the adjacent lanes. The other two antennas may extend the field of view so that vehicles entering or leaving the lane of vehicle 1500 may be quickly detected.For example, mid range RADAR systems may include a range of up to 1560m (forward) or 80m (rearward) and a field of view of up to 42 degrees (forward) or 1550 degrees (rearward). Short range RADAR systems may include, among other things, RADAR sensors configured for installation at both ends of the rear bumper. When such a RADAR sensor system is installed at both ends of the rear bumper, it can generate two beams that constantly monitor blind spot in the rear area and adjacent to the vehicle.Short range RADAR systems may be used in an ADAS blind spot detection system and / or as a lane change assist.The vehicle 1500 may also include one or more ultrasonic sensors 1562. The one or more ultrasonic sensors 1562 that may be mounted to the front, rear, and / or sides of the vehicle 1500 may be used for park assist and / or for creating and updating an occupancy grid. A plurality of ultrasonic sensors 1562 may be used, and different ultrasonic sensors 1562 may be used for different sensing ranges (e.g., 2.5 m, 4 m). The one or more ultrasonic sensors 1562 may operate with functional safety levels of ASIL B.The vehicle 1500 may include one or more LiDAR sensors 1564. The one or more LiDAR sensors 1564 may be used for object and pedestrian detection, emergency braking, crash avoidance, and / or other functions. The one or more LiDAR sensors 1564 may correspond to the functional safety plane ASIL B. In some examples, the vehicle 1500 may include multiple LIDAR sensors 1564 (e.g., two, four, six, etc.) that may use Ethernet (e.g., to provide data to a gigabit Ethernet switch).In some examples, the one or more LiDAR sensors 1564 may be capable of providing a list of objects and their distances for a 360 degree field of view. For example, commercially available LiDAR sensors 1564 may have an advertised range of about 1500 m, with an accuracy of 2 cm to 3 cm, and with support for a 1200 Mbit / s Ethernet connection. In some examples, one or more non-foregoing LiDAR sensors 1564 may be used. In such examples, the one or more LiDAR sensors 1564 may be implemented as a small device that may be embedded in the front, rear, sides, and / or corners of the vehicle 1500. The one or more LiDAR sensors 1564 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees, with a range of 200 m, even for low reflectivity objects, in such examples. The one or more front-mounted LiDAR sensors 1564 may be configured for a horizontal field of view between 45 degrees and 135 degrees.In some examples, LiDAR technologies such as 3D flash LIDAR may also be used. 3D flash LiDAR uses a laser flash as a transmission source to illuminate the environment of the vehicle up to about 200 meters. A flash LiDAR unit includes a receptor that records the time of flight of the laser pulse and the reflected light on each pixel, which in turn corresponds to the distance between the vehicle and the objects. Flash LiDAR may allow highly accurate and distortion-free images of the environment to be generated with each laser flash. In some examples, four flash LiDAR sensors may be employed, one on each side of the vehicle 1500. Available 3D flash LiDAR systems include a solid state 3D focal plane array LiDAR camera that does not include moving parts other than a blower (e.g., a non-scanning LiDAR device). The flash LiDAR device may use a 5 nanosecond Class I (eye proof) laser pulse per frame and acquire the reflected laser light in the form of 3D range point clouds and co-registered intensity data. By using flash LiDAR, and because flash LiDAR is a solid state device without moving parts, the one or more LiDAR sensors 1564 may be less susceptible to motion blur, vibration, and / or shock.The vehicle may also include one or more IMU sensors 1566. The one or more IMU sensors 1566 may be disposed in the center of the rear axle of the vehicle 1500, in some examples. The one or more IMU sensors 1566 may include, for example and without limitation, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other types of sensors. In some examples, such as in six-axis applications, the one or more IMU sensors 1566 may include accelerometers and gyroscopes, while in nine-axis applications the one or more IMU sensors 1566 may include accelerometers, gyroscopes, and magnetometers.In some embodiments, the one or more IMU sensors 1566 may be implemented as a miniaturized, high performance GPS-based inertial navigation system (GPS) combining microelectromechanical system (MEMS) inertial sensors, a high sensitivity GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and attitude. Thus, in some examples, the one or more IMU sensors 1566 may enable the vehicle 1500 to estimate the heading without requiring input from a magnetic sensor by observing and correlating the changes in speed from the GPS directly with the one or more IMU sensors 1566. In some examples, the one or more IMU sensors 1566 and the one or more GNSS sensors 1558 may be combined into a single integrated unit.The vehicle may include one or more microphones 1596 mounted in and / or around the vehicle 1500. The one or more microphones 1596 may be used to capture and identify emergency vehicles, among other things.The vehicle may further include any number of camera types, including one or more stereo cameras 1568, one or more wide-angle cameras 1570, one or more infrared cameras 1572, one or more environmental cameras 1574, one or more long-range and / or mid-range cameras 1598, and / or other camera types. The cameras may be used to capture image data around the entire periphery of the vehicle 1500. Which types of cameras are used will depend on the embodiments and requirements of the vehicle 1500, and any combination of camera types may be used to ensure the necessary coverage around the vehicle 1500. Moreover, the number of cameras may be different depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. The cameras may support gigabit multimedia serial link (GMSL) and / or gigabit Ethernet, by way of example and without limitation. Each of the one or more cameras will be described in more detail herein with reference to FIGS. 15A and 15B.The vehicle 1500 may further include one or more vibration sensors 1542. The one or more vibration sensors 1542 may measure vibrations of components of the vehicle, such as the one or more axles. For example, changes in vibrations may be indicative of a change in road surface. In another example, when two or more vibration sensors 1542 are used, the differences between the vibrations may be used to determine friction or slip on the road surface (e.g., when the difference in vibration exists between a driven axle and a free-rotating axle).The vehicle 1500 may include an ADAS system 1538. The ADAS system 1538 may include a SoC in some examples. The ADAS system 1538 may include autonomous / adaptive / automatic cruis...

Claims

A system, comprising: one or more processors to: obtain, using a direct memory access (DMA) system, data representing a frame format comprising a set of DMA transfers to be performed in a sequence according to a frame type, wherein the frame format includes a set of descriptor identifiers corresponding to descriptors; determine, using the DMA system, the frame type of the frame format from a set of frame types based at least on the frame format; Obtaining, using the DMA system, data associated with descriptors based at least on the descriptor identifiers of the frame format and the frame type of the frame format; and causing, using the DMA system, execution of the DMA transfers according to the sequence between a source memory and a destination memory based at least on the frame format and the descriptors, wherein the DMA system is configured to process frame formats associated with each frame type of the set of frame types.The system of claim 1, wherein the frame format is associated with a frame addressing frame format, and wherein the one or more processors are to: configure, using the DMA system, the set of DMA transfers to be performed based at least on the frame addressing frame format indicating that the set of DMA transfers are to be performed according to a raster scan sequence, wherein the raster scan sequence is associated with at least one pass order of a plurality of pass orders.The system of any preceding claim, wherein the frame format is associated with a descriptor addressing frame format, and wherein the one or more processors are to: cause, based at least on the configuration of the DMA system, the DMA transfers to be performed based at least on the descriptor addressing frame format indicating that the DMA transfers are to be performed based at least on a configuration of the DMA transfers by an accelerator.The system of any preceding claim, wherein the frame format is associated with a frame format for addressing random regions, and wherein the one or more processors are to: cause, using the DMA system, the set of DMA transfers to be performed based at least on the set of descriptors corresponding to the regions of interest identified by the frame format within a frame.The system of any preceding claim, wherein the one or more processors are to: determine the frame type based at least on the frame format, wherein the frame type indicates that one or more byte fields of the frame format are reserved byte fields; and obtain the data associated with the descriptors based at least on the frame format type.The system of claim 5, wherein the frame type includes: a descriptor addressing frame type associated with one or more updated descriptors generated using an accelerator, or a random region frame type associated with one or more descriptors indicating at least one offset and at least one descriptor associated with a tile bounding a region of interest within a frame, wherein the frame is specified by the frame format.The system of any preceding claim, wherein the one or more processors are to: cause the set of DMA transfers to be performed in a single channel based at least on the frame format type and descriptors.The system of any preceding claim, wherein the frame format comprises one or more reserved byte fields; and wherein the one or more processors are to: obtain the data associated with the descriptors, wherein the descriptors comprise one or more descriptor byte fields corresponding to one or more of the reserved byte fields of the frame format.The system of any preceding claim, wherein the one or more processors are to: determine one or more aspects of the set of DMA transfers based at least on the frame type; and determine the sequence of the set of DMA transfers based at least on one or more aspects of the set of DMA transfers.The system of any preceding claim, wherein the one or more processors comprise at least one of: an autonomous or semi-autonomous machine control system; an autonomous or semi-autonomous machine perception system; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative creation of content for 3D assets; a system for performing deep learning operations; a system for generating or presenting at least one of augmented reality content, virtual reality content, or mixed reality content; a system for hosting one or more real-time streaming applications; a system for implementing large language models (LLMs); a system for implementing vision language models (VLMs); a system for implementing multi-modal language models; a system implemented using an edge device; a system implemented using a robot; a system for performing operations with conversational AI; a system for performing operations with generative AI; a system for generating synthetic data; a system including one or more virtual machines (VMs); a system at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources.One or more processors, comprising: one or more circuitry to: obtain data representing a frame format including a set of DMA transfers to be performed in a sequence; determine a frame type of the frame format from a set of frame types based at least on the frame format; obtain data associated with the descriptors based at least on the descriptor identifiers of the frame format and the frame type of the frame format; and cause the set of DMA transfers to be performed according to the sequence between a source memory and a destination memory based at least on the frame format and the descriptors.The one or more processors of claim 11, wherein the frame format is associated with a frame addressing frame format, and wherein the one or more circuits are to: cause the set of DMA transfers to be performed based at least on the frame addressing frame format indicating that the set of DMA transfers is to be performed according to a raster scan sequence, wherein the raster scan sequence is associated with at least one pass order of a plurality of pass orders.The one or more processors of any of claims 11 or 12, wherein the frame format is associated with a descriptor addressing frame format, and wherein the one or more circuits are to: cause the set of DMA transfers to be performed based at least on the descriptor addressing frame format indicating that the set of DMA transfers is to be performed based at least on a configuration of the DMA transfers using an accelerator.The one or more processors of any of claims 11 to 13, wherein the frame format is associated with a frame format for addressing random regions, and wherein the one or more circuits are to cause the set of DMA transfers to be performed based at least on the set of descriptors corresponding to the regions of interest identified by the frame format within a frame.The one or more processors of any of claims 11 to 14, wherein the one or more circuits are to: determine the frame type based at least on the frame format, the frame type indicating that one or more byte fields of the frame format are reserved byte fields; and obtain the data associated with the descriptors based at least on the frame format type.The one or more processors of claim 15, wherein the frame type includes: a descriptor addressing frame type associated with one or more updated descriptors generated by an accelerator, or a random region frame type associated with one or more descriptors indicating at least one offset and at least one descriptor associated with a tile bounding a region of interest within a frame, the frame being specified by the frame format.The one or more processors of any of claims 11 to 16, wherein the one or more circuits are to cause the set of DMA transfers to be performed in a single channel based at least on the frame format type and descriptors.The one or more processors of any of claims 11 to 17, wherein the frame format comprises one or more reserved byte fields; and wherein the one or more circuits are to: obtain the data associated with the descriptors, wherein the descriptors comprise one or more descriptor byte fields corresponding to one or more of the reserved byte fields of the frame format.The one or more processors of any of claims 11 to 18, wherein the one or more processors comprise at least one of: an autonomous or semi-autonomous machine control system; an autonomous or semi-autonomous machine perception system; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for generating or presenting at least one of augmented reality content, virtual reality content, or mixed reality content; a system for hosting one or more real-time streaming applications; a system for implementing large language models (LLMs); a system for implementing vision language models (VLMs); a system for implementing multi-modal language models; a system implemented using an edge device; a system implemented using a robot; a system for performing operations with conversational AI; a system for performing operations with generative AI; a system for generating synthetic data; a system including one or more virtual machines (VMs); a system at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources.A method comprising: obtaining data representing a frame format comprising a set of DMA transfers to be performed in a sequence, the frame format including a set of descriptor identifiers corresponding to the descriptors; determining the frame type of the frame format from a set of frame types based at least on the frame format; obtaining data associated with the descriptors based at least on the descriptor identifiers of the frame format and the frame type of the frame format; and causing the set of DMA transfers to be performed according to the sequence between a source memory and a destination memory based at least on the frame format and the descriptors.

Citation Information

Patent Citations

  • US-PATENTANMELDUNGNR.16/101,232

  • US-PATENTANMELDUNGNR.17/391,491