Systems, devices and / or methods for servicing redirected memory access requests using direct memory access scatter and / or aggregate operations

By optimizing memory access through DMA controllers and buffers, the problem of low efficiency in signal fusion from multiple data sources is solved, enabling efficient signal processing in self-driving and autonomous driving applications.

CN120660079AActive Publication Date: 2025-09-16MERCEDES BENZ GRP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380093577.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-04
Filing Date
2023-12-19
Publication Date
2025-09-16
Estimated Expiration
2043-12-19

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently process signal fusion from multiple data sources, especially in self-driving and autonomous driving applications, where sensor-generated signals are processed inefficiently and computational resources and memory access are suboptimal.

Method used

It uses a direct memory access (DMA) controller to receive redirected write and read requests, perform gather and scatter operations, and combine buffers and computing node pools to optimize memory access and computing pipelines for efficient signal confluence and processing.

Benefits of technology

It improves the efficiency of signal processing and computing resource utilization, reduces latency and computing requirements, and is suitable for efficient data processing in self-driving and autonomous driving applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120660079A_ABST
    Figure CN120660079A_ABST
Patent Text Reader

Abstract

Example methods, apparatuses, and / or articles of manufacture are disclosed that may be implemented in whole or in part in conjunction with processing a signal stream. In one application, a direct memory access (DMA) controller may perform an aggregation operation based at least in part on one or more redirected write requests to obtain one or more data items from memory; transitioning the obtained data items to one or more addresses in the memory; and performing a scatter operation based on the one or more addresses.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art 1. Technical Field

[0001] The subject matter disclosed herein relates to processing signals received in streams from multiple data sources.

[0002] 2. Information

[0003] Self-driving and / or autonomous driving applications and other automotive and robotic applications can rely on the fusion of signals, measurements and / or observations generated by multiple sensors. The processing for such applications can involve manipulating arrays of data elements. Such applications can be performed and / or implemented by commercially available central processing units (CPUs) and / or graphics processing units (GPUs). Such commercially available processing units can be configured to manipulate the elements of an input array to generate elements of an output array. Summary of the Invention

[0004] One embodiment disclosed herein relates to a system comprising: a memory comprising one or more memory devices; and a direct memory access (DMA) controller coupled to the memory by a bus, the DMA controller configured to: receive one or more redirected write requests; perform a gather operation based at least in part on the received one or more redirected write requests to obtain one or more words from the memory; convert the one or more words obtained by performing the gather operation to one or more addresses in the memory; and perform a scatter operation to write data items to the one or more addresses in the memory. In a specific implementation, the memory is a first memory, and the system further comprises a second memory operable / functioning as a buffer, wherein the buffer is configured to store values ​​and / or states retrieved from the first memory. The DMA controller may also be configured to parse the values ​​and / or states in the buffer to determine at least one of the one or more addresses in the first memory. In one example, the DMA controller may be further configured to apply one or more arithmetic operations to at least one of the parsed values ​​and / or states to determine the at least one address.

[0005] In another specific embodiment, the DMA controller is further configured to interpret the one or more redirected write requests as a request for a gather operation following the scatter operation. In yet another specific embodiment, the DMA controller is further configured to form a request for the scatter operation for the redirected write request based at least in part on interpreting and / or converting data from the first gather operation into the one or more addresses in the memory to which the data is to be written. In yet another specific embodiment, the system further includes an initiator configured to initiate the one or more redirected write requests, and wherein the DMA controller is configured to perform the first gather operation in response to a signal from the initiator. The initiator of the one or more redirected read requests may include one or more registers of a local memory or a compute node, the one or more registers of the local memory or the compute node being configured to receive sensor measurements and / or observations as data items in a signal stream.

[0006] Another specific embodiment disclosed herein relates to a method at a direct memory access (DMA) controller, the method comprising: performing a first gather operation based at least in part on one or more redirected read requests; transforming one or more words obtained by performing the first gather operation to one or more addresses; and performing a second gather operation to forward data items located at the one or more addresses to a destination determined at least in part based on the one or more redirected read requests. In one example, transforming the one or more words obtained by performing the first gather operation to the one or more addresses further comprises parsing values ​​and / or states in a buffer to determine a memory address. In another example, the method further comprises applying one or more arithmetic operations to the parsed values ​​and / or states to determine the memory address.

[0007] In one particular implementation, the method further includes interpreting the one or more redirected read requests as a first aggregate operation request of two aggregate operation requests. In another particular implementation, the method further includes forming a second request of the two aggregate operation requests based at least in part on the one or more addresses determined from a result of the first aggregate operation. In yet another particular implementation, the method further includes performing the first aggregate operation in response to a signal from an initiator of the one or more redirected read requests. In one example, the initiator of the one or more redirected read requests includes a process executed by a compute node to perform one or more sensor fusion operations.

[0008] Another specific embodiment disclosed herein is directed to a system comprising: a memory comprising one or more memory devices; and a direct memory access (DMA) controller coupled to the memory by a bus. The DMA controller may be configured to: receive one or more redirected read requests; perform a first gather operation based at least in part on the received one or more redirected read requests to obtain one or more words from the memory; convert the one or more words obtained by performing the first gather operation to one or more addresses; and perform a second gather operation to forward data items located at the one or more addresses to a destination determined at least in part on the one or more redirected read requests. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The subject matter claimed is particularly pointed out and distinctly claimed in the concluding portion of the specification. However, as to the organization and / or method of operation, together with objects, features and / or advantages thereof, if used in conjunction with the accompanying Figure 1 The foregoing may be best understood by reference to the following detailed description, read together with the accompanying drawings:

[0010] Figure 1 is a schematic diagram of a computing device according to an embodiment;

[0011] Figure 2A and Figure 2B is a schematic diagram of a computing device including a direct memory access (DMA) controller and / or engine including a buffer according to one embodiment;

[0012] Figure 3A and Figure 3E is a schematic diagram of a computing device including a DMA controller and / or engine including a buffer for facilitating scatter-gather operations according to one embodiment;

[0013] Figure 3C is a schematic diagram illustrating a non-addressable portion of an addressable row according to one embodiment;

[0014] Figure 3B and Figure 3D is a flow chart of a process for facilitating scatter operations and gather operations according to one embodiment;

[0015] Figure 4A and Figure 4D is a schematic diagram of a computing device including a DMA controller including a buffer for facilitating redirection of DMA transactions according to one embodiment;

[0016] Figure 4B and Figure 4C is a flow chart of a process for facilitating redirection of DMA transactions according to one embodiment;

[0017] Figure 5A and Figure 5C is a schematic diagram of a computing device for facilitating processing of multiple signal streams from an associated plurality of sources, according to one embodiment;

[0018] Figure 5B is a flow chart of a method for processing multiple data streams from associated multiple sources according to one embodiment;

[0019] Figure 6 is an illustration depicting example sensor signal aggregation for an example vehicle according to one embodiment; and

[0020] Figure 7 An example schematic block diagram depicts features of an example vehicle in a self-driving / autonomous driving application according to one embodiment.

[0021] In the following detailed description, reference is made to the accompanying drawings, which form part of this description, wherein similar reference numerals may refer to corresponding and / or similar similar parts throughout the text. It should be understood that, such as for the simplicity and / or clarity of illustration, the drawings are not necessarily drawn to scale. For example, the dimensions of some aspects may be exaggerated relative to other aspects. In addition, it should be understood that other embodiments may be utilized. In addition, structural and / or other changes may be made without departing from the claimed subject matter. Throughout this specification, references to "claimed subject matter" refer to the subject matter intended to be covered by one or more claims or any part thereof, and are not necessarily intended to refer to a complete set of claims, a specific combination of claim sets (e.g., method claims, device claims, etc.), or a specific claim. It should also be noted that directions and / or references (e.g., such as upward, downward, top, bottom, etc.) can be used to facilitate discussion of the drawings and are not intended to limit the application of the claimed subject matter. Therefore, the following specific embodiments should not be considered to limit the claimed subject matter and / or equivalents. DETAILED DESCRIPTION

[0022] Throughout this specification, references to an implementation, an implementation, an embodiment, an embodiment, etc. mean that the specific features, structures, characteristics, and / or similar items described with respect to a particular implementation and / or embodiment are included in at least one implementation and / or embodiment of the claimed subject matter. Thus, the appearance of such phrases, for example, in various places throughout this specification is not necessarily intended to refer to the same implementation and / or embodiment or any one specific implementation and / or embodiment. Furthermore, it should be understood that the specific features, structures, characteristics, and / or similar items described can be combined in various ways in one or more implementations and / or embodiments and are therefore within the intended scope of the claims. Generally speaking, of course, just like the description of a patent application, these and other issues are likely to vary in a particular context of use. In other words, throughout this disclosure, specific description situations and / or usage situations provide useful guidance on reasonable inferences to be drawn; however, again, "in this context," without further qualification, generally refers to at least the context of this patent application.

[0023] To facilitate efficient processing of data items in one or more signal streams, a direct memory access (DMA) controller coupled to a memory by a bus may be configured to receive one or more redirected write requests. The DMA controller may, based at least in part on the received one or more redirected write requests, perform a gather operation to obtain one or more words from the memory; transform the one or more words obtained from performing the gather operation into one or more addresses in the memory; and perform a scatter operation to write the data items at the one or more addresses in the memory determined by the transformation. In another embodiment, the DMA controller may process the one or more redirected read requests by performing a first gather operation based at least in part on the one or more redirected read requests; transform the one or more words obtained from performing the first gather operation into one or more addresses; and perform a second gather operation to forward the data items at the one or more addresses determined by the transformation to a destination determined at least in part on the one or more redirected read requests.

[0024] In some cases, the host central processing unit (CPU) can be part of a computing device in the vehicle and can process data items in a memory for automotive applications. As an example, automotive applications such as self-driving / autonomous driving applications (e.g., fully autonomous, semi-autonomous, driver assistance systems, etc.) can use particle filters to fuse sensor signals and / or observed signal streams, for example, to update particle filter states. Such applications can be implemented, for example, in systems such as automated machines (cars, trucks, etc.). In this context, a "signal stream" as mentioned herein refers to the time-varying progression of a sequence of coded data items to be delivered to a recipient device via a signal transmission medium. The coded data items (also referred to as data values) delivered in the signal stream can express attributes indicating conditions and / or events, subject identifiers, timestamps indicating the time of the event, metadata, to name just a few examples of attributes that can be expressed by the coded data items delivered in the signal stream. In a specific implementation, the signal stream can deliver sensor measurements and / or observations together with associated timestamps to indicate the time when such measurements and / or observations were obtained.

[0025] In one aspect of one embodiment, information originating from different sources and arriving in corresponding signal streams can be processed as a "sink" of information to determine a computational result. Such a sink of information can be processed by correlating and / or correlating information items from different sources according to specific attributes (e.g., time, space, reliability, confidence, etc.). Processing of the sink of information can then include performing one or more operations on the items based on one or more attributes to produce a result. In one specific implementation, data items in the sink of information can be processed by updating one or more states of a particle filter. For example, such a particle filter can implement processing of a sink of arrays of sensor signals / observations (e.g., received from a signal stream) to update the states of measurement particles, filter particles, static particles, and / or dynamic particles. In one specific implementation, the measurements and / or observations in the sink of measurement arrays can be generated from different sensors. Such measurements and / or observations generated by different sensors can still be correlated and / or correlated in time and space. According to one embodiment, correlating data items from different sources in the sink of arrays can be implemented, at least in part, using a radix sort applied to the array of keys. In one particular implementation, sensor signals and / or observations may be formatted into arrays to be processed to generate a sink of arrays. An example process for generating such a sink of arrays may be performed according to the following pseudo code:

[0026]

[0027] According to one embodiment, the confluence of multiple signal streams may include a mapping of the data items in different input signal streams to the data items in one or more output signal streams. For example, the data items in this type of input signal stream may include sensor measurements and / or observations of the sensor associated with this input signal stream. Therefore, the confluence of this type of input signal stream may include a mapping of the data items in sensor measurements and / or observations (for example, from the different / distinguished sensors associated with this input signal stream) to the data items in output signal streams. This type of data items in this output signal stream may include the value inferred / calculated based on this sensor measurements and / or observations. In the pseudocode example provided above, the quantity of input signal streams may be defined as S_in[] (comprising data item value_in[]), and the quantity of output signal streams may be defined as S_out[] (comprising data item value_out[]). Here, the expression "(value_out[], S_out_enable[]) = f(value_in[], parameters) for (i: 0.. (l-1))" can map the data item value_in[] in the input signal stream S_in[] to the data item value_out[] in the output signal stream S_out[] according to the function f().

[0028] According to one embodiment, computing devices, circuits and / or logic may form a "convergence engine" (CE), also referred to as a "convergence" and / or a "convergence processor" (CP), to handle the confluence of data items as discussed above. Such a confluence of data items may include, for example, a confluence of signal streams and / or a sequence of confluences of signal streams, wherein the confluence of signal streams and / or the sequence of confluences of signal streams have reduced latency and / or reduced computing resources (e.g., power, memory, etc.). In a particular embodiment, the output of a confluence operation of such a CE may provide all or part of the input to a subsequent confluence operation. Characteristics of the confluence, such as functions like k, f(), read_next() as shown in the pseudo-code example above, may be part of the runtime programming of the CE.

[0029] According to one embodiment, the CE may employ a direct memory access (DMA) subsystem that may include a DMA controller (also referred to as a DMA "engine") that is configurable to, for example, initiate read operations and write operations between row accessible memory and transfer memory. The specific process by which the CE accesses memory may be determined in parameter form at compile time and in physical form at runtime. Thus, in one particular implementation, the events that trigger the execution of the DMA controller may not be limited to events occurring at the arithmetic logic unit (ALU) (e.g., loads and stores from the ALU). The DMA controller may be triggered to execute a DMA transaction by the start of a merge operation. The DMA controller may then execute such DMA transactions independently of the arithmetic logic unit (ALU) (e.g., depending solely on the availability of read and write access at the valid endpoints of such DMA transactions).

[0030] Figure 1 1 is a schematic diagram of a system 100 for performing DMA transactions. System 100 includes multiple components that communicate via bus 101. The components include a host CPU 102, a memory controller 106, RAM 108 (which may form main system memory), peripheral devices 114, and a DMA controller 112. In this context, "direct memory access," as referred to herein, refers to a process performed by one or more hardware subsystems and / or circuits to access a particular memory independently of a particular processing unit and / or central processing unit (e.g., independent of host CPU 102). According to one embodiment, DMA controller 112 may initiate operations to access (e.g., read access or write access) random access memory (RAM) 108 via bus 101 independently of host CPU 102. For example, DMA controller and / or engine 112 may control and / or execute transactions to transfer data items (also referred to as data values) between peripheral devices 114 and RAM 108 independently of the actions of host CPU 102 (via memory controller 106). For example, such DMA transactions may be triggered by signals, conditions, and / or events (eg, interrupt signals).

[0031] As discussed above, the processing of data items for automotive applications or other computing applications may be enhanced by a Convergence Engine (CE). Figure 2A and Figure 2Bis a schematic diagram of a computing device 200 according to one embodiment that includes a direct memory access (DMA) controller 212 (also referred to as a DMA engine 212) and a buffer 216 to implement one or more aspects of CE. The computing device 200 may be or may form, for example, a system on a chip (SoC), a microchip, a control circuit, or some other computing device. In some implementations, the computing device 200 may form a vehicle controller, such as an advanced driver assistance system (ADAS) device, a telematics control unit (TCU), an electronic control unit (ECU), a centralized vehicle computer, or some other vehicle controller. Figure 2A and Figure 2B As illustrated, computing device 200 may also include a CPU 202 (also referred to as a host CPU), a transfer memory 218 , a line-accessible memory 208 , a buffer 216 , and a compute node (CN) pool 220 (also referred to as CN pool 220 ).

[0032] In one embodiment, a CN may include a single processing circuit core capable of performing operations to map input operands to output computation results. In another embodiment, a CN may include multiple distinct processing cores configured to perform operations to map input operands to output computation results. In another embodiment, two or more non-concurrently executing CNs may be implemented on the same processing circuit core. For example, a processing circuit may execute a first CN to generate an output result (e.g., stored in transfer memory 218), which is then input to a second CN executed later on the same processing circuit core.

[0033] The CNs in CN pool 220 may include dedicated local memory (e.g., static random access memory (SRAM)) and general-purpose registers to receive operands for operations to be performed and / or provide results of the operations. In this example, host CPU 202 may use transfer memory 218 to store data items and may control the CNs in CN pool 220 to perform operations on these data items. Transfer memory 218 may be physically closer to the CNs and / or may operate with lower access latency, and thus may be used as a (high-speed) cache to store data items. Line-accessible memory 208 may be external to transfer memory 218 and may provide a larger amount of memory space than transfer memory 218, but may be physically further away from CN pool 220 and may operate with longer access latency. According to one embodiment, transfer memory 218 may include one or more synchronization mechanisms to facilitate inter-CN communication between and / or among CNs in CN pool 220 (e.g., for synchronizing communication between CNs with different execution latencies). In one particular implementation, the row-accessible memory 208 may be separated from the transfer memory 218, the CPU 202, and the buffer 216 by a bus (not shown). According to one embodiment, the buffer 216 may be formed in circuitry used to implement the core circuitry of the DMA controller and / or engine 212, such that the buffer 216 is distinct and separate from the circuitry used to form the transfer memory 218. Such formation of the buffer 216 in the core circuitry of the DMA controller and / or engine 212 may reduce and / or minimize the latency associated with loading data items into and storing data items from the buffer 216 in the course of performing DMA operations.

[0034] In one embodiment, the DMA controller and / or engine 212 may be configured to interface / dock with the transfer memory 218 and the row accessible memory 208. The transfer memory 218 and / or the row accessible memory 208 may also be cache line addressable. In other words, in this example, the row accessible memory 208 may be a cache line addressable memory. The transfer memory 218 and / or the row accessible memory 208 may provide data items (also referred to as data values) that the CNs in the CN pool 220 may manipulate, operate on, or otherwise process. According to one embodiment, all or a portion of the transfer memory 218 may be organized as a cache that may be integrated with commercially available components. It should also be noted that the cache in the transfer memory 218 or in the buffer 216 may mitigate manufacturing defects and / or enable various embodiments to be applicable to application scales beyond expectations.

[0035] In one embodiment, the DMA controller and / or engine 212 can be configured to handle cache line sized data items (e.g., 64 bytes or 128 bytes), wherein such data items can be located and accessed by cache line addresses even in scatter-gather operations. Such cache line addresses for 64-byte cache lines can be expressed in binary notation ending with six zeros. Similarly, cache line addresses for 128-byte cache lines can be expressed in binary notation ending with seven zeros. According to one embodiment, the DMA controller and / or engine 212 can be configured to handle word sized data items, even when the row accessible memory 208 can continue to be addressable only at cache lines. The transfer memory 218 can be organized as a (high-speed) cache, a word addressable register, or a collection of transfer memory buffers such as word-width first-in-first-out (FIFO) buffers, or a combination thereof. In one embodiment, such transfer memory buffers in the transfer memory 218 can include circuits and / or devices specifically designed to act as FIFO buffers. In other specific implementations, such a transfer memory buffer in the transfer memory 218 may include a static random access memory (SRAM) device or a network-on-chip (NOC) device specifically configured to act as a FIFO buffer. To facilitate scatter operations or gather / gather operations to transfer word-sized data items to and from the row-accessible memory 208, the DMA controller and / or engine 212 may implement a buffer 216 (e.g., located between the row-accessible memory 208 and the transfer memory 218). The buffer 216 may be configured to store bytes from multiple cache lines together for placement into a destination word in the transfer memory 218. The DMA controller and / or engine 212 may also be capable of performing multicast write operations in either direction (e.g., from the transfer memory 218 to the row-accessible memory 208 or from the row-accessible memory 208 to the transfer memory 218). The buffer 216 may be distinct / different from the transfer memory 218 and the memory DMA controller and / or engine 212. As discussed herein, buffer 216 may be formed in the client core to implement DMA controller and / or engine 212 .

[0036] In this context, "transfer memory," as referred to herein, means circuitry for facilitating the communication of data items between and / or among CNs, such as the CNs in the CN pool 220. In one particular implementation, such a transfer memory may transmit the result of executing a first operation at a first CN as an input operand for a second operation to be executed at a second CN (e.g., in a computation pipeline). In one particular implementation, the transfer memory 218 may be organized as a static random access memory (SRAM) device used as access-controlled memory or a shared memory used as cache memory, a word-addressable or SIMD vector-addressable register file, or circuitry and / or devices specifically configured to act as a first-in-first-out (FIFO) buffer. Such circuitry and / or devices specifically configured to act as a FIFO buffer may have a word width or a single instruction, multiple data (SIMD) vector width, or circuitry or a network-on-chip (NOC) device coupled between endpoints, or a combination thereof, to provide just a few examples.

[0037] As described above, the CNs in CN pool 220 (e.g., CN pool) can operate on data items read from memory. In this context, "computing node" as mentioned herein means an identifiable and distinguishable set of computing resources (e.g., hardware and executable instructions) that can be configured to perform operations to process input values ​​to provide output values. The CNs in CN pool 220 can include scalar CNs and / or processing circuit cores to implement arithmetic logic units (ALUs), digital signal processors (DSPs), vector CNs, VLIW engines, or field programmable gate arrays (FPGAs) units, or combinations thereof, providing only a few examples of specific circuit cores that can be used to implement the CNs in CN pool 220. CN pool 220 can be implemented according to various architectures. For example, CN pool 220 can include one or more CNs implemented in full-featured or simplified form according to a reduced instruction set computing (RISC) architecture, complex instruction set computing (CISC) or very long instruction word (VLIW) architecture, or some combination of these types. The CNs in CN pool 220 can also include a combination of scalar, SIMD, or multiple instruction and single data stream (MISD) ALUs.

[0038] According to one embodiment, features of computing device 200 may include commercially available features such as, for example, a transport ring, a Clos network, a shuffle circuit, etc. CN pool 220 may facilitate multithreading in CNs, CN clustering, mailboxes, interrupts, inter-CN synchronization features, atomic operations at locations in transport memory 218, features added to meet safety and security requirements, to provide just a few examples.

[0039] In certain implementations, CN clusters within CN pool 220 may be formed based, at least in part, on resource tradeoffs and may be permanently defined within an integrated circuit (IC) device. In certain embodiments, two devices within the same integrated circuit (IC) device may implement different clusters of associated CNs, depending on how CN pool 220 is configured. For example, a processor configuration may define the processing of multiple sinks with a CN cluster based, at least in part, on the associated sinks to be executed. Each associated cluster may, for example, process an array of sinks, such that the output of one cluster provides input to one or more other clusters.

[0040] According to one embodiment, transmit memory 218 may form one or more buffers, including one or more first-in, first-out (FIFO) buffers. The one or more FIFO buffers may include a "vertical" FIFO buffer. Such a vertical FIFO buffer may be a buffer having an endpoint that interfaces with the DMA controller and / or engine 212 (such an endpoint may be referred to as an "external" endpoint) and another endpoint that forms a register (such an endpoint may be referred to as an "internal" endpoint) that provides data items that can be used as operands for CNs in CN pool 220. The internal endpoint of the vertical FIFO buffer may be shared by multiple CNs in CN pool 220. If such an internal endpoint includes a FIFO outbound endpoint (i.e., the outbound end of the FIFO buffer), broadcasting to multiple CNs in CN pool 220 may be achieved, for example. If such an internal endpoint includes a FIFO inbound endpoint (i.e., the inbound end of the FIFO buffer), hardware locks and / or instructions executed on CNs in CN pool 220 may prevent race conditions. In some implementations, the FIFO buffer formed in transmit memory 218 may also include a "horizontal" FIFO buffer. This horizontal FIFO buffer can be a buffer with two endpoints that provide data items as operands for different CNs in CN pool 220, or simply provide data items to locations in transfer memory 218. It should be noted that the FIFO buffer formed in transfer memory 218 can provide a transparent blocking and releasing mechanism, for example, when input and output speeds differ. If the external endpoints of the FIFO buffer in transfer memory 218 are registers or operands consumed / processed by some CNs in CN pool 220, and if CN processing is slow, the DMA controller and / or engine 212 may ultimately implement a blocking mechanism. Conversely, if the DMA controller and / or engine 212 writes slowly to such a FIFO buffer in transfer memory 218, the CNs in CN pool 220 may ultimately implement a blocking mechanism. A similar transparent blocking and releasing mechanism can exist if FIFO buffers are implemented between and / or within CNs and / or between and / or within locations in transfer memory 218.

[0041] In an IC device, if transfer memory 218 forms a FIFO buffer, the circuitry at the endpoints of the FIFO buffer may be permanent or configurable (e.g., via internal FPGA circuitry). A specific implementation may include segments of a vertical FIFO buffer and a horizontal FIFO buffer formed within the IC device circuitry. Such buffer circuitry may have endpoints that are configurable at runtime to be at least one of: operands or registers of a CN in CN pool 220, an interface to buffer 216, a location in transfer memory 218, and endpoints of other FIFO buffers in transfer memory 218. If the endpoints of the FIFO buffer of transfer memory 218 are operands or registers of a CN in CN pool 220, such FIFO buffer endpoints may be shared among multiple CNs in CN pool 220.

[0042] According to one embodiment, the CNs in the CN pool 220 may receive operands for computational operations, for example, from registers in a register file, from a FIFO buffer in the transfer memory 218, and / or from special constant and parameter registers (which may include shared registers). In some cases, a portion of the register file may be shared between multiple CNs in the CN pool 220. Within the computing device 200, the host CPU 202 may also be able to access the constant and parameter registers. The host CPU 202 may provide host functions, such as, for example, initiation of a busing operation to be performed after configuration tasks have been completed. Such configuration tasks may include defining clusters of CNs in the CN pool 220, configuring inter-CN communication, configuring the width and depth of the FIFO buffer in the transfer memory 218, configuring endpoints of the FIFO buffer in the transfer memory 218, defining a manager CN in the CN pool 220 and the CN clusters in the CN pool 220 that the manager CN is to manage, establishing a DMA controller and / or engine 212, establishing atomics, establishing communication and synchronization resources, and monitoring the termination of a bus, to provide just a few examples.

[0043] In one embodiment, the buffer 216 can facilitate the processing of the bus in multiple aspects. In one such non-limiting aspect, the input signal stream in the bus can include an "indirect stream" (e.g., a set of addresses in the row-accessible memory 208 to be read). In this case, such an indirect signal stream may not be directly fed to the CN in the CN pool 220, but instead may be provided to the DMA controller and / or engine 212, which may transmit a signal stream of read data items to the CN. In other words, the latency of accesses through random read operations (random reads of rows or random reads of words or random indirect reads of rows or words, or a combination thereof) can be hidden by extracting the pattern of random accesses, establishing an access list of significant length, and performing the necessary extraction and indirection at the DMA controller and / or engine 212 itself. In other cases, the input signal stream may include a "double indirect stream," in which the signal stream may contain addresses, and after some processing, the data items at these addresses provide further indirection. In a specific embodiment, the first level of indirection of the double indirect stream can be read and used multiple times. In sensor fusion applications, multiple sinks may be configured, where an initial sink may be constructed using a lookup table (LUT) for the first level indirection data items and the second level indirection data items (e.g., in the transmit memory 218 itself). However, it should be understood that these examples are not limiting.

[0044] According to one embodiment, some processing within the sink of the signal flow (e.g., within f() and read_next() in the pseudo-code example above) can be viewed as extracting information from data items distributed across the CNs in the CN pool 220. An example of such extraction is provided in Table 1 below.

[0045]

[0046] Table 1

[0047] According to one embodiment, results may be provided by CNs in CN pool 220 that operate cooperatively using operations such as shifting, shuffling, broadcasting, and multicasting of operands between participating CNs, to provide just a few examples. Such features may exist within a CN, for example, word operands of a SIMD vector CN present in CN pool 220. CN pool 220 may be configured with such features for inter-CN communication via configurable bridge circuits between CNs. Such circuits may be hardwired or configurable at runtime.

[0048] According to one embodiment, CNs in CN pool 220 (e.g., arranged in CN clusters) may be configured for specialized processing functions. In one specific implementation, such specialized CNs in CN pool 220 may facilitate the management of application processing flows. For example, CN pool 220 may include one or more processing CNs 224 and one or more manager CNs 222. Processing CNs 224 may perform processing operations on, for example, sensor observations, measurements, and / or other signals. In this example, manager CNs 222 may manage different sets of processing CNs 224. Processing CNs 224 may communicate with one or more manager CNs 222. In some cases, manager CNs 222 may provide information to processing CNs 224 based on communications from processing CNs 224. Processing CNs 224 may continue their processing as determined by the information provided by manager CNs 222. In some implementations, manager CNs 222 may include physically separate processing CNs 224, or may simply include specialized circuitry formed within processing CNs 224.

[0049] It should be noted that the rates at which different individual signal streams (e.g., from different sources such as different sensors) of a sink are generated and consumed (e.g., processed) may not necessarily be equal. The rate at which such signal streams are consumed or generated may be determined by, for example, operational characteristics of a function such as f() (in the pseudocode example above). For example, the rate at which a signal stream of measurements and / or observations from a sensor used in a sensor fusion operation may be calculated / generated by function f(). Such sensor fusion operations may involve an inverse sensor model that results in an output signal stream that is longer than the input signal stream. To process sinks of longer signal streams, for example, the output rate may be matched to the bandwidth / throughput of the row-accessible memory 208. For example, configuring the output of one CN to be the input to another CN may assist in load balancing.

[0050] According to one embodiment, the computing device 200 can enable the deployment of advanced sensor fusion operations to update particle filter states (e.g., in autonomous driving or other motor vehicle applications) while consuming very little power. In one application, an instance of the computing device 200 can be implemented as, for example, a sequence of pipeline stages. For such applications, features of the computing device 200 can be configured to have different amounts of resources allocated to different pipeline stages when used. By using FIFO buffers (e.g., FIFO buffers formed in the transfer memory 218), the exchange of data items between and / or within the pipeline stages and the row-accessible memory 208 can be synchronized transparently (e.g., without mutexes, spin locks, etc.).

[0051] In another embodiment, the computing device 200 may be configured to optimize for power, space, and / or performance (e.g., accuracy and / or latency). Although the features of the computing device 200 may be suitable for implementing CE, the features of the computing device 200 may be suitable for other applications, including, for example, applications that rely on random access to the row-accessible memory 208. The features of the computing device 200 may also be implemented in so-called "supercomputers." In the context of supercomputers, the low-power features of the computing device may help overcome power constraints that may prevent the implementation of, for example, exascale supercomputers. In addition, the circuitry used to implement the computing device 200 may incorporate safety and security features in a manner that meets the requirements of embedded computing devices. Given that the computing device 200 has a small physical size and low power consumption, the use of the computing device 200 may not necessarily be limited to use as an external accelerator integrated circuit (IC) device, but may also be incorporated into a subsystem within an automotive-grade system-on-chip (SOC) IC device.

[0052] In one aspect, for example, the computing device 200 may include a specific arrangement and / or configuration of CNs (e.g., CN pool 220), transfer memory 218, and / or DMA controller and / or engine 212. The computing device 200 may be configured to provide a network of CNs and memory elements suitable for a specific type of computation, such as a specific application for processing signal streams (e.g., carrying measurements and / or observations from sensors). In one embodiment, such a network of CNs and memory elements may enable simultaneous processing of multiple signal streams with high throughput and low latency. Such a network of CNs and memory elements may be implemented, at least in part, using a specific implementation of an intra-device communication protocol (e.g., AXI) via physical connections and FIFO buffers. As noted above, the endpoints of the FIFO buffers may include addressable memory locations or registers to receive, for example, operands (e.g., general purpose registers configured as ALUs of a compute node as operands for computational operations) or results computed by the CNs. For example, a FIFO buffer pool may include endpoints that can be configured to be associated with various CNs or memories.

[0053] In one aspect, certain embodiments disclosed herein relate to so-called vectored input / output (I / O) operations, including "scatter" and "gather" operations. Such vectored operations can, for example, enable a large amount of data to be transferred to or from a physical memory (e.g., multiple addressable rows of memory in row-accessible memory 208) with high throughput using a single request or command to enhance efficiency and convenience. For example, a gather operation may require reading data from multiple memory locations (e.g., buffers) in a single transaction and writing the read data to a signal stream or a continuous portion of memory. In one embodiment, the DMA controller and / or engine 212 may perform a gather operation to service a gather request (e.g., from an application) that specifies multiple memory locations, not necessarily row-aligned, from which data items are to be read and a destination (e.g., memory address) for storing the read items. On the other hand, a scatter operation may require reading data items from a signal stream or continuous memory and writing the read data items to multiple different memory locations, not necessarily row-aligned. In one specific implementation, the DMA controller and / or engine 312 may perform a gather operation to service a scatter request (e.g., from an application) that may specify locations (e.g., consecutive memory addresses) of data items to be read and may also specify locations to which the read data items are to be written indirectly (by requiring reading locations that enable the determination of the locations to which the read data items are to be written). The DMA controller and / or engine 312 may also perform the indirectly specified gather operation.

[0054] According to one embodiment, the use of DMA transactions to assist in processing data items in a signal stream can be enhanced by using scatter and gather operations. In one particular implementation, the contents loaded into a buffer from a gather operation can be used to determine one or more addresses for a subsequent gather or scatter operation. Figure 3A and Figure 3E is a schematic diagram of a computing device 300 including a DMA controller and / or engine 312 that includes or communicates with a buffer 316 for facilitating scatter and / or gather operations, according to one embodiment. In one particular implementation, computing device 300 may include computing device 200 ( Figure 2A and Figure 2B The DMA controller and / or engine 312 may be configured as a scatter-gather multicast DMA engine (SGM-DMA), but with the ability to address words within an addressable row, for example.

[0055] According to one embodiment, the DMA controller and / or engine 312 may receive input signal streams and / or may feed such input signal streams to an initial cluster of CNs as blocks on virtual channels through FIFO buffers. Subsequently, the output signal streams from the initial cluster of CNs may be fed to subsequent downstream clusters of CNs as blocks on virtual channels or as data through FIFO buffers. For example, transmissions between clusters of CNs may be multicast transmissions as determined by a particular application. At any stage, some or all of the output signal streams from some CN clusters may be returned to the DMA controller and / or engine 312 (e.g., during the DMA execution of a scatter or gather operation).

[0056] According to one embodiment, the DMA controller and / or engine 312 may perform certain scatter and / or gather operations to transfer data items from one non-contiguous block of memory to another non-contiguous block of memory using a series of smaller, contiguous block transfers. Here, obtaining such data items from a non-contiguous block of source memory may be performed in a gather operation. Similarly, writing data items to a non-contiguous block of destination memory may be performed in a scatter operation. In one specific implementation, the smallest memory unit that can be accessed in such source or destination memory may be a single addressable row of values ​​and / or states (e.g., a single cache line or word in row-accessible memory). For example, the DMA controller and / or engine 312 may communicate with a row-accessible memory (LAM) 308, which may be accessible on a row-by-row basis.

[0057] According to one embodiment, a physical memory (such as LAM 308) may include bit cells that define values ​​and / or states to express information such as one or zero. Such physical memory may also organize the bit cells into words containing an integer number of 8-bit bytes (e.g., four-byte words on 32 bits or eight-byte words on 64 bits). In addition, such physical memory may define row addresses (e.g., word row addresses) associated with consecutive bits of an "addressable row" that define values ​​and / or states. For example, in response to a read or write request (e.g., from a host processor), a memory controller may access a portion of the memory in a read or write transaction anchored according to a word row address specified in the request. To service a read request, for example, the memory controller may retrieve the values ​​and / or states of all bytes in the row associated with the row address specified in the read request. Similarly, to service a write request, the memory controller may write the values ​​and / or states of all bytes in the addressable row associated with the row address specified in the write request. Although a row address may specify a memory location that includes all consecutive bytes of an addressable row, such a row address does not specify the location of individual sub-portions of such an addressable row, such as individual bytes or consecutive bytes that are less than the entirety of the addressable row, or bytes that span across addressable rows. Such sub-portions of an otherwise addressable row are referred to herein as "non-addressable portions."

[0058] According to one embodiment, an addressable row may define the smallest memory unit that can be located and / or accessed according to a memory addressing scheme. Figure 3C In a specific example embodiment, such addressable rows 360 may be composed of smaller memory units such as bits, bytes, or words. In a specific example embodiment, the addressable rows 360 include n+1 bytes 3620 to 362 n As noted above, certain specific implementations may involve updating an unaddressable portion and / or a portion that is less than the entirety of an addressable row and / or an unaddressable portion that spans multiple rows (whether or not including the addressable row). In the illustrated example embodiment, bytes 3622 and 3623 may define the unaddressable portion of addressable row 360. Here, while the entirety of the addressable row may be locatable and / or accessible via a unique address according to a memory addressing scheme, bytes 3622 and 3623 (the lesser portion) are not addressable according to the memory addressing scheme.

[0059] In one embodiment, the LAM 308 may have a LAM controller 306 that is configured to receive a request specifying an address for a row of data items stored in the LAM 308. Such a row of data items may include multiple words or multiple bytes and may be the smallest unit of data items that the LAM 308 can retrieve and return to another device. According to one embodiment, the computing device 300 may use a buffer 316 to enable access to smaller, non-addressable portions of a row of data items (e.g., single or multiple bytes in a word or group of words). According to one embodiment, the circuitry forming the buffer 316 may be integrated with the circuitry forming the DMA controller and / or engine 312 so that accesses to the buffer 316 initiated by the DMA controller and / or engine 312 have minimal latency. For example, buffer 316 may be formed as a static random access memory (SRAM) device that can be accessed by DMA controller and / or engine 312 without initiating requests and / or transactions on a main memory bus (e.g., a bus coupled to LAM 308 or a host computer / processor).

[0060] According to one embodiment, the DMA controller and / or engine 312 can communicate with an initiator 322. The initiator 322 can include a device (e.g., implemented at least in part by circuitry and / or logic) that implements a particular state to trigger one or more DMA transactions to be performed by the DMA controller. For example, the initiator 322 can include an output register of an ALU, a buffer, or a hardware interrupt handler, to provide just a few examples of devices that can initiate a DMA transaction.

[0061] According to one embodiment, the DMA controller and / or engine 312 may obtain a list of gather requests in response to a signal from the initiator 322. In one particular implementation, the initiator 322 may trigger a DMA transaction in response to an event or condition during the execution of a particle filtering process. For example, the particle filtering process may identify data items in memory that are expected to be retrieved for processing in a future execution cycle. Once a large number of such data items have been identified, a list of gather requests (e.g., as redirected gather requests) identifying such data items may be forwarded to the DMA controller and / or engine 312. In one embodiment, such a list of gather requests may be provided to the DMA controller and / or engine 312 in shared memory or a network on chip (NOC), to provide a few examples. Once such a list is known to be available for processing by the DMA controller and / or engine 312, the process for generating the list (e.g., the execution of computer-readable instructions) may trigger the DMA controller and / or engine 312 via an interrupt or a posted message. Such a trigger may enable the DMA controller and / or engine 312 to initiate one or more gather operations (e.g., redirected gather operations).

[0062] In response to a signal from the initiator 322, the DMA controller and / or engine 312 may obtain a list of gather requests in the form of a linked list. For example, such a linked list may be locatable in a memory (e.g., the reassembly buffer 316 or the line accessible memory (LAM) 318) based on an address provided by the initiator 322. According to one embodiment, such a list of gather requests may include individual gather requests that may be used as independent gather requests independent of other gather requests in the list of gather requests. The DMA controller and / or engine 312 may compare the addresses in such gather requests to be sent by the memory controller (e.g., Figure 1 316 ). DMA controller and / or engine 312 may form a packet from the extracted items to be forwarded to one or more requesting entities (e.g., processes executing on a host CPU 202 and / or a CN of a CN pool 220). In one aspect, a data item extracted from a row stored in buffer 316 may include two or more non-addressable portions of the row (e.g., selected bytes and / or fields within an addressable row / word such as an addressable row in memory containing the data item).

[0063] Figure 3B A flow chart illustrating a process 350 for a gather operation according to one aspect of the present disclosure is illustrated. In one embodiment, process 350 may include operations 352, 354, 356, and 358, which may be performed by one or more circuits such as a DMA controller and / or engine 312 and / or buffer 316. Operation 352 may include processing one or more gather requests to determine one or more addressable rows of data items to be retrieved from memory. For example, operation 352 may map parameters in a received request to addresses in LAM 308. Operation 354 may be implemented by circuitry to load values ​​and / or states to memory and may include loading signals and / or states (e.g., representing data items) stored in one or more addressable rows of memory such as LAM 308 into buffer 316. Figure 3BAs illustrated, operation 356 may include parsing one or more unaddressable portions of the line loaded into buffer 316 at operation 354 into portions such as, for example, single or multiple bytes in a word or word groups. Operation 358 may include completing processing of the two or more aggregate requests by, for example, returning the unaddressable portions parsed at operation 356.

[0064] like Figure 3B As illustrated, operation 358 may include responding to multiple gather requests presented to the DMA controller and / or engine 312 in the form of a gather request list. Loading one or more addressable lines into the buffer 316 at operation 354 may occur in response to the gather request list. Operation 354 may resolve two or more non-addressable portions based on the data items specified in the gather request list. For example, operation 356 may decode the parameters in the gather request to be mapped into byte offsets from the row address to the corresponding non-addressable portions. The resolved portions may then be forwarded to the initiator of the gather request (e.g., initiator 322). Here, for example, if multiple gather requests specify resolved portions within a single addressable line, process 350 may enable servicing the multiple gather requests by accessing the same single addressable line loaded into the buffer 316. This may eliminate the need for the DMA controller and / or engine 312 to access the same addressable line multiple times for separate gather requests for data items in the same addressable line (e.g., in the LAM 308).

[0065] To service one or more gather requests, process 350 may gather less than the entirety of an addressable line in memory by loading the addressable line into a buffer and resolving the non-addressable portion to be provided to the requester. While some gather requests may require gathering less than the entirety of an addressable line, one or more received gather requests may require gathering of the entirety of an addressable line and / or multiple lines and / or bytes across lines. For gather requests that require gathering less than the entirety of an addressable line, the DMA controller and / or engine 312 may perform process 350. According to one embodiment, for gather requests that require gathering of the entirety of an addressable line, the DMA controller and / or engine 312 may bypass operations 354, 356, and 358 and perform a gather operation without loading the addressable line into buffer 316.

[0066] In another specific implementation, the DMA controller and / or engine 312 may obtain a scatter request list in response to a signal from the initiator 322. The DMA controller and / or engine 312 may obtain the scatter request list in the form of a linked list that can be located in memory based on the address provided by the initiator 322. The DMA controller and / or engine 312 may then combine the addresses to be accessed by such scatter requests with a (potentially smaller) list of row read requests to be executed by a memory controller (e.g., memory controller 106). For example, scatter requests in a scatter request list that reference data items in the same addressable row of memory may be combined so that only a single row read of the request is required (to access the data items to service multiple scatter requests). The requested rows read by such a memory controller may be loaded into the buffer 316.

[0067] According to one embodiment, the obtained scatter request list may indicate specific non-addressable portions (e.g., individual bytes or fields) of an addressable row to be read and loaded into the buffer 316. The non-addressable portion of the addressable row loaded into the buffer 316 may then be modified and / or overwritten. When a requested row read by the memory controller arrives at the buffer 316, the DMA controller and / or engine 312 may reference the original list of scatter requests to determine the specific non-addressable portions of the read row arriving at the buffer 316 to be modified and / or overwritten. The DMA controller and / or engine 312 may form packets from the modified row in the buffer 316 to be written back to the memory via the memory controller.

[0068] Figure 3DA flow chart illustrating a process 370 for a scatter operation according to one embodiment is shown. Process 370 may include operations 372, 374, 376, and 378. Operation 372 may include, for example, receiving one or more scatter requests from initiator 322. Operation 374 may include loading signals and / or states (e.g., representing data items) of one or more addressable rows in a memory such as LAM 308 into buffer 316. In one particular implementation, DMA controller and / or engine 312 may determine the addressable rows to be fetched from LAM 308 at operation 374 by processing one or more gather requests. Operation 376 may include writing values ​​and / or states to at least one non-addressable portion of the row loaded into buffer 316 at operation 374 to at least partially modify the one or more addressable rows stored in buffer 316. For example, operation 374 may write values ​​and / or states to two or more non-addressable portions based on the multiple scatter requests received at block 372. Operation 378 may complete servicing of such a scatter request by initiating a write operation to write at least one of the one or more addressable rows modified at operation 374 back to memory (e.g., LAM 308). For example, operation 378 may initiate an operation to write back the row address in the LAM 308 addressable row loaded at operation 374 and modified at operation 376.

[0069] In one particular implementation, operation 374 may be initiated by a scatter request received at operation 372. Here, the one or more addressable rows of values ​​and / or states loaded into buffer 316 at operation 374 may be obtained by servicing one or more row read requests by a memory controller (e.g., memory controller 106). To write the modified addressable rows, operation 378 may include initiating the memory controller to perform one or more operations to write the one or more modified addressable rows. Here, process 370 may enable servicing multiple scatter requests by accessing a single addressable row loaded into buffer 316. For example, the multiple scatter requests received at operation 372 may specify words, bytes, fields, etc. within the same single addressable row. This may eliminate the need for the DMA controller and / or engine 312 to access the same addressable row multiple times for separate scatter requests for data items in the same addressable row (e.g., in LAM 308). Process 370 may also include converting the multiple scatter requests received at operation 372 into a list of row read requests (a list of row read requests for one or more addressable rows specifying values ​​and / or states) to be issued to the memory controller. This conversion of the multiple scatter requests into a list of row read requests may also include constructing at least one single row read request for an addressable row of memory containing the data items requested by the at least two scatter requests received at operation 372.

[0070] To service one or more scatter requests, process 370 may update less than the entirety of an addressable row in memory by loading the addressable row into buffer 316 and updating some portions of the loaded addressable row while leaving other portions unchanged. While some scatter requests may require an update of less than the entirety of the addressable row, one or more received scatter requests may require an update of the entirety of the addressable row, and the DMA will only write to that addressable row rather than performing a read-modify-write operation. For scatter requests that require an update of less than the entirety of the addressable row, the DMA controller and / or engine 312 may perform process 370. The DMA controller and / or engine 312 may also be configured to service scatter requests that require an update of the entirety of the addressable row by bypassing loading the addressable row into buffer 316. Here, to accomplish such an update of the entirety of the addressable row, the DMA controller and / or engine 312 may initiate a write operation to the addressable row in LAM 308 without loading the addressable row into buffer 316.

[0071] According to one embodiment, the DMA controller and / or engine 312 may receive multiple scatter requests that collectively request updates to the same overlapping portion of an addressable row in the LAM 308. This may create conflicts regarding how the overlapping portion should be updated to service the multiple scatter requests. For example, such multiple scatter requests may be sorted based on creation time or receipt time. According to one embodiment, for example, conflicts regarding updating a portion of an addressable row by multiple scatter requests may be resolved based on the most recent scatter request created or received.

[0072] In certain implementations of processes 350 and 370, the non-addressable portion of a line stored in buffer 316 may be a byte, a collection of bytes or fields, etc. Although the specific actions in the scatter and gather operations described above are described as occurring in a particular order, certain actions may be performed concurrently, and certain actions may be performed concurrently or in a particular order as a matter of engineering choice. Additionally, physical optimizations, such as the number and type of processing cores to be used to implement the features of the DMA controller and / or engine 312 and the interface engine, the number and type of memory blocks to be used for the associated memory elements and buffer 316, the number of ports and addressability characteristics of the memory forming LAM 308, may be selected as a matter of engineering choice. For example, buffer 316 may or may not be byte addressable, and the DMA controller and / or engine 312 may include, for example, a scalar or vector engine.

[0073] Figure 4A and Figure 4D4 is a schematic diagram of a computing device 400 including a DMA controller and / or engine 412 that communicates with a buffer 416 and an initiator 422 to facilitate redirection of DMA transactions, according to one embodiment. In one particular implementation, the computing device 400 may include the computing device 200 ( Figure 2A and Figure 2B ) one or more characteristics. According to one embodiment, the DMA controller and / or engine 412 may obtain a list of redirected read requests in response to a signal from the initiator 422. According to one embodiment, a read request may include a message and / or signal specifying one or more target memory addresses to be accessed in a read transaction to service the read request. For example, the read request may specify one or more target addresses as word row addresses of a location in memory containing content to be retrieved in a memory read transaction to service the read request. As referred to herein, a "redirected read request" means a read request that has been transformed or altered such that the content at the original target memory address is modified and / or shifted to a different memory address. Here, the different memory address specifies a location in memory containing content to be retrieved in a read transaction to service the redirected read request. For example, the DMA controller and / or engine 412 interprets the address specified in such a redirected read request as a word aggregation request to store the aggregated words in the buffer 416. For example, the size of the associated words of the word aggregation request portion to be read first to service the redirected read request may depend at least in part on how the associated words are interpreted to form the target address for redirection. In one embodiment, the word to be read may include, for example, an address or an index into an array that can be converted to an address. In one embodiment, the DMA controller and / or engine 412 may then convert the aggregated word stored in the buffer 416 into an address and merge the address (i.e., the converted aggregated word) with the redirected read request to form an aggregate request. The DMA controller and / or engine 412 may then service the formed aggregate request to send the resulting data item to the destination, as described in the redirected read request obtained in response to the signal from the initiator 422.

[0074] Figure 4B4 is a flow diagram of a process 450 for facilitating redirection of DMA transactions according to one embodiment. Process 450 may include operations 452, 454, 456, and 458. Operation 454 may include performing a gather operation (e.g., a word gather operation) based at least in part on one or more redirected read requests received at operation 452. In this operation, for example, DMA controller and / or engine 412 may obtain or otherwise receive such redirected read requests in response to a signal from initiator 422. In a particular implementation, DMA controller and / or engine 412 may interpret the one or more redirected read requests as a request for a first of two gather operations to be performed. During the first gather operation performed at operation 454, the "gathered" data items may be loaded into buffer 416.

[0075] Operation 456 may include converting one or more aggregated data items stored in buffer 416 to one or more addresses to specify a subsequent aggregate operation. For example, operation 456 may include converting one or more words obtained by performing the first word aggregate operation (performed at operation 454) to one or more addresses (also referred to as one or more address values). In one particular implementation, operation 456 may include parsing the values ​​and / or states in buffer 416 (from the aggregate operation) to determine one or more memory addresses in LAM 408. Operation 456 may also include applying one or more arithmetic operations to the parsed values ​​and / or states to determine the one or more memory addresses in LAM 408. For example, operation 456 may apply one or more arithmetic operations to the parsed values ​​and / or states to be stored in the buffer, which will form a memory address to a memory location in LAM 408. Such a formed address to a memory location in LAM 408 may form a basis for a subsequent aggregate operation.

[0076] According to one embodiment, the arithmetic operation applied at operation 456 may be defined according to the following expression (1):

[0077] address = base + x × element_size, (1)

[0078] in:

[0079] address is the target address (e.g., for the gather operation determined at operation 456 or for the scatter operation at operation 476);

[0080] x is a value obtained from an aggregate operation (e.g., at operation 454 or 474); and

[0081] base and element_size are parameters provided in the redirected request (eg, the redirected read request received at operation 452 or the redirected write request received at operation 472).

[0082] Operation 458 may include performing a second (e.g., subsequent) gather operation to forward the data item located at the one or more determined addresses to a destination. Such a destination may be determined based at least in part on the one or more redirected read requests. In another specific implementation, operation 458 may perform two or more gather operations based on the one or more addresses obtained at operation 456. According to one embodiment, when performing the gather operation, operation 458 may interpret the one or more redirected read requests as two requests for a word gather operation.

[0083] According to one embodiment, a write request may include a message and / or signal specifying one or more target memory addresses to be accessed in a memory write transaction to service the write request. For example, a write request may specify one or more target addresses as the word row addresses of a location in memory to be written in a memory write transaction (to service the write request). As referred to herein, a "redirected write request" means a write request that has been transformed or altered so that the original target memory address is modified and / or transformed to a different target memory address. Here, the different target memory address specifies a location in memory to be written in a write transaction to service the redirected write request.

[0084] In another specific embodiment, the DMA controller and / or engine 412 may obtain a list of redirected write requests in response to a signal from the initiator 422. The DMA controller and / or engine 412 may interpret the addresses specified in such redirected write requests as gather requests. For example, such gather requests may require that gathered words be loaded into the buffer 416. For example, the associated words to be read may have a size that depends at least in part on how the associated words are interpreted to form the target address for redirection. One or more such words loaded into the buffer 416 may be interpreted to form the target address for redirection. In one embodiment, such words loaded into the buffer 416 may include, for example, addresses or indices into an array that can be converted to addresses. In one embodiment, the DMA controller and / or engine 412 may then convert the gathered words stored in the buffer 416 into addresses and merge the addresses with the redirected write requests to form scatter requests. The DMA controller and / or engine 412 may then service the formed scatter requests, resulting in updates to the rows performed according to the redirected write requests obtained in response to the signal from the initiator 422.

[0085] Figure 4C4 is a flow chart of a process 470 for facilitating redirection of DMA transactions according to one embodiment. Operation 472 may include performing a word gather operation based at least in part on one or more redirected write requests received at operation 472. For example, the DMA controller and / or engine 412 may receive such redirected write requests at operation 472 in response to a signal from initiator 422. In a particular implementation, the DMA controller and / or engine 412 may interpret the one or more redirected write requests as a request for a word gather operation to be performed after the execution of the word scatter operation. During the word gather operation performed at operation 474, the "gathered" data items may be loaded into buffer 416. Operation 476 may include converting one or more gathered data items (from the gather operation) stored in buffer 416 to one or more addresses to specify a subsequent word scatter operation. In a particular implementation, operation 476 may include parsing the values ​​and / or states in buffer 416 to determine one or more memory addresses in LAM 408. For example, operation 476 may be performed at least in part by circuitry for parsing the values ​​and / or states in the buffer (e.g., a DMA controller and / or engine adapted to parse the values ​​and / or states) to determine at least one of the one or more addresses. For example, operation 476 may apply one or more arithmetic operations to the parsed values ​​and / or states to be stored in the buffer, which will form memory addresses in LAM 408. Operation 476 may also include applying one or more arithmetic operations to the parsed values ​​and / or states to determine the one or more memory addresses to the memory locations in LAM 408.

[0086] Operation 478 may include performing a scatter to write specific data items to one or more addresses obtained at operation 476 based at least in part on the one or more redirected read requests. According to one embodiment, operation 476 may apply an arithmetic operation to calculate a target address for the scatter operation to be performed at operation 478 according to expression (1). For example, the DMA controller and / or engine 412 may form such a scatter request based at least in part on the contents at the one or more addresses determined at operation 476. Such specific data items to be written from such a scatter request may be specified, for example, in one or more redirected write requests received at operation 472 in response to a signal from the initiator 422. In one particular implementation, operation 478 may include interpreting the contents at the addresses determined at operation 476 as addresses for the scatter operation.

[0087] One particular implementation of a computing device for processing a plurality of signal streams from an associated plurality of associated sources is provided by Figure 5A and Figure 5CIn one particular implementation, the computing device 500 may include the computing device 200 ( Figure 2A and Figure 2B ) one or more features. The computing device 500 may include a pool of compute nodes (CNs) 520 with associated register files and local memories, a pool of block memories and FIFO buffers, and a NOC that interconnects the various components. As noted above, the CNs in the CN pool 520 may include scalar CNs and / or processing circuit cores to implement an arithmetic logic unit (ALU), a digital signal processor (DSP), a vector CN, a VLIW engine, or a field programmable gate array (FPGA) unit, or a combination thereof, to provide just a few examples of specific circuit cores that may be used to implement the CNs in the CN pool 520. In one embodiment, the CN may include a single processing circuit core that is capable of performing operations to map input operands to output computation results. In another embodiment, the CN may include multiple distinct processing cores that are used to perform operations to map input operands to output computation results. The CNs in the CN pool 520 may include dedicated local memories (e.g., SRAM) and general registers to receive operands for operations to be performed and / or provide results from the execution of the operations. In one particular implementation, the CNs in CN pool 520 may be formed from one or more processing circuit cores integrated / configured with elements of memory pool 508 (e.g., transfer memory 218), such as FIFO buffers. In one embodiment, memory pool 508 may provide memory resources to be shared between the CNs in CN pool 520 and may facilitate communication between and / or among the CNs in CN pool 520. For example, endpoints of the FIFO buffers integrated / configured with CN pool 520 may include memory blocks local to or isolated from the CNs in CN 520 and / or their register files. The CNs in CN pool 520 may function as a host processor, a process manager, or a signal processor, or a combination thereof, to provide a few examples. According to one embodiment, the CNs in CN pool 520 may support atomic operations (atoms), interrupts, and other inter-processor communication (IPC) features to facilitate signal communication between and / or among the CNs in CN pool 520.

[0088] According to one embodiment, the computing device 500 may receive a signal stream containing data items (e.g., sensor signals, observations and / or measurements, timestamps, metadata, etc.) fed from an external source such as a sensor and / or memory. The external memory (not shown) may be coupled to a DMA controller and / or engine (e.g., DMA controller and / or engine 312 and / or 412). The CNs in the CN pool 520 may also provide a source of signal streams containing data items. For example, the CNs in the CN pool 520 may feed data items in the signal stream to be loaded into a FIFO buffer. The CNs in the CN pool 520 may also feed data items in the signal stream by transmitting groups of output parameters as blocks of data items within the IC devices in the NOC. According to one embodiment, a single logical signal stream may be sent in a multicast manner to multiple CNs in the CN pool 520. In addition, a single logical signal stream may be segmented into sub-signal streams to be fed to a subset of the CNs in the CN pool 520. In another embodiment, some CNs in the CN pool 520 may process data items from various signal streams to provide output signal streams as input streams to other CNs in the CN pool 520. The final output of the CNs in the CN pool 520 may include a signal stream output from the computing device 500 to a sink such as an actuator, memory, storage device, or display device, to provide just a few examples. Processing between and / or among the CNs in the CN pool 520 may be controlled and / or orchestrated via a combination of interrupts, polling of status flags, and periodic detection of work, to provide just a few examples.

[0089] Figure 5BFIG5 is a flow chart of a method 500 (also referred to as process 500) for processing multiple signal streams from multiple associated sources, according to one embodiment. According to one embodiment, process 500 may combine data items received in signal streams from multiple sources to update the state of a particle filter (e.g., to support automated operation of a motor vehicle). Such multiple sources may include, for example, multiple sensor devices (e.g., deployed in a motor vehicle), such as cameras, speedometers, active sensing devices (e.g., radar / lidar), environmental sensors (e.g., thermometers, light sensors, altimeters, etc.), and microphones, to provide just a few examples. Such signal streams from multiple sources may be received in signal packets from a communication network (e.g., signal packets including sensor measurements and / or observations obtained from a remote vehicle). In one specific implementation, the data items received in the signal streams from the multiple sources may include sensor measurements and / or observations having common attributes that can be identified by process 500. In one specific implementation, such common attributes may be identified by metadata co-located with the measurements and / or observations received from the multiple sources. For example, such common attributes may be associated by time (e.g., a timestamp indicating the time of the measurement and / or observation), space (e.g., a location with reference to an origin such as a point on a motor vehicle), and source (e.g., which specific physical sensor on the motor vehicle or which specific type of sensor among sensors mounted on the motor vehicle).

[0090] Operation 552 may include associating the data items at least in part based on the common attributes of the data items received from a plurality of data streams. Such multiple data streams may be provided as the output of the read / aggregation operation of CN and / or DMA aggregation or word aggregation or redirection. For example, operation 552 may sort and / or correlate (e.g., "bucketing") the measurements and / or observations received from different signal streams by time (e.g., according to timestamp) and / or space (e.g., the location of the object being observed relative to a reference point). In a specific implementation, operation 552 may associate the measurement and / or observations received from different sources (e.g., different sensors) at roughly the same time with the location of the particle limited in the current state of the particle filter. In a specific embodiment, operation 552 may associate the measurement and / or observations from different sources according to the location of the specific object observed and / or measured by such associated observations and / or measurements. In another particular implementation, data items from each of the plurality of signal streams may be loaded into a buffer (e.g., buffer 216) associated with a direct memory access (DMA) controller associated with the signal stream. Operation 552 may then identify at least one common attribute associated with the data items loaded into the buffer based at least in part on the contents of the data items loaded into the buffer.

[0091] Operation 554 may include simultaneously loading the data items associated at operation 552 into one or more registers of the compute node (e.g., general purpose registers of an ALU or other processing core forming the compute node) without storing the data items in row-accessible memory. According to one embodiment, the registers of the compute node may be loaded with data items to be retrieved on an execution cycle of the compute node. For example, the data items loaded into the registers of the compute node in one execution cycle (e.g., at the endpoints of a FIFO buffer) may provide operands for a computational operation to be performed by the compute node in the next execution cycle. In a specific implementation, the one or more registers of the compute node may include the endpoints of an associated FIFO buffer formed by internal memory (e.g., memory pool 508). Multiple such FIFO buffers having endpoints at registers of the compute node may be synchronized to apply data items from multiple sources (e.g., loaded data items from different sensors) as operands for compute nodes having common properties (e.g., temporal and spatial properties). Operation 556 may include execution of a compute node to process the data items simultaneously loaded at operation 554 to process the loaded data items as operands for one or more compute operations (e.g., to perform one or more functions, such as, for example, updating the state of a particle filter). The data items output from execution of the one or more compute operations at operation 556 may form data items for additional signal streams to be processed by additional compute nodes and / or for storage in memory.

[0092] In one embodiment, operation 552 may be performed by a first computing node to sort sensor observations and / or measurements received in a plurality of signal streams based at least in part on associated timestamps and locations of objects observed and / or measured by the sensor observations and / or measurements. The sorted sensor observations and / or measurements may then be loaded into one or more registers of a second computing node at block 554. Execution of the second node may then combine the sorted sensor observations and / or measurements at block 556.

[0093] In another embodiment, the data item of the additional signal stream as the output of operation 556 can be loaded into one or more registers of a subsequent compute node as an operand for one or more additional compute operations. For example, one or more direct memory access transactions can be performed to store the data item of the additional signal stream to external memory, a word scatter operation can be performed to write the data item of the additional signal stream, a redirected write operation can be performed to write the data item of the additional signal stream, or the data item of the additional signal stream can be provided as a control signal to one or more actuators, or a combination thereof.

[0094] In this context, "simultaneous loading" as mentioned herein means the loading of data items to be processed by a computing node in the same execution cycle. If data items from a synchronous signal stream are to be loaded simultaneously into a register of a computing node, such data items may be loaded into the register in the same execution cycle of the computing node (e.g., as operands of a computational operation performed in the next execution cycle). If data items from corresponding asynchronous signal streams are to be loaded simultaneously into a register of a computing node, such data items may be loaded into the register in different (e.g., adjacent) execution cycles of the computing node. For example, the execution of the computing node may be suspended for one or more execution cycles to allow multiple data items from different asynchronous signal streams to be loaded into the register as operands of a computational operation in the execution cycle of the computing node. In another embodiment, a transparent blocking and releasing mechanism may be applied to a computing node to facilitate simultaneous loading of data items from asynchronous signal streams into the register of the computing node.

[0095] According to one embodiment, the data items in the at least two signal streams associated at operation 552 may include the sensor observations and / or measurements associated at operation 554. The data items loaded simultaneously at operation 554 may then include associated sensor observations from multiple signal streams (e.g., from multiple different sources). In one particular implementation, operation 552 may include combining the sensor observations and / or measurements in the at least two signal streams to provide a combined sensor observation and / or measurement. Operation 552 may then sort and / or correlate (e.g., bucket) the combined sensor observations and / or measurements based, at least in part, on the associated timestamps and locations of the objects observed and / or measured by the sensor observations and / or measurements.

[0096] In one particular implementation, operation 554 may utilize, at least in part, process 350 ( Figure 3B ) and / or process 450( Figure 4B) is implemented. For example, the DMA controller and / or engine may perform the aggregation operations described in processes 350 and / or 450 to fill the queues maintained by the FIFO buffers (e.g., with associated measurements and / or observations) to load the registers of the compute node. Such execution by the DMA controller and / or engine may selectively load the data items into one or more registers based at least in part on an indication of a common attribute in the contents of the data items loaded into the FIFO buffers. In one specific implementation, multiple FIFO buffers may have endpoints at corresponding registers of the compute node, wherein the queues of different FIFO buffers are to be filled with data items from different signal streams (e.g., measurements and / or observations from different sensors). The queues of the multiple FIFO buffers may be filled so that data items from different signal streams that are associated by specific attributes (e.g., temporal and spatial attributes) are simultaneously loaded into the endpoints of the FIFO buffers (e.g., at the general registers of the compute node). According to one embodiment, operations 552 and 554 may be performed by a first compute node to simultaneously load the associated data items into one or more registers of the second compute node during one execution cycle of the second compute node. In a subsequent execution cycle of the second compute node, the second compute node can use the associated data items loaded simultaneously at block 554 as operands to perform a computational operation. In one specific implementation, the simultaneous loading of the data items at block 554 can be facilitated by a FIFO buffer having endpoints at registers of the second compute node. Here, the first compute node can fill the queue of the FIFO buffer with associated data items from different signal streams so that the corresponding associated data items reach the endpoints of the FIFO buffer in the same or similar execution cycles. Such filling of the queue of the FIFO buffer with associated data items can eliminate any need for the first compute node to store the associated data items in a row-accessible memory (e.g., row-accessible memory 208 or RAM 108).

[0097] In another specific implementation, the results of the execution of the compute node at operation 556 may provide data items for one or more additional signal flows. In one implementation, such data items for the additional signal flows may be loaded into one or more registers of subsequent compute nodes and / or downstream compute nodes. In one example, the data items at the output registers of the compute nodes (e.g., loaded from the execution of the compute operation) may be loaded via DMA write transactions, word scatter DMA transactions, and / or redirected scatter DMA transactions (e.g., based on the Figure 4C470) is transferred to the row-accessible memory. For example, such a DMA transaction may store data items of an additional signal stream to the row-accessible memory. In another example, such data items provided at the output register of a compute node may be applied as data items of an input signal stream to one or more other compute nodes. In one particular implementation, the result from the execution of the first compute node at operation 556 may load the result into an output register defined as a first endpoint of an associated FIFO buffer (e.g., from the memory pool 508). The associated FIFO buffer may then have a second endpoint defined as an input register of a second compute node to receive the result determined by the execution of the first compute node as an operand.

[0098] According to one embodiment, operations 552 and 554 may be performed by a first CN in CN pool 520, while operation 556 may be performed by a second CN in CN pool 520. At operation 554, the first compute node may simultaneously load associated data items (e.g., data items containing sensor observations and / or measurements associated at least in part based on spatial and temporal attributes, and the simultaneously loaded data items include the associated sensor observations and / or measurements) into one or more registers of a second CN in CN pool 520. At operation 556, the second CN may then process the data items simultaneously loaded by the first CN as operands for one or more computational operations. In one specific implementation, the first FIFO buffer may define a first endpoint as a register of the first CN (e.g., an output register of the first CN). The second endpoint of the first FIFO buffer may define a first register of the one or more registers of the second CN (e.g., an input register of the second CN). As can be observed, at operation 554, the first FIFO may enable the data items to be simultaneously loaded into the one or more registers of the second CN without storing the associated data items in row-accessible memory, as noted above. In another specific implementation, the FIFO buffer may define a first endpoint as a register of a third CN in the CN pool 520, and may define a second endpoint as at least a second register of one or more registers of the second CN. Here, at operation 554, both the first CN and the third CN may simultaneously load associated data items into one or more registers of the second CN node without requiring the associated data items to be stored in the row-accessible memory.

[0099] According to one embodiment, all or a portion of computing devices 200, 300 (e.g., including features implementing processes 350 and / or 370, such as circuitry for forming a DMA controller), 400 (e.g., including features implementing processes 450 and / or 470, such as circuitry for forming a DMA controller), and / or 500 (e.g., including features implementing process 550) may be formed and / or expressed in transistors and / or lower metal interconnects (not shown) in whole or in part in a process (e.g., front-end-of-line and / or back-end-of-line process), such as a process for forming complementary metal oxide semiconductor (CMOS) circuits, by way of example only. However, it should be understood that this is merely an example of how circuitry may be formed in a device in a front-end process, and claimed subject matter is not limited in this respect.

[0100] It should be noted that the various circuits disclosed herein can be described using computer-aided design tools and expressed (or represented) in terms of their behavior, register transfers, logic components, transistors, layout geometry, and / or other characteristics as data and / or computer-readable instructions embodied in various computer-readable media (e.g., non-transitory storage media). The formats of files and other objects in which such circuit representations can be implemented (e.g., in circuit devices) include, but are not limited to, formats supporting behavioral languages ​​such as C, Verilog, and Very High Speed ​​Integrated Circuit Hardware Description Language (VHDL), formats supporting register-level description languages ​​such as Register Transfer Language (RTL), formats supporting geometric description languages ​​such as Graphic Design System II (GDSII), Graphic Design System III (GDSIII), Graphic Design System IV (GDSIV), California Institute of Technology Intermediate Format (CIF), Manufacturing Electron Beam Exposure System (MEBES), and any other suitable formats and languages. Storage media in which such formatted data and / or instructions can be embodied include, but are not limited to, various forms of non-volatile storage media (e.g., optical, magnetic, or semiconductor storage media) and carrier waves that can be used to transmit such formatted data and / or instructions via wireless, optical, or wired signaling media, or any combination thereof. Examples of transmission of such formatted data and / or instructions by carrier waves include, but are not limited to, transmission (upload, download, email, etc.) via the Internet and / or other computer networks via one or more electronic communication protocols (e.g., Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Simple Mail Transfer Protocol (SMTP), etc.).

[0101] If such a data and / or instruction-based representation of the aforementioned circuitry is received within a computer system via one or more machine-readable media, such a data and / or instruction-based representation of the aforementioned circuitry may be processed by a processing entity (e.g., one or more processors) within the computer system in conjunction with the execution of one or more other computer programs to generate a representation or image of the physical manifestation of such circuitry, such one or more other computer programs including, but not limited to, a netlist generation program, a place and route program, etc. Such a representation or image may then be used in device fabrication, for example, by enabling the generation of one or more masks used to form various components of the circuitry in a device fabrication process (e.g., a wafer fabrication process).

[0102] In the context of this patent application, the term "between" and / or similar terms are understood to include "among" (if appropriate for a particular purpose), and vice versa. Similarly, in the context of this patent application, the terms "compatible with," "conform to," and / or similar terms are understood to include substantial compatibility and / or substantial conformity, respectively.

[0103] Unless otherwise indicated, in the context of this patent application, the term "or" (when used in an associative list such as A, B, or C) is intended to mean "A, B, and C" (used herein in an inclusive sense), as well as "A, B, or C" (used herein in an exclusive sense). With this understanding, "and" is used in an inclusive sense and is intended to mean A, B, and C; while "and / or" may be used with caution to clearly indicate that all of the foregoing meanings are intended, although such usage is not required. In addition, the terms "one or more" and / or similar terms are used to describe any feature, structure, characteristic, etc. in the singular, and "and / or" is also used to describe multiple features, structures, characteristics, and / or similar items and / or some other combination of features, structures, characteristics, and / or similar items. Likewise, the term "based on" and / or similar terms are understood to not necessarily be intended to convey an exhaustive list of factors, but rather to allow for the presence of additional factors that are not necessarily explicitly described.

[0104] Algorithmic descriptions and / or symbolic representations are examples of techniques used by those skilled in the art of signal processing and / or related arts to convey the substance of their work to others skilled in the art. In the context of this patent application, an algorithm is generally considered to be a self-consistent sequence of operations and / or similar signal processing leading to a desired result. In the context of this patent application, the operations and / or processing involve physical manipulations of physical quantities. Typically, although not necessarily, such quantities may take the form of electrical and / or magnetic signals and / or states capable of being stored, transferred, combined, compared, processed, and / or otherwise manipulated, such as electronic signals and / or states that constitute components of various forms of digital content (such as signal measurements, text, images, video, audio, etc.).

[0105] Primarily for reasons of commonality, it has proven sometimes convenient to refer to such physical signals and / or physical states as bits, values, elements, parameters, symbols, characters, terms, samples, observations, weights, numbers, numerical values, measurements, contents, and the like. However, it will be understood that all of these and / or similar terms are associated with appropriate physical quantities and are merely convenient labels. Unless otherwise specifically stated, as will be apparent from the foregoing discussion, it will be understood that throughout this specification, discussions utilizing terms such as "processing," "calculating," "calculating," "determining," "establishing," "obtaining," "identifying," "selecting," "generating," and the like may refer to actions and / or processes of specific devices such as special-purpose computers and / or similar special-purpose computing and / or network devices. Thus, in the context of this specification, special-purpose computers and / or similar special-purpose computing and / or network devices are capable of processing, manipulating, and / or transforming signals and / or states, typically in the form of physical electronic and / or magnetic quantities, within the memory, registers, and / or other storage devices, processing devices, and / or display devices of the special-purpose computers and / or similar special-purpose computing and / or network devices. In the context of this particular patent application, as mentioned, the term "specific apparatus" therefore includes a general-purpose computing and / or networking device (once it is programmed to perform particular functions (e.g., according to program software instructions)), such as a general-purpose computer.

[0106] In some cases, the operation of a memory device (such as a change of state from binary one to binary zero or a change of state from binary zero to binary one) may, for example, include a transformation, such as a physical transformation. For certain types of memory devices, such a physical transformation may include a physical transformation of an article to a different state or thing. For example, but not limitation, for some types of memory devices, the change of state may involve the accumulation and / or storage of charge or the release of stored charge. Similarly, in other memory devices, the change of state may include a physical change, such as a change in magnetic orientation. Similarly, the physical change may include a change in molecular structure, such as a change from a crystalline form to an amorphous form or from an amorphous form to a crystalline form. In yet other memory devices, the change in physical state may involve quantum mechanical phenomena, such as superposition, entanglement, and / or the like, for example, which may involve quantum bits (qubits). The foregoing is not intended to be an exhaustive list of all examples in which a change of state from binary one to binary zero or a change of state from binary zero to binary one in a memory device may include a transformation, such as a physical but non-transient transformation. Instead, the foregoing is intended to serve as an illustrative example.

[0107] Figure 6 is an illustration of an example sensor signal aggregation pattern for one embodiment of a vehicle 2200 operable in one or more autonomous driving modes (e.g., a fully autonomous mode, a semi-autonomous mode, a driver assistance mode). As depicted, an automated vehicle such as vehicle 2200 may include multiple sensors to provide measurements and / or observations to be processed, such as according to processes 350, 370, 450, 470, and / or 550. Although in Figure 6 Specific patterns and / or specific numbers of sensors are depicted in FIG, but the scope of the subject matter is not limited in these respects. For example, a system or device such as vehicle 2200 may include any number of sensors in any of a wide range of arrangements and / or configurations. Additionally, although Figure 6 Depicted as a two-dimensional representation, sensors of a vehicle capable of operating in one or more automated modes may, for example, generate signals and / or signal groups representing conditions in three-dimensional space surrounding vehicle 2200. In particular implementations, one-dimensional and / or two-dimensional sensor measurements may be combined and / or otherwise processed to produce, for example, a three-dimensional representation of conditions surrounding vehicle 2200.

[0108] In a particular implementation, various sensors may be installed in vehicle 2200, for example, to capture observations and / or measurements of different portions of the environment surrounding and / or near the vehicle. In a particular implementation, vehicle 2200 may include a plurality of different sensors capable of detecting incoming signals such as, for example, optical signals, electromagnetic signals, and / or acoustic signals. Each sensor may have a different field of view of the environment surrounding vehicle 2200. While example fields of view 2210a through 2210h are depicted, the subject matter is, of course, not limited in scope in these respects.

[0109] In a particular implementation, the sensor signals and / or signal packets may be utilized by at least one processor of the vehicle 2200, for example, to identify objects and / or other environmental conditions near the vehicle 2200, which may be utilized by the processing system of the vehicle 2200 to autonomously guide the vehicle through the environment, for example. Example objects that may be detected in the environment around a vehicle such as the vehicle 2200 may include other vehicles, trucks, cyclists, pedestrians, animals, rocks, trees, lampposts, guardrails, painted lines, signal lights, buildings, road signs, etc. Some objects may be stationary, and other objects, such as pedestrians, may move through the environment.

[0110] In one specific implementation, one or more sensors of the example vehicle 2200 may generate signals and / or signal groups that may represent at least a portion of the environment surrounding and / or near the vehicle 2200. Other sensors may provide signals and / or signal groups that represent the speed, acceleration, orientation, position (e.g., via a global navigation satellite system (GNSS)), etc. of the vehicle 2200. As described more fully below, the sensor signals and / or signal groups may be processed, such as via a particle filter, to generate a plurality of particles. Such particles may be utilized, at least in part, to influence the operation of the vehicle 2200. In a specific implementation, as the vehicle 2200, for example, moves through the environment, the sensor signals and / or signal states may be utilized, at least in part, to update the particle filter, e.g., to further influence the operation of the vehicle. As discussed more fully below, for particle filters and the like that utilize a relatively wide range of signals and / or signal groups generated by a variety of sensors, any of a plurality of sorting operations may be performed on the sensor signals and / or signal groups.

[0111] In this context, a "particle" refers to a digital representation of an environmental condition at a particular point in a particular coordinate system at a particular point in time that is derived at least in part from a sensor signal and / or signal grouping. For example, a particular particle may include an array of parameters describing a particular point within the environment surrounding vehicle 2200 at a particular point in time. In a specific implementation, a "particle filter" or the like may be utilized to process sensor signals and / or signal groups to generate a plurality of particles describing an environment, such as the environment surrounding vehicle 2200. Of course, a particle filter is merely an example type of processing that may be performed on sensor signals and / or signal groups, and the scope of the subject matter is not limited in this respect.

[0112] In certain implementations, a particular coordinate system may be specified, although the scope of the subject matter is not limited to any particular coordinate system. In certain implementations, such a coordinate system may include a three-dimensional parameter space, although other implementations may specify other numbers of dimensions. In certain implementations, each particle may belong to a particular location within a particular three-dimensional space (e.g., the X, Y, and Z axes).

[0113] Figure 7 An example schematic block diagram of an example vehicle 2200 is depicted. As mentioned, in a specific implementation, vehicle 2200 can include a plurality of sensors 2210. Such sensors can include, for example, image capture (e.g., cameras), radar, lidar, and / or ultrasonic sensors, to name a few non-limiting examples. In a specific implementation, sensors 2210 can generate sensor signals and / or signal packets 2215 that can be provided to and / or otherwise obtained by a control system such as control system 2220.

[0114] In specific implementations, the control system 2220 may include, for example, at least one processor, at least one memory device, and / or at least one communication interface. In specific implementations, the control system 2220 may include, for example, one or more central processing units (CPUs), neural network processors (NNPs), and / or graphics processing units (GPUs). In specific implementations, the control system 2220 may process sensor signals and / or signal packets to generate signals and / or signal packets that may affect the operation of the vehicle 2200. For example, the signals and / or signal packets may be generated by the control system 2220 and provided to and / or otherwise obtained by the drive system 2230. In specific implementations, the processing of the sensor signals and / or signal packets by the control system 2220 may include, for example, the use of a particle filter, although other implementations may utilize other signal processing algorithms, techniques, methods, etc., and the scope of the subject matter is not limited in this respect. In specific implementations, the drive system 2230 may include, for example, devices, mechanisms, systems, etc. for affecting the operation of the vehicle 2200. As mentioned, as the vehicle 2200 traverses traffic, additional sensor signals and / or signal packets may be acquired and processed so that the operation of the vehicle 2200 may be updated over time.

[0115] In the foregoing description, various aspects of the claimed subject matter have been described. For purposes of explanation, exemplary details, such as quantities, systems, and / or configurations, have been set forth. In other cases, well-known features have been omitted and / or simplified so as not to obscure the claimed subject matter. Although certain features have been illustrated and / or described herein, numerous modifications, substitutions, variations, and / or equivalents will now occur to those skilled in the art. Therefore, it should be understood that the appended claims are intended to cover all modifications and / or variations that fall within the claimed subject matter.

Claims

1. A system, comprising: memory, said memory comprising one or more memory devices; and a direct memory access (DMA) controller coupled to the memory by a bus, the DMA controller being configured to: receiving one or more redirected write requests; performing a gather operation based at least in part on the received one or more redirected write requests to obtain one or more words from the memory; converting the one or more words obtained by the gather operation into one or more addresses in the memory; as well as A scatter operation is performed to write data items to the one or more addresses in the memory.

2. The system of claim 1 , wherein the memory is a first memory, wherein the system further comprises a second memory operable as / used as a buffer, wherein the buffer is configured to store values ​​and / or status retrieved from the first memory, wherein the DMA controller is further configured to parse the values ​​and / or status in the buffer to determine at least one of the one or more addresses in the memory. 3 . The system of claim 2 , wherein the DMA controller is further configured to apply one or more arithmetic operations to at least one of the resolved values ​​and / or states to determine the at least one address. 4 . The system of claim 1 , wherein the DMA controller is further configured to interpret the one or more redirected write requests as a request for the gather operation. 5 . The system of claim 1 , wherein the DMA controller is further configured to form a request for the scatter operation based at least in part on the one or more addresses in the memory. 6 . The system of claim 1 , further comprising an initiator configured to initiate the one or more redirected write requests, wherein the DMA controller is configured to perform a first gather operation in response to a signal from the initiator.

7. The system of claim 6, wherein the initiator of the one or more redirected write requests comprises a local memory or one or more registers of a computing node, wherein the local memory or one or more registers of the computing node are configured to receive sensor measurements and / or observations as data items in a signal stream.

8. A method at a direct memory access (DMA) controller, the method comprising: performing a first aggregation operation based at least in part on the one or more redirected read requests; converting one or more words obtained by performing the first gather operation to one or more addresses; as well as A second gather operation is performed to forward the data items located at the one or more addresses to a destination determined based at least in part on the one or more redirected read requests.

9. The method of claim 8, wherein translating the one or more words obtained by performing the first gather operation to the one or more addresses further comprises: The values ​​and / or states in the buffer are parsed to determine the memory address.

10. The method according to claim 9, further comprising: One or more arithmetic operations are applied to the resolved value and / or state to determine the memory address.

11. The method according to claim 8, further comprising: The one or more redirected read requests are interpreted as a first request for an aggregate operation of two requests.

12. The method according to claim 8, further comprising: A request for a second of the two requests for an aggregate operation is formed based at least in part on the one or more addresses.

13. The method according to claim 8, further comprising: The first aggregate operation is performed in response to a signal from an initiator of the one or more redirected read requests.

14. The method of claim 13, wherein the initiator of the one or more redirected read requests comprises a process executed by an ALU to perform one or more sensor fusion operations.

15. A system comprising: memory, said memory comprising one or more memory devices; and a direct memory access (DMA) controller coupled to the memory by a bus, the DMA controller being configured to: receiving one or more redirected read requests; performing a first gather operation based at least in part on the received one or more redirected read requests to obtain one or more words from the memory; converting the one or more words obtained by performing the first gather operation into one or more addresses; as well as A second gather operation is performed to forward the data items located at the one or more addresses to a destination determined based at least in part on the one or more redirected read requests.

16. The system of claim 15, wherein the DMA controller is further configured to parse values ​​and / or status in a buffer to determine the one or more addresses.

17. The system of claim 15, wherein the DMA controller is further configured to: parsing the values ​​and / or states in the buffer; and One or more arithmetic operations are applied to the resolved values ​​and / or states to determine the one or more memory addresses.

18. The system of claim 15, wherein the DMA controller is further configured to: The one or more redirected read requests are interpreted as a first request for an aggregate operation of two requests.

19. The system of claim 15, wherein the DMA controller is further configured to form a request for a second of two requests for a gather operation based at least in part on the one or more addresses.

20. The system of claim 15, wherein the DMA controller is further configured to: The first aggregate operation is performed in response to a signal from an initiator of the one or more redirected read requests.

Citation Information

Patent Citations

  • DMA controller

    JP2006323579A

  • Multidimensional address generation for direct memory access

    US20200371978A1

  • Logical address direct memory access with multiple concurrent physical ports and internal switching

    US7877524B1