Systems, devices, and / or methods for processing signal streams

The system addresses inefficiencies in processing multiple signal streams by using computational nodes and DMA controllers to associate and process data items, achieving reduced latency and resource usage for enhanced data fusion in self-driving applications.

JP2026505156APending Publication Date: 2026-02-12MERCEDES BENZ GROUP AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025539630
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-04
Filing Date
2023-12-19
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing systems struggle to efficiently process and fuse multiple signal streams from diverse sensors in automotive and robotic applications, particularly in self-driving and autonomous driving, due to limitations in computational resources and latency in data manipulation.

Method used

A system comprising multiple computational nodes that associate and process data items from multiple signal streams based on common attributes, utilizing direct memory access (DMA) controllers to load and process data without external memory storage, and employing a Confluence Engine (CE) for reduced latency and resource usage.

Benefits of technology

Enhances the processing of sensor observations and measurements by reducing latency and computational resources, enabling efficient fusion of data streams for applications like self-driving vehicles with improved power efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026505156000001_ABST
    Figure 2026505156000001_ABST
Patent Text Reader

Abstract

Exemplary methods, apparatus, and / or articles of manufacture that may be implemented in whole or in part in connection with processing signal streams are disclosed. In one application, multiple computational nodes may be configured to process data items transported in signal streams from multiple sources.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The subject matter disclosed herein relates to processing signals received in streams from multiple data sources. [Background technology]

[0002] Self-driving and / or autonomous driving applications, as well as other automotive and robotic applications, may rely on the fusion of signals, measurements, and / or observations generated by multiple sensors. Processing for such applications may involve the manipulation of arrays of data elements. Such applications may be performed and / or implemented by commercially available central processing units (CPUs) and / or graphics processing units (GPUs). Such commercially available processing units may be configured to manipulate elements of input arrays to generate elements of output arrays. Summary of the Invention

[0003] One embodiment disclosed herein is directed to a system including a plurality of computational nodes, each computational node of the plurality of computational nodes being a respective circuit adapted to perform a computational operation on data, the plurality of computational nodes including at least a first computational node and a second computational node. A first computational node of the plurality of computational nodes may be configured to receive a plurality of signal streams from a plurality of sources, each signal stream including a respective set of data items, identify data items from the respective sets of data items having a common attribute, and simultaneously load the data items received from one or more of the plurality of signal streams into one or more registers of a second computational node. The simultaneously loaded data items may be associated based on the attribute common to the simultaneously loaded data items. The second computational node may be configured to process the simultaneously loaded data items as operands of one or more computational operations.

[0004] In one particular implementation, the system further comprises a transport memory for transporting data and / or values ​​in the form of a first transport memory buffer, the first transport memory buffer having a first endpoint and a second endpoint, the plurality of computing nodes are configured to communicate with an external memory external to the system via a bus, the first computing node comprises a register forming the first endpoint of the first transport memory buffer, a first register of the one or more registers of the second computing node forms the second endpoint of the first transport memory buffer, and the first computing node is configured to load data items into one or more registers of the second computing node without storing the data items in the external memory. The system may also include a second transport memory buffer, the second transport memory buffer having a first endpoint and a second endpoint, the first endpoint of the second transport memory buffer formed by a register of a third computing node of the plurality of computing nodes, the second endpoint of the second transport memory buffer formed by a second register of the one or more registers of the second computing node, and the first computing node and the third computing node configured to simultaneously load data items into the one or more registers of the second computing node without storing the data items in external memory.

[0005] In another particular implementation, the data items in at least two of the plurality of signal streams include sensor observations and / or measurements, the first computing node is configured to associate the sensor observations and / or measurements based at least in part on spatial and temporal attributes, and the simultaneously loaded data items include the associated sensor observations and / or measurements. In one example, the first computing node is further configured to sort the sensor observations and / or measurements based at least in part on associated timestamps and locations of objects observed and / or measured by the sensor observations and / or measurements, and the second computing node is further configured to combine the sorted sensor observations and / or measurements to provide combined sensor observations and / or measurements.

[0006] Another embodiment disclosed herein is directed to a method including associating data items of multiple signal streams from multiple sources based at least in part on a common attribute of the data items; simultaneously loading the associated data items from the multiple signal streams into one or more registers of a computational node; and executing the computational node to process the simultaneously loaded associated data items as operands of one or more computational operations. In one particular implementation, the method further includes providing results of processing the simultaneously loaded associated data items as operands to one or more computational instructions as data items of additional signal streams. In one example, the method may further include loading the data items of the additional signal streams into one or more registers of a subsequent computational node as operands for one or more additional computational operations. In another example, the method further includes performing one or more direct memory access transactions to store the data items of the additional signal streams in external memory, performing a word scatter operation, performing a redirected write operation, or providing control signals to one or more actuators, or a combination thereof.

[0007] In another particular implementation, associating the data items of the multiple signal streams further includes executing a direct memory access (DMA) controller to load a data item from each of the multiple signal streams into a buffer associated with the signal stream, and identifying at least one common attribute between the loaded data item and at least one other data item based at least in part on the content of the data item loaded into the buffer. In one example, executing the DMA controller may further include loading one or more addressable lines of values ​​and / or states stored in memory into the buffer, analyzing one or more non-addressable portions of at least one of the loaded one or more addressable lines of values ​​and / or states, and processing one or more collection requests based at least in part on the analyzed one or more non-addressable portions. In another example, executing the DMA controller may further include performing a first word gather operation based at least in part on the one or more redirected read requests, converting one or more words obtained from performing the first word gather operation into one or more addresses, and performing a second word gather operation to transfer data items located at the one or more addresses to a destination determined at least in part based on the one or more redirected read requests.

[0008] In another particular implementation, simultaneously loading data items associated with one or more registers of the compute node from the multiple signal streams includes loading the data items of the multiple signal streams into buffers associated with the multiple signal streams, and executing a direct memory access (DMA) controller associated with the multiple signal streams to selectively load the data items into the one or more registers based at least in part on an indication of a common attribute in the content of the data items loaded into the buffer. In yet another particular implementation, the multiple sources include at least a first sensor integrated with the vehicle and a second sensor external to the vehicle.

[0009] In yet another particular implementation, the data items in at least two of the multiple signal streams include sensor observations and / or measurements, the sensor observations and / or measurements in at least two of the multiple signal streams are associated based at least in part on spatial and temporal attributes, and the simultaneously loaded associated data items include the sensor observations and / or measurements associated based at least in part on the spatial and temporal attributes. In one example, associating the data items of the multiple signal streams from the multiple sources includes sorting, at a previous computational node, the sensor observations and / or measurements based at least in part on associated timestamps and locations of objects observed and / or measured by the sensor observations and / or measurements, and combining, at the computational node, the sorted sensor observations and / or measurements in the at least two or more signal streams to provide combined sensor observations and / or measurements. In another example, the method further includes executing the computational node to update a state of a particle filter based at least in part on the simultaneously loaded associated data items.

[0010] Another embodiment disclosed herein is directed to a system comprising: a plurality of sensors for generating a plurality of associated signal streams; and a plurality of computational nodes coupled to the plurality of sensors, wherein at least a first computational node of the plurality of computational nodes is configurable to concurrently load data items from two or more of the associated plurality of signal streams originating from two or more of the sensors into one or more registers of a second computational node, wherein the concurrently loaded data items are associated based at least in part on attributes common to the concurrently loaded data items, and the second computational node is configured to process the concurrently loaded data items as operands of one or more computational operations. In one particular implementation, the data items in at least two of the associated plurality of signal streams include sensor observations and / or measurements, and the first computational node is configured to associate the sensor observations and / or measurements based at least in part on spatial and temporal attributes, and the concurrently loaded data items include the associated sensor observations and / or measurements. In another particular implementation, the second computing node is further configured to update a state of the particle filter based at least in part on the concurrently loaded data items.

[0011] The claimed subject matter is particularly pointed out and distinctly claimed in the concluding portion of this specification, however, both as to organization and / or method of operation, together with its objects, features, and / or advantages, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a schematic diagram of a computing device according to an embodiment. [Figure 2A] FIG. 1 is a schematic diagram of a computing device including a direct memory access (DMA) controller and / or engine including a buffer according to one embodiment. [Figure 2B]FIG. 1 is a schematic diagram of a computing device including a direct memory access (DMA) controller and / or engine including a buffer according to one embodiment. [Figure 3A] FIG. 1 is a schematic diagram of a computing device including a DMA controller and / or engine including buffers for facilitating scatter-gather operations according to one embodiment. [Figure 3B] FIG. 1 is a flow diagram of a process for facilitating scatter and gather operations according to one embodiment. [Figure 3C] FIG. 1 is a schematic diagram illustrating non-addressable portions of addressable lines according to one embodiment. [Figure 3D] FIG. 1 is a flow diagram of a process for facilitating scatter and gather operations according to one embodiment. [Figure 3E] FIG. 1 is a schematic diagram of a computing device including a DMA controller and / or engine including buffers for facilitating scatter-gather operations according to one embodiment. [Figure 4A] FIG. 1 is a schematic diagram of a computing device including a DMA controller including a buffer for facilitating redirection of DMA transactions according to one embodiment. [Figure 4B] FIG. 1 is a flow diagram of a process for facilitating redirection of DMA transactions according to one embodiment. [Figure 4C] FIG. 1 is a flow diagram of a process for facilitating redirection of DMA transactions according to one embodiment. [Figure 4D] FIG. 1 is a schematic diagram of a computing device including a DMA controller including a buffer for facilitating redirection of DMA transactions according to one embodiment. [Figure 5A] FIG. 1 is a schematic diagram of a computing device for facilitating processing of multiple signal streams from multiple associated sources according to one embodiment. [Figure 5B] FIG. 1 is a flow diagram of a method for processing multiple data streams from associated multiple sources according to one embodiment. [Figure 5C] FIG. 1 is a schematic diagram of a computing device for facilitating processing of multiple signal streams from multiple associated sources according to one embodiment. [Figure 6] FIG. 2 illustrates an exemplary sensor signal collection for an exemplary vehicle according to one embodiment. [Figure 7] FIG. 1 illustrates an exemplary schematic block diagram of exemplary vehicle features in a self-driving / autonomous driving application according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] In the following detailed description, references are made to the accompanying drawings, which form a part hereof, in which like numerals may indicate corresponding and / or similar like parts throughout. It is understood that the figures have not necessarily been drawn to scale, such as for simplicity and / or clarity of illustration. For example, the dimensions of some aspects may be exaggerated relative to other aspects. It is further to be understood that other embodiments may be utilized. Furthermore, structural and / or other changes may be made without departing from the claimed subject matter. References throughout this specification to "claimed subject matter" refer to subject matter intended to be covered by one or more claims, or any portion thereof, and are not necessarily intended to refer to a complete set of claims, a specific combination of a set of claims (e.g., method claims, apparatus claims, etc.), or a particular claim. It should also be noted that directions and / or references, such as up, down, top, bottom, etc., may be used to facilitate description of the drawings and are not intended to limit the application of the claimed subject matter. Therefore, the following detailed description should not be construed as limiting the claimed subject matter and / or equivalents.

[0014] Throughout this specification, references to an implementation, a certain implementation, an embodiment, a certain embodiment, etc., mean that a particular feature, structure, characteristic, etc. described with respect to a particular implementation and / or embodiment is included in at least one implementation and / or embodiment of the claimed subject matter. Thus, for example, the appearance of such phrases in various places throughout this specification is not necessarily intended to refer to the same implementation and / or embodiment, or to any one particular implementation and / or embodiment. Furthermore, it is to be understood that particular features, structures, characteristics, and / or the like that are described can be combined in various ways in one or more implementations and / or embodiments and thus fall within the scope of the intended claims. Generally, of course, as always with the specification of a patent application, these and other issues have the potential to vary in the particular context of use. In other words, throughout this disclosure, the particular context of description and / or use provides useful guidance regarding reasonable inferences to be drawn. Similarly, however, "in this context" generally refers, without further limitation, to at least the context of this patent application.

[0015] According to one embodiment, multiple signal streams may be processed using multiple computational nodes, with each computational node configured to perform a specific operation on data items. In one implementation, a first computational node of the multiple computational nodes may be configured to receive multiple signal streams from multiple sources (e.g., sensors) and identify among respective sets of data items in the different received signal streams that have a common attribute. The first computational node may then simultaneously load data items received from two or more of the multiple signal streams that have the common identified attribute into one or more registers of a second computational node. The second computational node may be configured to process the loaded data items as operands of one or more computational operations.

[0016] In some examples, the host central processing unit (CPU) may be part of a computing device within a vehicle and may process data items in memory for automotive applications. As an example, automotive applications such as self-driving / autonomous driving applications (e.g., fully autonomous, semi-autonomous, driver assistance systems, etc.) may employ, for example, particle filters to fuse signal streams of sensor signals and / or observations to, for example, update particle filter states. Such applications may be implemented in systems such as, for example, automated machines (cars, trucks, etc.). In this context, a "signal stream" as referred to herein refers to a time-varying progression of a sequence of coded data items delivered to a receiving device via a signal transmission medium. The coded data items (also referred to as data values) delivered in a signal stream may represent attributes indicating conditions and / or events, object identifiers, timestamps indicating the time of an event, and metadata, to name a few attributes that may be represented in coded data items delivered in a signal stream. In certain implementations, a signal stream may deliver sensor measurements and / or observations in combination with associated timestamps to indicate the time at which such measurements and / or observations were obtained.

[0017] In one aspect of an embodiment, information originating from different sources and arriving in respective signal streams may be processed as a “confluence” of information to determine a computational result. Such a confluence of information may be processed by associating and / or correlating information items from the different sources by particular attributes (e.g., time, space, reliability, trustworthiness, etc.). Processing the confluence of information may then include performing one or more operations on the items based on the one or more attributes to generate a result. In particular implementations, data items in the confluence of information may be processed by updating one or more states of a particle filter. For example, such a particle filter may implement processing of a confluence of arrays of sensor signals / observations (e.g., received from signal streams) to update the states of measurement particles, filter particles, static particles, and / or dynamic particles. In one implementation, measurements and / or observations in the confluence of arrays of measurements may be generated from different sensors. Nevertheless, such measurements and / or observations generated by different sensors may be associated and / or correlated by time and space. According to one embodiment, associating data items from different sources in a union of arrays may be implemented at least in part using a radix sort applied to an array of keys. In particular implementations, sensor signals and / or observations may be formatted into arrays that are processed to generate the union of arrays. An exemplary procedure for generating such a union of arrays may be performed according to the following pseudocode: [Table 1] TIFF2026505156000003.tif155150

[0018] According to one embodiment, the merging of multiple signal streams may include mapping data items in different input signal streams to data items in one or more output signal streams. For example, data items in such input signal streams may include sensor measurements and / or observations of sensors associated with the input signal streams. Accordingly, the merging of such input signal streams may include mapping sensor measurements and / or observations (e.g., from different / separate sensors associated with the input signal streams) to data items of the output signal stream. Such data items of the output signal stream may include values ​​inferred / calculated based on the sensor measurements and / or observations. In the pseudocode example provided above, the number of input signal streams may be defined as S_in[] (which includes the data item value_in[]), and the number of output signal streams may be defined as S_out[] (which includes the data item value_out[]). Here, the expression "(value_out[],S_out_enable[])=f(value_in[],parameters)for(i:0..(l-1))" may map the data item value_in[] in the input signal stream S_in[] to the data item value_out[] in the output signal stream S_out[] according to the function f().

[0019] According to one embodiment, computing devices, circuits, and / or logic may form a "Confluence Engine" (CE), also referred to as a "Confluencer" and / or "Confluence Processor" (CP), to process the confluence of data items as described above. Such confluence of data items may include, for example, a confluence of signal streams and / or a sequence of confluences of signal streams with reduced latency and / or reduced computing resources (e.g., power, memory, etc.). In particular implementations, the output of a confluence operation of such a CE may provide all or part of the input to a subsequent confluence operation. Confluence properties, such as functions like k, f(), and read_next() as shown in the pseudocode example above, may be part of the runtime programming of the CE.

[0020] According to one embodiment, a CE may employ a direct memory access (DMA) subsystem, which may include, for example, a DMA controller (also referred to as DMA "engine(s)") configurable to initiate read and write operations between line-accessible memory and transport memory. The particular process by which the CE accesses memory may be determined parametrically at compile time and physically at runtime. Thus, in particular implementations, events that trigger the execution of the DMA controller may not be limited to events occurring in the arithmetic logic unit (ALU) (e.g., from loads and stores in the ALU). The DMA controller may be triggered to perform DMA transactions by the initiation of a join operation. The DMA controller may then perform such DMA transactions independently of the arithmetic logic unit (ALU) (e.g., dependent only on the availability of read and write accesses at valid endpoints for such DMA transactions).

[0021] 1 is a schematic diagram of a system 100 that performs DMA transactions. System 100 includes multiple components that communicate via bus 101. The components include a host CPU 102, a memory controller 106, RAM 108 (which may form main system memory), peripheral devices 114, and a DMA controller 112. In this context, "direct memory access," as referred to herein, refers to a process performed by one or more hardware subsystems and / or circuits to access a particular memory independently of a particular processing unit and / or central processing unit (e.g., independently of host CPU 102). According to one embodiment, DMA controller 112 may initiate an operation to access (e.g., read or write) random access memory (RAM) 108 via bus 101 independently of host CPU 102. For example, the DMA controller and / or engine 112 may control and / or execute transactions to transfer data items (also called data values) between the peripheral devices 114 and RAM 108 (through the memory controller 106) independent of any action by the host CPU 102. Such DMA transactions may be triggered, for example, by a signal, a condition, and / or an event (e.g., an interrupt signal).

[0022] As described above, processing of data items for automotive or other computing applications may be enhanced through a merging engine (CE). Figures 2A and 2B are schematic diagrams of a computing device 200 including a direct memory access (DMA) controller 212 (also referred to as a DMA engine 212) and a buffer 216 for implementing one or more aspects of a CE, according to one embodiment. Computing device 200 may be or form, for example, a system-on-chip (SoC), a microchip, a control circuit, or some other computing device. In some implementations, computing device 200 may form a vehicle controller, such as an Advanced Driver Assistance System (ADAS) device, a telematics control unit (TCU), an electronic control unit (ECU), a centralized vehicle computer, or some other vehicle controller. As shown in Figures 2A and 2B, the computing device 200 may further include a CPU 202 (also referred to as a host CPU), a transport memory 218, a line-accessible memory 208, a buffer 216, and a pool of compute nodes (CNs) 220 (also referred to as a CN pool 220).

[0023] In one implementation, a CN may comprise a single processing circuit core capable of performing operations for mapping input operands to output computation results. In another implementation, a CN may comprise multiple separate processing cores for performing operations for mapping input operands to output computation results. In another implementation, two or more non-concurrently executing CNs may be implemented on the same processing circuit core. For example, a processing circuit may implement a first CN to generate an output result (e.g., stored in transport memory 218) that becomes input to a second, subsequently executing CN implemented on the same processing circuit.

[0024] The CNs in pool of CNs 220 may include dedicated local memory (e.g., static random access memory (SRAM)) and general-purpose registers to receive operands for operations to be performed and / or to provide results from the execution of operations. In this example, host CPU 202 may use transport memory 218 to store data items and control the CNs in CN pool 220 to perform operations on these data items. Transport memory 218 may be physically closer to the CNs and / or operate with lower access latency and may therefore be used as a cache for storing data items. Line-accessible memory 208 may be external to transport memory 218 and may provide a larger amount of memory space for transport memory 218, but may be physically farther from CN pool 220 and operate with longer access latency. According to one embodiment, transport memory 218 may include one or more synchronization mechanisms to facilitate inter-CN communication between CNs in CN pool 220 (e.g., for synchronizing communication between CNs with different execution latencies). In particular implementations, line-accessible memory 208 may be separated from transport memory 218, CPU 202, and buffer 216 by a bus (not shown). According to one embodiment, buffer 216 may be formed in circuitry for implementing core circuitry of DMA controller and / or engine 212 such that buffer 216 is separate and isolated from circuitry for forming transport memory 218. Such formation of buffer 216 in core circuitry of DMA controller and / or engine 212 may reduce and / or minimize latency associated with loading data items into and storing data items from buffer 216 in the course of performing DMA operations.

[0025] In one embodiment, DMA controller and / or engine 212 may be configured to interface with transport memory 218 and line-accessible memory 208. Transport memory 218 and / or line-accessible memory 208 may be cache-line addressable. In other words, line-accessible memory 208 in this example may be cache-line addressable memory. Transport memory 218 and / or line-accessible memory 208 may provide data items (also referred to as data values) that CNs in CN pool 220 can manipulate, operate on, or otherwise process. According to one embodiment, all or a portion of transport memory 218 may be configured as a cache, which may be integrated with commercially available components. It should also be noted that caching either within transport memory 218 or buffer 216 may mitigate manufacturing defects and / or enable the use of an embodiment larger than the anticipated application size.

[0026] In one implementation, DMA controller and / or engine 212 may be configured to handle cache-line-sized data items (e.g., 64 bytes or 128 bytes), and even in scatter-gather operations, such data items are locatable and accessible by cache-line addresses. Such cache-line addresses for 64-byte cache lines may be represented in binary notation ending with six zeros. Similarly, cache-line addresses for 128-byte cache lines may be represented in binary notation ending with seven zeros. According to one embodiment, DMA controller and / or engine 212 may be configured to handle word-sized data items, even while line-accessible memory 208 may remain addressable only in cache lines. To facilitate scatter or gather operations to transfer word-sized data items to or from line-accessible memory 208, DMA controller and / or engine 212 may implement buffer 216 (e.g., located between line-accessible memory 208 and transport memory 218). Buffer 216 may be configurable to store together bytes from multiple cache lines that are placed into a destination word in transport memory 218. DMA controller and / or engine 212 may also be capable of performing multicast write operations in either direction (e.g., from transport memory 218 to line-accessible memory 208 or from line-accessible memory 208 to transport memory 218). Buffer 216 may be separate / different from transport memory 218. As described herein, buffer 216 may be formed within a client core to implement DMA controller and / or engine 212.

[0027] In this context, a "transport memory" as referred to herein means circuitry for facilitating communication of data items between CNs, such as CNs in pool of CNs 220. In one particular implementation, such a transport memory may transport results from the execution of a first operation in a first CN to become input operands of a second operation executed in a second CN (e.g., in a computing pipeline). In particular implementations, transport memory 218 may be configured as a static random access memory (SRAM) device used as access control memory, or a shared memory as cache memory, a word-addressable or SIMD vector-addressable register file, or a circuit and / or device specifically structured to function as a first-in, first-out (FIFO) buffer. Such a circuit and / or device specifically structured to function as a FIFO buffer may have a width that is word-wide (word width) or single instruction, multiple data (SIMD) vector-wide (vector width) or a circuit or network-on-chip (NOC) device coupled between endpoints, or a combination thereof, to name a few.

[0028] As described above, CNs in CN pool 220 (e.g., a pool of CNs) may operate on data items read from memory. In this context, a "computation node," as referred to herein, means a set of identifiable and distinct computing resources (e.g., hardware and executable instructions) configurable to perform operations to process input values ​​and provide output values. CNs in CN pool 220 may comprise scalar CNs and / or processing circuit cores for implementing arithmetic logic units (ALUs), digital signal processors (DSPs), vector CNs, VLIW engines, or field programmable gate array (FPGA) cells, or combinations thereof, to name a few examples of specific circuit cores that may be used to implement CNs in CN pool 220. CN pool 220 may be implemented according to a variety of architectures. For example, CN pool 220 may comprise one or more CNs implemented in full-featured or simplified form according to a reduced instruction set computing (RISC) architecture, a complex instruction set computing (CISC) architecture, or a very long instruction word (VLIW) architecture, or some combination of these types. The CNs in the CN pool 220 may also include a combination of scalar, SIMD, or multiple instruction and single data stream (MISD) ALUs.

[0029] According to one embodiment, the functionality of computing device 200 may include, for example, commercially available functionality such as transport rings, clos networks, shuffle circuits, etc. Pool of CNs 220 may facilitate multithreading in CNs, clustering of CNs, mailboxes, interrupts, functionality for synchronization between CNs, atomic operations at locations in transport memory 218, and functionality added to meet safety and security requirements, to name a few.

[0030] In particular implementations, the clustering of CNs in CN pool 220 may be formed based at least in part on resource tradeoffs and may be persistently defined in an integrated circuit (IC) device. In particular embodiments, depending on how CN pool 220 is configured, two of the same integrated circuit (IC) devices may implement different clusterings of related CNs. For example, a processor configuration may define the processing of multiple joins using clustering of CNs based at least in part on the associated joins to be performed. Each associated cluster may process a join of arrays, for example, such that the output of one cluster provides input to one or more other clusters.

[0031] According to one embodiment, transport memory 218 may form one or more buffers, including one or more first-in, first-out (FIFO) buffers. The one or more FIFO buffers may comprise a “vertical” FIFO buffer. Such a vertical FIFO buffer may have an endpoint that interfaces with the DMA controller and / or engine 212 (such endpoint may be referred to as an “external” endpoint) and another endpoint that forms a register that may provide data items usable as operands for CNs in pool of CNs 220 (such endpoint may be referred to as an “internal” endpoint). The internal endpoint of the vertical FIFO buffer may be shared by multiple CNs in pool of CNs 220. If such internal endpoint includes, for example, a FIFO out-end (i.e., an outbound end of a FIFO buffer), broadcasting to multiple CNs in pool of CNs 220 may be achieved. If such internal endpoint includes a FIFO in-end (i.e., an inbound end of a FIFO buffer), hardware locks and / or instructions executed on CNs in pool of CNs 220 may prevent race conditions. In some implementations, the FIFO buffers formed in transport memory 218 may also include “horizontal” FIFO buffers. A horizontal FIFO buffer may be a buffer with two endpoints that either provide data items as operands for different CNs in pool of CNs 220 or only provide data items to locations in transport memory 218. Note that the FIFO buffers formed in transport memory 218 can provide a transparent blocking and releasing mechanism when input and output rates differ. If the outer endpoints of the FIFO buffers in transport memory 218 are registers or operands that some CNs in pool of CNs 220 consume / process, and if the CNs are slow to do so, the DMA controller and / or engine 212 may eventually implement a blocking mechanism.Conversely, if the DMA controller and / or engine 212 is slow to write to such FIFO buffers in transport memory 218, the CNs in pool of CNs 220 may eventually implement a blocking mechanism. If FIFO buffers are implemented between CNs and / or between locations in transport memory 218, a similar transparent blocking and releasing mechanism may exist.

[0032] In an IC device, when transport memory 218 forms a FIFO buffer, the circuitry at the FIFO buffer endpoints may be permanent or configurable (e.g., via internal FPGA circuitry). Particular implementations may include vertical and horizontal FIFO buffer segments formed in the IC device circuitry. Such buffer circuitry may have endpoints that are run-time configurable from among at least one of operands or registers of a CN in pool of CNs 220, an interface to buffer 216, a location in transport memory 218, and endpoints of other FIFO buffers in transport memory 218. When the endpoints of FIFO buffers in transport memory 218 are operands or registers of a CN in pool of CNs 220, such FIFO buffer endpoints may be shared among multiple CNs in pool of CNs 220.

[0033] According to one embodiment, CNs in pool of CNs 220 may receive operands for computational operations, for example, from registers in a register file, from FIFO buffers in transport memory 218, and / or from special constant and parameter registers (which may include shared registers). In some examples, portions of a register file may be shared among multiple CNs in pool of CNs 220. Within computing device 200, host CPU 202 may also be able to access constant and parameter registers. Host CPU 202 may provide host functionality, such as initiating join operations that are executed after completing configuration tasks. Such configuration tasks may include, to name a few, defining clusters of CNs in the pool of CNs 220, configuring inter-CN communications, configuring the width and depth of FIFO buffers in transport memory 218, configuring endpoints of FIFO buffers in transport memory 218, defining manager CNs in the pool of CNs 220 and CN clusters that the manager CNs in the pool of CNs 220 manage, setting up DMA controllers and / or engines 212, setting up atomics, setting up communication and synchronization resources, and monitoring for merge completion.

[0034] In one embodiment, buffer 216 can facilitate the processing of the merge in multiple ways. In one such non-limiting aspect, the input signal stream at the merge may comprise an “indirection stream” (e.g., a collection of addresses in line-accessible memory 208 to be read). In such a case, such an indirection signal stream may not be directly fed to a CN in pool of CNs 220, but may instead be provided to DMA controller and / or engine 212, which may in turn transport the signal stream of read data items to the CN. In other words, the access latency due to random read operations (random line reads, or random word reads, or random indirect line or word reads, or combinations thereof) may be hidden by extracting the pattern of the random accesses, building an access list of significant length, and performing the necessary extraction and indirection in DMA controller and / or engine 212 itself. In other cases, the input signal stream may comprise a “double indirection stream” in which the signal stream contains addresses, and the data items at these addresses provide further indirection after some processing. In certain implementations, the first level of indirection of a double indirection stream may be read and used multiple times. In sensor fusion applications, multiple mergings may be configured, and an initial merging may be built (e.g., within the transport memory 218 itself) using lookup tables (LUTs) for the first level of indirection data items and the second level of indirection data items. However, it should be understood that these examples are not limiting.

[0035] According to one embodiment, some operations within the merging of signal streams (e.g., within f() and read_next() in the pseudocode example above) can be viewed as extracting information from data items distributed across CNs in pool of CNs 220. An example of such an extraction is shown in Table 1 below. [Table 2]

[0036] According to one embodiment, results may be provided by CNs in pool of CNs 220 operating cooperatively using operations such as shifting, shuffling, broadcasting, and multicasting operands among participating CNs, to name a few. Such features may be present within the CNs, for example, for word operands of SIMD vector CNs in pool of CNs 220. Pool of CNs 220 may be configured with such features for intra-CN communication via configurable bridge circuitry between CNs. Such circuitry may be hardwired or may be configurable at runtime.

[0037] According to one embodiment, CNs in a pool of CNs 220 (e.g., configured in a CN cluster) may be configured for specialized processing functions. In particular implementations, such specialized CNs in a pool of CNs 220 may facilitate management of an application's processing flow. For example, a pool of CNs 220 may include one or more processing CNs 224 and one or more manager CNs 222. The processing CNs 224 may perform processing operations, for example, on sensor observations, measurements, and / or other signals. In this example, the manager CN 222 may manage different sets of processing CNs 224. The processing CNs 224 may communicate with one or more manager CNs 222. In some examples, the manager CN 222 may provide information to the processing CNs 224 based on communications from the processing CNs 224. The processing CNs 224 may continue their processing as qualified by the information provided by the manager CN 222. In some implementations, the manager CN 222 may comprise a physically separate processing CN 224 or only dedicated circuitry formed within the processing CN 224.

[0038] Note that the rates at which different individual signal streams of a fusion (e.g., from different sources, such as different sensors) are generated and consumed (e.g., processed) may not necessarily be equal. The rate at which such signal streams are consumed or generated may be determined, for example, by characteristics of an operation such as function f() (in the pseudocode example above). For example, the rate at which a signal stream of sensor measurements and / or observations for a sensor fusion operation may be calculated / generated by function f(). Such a sensor fusion operation may involve an inverse sensor model that results in outputting a signal stream that is longer than the input signal stream. To process the fusion of longer signal streams, for example, the output rate may be matched to the bandwidth / throughput of the line-accessible memory 208. Configuring the output of one CN to be the input to another CN may assist, for example, in load balancing.

[0039] According to one embodiment, computing device 200 may enable the deployment of advanced sensor fusion operations to update particle filter states (e.g., in autonomous driving or other automotive applications) while consuming little power. In one application, an instance of computing device 200 may be implemented, for example, as a sequence of pipeline stages. For such an application, features of computing device 200 may be configured in use to have different amounts of resources allocated to different pipeline stages. The exchange of data items between pipeline stages and line-accessible memory 208 may be synchronized transparently (e.g., without mutexes, spinlocks, etc.) through the use of FIFO buffers (e.g., FIFO buffers formed in transport memory 218).

[0040] In another embodiment, computing device 200 may be configurable to optimize power, space, and / or performance (e.g., accuracy and / or latency). While features of computing device 200 may be adapted to implement CE, features of computing device 200 may also be adapted for other applications, including, for example, applications that rely on random access to line-accessible memory 208. Features of computing device 200 may also be implemented in so-called “supercomputers.” In the context of supercomputers, the low-power features of computing device 200 may help overcome power constraints that may prevent the realization of, for example, exascale supercomputers. Furthermore, circuitry for implementing computing device 200 may incorporate safety and security features to meet the requirements of embedded computing devices. Due to the small physical size and low power consumption of computing device 200, the use of computing device 200 need not necessarily be limited to use as an external accelerator integrated circuit (IC) device, but may also be incorporated into a subsystem within an automotive-grade system-on-chip (SOC) IC device.

[0041] In one aspect, for example, computing device 200 may comprise a particular arrangement and / or configuration of CNs (e.g., pool of CNs 220), transport memory 218, and / or DMA controller and / or engine 212. Computing device 200 may be configurable to provide a network of CNs and memory elements adapted to a particular type of computation, such as a particular application for processing signal streams (e.g., conveying measurements and / or observations from sensors). In one implementation, such a network of CNs and memory elements may enable simultaneous processing of multiple signal streams with high throughput and low latency. Such a network of CNs and memory elements may be implemented, at least in part, using an implementation of an intra-device communication protocol (e.g., AXI) via physical connections and FIFO buffers. As noted above, the endpoints of a FIFO buffer may include, for example, addressable memory locations or registers for receiving operands (e.g., general-purpose registers of an ALU configured as a computation node as operands for a computation operation) or results computed by a CN. A pool of FIFO buffers may comprise, for example, configurable endpoints to be associated with various CNs or memories.

[0042] In one aspect, certain embodiments disclosed herein are directed to so-called vectorized input / output (I / O) operations, including “scatter” and “gather” operations. Such vectorized operations may enable, for example, high-throughput transfers of large amounts of data into or out of physical memory (e.g., multiple addressable lines of memory in line-accessible memory 208) using a single request or command to improve efficiency and convenience. For example, a gather operation may involve sequentially reading data from multiple memory locations (e.g., buffers) and writing the read data to a signal stream or contiguous portion of memory in a single transaction. In one implementation, DMA controller and / or engine 212 may perform gather operations to service gather requests (e.g., arising in an application) that specify multiple, not necessarily line-aligned, memory locations from which data items are to be read and destinations (e.g., memory addresses) for storing the read items. On the other hand, a scatter operation may include (or involve) reading data items from a signal stream or contiguous memory and writing the read data items to multiple different memory locations that are not necessarily line-aligned. In one implementation, the DMA controller and / or engine 312 may perform gather operations to service scatter requests (e.g., generated by an application) that may specify locations (e.g., contiguous memory addresses) of data items to be read. Such scatter requests may also indirectly specify locations to which the read data items are to be written (by requesting the reading of specific locations that allow the locations to which the read data items will be written to be determined). The DMA controller and / or engine 312 may similarly perform indirectly specified gather operations.

[0043] According to one embodiment, the use of DMA transactions to assist in the processing of data items in a signal stream may be enhanced through the use of scatter and gather operations. In particular implementations, the contents loaded into a buffer from a gather operation may be used to determine one or more addresses for a subsequent gather or scatter operation. FIGS. 3A and 3E are schematic diagrams of a computing device 300 according to one embodiment, including a DMA controller and / or engine 312 that includes or communicates with a buffer 316 to facilitate scatter and / or gather operations. In particular implementations, the computing device 300 may include one or more features of the computing device 200 ( FIGS. 2A and 2B ). The DMA controller and / or engine 312 may be configured as a scatter-gather multicast DMA engine (SGM-DMA), e.g., with the ability to address words within an addressable line.

[0044] According to one embodiment, the DMA controller and / or engine 312 may receive an input signal stream and / or provide such input signal stream as blocks on a virtual channel through a FIFO buffer to an initial cluster of CNs. The output signal stream from the initial cluster of CNs may then be provided as blocks on a virtual channel or as data through a FIFO buffer to subsequent downstream clusters of CNs. Transfers between clusters of CNs may be, for example, multicast transfers dictated by a particular application. At any stage, some or all of the output signal streams from some CN clusters may be returned to the DMA controller and / or engine 312 (e.g., in the course of the DMA performing a scatter or gather operation).

[0045] According to one embodiment, the DMA controller and / or engine 312 may perform certain scatter and / or gather operations to transfer data items from one non-contiguous memory block to another non-contiguous memory block using a series of smaller contiguous block transfers. Here, obtaining such data items from non-contiguous blocks of a source memory may be performed in a gather operation. Similarly, writing data items to non-contiguous blocks of a destination memory may be performed in a scatter operation. In one implementation, the smallest unit of memory that can be accessed in such a source or destination memory may be a single addressable line of values ​​and / or state (e.g., a single cache line or word inline-accessible memory). For example, the DMA controller and / or engine 312 may communicate with a line-accessible memory (LAM) 308, which may be accessible line-by-line.

[0046] According to one embodiment, a physical memory (such as LAM 308) may comprise bit cells for defining values ​​and / or states to represent information, such as 1s or 0s. Such a physical memory may further organize the bit cells into words containing integer numbers of 8-bit bytes (e.g., 4-byte words spanning 32 bits or 8-byte words spanning 64 bits). Furthermore, such a physical memory may define line addresses (e.g., word line addresses) associated with consecutive bits that define an “addressable line” of values ​​and / or states. For example, in response to a read or write request (e.g., originating from a host processor), a memory controller may access a portion of memory in a read or write transaction targeted according to the word line address specified in the request. To service a read request, for example, the memory controller may retrieve the values ​​and / or states of all bytes of the line associated with the line address specified in the read request. Similarly, to service a write request, the memory controller may write the values ​​and / or states of all bytes of the addressable line associated with the line address specified in the write request. Although a line address may specify a memory location that includes all contiguous bytes of an addressable line, such a line address does not specify the location of individual sub-portions of such an addressable line, such as individual or contiguous bytes of less than the entire addressable line, or bytes that span the addressable line. Such sub-portions of an otherwise addressable line are referred to herein as "non-addressable portions."

[0047] According to one embodiment, an addressable line may define the smallest unit of memory that may be locatable and / or accessible according to a memory addressing scheme. In the particular exemplary embodiment of FIG. 3C, such addressable line 360 ​​may be made up of smaller units of memory, such as bits, bytes, or words. In the particular illustrated embodiment, addressable line 360 ​​is made up of n+1 bytes 3620-3622. n As noted above, certain implementations may be directed to updating non-addressable portions and / or portions of less than an entire addressable line and / or non-addressable portions that span multiple lines, with or without an addressable line. In the illustrated exemplary embodiment, bytes 3622 and 3623 may define the non-addressable portion of addressable line 360, where the entire addressable line may be locatable and / or accessible via a unique address according to a memory addressing scheme, but bytes 3622 and 3623 (less than the entire) are not addressable according to the memory addressing scheme.

[0048] In one implementation, the LAM 308 may have a LAM controller 306 configured to receive requests specifying the address of a line of data items stored in the LAM 308. Such a line of data items may include multiple words or multiple bytes and may be the smallest unit of a data item that the LAM 308 can retrieve and return to another device. According to one embodiment, the computing device 300 may use a buffer 316 to enable access of smaller, non-addressable portions of a line (e.g., single or multiple bytes within a word or group of words). According to one embodiment, the circuitry forming the buffer 316 may be integrated with the circuitry forming the DMA controller and / or engine 312 to enable minimal latency for accesses of the buffer 316 initiated by the DMA controller and / or engine 312. For example, buffer 316 may be formed as a static random access memory (SRAM) device that is accessible by the DMA controller and / or engine 312 without initiating a request and / or transaction on the main memory bus (e.g., a bus coupled to LAM 308 or a host computer / processor).

[0049] According to one embodiment, DMA controller and / or engine 312 may communicate with initiator 322. Initiator 322 may comprise a device (e.g., implemented at least in part by circuitry and / or logic) that achieves a particular state to trigger one or more DMA transactions performed by the DMA controller. For example, initiator 322 may comprise an output register of an ALU, a buffer, or a hardware interrupt handler, to name a few examples of devices that can initiate a DMA transaction.

[0050] According to one embodiment, the DMA controller and / or engine 312 may obtain a list of collection requests in response to a signal from the initiator 322. In one particular implementation, the initiator 322 may trigger a DMA transaction in response to an event or condition in the execution of a particle filter process. For example, the particle filter process may identify data items in memory that are expected to be retrieved for processing in a future execution cycle. Once a substantial amount of such data items is identified, a list of collection requests identifying such data items (e.g., as redirected collection requests) may be forwarded to the DMA controller and / or engine 312. In one implementation, such a list of collection requests may be provided to the DMA controller and / or engine 312 in shared memory or a network-on-chip (NOC), to name a few examples. Once such a list is known to be available for processing by the DMA controller and / or engine 312, a process (e.g., execution of computer-readable instructions) to generate the list may trigger the DMA controller and / or engine 312 via an interrupt or a post message. Such a trigger may initiate the DMA controller and / or engine 312 to initiate one or more gather operations (eg, redirected gather operations).

[0051] In response to a signal from the initiator 322, the DMA controller and / or engine 312 may obtain a list of collection requests in the form of a linked list. Such a linked list may be locatable in memory (e.g., the reassembly buffer 316 or the line-accessible memory (LAM) 318), for example, according to an address provided by the initiator 322. According to one embodiment, such a list of collection requests may include individual collection requests that are serviceable as standalone collection requests independent of other collection requests in the list of collection requests. The DMA controller and / or engine 312 may combine the addresses in such collection requests with a (potentially smaller) list of line read requests to be executed by a memory controller (e.g., the memory controller 106 of FIG. 1). Individual line read requests in such a list of line read requests may indicate physical memory addresses (e.g., in the LAM 308) that specify memory locations from which the memory controller reads to service the individual line read requests. Such a memory controller may service such line read requests by loading the requested lines into the buffer 316. When a requested line read by the memory controller arrives at buffer 316, DMA controller and / or engine 312 may refer to the original list of gather requests (used to form the list of line read requests) to extract the requested data items from the read line arriving at buffer 316. DMA controller and / or engine 312 may form packets from the extracted items to be forwarded to one or more requesting entities (e.g., host CPU 202 and / or processes executing on the CN of CN pool 220). In one aspect, a data item extracted from a line stored in buffer 316 may include two or more non-addressable portions of the line (e.g., selected bytes and / or fields within an addressable line / word, such as an addressable line in memory that contains the data items).

[0052] 3B illustrates a flow diagram of a process 350 for a gather operation according to one aspect of the present disclosure. In one embodiment, the process 350 may include operations 352, 354, 356, and 358, which may be performed by one or more circuits, such as the DMA controller and / or engine 312 and / or the buffer 316. The operation 352 may include processing one or more gather requests to determine one or more addressable lines of a data item to fetch from memory. For example, the operation 352 may map parameters in the received request to addresses in the LAM 308. The operation 354 may be implemented by a circuit for loading values ​​and / or states into memory and may include loading signals and / or states of one or more addressable lines (e.g., representing a data item) stored in a memory, such as the LAM 308, into the buffer 316. 3B, operation 356 may include parsing one or more non-addressable portions of the line loaded into buffer 316 in operation 354 into portions, such as, for example, single or multiple bytes within a word or group of words. Operation 358 may include completing processing of the two or more collection requests, for example, by returning the non-addressable portions parsed in operation 356.

[0053] As shown in FIG. 3B , operation 358 may include responding to multiple collection requests presented to the DMA controller and / or engine 312 in a list of collection requests. Loading one or more addressable lines into buffer 316 in operation 354 may occur in response to the list of collection requests. Operation 354 may parse two or more unaddressable portions according to data items specified in the list of collection requests. For example, operation 356 may decode parameters in the collection requests that map from line addresses to byte offsets to the corresponding unaddressable portions. The parsed portions may then be forwarded to the initiator of the collection requests (e.g., initiator 322). Here, process 350 may enable servicing multiple collection requests using access of a single addressable line loaded into buffer 316, for example, if the multiple collection requests specify parsed portions within the same single addressable line. This may eliminate the need for the DMA controller and / or engine 312 to access the same addressable line (e.g., LAM 308) multiple times for separate collection requests for data items within the same addressable line.

[0054] To service one or more collection requests, process 350 may collect less than an entire addressable line in memory by loading the addressable line into a buffer and parsing the non-addressable portion provided to the requester. While some collection requests may request collection of less than an entire addressable line, one or more received collection requests may request collection of an entire addressable line and / or multiple lines and / or bytes spanning a line. For collection requests that request collection of less than an entire addressable line, DMA controller and / or engine 312 may perform process 350. According to one embodiment, for collection requests that request collection of an entire addressable line, DMA controller and / or engine 312 may bypass operations 354, 356, and 358 and perform the collection operation without loading the addressable line into buffer 316.

[0055] In another implementation, the DMA controller and / or engine 312 can obtain a list of scattered requests in response to a signal from the initiator 322. The DMA controller and / or engine 312 may obtain such a list of scattered requests in the form of a linked list locatable in memory according to an address provided by the initiator 322. The DMA controller and / or engine 312 may then combine addresses accessed by such scattered requests with a (potentially smaller) list of line read requests to be executed by a memory controller (e.g., memory controller 106). For example, scattered requests in the list of scattered requests that reference data items in the same addressable line of memory may be combined such that only a single line read is required (to access the data item to service multiple scattered requests). The requested lines read by such a memory controller may be loaded into the buffer 316.

[0056] According to one embodiment, the obtained list of scatter requests may indicate specific non-addressable portions (e.g., individual bytes or fields) of the addressable lines to be read and loaded into buffer 316. The non-addressable portions of the addressable lines loaded into buffer 316 may then be modified and / or overwritten. When the requested lines read by the memory controller arrive at buffer 316, DMA controller and / or engine 312 may refer to the original list of scatter requests to determine the specific non-addressable portions of the read lines that arrived at buffer 316 to be modified and / or overwritten. DMA controller and / or engine 312 may form packets from the modified lines in buffer 316 and write them back to memory via the memory controller.

[0057] 3D shows a flow diagram of a process 370 for a scatter operation, according to one embodiment. Process 370 may include operations 372, 374, 376, and 378. Operation 372 may include, for example, receiving one or more scatter requests from initiator 322. Operation 374 may include loading signals and / or states (e.g., representing data items) of one or more addressable lines in a memory, such as LAM 308, into buffer 316. In particular implementations, DMA controller and / or engine 312 may determine the addressable lines to be fetched from LAM 308 in operation 374 by processing one or more gather requests. Operation 376 may include writing values ​​and / or states to at least one non-addressable portion of the lines loaded into buffer 316 in operation 374 to at least partially modify the one or more addressable lines stored in buffer 316. For example, operation 374 may write values ​​and / or states to two or more unaddressable portions based on the multiple scattered requests received in block 372. Operation 378 may complete servicing such scattered requests by initiating a write operation to write back to memory (e.g., LAM 308) at least one of the one or more addressable lines modified in operation 374. For example, operation 378 may initiate a write back operation to a line address in an addressable line of LAM 308 that was loaded in operation 374 and modified in operation 376.

[0058] In particular implementations, operation 374 may be initiated by the scatter request received in operation 372. Here, the one or more addressable lines of values ​​and / or states loaded into buffer 316 in operation 374 may be obtained by servicing one or more line read requests by a memory controller (e.g., memory controller 106). Operation 378 may include initiating the memory controller to perform one or more operations to write the one or more modified addressable lines to write the modified addressable lines. Here, process 370 may enable servicing multiple scatter requests using accesses of a single addressable line loaded into buffer 316. The multiple scatter requests received in operation 372 may specify, for example, words, bytes, fields, etc., within the same single addressable line. This may eliminate the need for the DMA controller and / or engine 312 to access the same addressable line multiple times for separate scatter requests for data items within the same addressable line (e.g., in LAM 308). Process 370 may further include converting the multiple scatter requests received in operation 372 into a list of line read requests to be issued to a memory controller (a list of line read requests specifying one or more addressable lines of values ​​and / or states). Such conversion of the multiple scatter requests into a list of line read requests may further include constructing at least one single line read request for an addressable line in memory that includes data items requested by the at least two scatter requests received in operation 372.

[0059] To service one or more scatter requests, process 370 may update less than an entire addressable line in memory by loading the addressable line into buffer 316 and updating some portions of the loaded addressable line while leaving other portions unchanged. While some scatter requests may request an update of less than an entire addressable line, one or more received scatter requests may request an update to an entire addressable line, which the DMA simply writes rather than performing a read-modify-write operation. For scatter requests that request an update to less than an entire addressable line, DMA controller and / or engine 312 may perform process 370. DMA controller and / or engine 312 may also be configured to service scatter requests that request an update of an entire addressable line by bypassing loading the addressable line into buffer 316. Here, to complete such an update to an entire addressable line, DMA controller and / or engine 312 may initiate a write operation to the addressable line in LAM 308 without loading the addressable line into buffer 316.

[0060] According to one embodiment, the DMA controller and / or engine 312 may receive multiple scatter requests that collectively request updates to the same overlapping portion of an addressable line within the LAM 308. This can cause conflicts over how to update the overlapping portion to service the multiple scatter requests. Such multiple scatter requests may be ordered, for example, according to creation time or reception time. According to one embodiment, conflicts to update a portion of an addressable line by multiple scatter requests may be resolved, for example, according to the most recent scatter request created or received.

[0061] In particular implementations of processes 350 and 370, the non-addressable portion of a line stored in buffer 316 may be a byte, a collection of bytes or fields, or the like. While certain actions in the scatter and gather operations above are described as occurring in a particular sequence, some actions may be performed simultaneously, and some actions may be performed simultaneously or in a particular sequence as a matter of engineering choice. Furthermore, physical optimizations such as the number and type of processing cores used to implement the DMA controller and / or engine 312 and interface engine features, the number and type of associated memory elements and memory blocks for buffer 316, the number of ports and addressability features of the memory forming LAM 308, and the like may be selected as a matter of engineering choice. For example, buffer 316 may or may not be byte-addressable, and DMA controller and / or engine 312 may comprise, for example, a scalar or vector engine.

[0062] 4A and 4D are schematic diagrams of a computing device 400 including a DMA controller and / or engine 412 in communication with a buffer 416 and an initiator 422 to facilitate redirection of DMA transactions, according to one embodiment. In particular implementations, the computing device 400 may include one or more features of the computing device 200 (FIGS. 2A and 2B). According to one embodiment, the DMA controller and / or engine 412 may obtain a list of redirected read requests in response to a signal from the initiator 422. According to one embodiment, the read request may include a message and / or signal specifying one or more target memory addresses to be accessed in a read transaction to service the read request. For example, the read request may specify one or more target addresses as word line addresses of locations in memory containing content to be retrieved in a memory read transaction to service the read request. As referred to herein, a "redirected read request" refers to a read request that has been transformed or altered such that content(s) at the original target memory address(es) are modified and / or translated to different memory address(es). Here, the different memory address(es) specify locations in memory containing content to be retrieved in a read transaction to service the redirected read request. The DMA controller and / or engine 412 will interpret the addresses specified in such redirected read requests as word gather requests, for example, to store the gathered words in buffer 416. The size of the associated word initially read to service the word gather request portion of the redirected read request may depend, at least in part, on how the associated word is interpreted to form, for example, a target address for the redirection. In one implementation, the word read may include, for example, an address or an index into an array that can be converted to an address.In one implementation, DMA controller and / or engine 412 may then convert the collected words stored in buffer 416 to addresses and merge the addresses (i.e., the converted collected words) with the redirected read request to form a collect request. DMA controller and / or engine 412 may then service the formed collect request to send the resulting data items to a destination as described in the redirected read request obtained in response to a signal from initiator 422.

[0063] 4B is a flow diagram of a process 450 for facilitating redirection of DMA transactions, according to one embodiment. Process 450 may include operations 452, 454, 456, and 458. Operation 454 may include performing a gather operation (e.g., a word gather operation) based at least in part on the one or more redirected read requests received in operation 452. In this operation, DMA controller and / or engine 412 may obtain or otherwise receive such redirected read requests, for example, in response to a signal from initiator 422. In particular implementations, DMA controller and / or engine 412 may interpret the one or more redirected read requests as a request for a first of two gather operations to be performed. During the first gather operation performed in operation 454, a “gathered” data item may be loaded into buffer 416.

[0064] Operation 456 may include converting one or more collected data items stored in buffer 416 into one or more addresses to specify a subsequent collection operation. For example, operation 456 may include converting one or more words obtained from performing a first word collection operation (performed in operation 454) into one or more addresses (also referred to as one or more address values). In one particular implementation, operation 456 may include parsing values ​​and / or states in buffer 416 (from the collection operation) to determine one or more memory addresses within LAM 408. Operation 456 may further include applying one or more arithmetic operations to the parsed values ​​and / or states to determine one or more memory addresses within LAM 408. For example, operation 456 may apply one or more arithmetic operations to the parsed values ​​and / or states stored in the buffer, which form memory addresses to memory locations within LAM 408. Such formed addresses to memory locations within LAM 408 may form the basis of a subsequent collection operation.

[0065] According to one embodiment, the arithmetic operation applied in operation 456 may be defined according to equation (1) as follows: address=base+x×element_size (1) During the ceremony, address is the target address (e.g., for a gather operation determined in operation 456 or for a scatter operation in operation 476); x is the value obtained from a collection operation (e.g., in operation 454 or 474), and The base and element_size are parameters provided in the redirected request (eg, the redirected read request received in operation 452 or the redirected write request received in operation 472).

[0066] Operation 458 may include performing a second (e.g., subsequent) gather operation to transfer data items located at the one or more determined addresses to a destination. Such destination may be determined based at least in part on the one or more redirected read requests. In another implementation, operation 458 may perform two or more gather operations based on the one or more addresses obtained in operation 456. According to one embodiment, in performing the gather operations, operation 458 may interpret the one or more redirected read requests as two requests for word gather operations.

[0067] According to one embodiment, a write request may include messages and / or signals that specify one or more target memory addresses to be accessed in a memory write transaction to service the write request. For example, a write request may specify one or more target addresses as word line addresses of locations in memory to be written in a memory write transaction (to service the write request). As referred to herein, a "redirected write request" means a write request that has been translated or modified such that the original target memory address(es) are modified and / or translated to different target memory address(es), where the different target memory address(es) specify locations in memory to be written in a write transaction to service the redirected write request.

[0068] In another particular implementation, the DMA controller and / or engine 412 may obtain a list of redirected write requests in response to a signal from the initiator 422. The DMA controller and / or engine 412 may interpret addresses specified in such redirected write requests as gather requests. Such gather requests may, for example, be to load gathered words into the buffer 416. The associated words that are read may have a size that depends, at least in part, on how the associated words are interpreted to form target addresses for redirection. One or more such words loaded into the buffer 416 may be interpreted to form target addresses for redirection. In one implementation, such words loaded into the buffer 416 may include, for example, addresses or indexes into an array that can be converted to addresses. In one implementation, the DMA controller and / or engine 412 may then convert the gathered words stored in the buffer 416 into addresses and merge the addresses with the redirected write requests to form a scatter request. The DMA controller and / or engine 412 may then service the formed distributed requests and effect updates of the lines according to the redirected write requests obtained in response to signals from the initiator 422.

[0069] 4C is a flow diagram of a process 470 for facilitating redirection of DMA transactions, according to one embodiment. Operation 472 may include performing a word gather operation based at least in part on the one or more redirected write requests received at operation 472. The DMA controller and / or engine 412 may receive such redirected write requests at operation 472 in response to a signal from the initiator 422, for example. In particular implementations, the DMA controller and / or engine 412 may interpret the one or more redirected write requests as a request to perform a word gather operation followed by a word scatter operation. During the word gather operation performed at operation 474, “gathered” data items may be loaded into buffer 416. Operation 476 may include translating the one or more gathered data items stored in buffer 416 (from the gather operation) into one or more addresses for specifying a subsequent word scatter operation. In one particular implementation, operation 476 may include analyzing values ​​and / or states in buffer 416 to determine one or more memory addresses within LAM 408. For example, operation 476 may be performed at least in part by circuitry (e.g., a DMA controller and / or engine adapted to analyze values ​​and / or states) that analyzes the values ​​and / or states in the buffer to determine at least one of the one or more addresses. For example, operation 476 may apply one or more arithmetic operations to the analyzed values ​​and / or states stored in the buffer to form memory addresses within LAM 408. Operation 476 may further include applying one or more arithmetic operations to the analyzed values ​​and / or states to determine one or more memory addresses to memory locations within LAM 408.

[0070] Operation 478 may include performing a scatter to write specific data items to one or more addresses obtained in operation 476 based at least in part on the one or more redirected read requests. According to one embodiment, operation 476 may apply an arithmetic operation to calculate target addresses of the scatter operation performed in operation 478 according to equation (1). For example, DMA controller and / or engine 412 may form such a scatter request based at least in part on the content at the one or more addresses determined in operation 476. Such specific data items to be written from such a scatter request may be specified in one or more redirected write requests received in operation 472 in response to a signal from initiator 422, for example. In one particular implementation, operation 478 may include interpreting the content at the addresses determined in operation 476 as addresses for the scatter operation.

[0071] One particular implementation of a computing device for processing multiple signal streams from multiple related sources is illustrated by computing device 500 of FIGS. 5A and 5C. In a particular implementation, computing device 500 may include one or more features of computing device 200 (FIGS. 2A and 2B). Computing device 500 may comprise a pool 520 of computation nodes (CNs) with associated register files and local memory, a pool of block memory and FIFO buffers, and a NOC interconnecting various components. As noted above, the CNs in pool 520 of CNs may comprise scalar CNs and / or processing circuit cores for implementing ALUs, digital signal processors (DSPs), vector CNs, VLIW engines, or field programmable gate array (FPGA) cells, or combinations thereof, to name a few examples of specific circuit cores that may be used to implement the CNs in pool 520 of CNs. In one implementation, a CN may comprise a single processing circuit core capable of performing operations to map input operands to output computation results. In another implementation, a CN may comprise multiple separate processing cores for performing operations to map input operands to output computation results. CNs in the pool of CNs 520 may comprise dedicated local memory (e.g., SRAM) and general-purpose registers for receiving operands for operations to be performed and / or for providing results from the execution of operations. In one particular implementation, a CN in the pool of CNs 520 may be formed from one or more processing circuit cores integrated / configured with elements of a pool of memory 508 (e.g., transport memory 218), such as a FIFO buffer. In one embodiment, the pool of memory 508 may provide memory resources shared among the CNs in the pool of CNs 520 and facilitate communication between the CNs in the pool of CNs 520. For example, the endpoints of FIFO buffers integrated / configured with the pool of CNs 520 may include memory blocks local to or separate from the CNs and / or the CN's register file of the CN 520.The functionality of a CN in pool of CNs 520 may be that of a host processor, a processing manager, or a signal processor, or a combination thereof, to name a few examples. According to one embodiment, CNs in pool of CNs 520 may support atomic, interrupt, and other inter-processor communication (IPC) functions to facilitate signal communication between CNs in pool of CNs 520.

[0072] According to one embodiment, computing device 500 can receive signal streams containing data items (e.g., sensor signals, observations and / or measurements, timestamps, metadata, etc.) provided from external sources such as sensors and / or memory. The external memory (not shown) can be coupled to a DMA controller and / or engine (e.g., DMA controller and / or engine 312 and / or 412). CNs in pool of CNs 520 can also provide sources of signal streams containing data items. For example, a CN in pool of CNs 520 can provide data items in a signal stream that are loaded into a FIFO buffer. A CN in pool of CNs 520 can also provide data items in a signal stream by transporting packets of output parameters within IC devices in the NOC as blocks of data items. According to one embodiment, a single logical signal stream can be transmitted to multiple CNs in pool of CNs 520 in a multicast manner. A single logical signal stream can also be segmented into sub-signal streams that are provided to subsets of CNs in pool of CNs 520. In another embodiment, some CNs in pool of CNs 520 may process data items from various signal streams and provide output signal streams as input streams to other CNs in pool of CNs 520. The final output of a CN in pool of CNs 520 may include a signal stream output from computing device 500 to a sink such as an actuator, memory, storage, or display device, to name a few examples. Processing between CNs in pool of CNs 520 may be controlled and / or orchestrated via a combination of interrupts, polling of status flags, and periodic checking of operations, to name a few examples.

[0073] 5B is a flow diagram of a method 500 (also referred to as process 500) for processing multiple signal streams from related multiple sources, according to one embodiment. According to one embodiment, process 500 can combine data items received in the signal streams from multiple sources to update the state of a particle filter (e.g., to support autonomous driving of a vehicle). Such multiple sources can include, for example, multiple sensor devices (e.g., mounted on a vehicle), such as, for example, a camera, a speedometer, active sensing devices (e.g., radar / lidar), environmental sensors (e.g., thermometers, light sensors, altimeters, etc.), microphones, etc., to name a few. Such signal streams from multiple sources can be received in signal packets (e.g., signal packets including sensor measurements and / or observations obtained from a remote vehicle) from a communication network. In certain implementations, the data items received in the signal streams from the multiple sources can include sensor measurements and / or observations having common attributes identifiable by process 500. In particular implementations, such common attributes may be identifiable by metadata co-located with measurements and / or observations received from multiple sources. For example, such common attributes may be associated by time (e.g., a timestamp indicating the time of the measurement and / or observation), space (e.g., location relative to an origin, such as a point on the vehicle), and source (e.g., which particular physical sensor on the vehicle or which particular type of sensor attached to the vehicle).

[0074] Operation 552 may include associating data items received from multiple data streams based at least in part on attributes common to the data items. Such multiple data streams may be provided as the output of a CN and / or DMA collection, word collection, or redirected read / collection operation. For example, operation 552 may sort and / or correlate (e.g., “bucketize”) measurements and / or observations received from different signal streams by time (e.g., according to timestamps) and / or space (e.g., position of observed objects relative to a reference point). In particular implementations, operation 552 may associate measurements and / or observations obtained approximately simultaneously and received from different sources (e.g., different sensors) with particle positions defined in the current state of the particle filter. In one particular embodiment, operation 552 may associate measurements and / or observations from different sources according to the locality of particular objects observed and / or measured by such associated measurements and / or observations. In another particular implementation, data items from each of the multiple signal streams may be loaded into a buffer associated with a direct memory access DMA controller associated with the signal stream (e.g., buffer 216). Operation 552 may then identify at least one of the common attributes associated with the data items loaded into the buffer based at least in part on the contents of the data items loaded into the buffer.

[0075] Operation 554 may include simultaneously loading the data items associated in operation 552 into one or more registers of a computational node (e.g., general-purpose registers of ALUs or other processing cores forming the computational node) without storing them in line-accessible memory. According to one embodiment, a register of a computational node may be loaded with data items retrieved in an execution cycle of the computational node. For example, a data item loaded into a register of a computational node (e.g., at an endpoint of a FIFO buffer) in one execution cycle may provide an operand for a computational operation performed by the computational node in a next execution cycle. In one particular implementation, one or more registers of a computational node may include an endpoint of an associated FIFO buffer formed by internal memory (e.g., pool of memory 508). Multiple such FIFO buffers with endpoints in the registers of a computational node may be synchronized to apply data items from multiple sources (e.g., data items loaded from different sensors) as operands of a computational node having common attributes (e.g., time and spatial attributes). Operation 556 may include execution of a computational node to process the data items simultaneously loaded in operation 554 to process the loaded data items as operands for one or more computational operations (e.g., to perform one or more functions, such as updating the state of a particle filter). Data items output from the execution of one or more computational operations in operation 556 may form data items for additional signal streams to be processed by additional computational nodes and / or for storage in memory.

[0076] In one embodiment, operation 552 may be performed by a first computational node sorting sensor observations and / or measurements received in multiple signal streams based at least in part on associated timestamps and localities of objects observed and / or measured by the sensor observations and / or measurements. The sorted sensor observations and / or measurements may then be loaded into one or more registers of a second computational node at block 554. Execution of the second node at block 556 may then combine the sorted sensor observations and / or measurements.

[0077] In another embodiment, the additional signal stream data item as output of operation 556 may be loaded into one or more registers of a subsequent compute node as an operand for one or more additional computational operations. For example, one or more direct memory access transactions may be performed to store the additional signal stream data item in external memory, a word scatter operation may be performed to write the additional signal stream data item, a redirected write operation may be performed to write the additional signal stream data item, or the additional signal stream data item may be provided as a control signal to one or more actuators, or any combination thereof.

[0078] In this context, "loading simultaneously" as referred to herein means loading data items to be processed by a compute node in the same execution cycle. When data items from synchronized signal streams are loaded simultaneously into registers of a compute node, such data items may be loaded into the register in the same execution cycle of the compute node (e.g., as operands of a computational operation executed in the next execution cycle). When data items from respective asynchronous signal streams are loaded simultaneously into registers of a compute node, such data items may be loaded into the register in different (e.g., adjacent) execution cycles of the compute node. For example, execution of a compute node may be paused for one or more execution cycles to allow multiple data items from different asynchronous signal streams to be loaded into registers as operands of a computational operation in an execution cycle of the compute node. In another embodiment, a transparent block and release mechanism may be applied to a compute node to facilitate the simultaneous loading of data items from asynchronous signal streams into registers of the compute node.

[0079] According to one embodiment, the data items in the at least two signal streams associated in operation 552 may include sensor observations and / or measurements associated in operation 554. The data items loaded simultaneously in operation 554 may then include related sensor observations from multiple signal streams (e.g., from multiple different sources). In particular implementations, operation 552 may include combining the sensor observations and / or measurements in the at least two signal streams to provide combined sensor observations and / or measurements. Operation 552 may then sort and / or correlate (e.g., bucket) the combined sensor observations and / or measurements based at least in part on associated timestamps and locations of objects observed and / or measured by the sensor observations and / or measurements.

[0080] In particular implementations, operation 554 may be implemented, at least in part, using process 350 (FIG. 3B) and / or process 450 (FIG. 4B). For example, a DMA controller and / or engine may perform a collection operation as described in process 350 and / or 450 to fill queues maintained by FIFO buffers (e.g., with associated measurements and / or observations) and load registers of the computational node. Such execution of the DMA controller and / or engine may selectively load data items into one or more registers based, at least in part, on an indication of common attributes in the contents of the data items loaded into the FIFO buffers. In one implementation, multiple FIFO buffers may have endpoints at corresponding registers of the computational node, with the queues of different FIFO buffers being filled with data items from different signal streams (e.g., measurements and / or observations from different sensors). The queues of multiple FIFO buffers may be filled such that data items from different signal streams related by particular attributes (e.g., temporal and spatial attributes) are loaded simultaneously into the endpoints of the FIFO buffers (e.g., into general-purpose registers of a compute node). According to one embodiment, operations 552 and 554 may be performed by a first compute node in one execution cycle of the second compute node to simultaneously load related data items into one or more registers of the second compute node. In a subsequent execution cycle of the second compute node, the second compute node may perform a computation operation using the related data items simultaneously loaded in block 554 as operands. In one implementation, the simultaneous loading of data items in block 554 may be facilitated by a FIFO buffer having endpoints at registers of the second compute node. Here, the first compute node may fill the queues of the FIFO buffers with related data items from different signal streams such that each related data item arrives at the endpoint of the FIFO buffer in the same or adjacent execution cycle.By filling the FIFO buffer queue with related data items in this manner, the first computing node does not need to store the related data items in line-accessible memory (e.g., line-accessible memory 208 or RAM 108).

[0081] In another particular implementation, the result of the execution of the computational node in operation 556 may provide data items of one or more additional signal streams. In one implementation, such data items of the additional signal streams may be loaded into one or more registers of a subsequent computational node and / or a downstream computational node. In one example, a data item in an output register of the computational node (e.g., loaded from execution of a computational operation) may be transferred to a line-accessible memory by a DMA write transaction, a word distributed DMA transaction, and / or a redirect distributed DMA transaction (e.g., according to process 470 of FIG. 4C ). Such a DMA transaction may, for example, store the data item of the additional signal stream in the line-accessible memory. In another example, such a data item provided in the output register of the computational node may be applied as a data item of an input signal stream to one or more other computational nodes. In one particular implementation, the result from the execution of the first computational node in operation 556 may load the result (e.g., from the pool of memory 508) into an output register defined as a first endpoint of an associated FIFO buffer. The associated FIFO buffer may then have a second endpoint defined as an input register of a second computational node to receive as an operand the result determined by the execution of the first computational node.

[0082] According to one embodiment, operations 552 and 554 may be performed by a first CN in pool of CNs 520, and operation 556 may be performed by a second CN in pool of CNs 520. In operation 554, the first computational node may simultaneously load related data items (e.g., data items including sensor observations and / or measurements associated based at least in part on spatial and temporal attributes, where the simultaneously loaded data items include the associated sensor observations and / or measurements) into one or more registers of a second CN in pool of CNs 520. The second CN may then process the data items simultaneously loaded by the first CN as operands to one or more computational operations in operation 556. In one implementation, the first FIFO buffer may define a first endpoint as a register of the first CN (e.g., an output register of the first CN). The second endpoint of the first FIFO buffer may define a first register of one or more registers of the second CN (e.g., an input register of the second CN). As can be observed, the first FIFO may allow data items to be loaded simultaneously into one or more registers of the second CN in operation 554 without storing the associated data items in line-accessible memory, as noted above. In another implementation, the FIFO buffer may define a first endpoint as a register of a third CN in the pool of CNs 520 and a second endpoint as at least a second register of the one or more registers of the second CN. Here, both the first CN and the third CN may simultaneously load associated data items into one or more registers of the second CN node in operation 554 without storing the associated data items in line-accessible memory.

[0083] According to one embodiment, all or a portion of computing devices 200, 300 (e.g., including features implementing processes 350 and / or 370, such as circuitry for forming a DMA controller), 400 (e.g., including features implementing processes 450 and / or 470, such as circuitry for forming a DMA controller), and / or 500 (e.g., including features implementing process 550) may be formed by and / or represented in whole or in part by transistors and / or underlying metal interconnects (not shown) in a process (e.g., a front-end-of-line and / or back-end-of-line process), such as, by way of example, a process for forming complementary metal-oxide-semiconductor (CMOS) circuits. It should be understood, however, that this is merely one example of how circuitry may be formed in a device in a front-end-of-line process, and that claimed subject matter is not limited in this respect.

[0084] It should be noted that the various circuits disclosed herein may be described using computer-aided design tools and represented in terms of their behavior, register transfers, logical components, transistors, layout geometry, and / or other characteristics as data and / or computer-readable instructions embodied in various computer-readable media (e.g., non-transitory storage media). File and other object formats in which such circuit representations may be implemented (e.g., in a circuit device) include, but are not limited to, formats that support behavioral languages ​​such as C, Verilog, and Very High Speed ​​Integrated Circuit Hardware Description Language (VHDL), formats that support register-level description languages ​​such as Register Transfer Language (RTL), formats that support geometry description languages ​​such as Graphic Design System II (GDSII), Graphic Design System III (GDSIII), Graphic Design System IV (GDSIV), Caltech Intermediate Format (CIF), Manufacturing Electron Beam Exposure System (MEBES), and any other suitable formats and languages. Storage media in which such formatted data and / or instructions may be embodied may include, but are not limited to, various forms of non-volatile storage media (e.g., optical, magnetic, or semiconductor storage media) and carrier waves that may be used to transfer such formatted data and / or instructions through wireless, optical, or wired signaling media, or any combination thereof. Examples of transfer of such formatted data and / or instructions via a carrier wave may include, but are not limited to, transfer (upload, download, email, etc.) over the Internet and / or other computer networks via one or more electronic communication protocols (e.g., Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Simple Mail Transfer Protocol (SMTP), etc.).

[0085] When received within a computer system via one or more machine-readable media, such data and / or instruction-based representations of the circuit described above may be processed by a processing entity (e.g., one or more processors) within the computer system in conjunction with the execution of one or more other computer programs, including, without limitation, netlist generators, place-and-route programs, etc., to generate a representation or image of a physical representation of such circuit. Such representation or image may then be used in device fabrication, for example, by enabling the generation of one or more masks used to form various components of the circuit in a device fabrication process (e.g., a wafer fabrication process).

[0086] In the context of this patent application, the term "between" and / or similar terms will be understood to include "among," where appropriate for the particular use, and vice versa. Similarly, in the context of this patent application, the terms "compatible with," "comply with," and / or similar terms will be understood to include substantial compatibility and / or substantial compliance, respectively.

[0087] Unless otherwise indicated, in the context of this patent application, the term "or" when used to relate a list such as A, B, or C is intended to mean A, B, and C used in an inclusive sense, as well as A, B, or C used in an exclusive sense. In this understanding, "and" is used in an inclusive sense and is intended to mean A, B, and C, while "and / or" may be used with due care to make clear that all of the foregoing meanings are intended, but such use is not required. Furthermore, the term "one or more" and / or similar terms are used to describe any feature, structure, characteristic, etc. in the singular, and "and / or" is also used to describe multiple features, structures, characteristics, etc. and / or any other combination thereof. Similarly, the term "based on" and / or similar terms are not necessarily intended to convey an exhaustive list of factors, but are understood to allow for the existence of additional factors not necessarily explicitly described.

[0088] Algorithmic descriptions and / or symbolic representations are examples of techniques used by those skilled in the signal processing and / or related arts to convey the substance of their work to others skilled in the art. An algorithm, in the context of this patent application, is generally considered to be a self-consistent sequence of operations and / or similar signal processing leading to a desired result. In the context of this patent application, operations and / or processing involve physical manipulations of physical quantities. Typically, although not necessarily, such quantities may take the form of electrical and / or magnetic signals and / or states capable of being stored, transferred, combined, compared, processed, and / or otherwise manipulated, e.g., as electronic signals and / or states constituting components of various forms of digital content, such as signal measurements, text, images, video, sound, etc.

[0089] It has proven convenient at times, principally for reasons of common usage, to refer to such physical signals and / or physical states as bits, values, elements, parameters, symbols, characters, terms, samples, observations, weights, numbers, numerals, measurements, contents, or the like. It should be understood, however, that all of these and / or similar terms are to be associated with the appropriate physical quantities and are merely convenient labels. Unless otherwise indicated, and as is clear from the foregoing description, it should be understood that throughout this specification, descriptions utilizing terms such as "processing," "computing," "calculating," "determining," "establishing," "obtaining," "identifying," "selecting," "generating," and the like can refer to actions and / or processes of particular devices, such as special purpose computers and / or similar special purpose computing devices and / or network devices. Thus, in the context of this specification, a special purpose computer and / or similar special purpose computing device and / or network device may process, manipulate, and / or transform signals and / or states, typically in the form of physical electronic and / or magnetic quantities, within the memory, registers, and / or other storage, processing, and / or display devices of the special purpose computer and / or similar special purpose computing device and / or network device. Thus, in the context of this particular patent application, as noted above, the term "particular device" includes a general purpose computing device, such as a general purpose computer, and / or network device once it has been programmed to perform a particular function, such as by following program software instructions.

[0090] In some situations, the operation of a memory device, such as a change of state from a binary 1 to a binary 0 or vice versa, may involve a transformation, such as a physical transformation. In certain types of memory devices, such a physical transformation may involve the physical transformation of a product to a different state or thing. For example, without limitation, in some types of memory devices, a change of state may involve the accumulation and / or storage of an electric charge or the release of a stored electric charge. Similarly, in other memory devices, a change of state may involve a physical change, such as a transformation of magnetic orientation. Similarly, a physical change may involve a transformation of molecular structure, such as from a crystalline form to an amorphous form or vice versa. In still other memory devices, a change of physical state may involve quantum mechanical phenomena, such as superposition, entanglement, etc., which may involve, for example, quantum bits (qubits). The above is not intended to be an exhaustive list of all examples in which a change of state from a binary 1 to a binary 0 or vice versa in a memory device may involve a transformation, such as a physical but non-transient transformation. Rather, the foregoing is intended as illustrative examples.

[0091] FIG. 6 illustrates an exemplary sensor signal collection pattern for one embodiment of a vehicle 2200 that can operate in one or more automated driving modes (e.g., fully autonomous mode, semi-autonomous mode, driver-assisted mode). As illustrated, an automated vehicle such as vehicle 2200 can include several sensors to provide measurements and / or observations that are processed according to processes 350, 370, 450, 470, and / or 550, for example. While a particular pattern and / or number of sensors is illustrated in FIG. 6, the subject matter is not limited in scope in these respects. For example, a system or device such as vehicle 2200 can include any number of sensors in any of a wide range of arrangements and / or configurations. Also, while FIG. 6 is illustrated as a two-dimensional representation, sensors in a vehicle that can operate in one or more automated modes can generate signals and / or signal packets that represent, for example, conditions surrounding vehicle 2200 in three-dimensional space. In implementations, one-dimensional and / or two-dimensional sensor measurements can be combined and / or otherwise processed to, for example, generate a three-dimensional representation of conditions surrounding vehicle 2200.

[0092] In implementations, various sensors may be attached to vehicle 2200 to capture observations and / or measurements of different portions of the environment surrounding and / or adjacent to the vehicle, for example. In implementations, vehicle 2200 may include multiple different sensors capable of detecting input signals, such as optical, electromagnetic, and / or audio signals. Each sensor may have a different observation field of view into the environment surrounding vehicle 2200. While exemplary fields of view 2210a-2210h are shown, it should be understood that the subject matter is not limited in scope in these respects.

[0093] In implementations, the sensor signals and / or signal packets may be utilized by at least one processor of vehicle 2200, for example, to identify objects and / or other environmental conditions in the vicinity of vehicle 2200, which may be utilized by a processing system of vehicle 2200 to autonomously guide the vehicle through the environment. Example objects that may be detected in the environment surrounding a vehicle such as vehicle 2200 may include other vehicles, trucks, cyclists, pedestrians, animals, rocks, trees, lampposts, guardrails, painted lines, traffic lights, buildings, road signs, etc. Some objects may be stationary, while other objects, such as pedestrians, may move within the environment.

[0094] In one implementation, one or more sensors of the exemplary vehicle 2200 may generate signals and / or signal packets that may represent at least a portion of the environment surrounding and / or adjacent to the vehicle 2200. Other sensors may provide signals and / or signal packets that represent the speed, acceleration, orientation, position (e.g., via a Global Navigation Satellite System (GNSS)), etc. of the vehicle 2200. As described more fully below, the sensor signals and / or signal packets may be processed through a particle filter or the like to generate a plurality of particles. Such particles may be utilized, at least in part, to affect operation of the vehicle 2200. In an implementation, as the vehicle 2200, for example, progresses through an environment, the sensor signals and / or signal states may be utilized, at least in part, to update a particle filter, for example, to further affect operation of the vehicle. As described more fully below, any of several sorting operations may be performed on the sensor signals and / or signal packets in order for the particle filter or the like to utilize a relatively wide range of signals and / or signal packets generated by the various sensors.

[0095] In this context, a "particle" refers to a digital representation, derived at least in part from a sensor signal and / or signal packet, of environmental conditions for a particular point in a particular coordinate system for a particular point in time. For example, a particular particle may include an array of parameters that describe a particular point in the environment surrounding vehicle 2200 at a particular point in time. In implementations, a "particle filter" or the like may be utilized to process the sensor signals and / or signal packets to generate a plurality of particles that describe an environment, such as the environment surrounding vehicle 2200. Of course, a particle filter is merely an exemplary type of processing that may be performed on sensor signals and / or signal packets, and the subject matter is not limited in scope in this respect.

[0096] In implementations, a particular coordinate system may be specified, although the subject matter is not limited in scope to any particular coordinate system. In implementations, such a coordinate system may include a three-dimensional parameter space, although other embodiments may specify other numbers of dimensions. In implementations, individual particles may be associated with particular positions within a particular three-dimensional space (e.g., X, Y, and Z axes).

[0097] 7 shows an example schematic block diagram of an example vehicle 2200. As mentioned above, in implementations, the vehicle 2200 may include a number of sensors 2210. Such sensors may include, for example, image capture (e.g., cameras), radar, lidar, and / or ultrasonic sensors, to name a few non-limiting examples. In implementations, the sensors 2210 may generate sensor signals and / or signal packets 2215 that may be provided to and / or otherwise acquired by a control system, such as control system 2220.

[0098] In implementations, control system 2220 may include, for example, at least one processor, at least one memory device, and / or at least one communication interface. In implementations, control system 2220 may include, for example, one or more central processing units (CPUs), neural network processors (NNPs), and / or graphics processing units (GPUs). In implementations, control system 2220 may process sensor signals and / or signal packets to generate signals and / or signal packets that may affect operation of vehicle 2200. For example, signals and / or signal packets may be generated by control system 2220, provided to drive system 2230, and / or acquired by drive system 2230. In implementations, the processing of sensor signals and / or signal packets by control system 2220 may include, for example, a particle filter, although other implementations may utilize other signal processing algorithms, techniques, approaches, etc., and the subject matter is not limited in scope in this respect. In implementations, drive system 2230 may include, for example, devices, mechanisms, systems, etc. for affecting the operation of vehicle 2200. As described above, as vehicle 2200 traverses an environment, additional sensor signals and / or signal packets may be acquired and processed so that the operation of vehicle 2200 may be updated over time.

[0099] In the foregoing description, various aspects of the claimed subject matter have been described. For purposes of explanation, details have been set forth by way of example, such as quantities, systems, and / or configurations. In other instances, well-known features have been omitted and / or simplified so as not to obscure the claimed subject matter. While certain features have been illustrated and / or described herein, many modifications, substitutions, changes, and / or equivalents will occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to cover all modifications and / or variations that fall within the scope of the claimed subject matter.

Claims

1. 1. A system comprising: a plurality of computational nodes, each computational node of the plurality of computational nodes being a respective circuit adapted to perform a computational operation on data, the plurality of computational nodes including at least a first computational node and a second computational node; The first computing node of the plurality of computing nodes is receiving a plurality of signal streams from a plurality of sources, each signal stream including a respective set of data items; identifying data items from each set of data items that have a common attribute; concurrently loading data items received from two or more of the plurality of signal streams into one or more registers of the second compute node, the concurrently loaded data items being associated based on an attribute common to the concurrently loaded data items; and configured to the second compute node is configured to process the simultaneously loaded data items as operands of one or more computational operations; system.

2. a first transport memory buffer, the first transport memory buffer having a first endpoint and a second endpoint; the plurality of computing nodes are configured to communicate with an external memory external to the system via a bus; the first computing node comprises a register forming a first endpoint of the first transport memory buffer; a first register of the one or more registers of the second compute node forms the second endpoint of the first transport memory buffer; the first computing node is configured to load the data item into the one or more registers of the second computing node without storing the data item in the external memory. The system of claim 1 .

3. a second transport memory buffer, the second transport memory buffer having a first endpoint and a second endpoint; the first endpoint of the second transport memory buffer is formed by a register of a third compute node of a plurality of compute nodes; the second endpoint of the second transport memory buffer is formed by a second register of the one or more registers of the second compute node; the first computing node and the third computing node are configured to simultaneously load the data item into the one or more registers of the second computing node without storing the data item in the external memory. The system of claim 2 .

4. the data items in at least two of the plurality of signal streams include sensor observations and / or sensor measurements; the first computing node is configured to associate the sensor observations and / or measurements based at least in part on spatial and temporal attributes; the simultaneously loaded data items include the associated sensor observations and / or measurements; The system of claim 1 .

5. the first computing node is further configured to sort the sensor observations and / or measurements based at least in part on associated timestamps and locations of objects observed and / or measured by the sensor observations and / or measurements; the second computing node is configured to combine the sorted sensor observations and / or measurements. The system of claim 4.

6. Associating data items of a plurality of signal streams from a plurality of sources based at least in part on attributes common to said data items; concurrently loading associated data items from said plurality of signal streams into one or more registers of a compute node; executing said compute nodes to process said simultaneously loaded associated data items as operands of one or more computational operations; A method comprising:

7. and providing the results of processing the simultaneously loaded associated data items as operands to one or more computational instructions as data items for an additional signal stream. The method of claim 6 further comprising:

8. loading data items of said additional signal stream into one or more registers of a subsequent compute node as operands for one or more additional computational operations; The method of claim 7 further comprising:

9. performing one or more direct memory access transactions to store data items of the additional signal stream in external memory, performing word scatter operations, performing redirected write operations, or providing control signals to one or more actuators, or any combination thereof; The method of claim 7 further comprising:

10. Associating the data items of the plurality of signal streams comprises: Executing a direct memory access (DMA) controller to load a data item from each of the plurality of signal streams into a buffer associated with the signal stream; identifying at least one common attribute between the loaded data item and at least one other data item based at least in part on the content of the data item loaded into the buffer; The method of claim 6 further comprising:

11. Executing the DMA controller loading one or more addressable lines of values ​​and / or states stored in memory into said buffer; analyzing one or more non-addressable portions of at least one of the one or more addressable lines of the loaded value and / or state; processing one or more collection requests based at least in part on the analyzed one or more non-addressable portions; The method of claim 10, comprising:

12. Executing the DMA controller performing a first word gather operation based at least in part on the one or more redirected read requests; converting one or more words obtained from performing the first word gathering operation into one or more addresses; performing a second word gather operation to transfer data items located at the one or more addresses to a destination determined at least in part based on the one or more redirected read requests; and b.

13. Concurrently loading the associated data items from the multiple signal streams into the one or more registers of the compute node includes: loading data items of the plurality of signal streams into buffers associated with the plurality of signal streams; executing direct memory access (DMA) controllers associated with the plurality of signal streams to selectively load data items into the one or more registers based at least in part on an indication of a common attribute in the contents of the data items loaded into the buffer; and b.

14. the data items in at least two of the plurality of signal streams include sensor observations and / or measurements; the sensor observations and / or measurements in the at least two of the plurality of signal streams are associated based at least in part on spatial and temporal attributes; the simultaneously loaded associated data items include sensor observations and / or measurements associated at least in part based on spatial and temporal attributes; The method of claim 6.

15. Associating the data items of the multiple signal streams from multiple sources comprises: sorting, at a previous computational node, the sensor observations and / or measurements based at least in part on associated timestamps and locations of objects observed and / or measured by the sensor observations and / or measurements; Combining the sorted sensor observations and / or measurements at the computational node and b.

16. executing the computational nodes to update a state of a particle filter based at least in part on the simultaneously loaded associated data items; The method of claim 14 further comprising:

17. the plurality of sources comprising at least a first sensor integrated with the vehicle and a second sensor external to the vehicle; The method of claim 6.

18. 1. A system comprising: a plurality of sensors for generating a plurality of associated signal streams; a plurality of computational nodes coupled to the plurality of sensors; At least a first computing node of the plurality of computing nodes concurrently loading data items from two or more of the associated plurality of signal streams originating from two or more of the sensors into one or more registers of a second computational node, the concurrently loaded data items being associated based at least in part on an attribute common to the concurrently loaded data items; and the second compute node is configured to process the simultaneously loaded data items as operands of one or more computational operations; system.

19. the data items in at least two of the associated signal streams include sensor observations and / or measurements; the first computing node is configured to associate the sensor observations and / or measurements based at least in part on spatial and temporal attributes; the simultaneously loaded data items include the associated sensor observations and / or measurements; 20. The system of claim 18.

20. the second computing node is further configured to update a state of a particle filter based at least in part on the simultaneously loaded data items.

20. The system of claim 18.

Citation Information

Patent Citations

  • Inter-processor transfer system for message information

    JP1989072255A

  • Inter-processor communication equipment

    JP1989240963A

  • Data management method, information processor, and program

    JP2014071495A

  • System parameter identification device, system parameter identification method, and computer program therefor

    JP2017083922A