Systems, devices, and / or methods for processing signal streams
Patent Information
- Application Number
- JP2025539630
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-01-04
- Filing Date
- 2023-12-19
- Publication Date
- 2026-09-14
- Estimated Expiration
- 2043-12-19
Smart Images

Figure 0007920463000004 
Figure 0007920463000005 
Figure 0007920463000006
Abstract
Description
TECHNICAL FIELD
[0001] The subject matter disclosed herein relates to processing signals received in streams from a plurality of data sources. BACKGROUND ART
[0002] Self-propelled driving and / or autonomous driving applications, as well as other automotive and robotics applications, may rely on the fusion of signals, measurements and / or observations generated by a plurality of sensors. Processing for such applications may involve manipulation of arrays of data elements. Such applications may be executed and / or implemented by commercially available central processing units (CPUs) and / or graphics processing units (GPUs). Such commercially available processing units may be configured to manipulate elements of an input array to generate elements of an output array. SUMMARY OF THE INVENTION
[0003] One embodiment disclosed herein includes a plurality of compute nodes and a transport memory for transporting data and / or values, and is directed to a system The transport memory includes a transport memory buffer. , wherein each compute node of the plurality of compute nodes is a respective circuit adapted to perform computational operations on data, and the plurality of compute nodes includes at least a first compute node and a second compute node. The first compute node among the plurality of compute nodes is configured to , receiving multiple signal streams from multiple sources, wherein the multiple signal streams are include a respective set of data items Furthermore, the data items in at least two of the plurality of signal streams include sensor observations and / or measured values. ; The aforementioned identify data items having a common attribute from among the respective set of data items; and Associating the sensor observations and / or measurements based at least partially on spatial and temporal attributes, and of the plurality of signal streams 2 received from one or more The identified data items , of the transport memory buffer, of the second compute node The aforementioned into one or more registers , in the same execution cycle load simultaneously and the above the simultaneously loaded data items are is, associated based on the common attribute and the above rela The system is configured to include linked sensor observations and / or measurements, and to sort the sensor observations and / or measurements at least partially based on the relevant timestamps and locations of the objects observed and / or measured by the sensor observations and / or measurements.The second computing node is, The aforementioned Simultaneously load data items and process them as operands for one or more calculation operations. This involves combining the sorted sensor observations and / or measured values, It may be configured to do so.
[0004] In one particular implementation, the system further comprises transport memory for transporting data and / or values in the form of a first transport memory buffer, the first transport memory buffer having a first endpoint and a second endpoint, and a plurality of compute nodes are configured to communicate with external memory located outside the system via a bus, the first compute node having a register that forms the first endpoint of the first transport memory buffer, the first register of one or more registers of a second compute node forming the second endpoint of the first transport memory buffer, and the first compute node is configured to load data items into one or more registers of the second compute node without storing the data items in external memory. The system may also include a second transport memory buffer having a first endpoint and a second endpoint, the first endpoint of the second transport memory buffer being formed by a register of a third compute node among a plurality of compute nodes, and the second endpoint of the second transport memory buffer being formed by a second register among one or more registers of the second compute node, and the first and third compute nodes being configured to load data items simultaneously into one or more registers of the second compute node without storing the data items in external memory.
[0005] In another specific implementation, data items in at least two of multiple signal streams include sensor observations and / or measurements, and a first compute node is configured to associate the sensor observations and / or measurements based at least partially on spatial and temporal attributes, with the simultaneously loaded data items including the associated sensor observations and / or measurements. In one example, the first compute node is further configured to sort the sensor observations and / or measurements based at least partially on the relevant timestamps and locations of the objects observed and / or measured by the sensor observations and / or measurements, and a second compute node is further configured to combine the sorted sensor observations and / or measurements to provide combined sensor observations and / or measurements.
[0006] Another embodiment disclosed herein relates to a method comprising: associating data items from multiple signal streams from multiple sources, at least in part, on attributes common to the data items; simultaneously loading the associated data items (associated data items) from the multiple signal streams into one or more registers of a compute node; and executing the compute node to process the simultaneously loaded associated data items as operands for one or more compute operations. In one particular implementation, the method further includes providing the results of processing the simultaneously loaded associated data items as operands to one or more compute instructions as data items from additional signal streams. In one example, the method may further include loading data items from additional signal streams into one or more registers of a subsequent compute node as operands for one or more additional compute operations. In another example, the method further includes executing one or more direct memory access transactions to store the additional signal stream data items in external memory, perform a word distribution operation, perform a redirected write operation, or provide control signals to one or more actuators, or a combination thereof.
[0007] In other specific implementations, associating data items from multiple signal streams further includes running a direct memory access (DMA) controller to load data items from each of the multiple signal streams into buffers associated with the signal streams, and identifying at least one common attribute between the loaded data items and at least one other data item, at least in part, based on the content of the data items loaded into the buffers. For example, running a DMA controller may further include loading one or more addressable lines of values and / or states stored in memory into a buffer, parsing one or more unaddressable portions of at least one of the one or more addressable lines of values and / or states loaded, and processing one or more retrieval requests, at least in part, based on the parsed one or more unaddressable portions. In another example, the execution of a DMA controller may further include: performing a first word collection operation at least partially based on one or more redirected read requests; translating one or more words obtained from the first word collection operation to one or more addresses; and performing a second word collection operation to transfer data items located at one or more addresses to destinations determined at least partially based on one or more redirected read requests.
[0008] In another specific implementation, simultaneously loading data items associated with one or more registers of a compute node from multiple signal streams includes loading data items from multiple signal streams into buffers associated with multiple signal streams, and executing a direct memory access (DMA) controller associated with multiple signal streams to selectively load the data items into one or more registers, at least in part, based on indications of common attributes in the content of the data items loaded into the buffers. In yet another specific implementation, the multiple sources include at least a first sensor integrated with the vehicle and a second sensor outside the vehicle.
[0009] In yet another specific implementation, data items in at least two of multiple signal streams include sensor observations and / or measurements, and the sensor observations and / or measurements in at least two of the multiple signal streams are associated at least partially on spatial and temporal attributes, and the associated data items loaded simultaneously include sensor observations and / or measurements associated at least partially on spatial and temporal attributes. In one example, associating data items from multiple signal streams from multiple sources includes, in a previous compute node, sorting the sensor observations and / or measurements at least partially on the relevant timestamps and locations of the objects observed and / or measured by the sensor observations and / or measurements, and in a compute node, combining the sorted sensor observations and / or measurements in at least two or more signal streams to provide combined sensor observations and / or measurements. In another example, the method further includes running a compute node to update the state of a particle filter at least partially on the associated data items loaded simultaneously.
[0010] Another embodiment disclosed herein relates to a system comprising a plurality of sensors for generating a plurality of associated signal streams, and a plurality of compute nodes coupled to the plurality of sensors, wherein at least a first compute node of the plurality of compute nodes is configurable to simultaneously load data items occurring at two or more of the sensors from two or more of the plurality of associated signal streams into one or more registers of a second compute node, wherein the simultaneously loaded data items are associated at least in part on attributes common to the simultaneously loaded data items, and the second compute node is configured to process the simultaneously loaded data items as operands of one or more compute operations. In one particular implementation, the data items in at least two of the plurality of associated signal streams include sensor observations and / or measurements, and the first compute node is configured to associate the sensor observations and / or measurements at least in part on spatial and temporal attributes, and the simultaneously loaded data items include associated sensor observations and / or measurements. In another specific implementation, a second computing node is further configured to update the particle filter state, at least partially, based on concurrently loaded data items.
[0011] The subject matter to be claimed is specifically pointed out and explicitly claimed in the concluding section of this specification. However, both the configuration and / or method of operation, along with their purpose, features, and / or advantages, can be best understood by referring to the following detailed description when read together with the accompanying drawings. [Brief explanation of the drawing]
[0012] [Figure 1] This is a schematic diagram of a computing device according to an embodiment. [Figure 2A] This is a schematic diagram of a computing device comprising a direct memory access (DMA) controller and / or engine including a buffer according to one embodiment. [Figure 2B]This is a schematic diagram of a computing device comprising a direct memory access (DMA) controller and / or engine including a buffer according to one embodiment. [Figure 3A] This is a schematic diagram of a computing device comprising a DMA controller and / or engine including a buffer for facilitating distributed data acquisition operations according to one embodiment. [Figure 3B] This is a flowchart of a process for facilitating distributed operation and collection operation according to one embodiment. [Figure 3C] This is a schematic diagram showing the non-addressable portion of an addressable line according to one embodiment. [Figure 3D] This is a flowchart of a process for facilitating distributed operation and collection operation according to one embodiment. [Figure 3E] This is a schematic diagram of a computing device comprising a DMA controller and / or engine including a buffer for facilitating distributed data acquisition operations according to one embodiment. [Figure 4A] This is a schematic diagram of a computing device equipped with a DMA controller that includes a buffer for facilitating the redirection of DMA transactions according to one embodiment. [Figure 4B] This is a flowchart of a process for facilitating the redirection of a DMA transaction according to one embodiment. [Figure 4C] This is a flowchart of a process for facilitating the redirection of a DMA transaction according to one embodiment. [Figure 4D] This is a schematic diagram of a computing device equipped with a DMA controller that includes a buffer for facilitating the redirection of DMA transactions according to one embodiment. [Figure 5A] This is a schematic diagram of a computing device for facilitating the processing of multiple signal streams from multiple associated sources according to one embodiment. [Figure 5B] This is a flowchart illustrating a method for processing multiple data streams from multiple associated sources according to one embodiment. [Figure 5C] It is a schematic diagram of a computing device for facilitating processing of a plurality of signal streams from a plurality of associated sources according to an embodiment. [Figure 6] It is a diagram illustrating exemplary sensor signal collection in an exemplary vehicle according to an embodiment. [Figure 7] An exemplary schematic block diagram of exemplary vehicle features in a self-driving / autonomous driving application according to an embodiment is shown. DETAILED DESCRIPTION OF EMBODIMENTS
[0013] In the following detailed description, reference is made to the accompanying drawings which form a part hereof, and in which like reference numerals may refer to corresponding and / or similar like parts throughout. It will be understood that the drawings are not necessarily drawn to scale, for example, for brevity and / or clarity of illustration. For example, dimensions of some aspects may be exaggerated relative to other aspects. Further, it is to be understood that other embodiments may be utilized. Furthermore, structural and / or other changes may be made without departing from the scope of the claimed subject matter. Throughout this specification, reference to "the claimed subject matter" refers to subject matter that is intended to be covered by one or more claims, or any portion thereof, and is not necessarily intended to refer to a complete set of claims, a particular combination of sets of claims (e.g., method claims, apparatus claims, etc.), or a particular claim. It should also be noted that, for example, directions and / or references such as top, bottom, topmost, and bottommost may be used to facilitate description of the drawings and are not intended to limit the application of the claimed subject matter. Accordingly, the following detailed description is not to be construed as limiting the claimed subject matter and / or equivalents.
[0014] Throughout this specification, references to a particular implementation, a given implementation, an embodiment, a given embodiment, etc., mean that the specific features, structures, characteristics, etc. described in relation to a particular implementation and / or embodiment are included in at least one implementation and / or embodiment of the claimed subject matter. Therefore, for example, the appearance of such phrases in various places throughout this specification is not necessarily intended to refer to the same implementation and / or embodiment, or any one particular implementation and / or embodiment. Furthermore, it should be understood that the specific features, structures, characteristics, and / or similar described can be combined in various ways in one or more implementations and / or embodiments, and are therefore within the scope of the intended claims. Generally, and naturally as is always the case with patent application specifications, these and other issues may vary in the context of a particular use. In other words, throughout this disclosure, the specific context of description and / or use provides useful guidance regarding reasonable inferences to be drawn. However, similarly, “in this context” generally refers, without further limitation, to at least the context of this patent application.
[0015] According to one embodiment, multiple signal streams may be processed using multiple compute nodes, each configured to perform specific operations on data items. In one implementation, a first compute node among the multiple compute nodes may be configured to receive multiple signal streams from multiple sources (e.g., sensors) and identify from each set of data items in different received signal streams that have common attributes. The first compute node may then simultaneously load the data items received from two or more of the multiple signal streams that have the common identified attributes into one or more registers of a second compute node. The second compute node may be configured to process the loaded data items as operands for one or more compute operations.
[0016] In some examples, a host central processing unit (CPU) may be part of the computing device within a vehicle and may process data items in memory for automotive applications. For example, automotive applications such as autonomous driving / autonomous driving applications (e.g., fully autonomous, semi-autonomous, driver assistance systems, etc.) may employ, for example, particle filters to merge signal streams of sensor signals and / or observations to update, for example, the particle filter state. Such applications may be implemented in systems such as automated machinery (cars, trucks, etc.). In this context, “signal stream” as referred to herein means the time-varying progression of a sequence of encoded data items delivered to a receiving device via a signal transmission medium. Encoded data items (also called data values) delivered in a signal stream may represent some attributes, such as attributes indicating conditions and / or events, object identifiers, timestamps indicating the time of an event, and metadata. In certain implementations, a signal stream may deliver sensor measurements and / or observations combined with relevant timestamps indicating the time when such measurements and / or observations were acquired.
[0017] In one aspect of one embodiment, information originating from different sources and arriving in their respective signal streams may be processed as an information “confluence” to determine a computation result. Such an information confluence may be processed by relating and / or correlating information items from different sources by specific attributes (e.g., time, space, reliability, confidence, etc.). Processing the information confluence may then involve performing one or more actions on the items based on one or more attributes to produce a result. In certain implementations, data items within the information confluence may be processed by updating one or more states of a particle filter. For example, such a particle filter may implement processing of an array of sensor signals / observations (e.g., received from a signal stream) to update the states of measured particles, filtered particles, static particles, and / or dynamic particles. In one implementation, the measured and / or observed values in the array of measured values may be generated from different sensors. Nevertheless, such measured and / or observed values generated by different sensors may be related and / or correlated by time and space. According to one embodiment, associating data items from different sources within an array confluence can be implemented using, at least in part, a radix sort applied to the key array. In a particular implementation, sensor signals and / or observations can be formatted into an array to be processed for generating the array confluence. An exemplary procedure for generating such an array confluence can be performed according to the following pseudocode. [Table 1] TIFF0007920463000002.tif155150
[0018] According to one embodiment, the merging of multiple signal streams may involve mapping data items in different input signal streams to data items in one or more output signal streams. For example, such data items in input signal streams may include sensor measurements and / or observations of sensors associated with the input signal streams. Therefore, such merging of input signal streams may involve, (For example, from different / separate sensors associated with the input signal stream) Sensor measurements and / or observations added to the data items of the output signal stream. value Mapping may be included. Such data items in the output signal stream may include values inferred / calculated based on sensor measurements and / or observations. In the pseudocode example provided above, the number of input signal streams may be defined as S_in[] (containing the data item value_in[]), and the number of output signal streams may be defined as S_out[] (containing the data item value_out[]). Here, the expression "(value_out[],S_out_enable[])=f(value_in[],parameters)for(i:0..(l-1))" may map the data item value_in[] in the input signal stream S_in[] to the data item value_out[] in the output signal stream S_out[] according to the function f().
[0019] According to one embodiment, computing devices, circuits, and / or logic can form a "Confluence Engine" (CE), also called a "Confluencer" and / or "Confluence Processor" (CP), to process the merging of data items as described above. Such merging of data items may include, for example, the merging of signal streams and / or sequences of signal streams having reduced latency and / or reduced computing resources (e.g., power, memory, etc.). In certain implementation forms, the output of a merging operation of such a CE may provide all or part of the input to a subsequent merging operation. Merging characteristics, such as functions like k, f(), and read_next() as shown in the pseudocode example above, may be part of the runtime programming of the CE.
[0020] In one embodiment, the CE may employ a direct memory access (DMA) subsystem, which may include a DMA controller (also called a DMA "engine") configurable to initiate read and write operations between line-accessible memory and transport memory. The specific processes through which the CE accesses memory may be determined parameterically at compile time and physically at runtime. Thus, in certain implementations, events that trigger the execution of the DMA controller may not be limited to events occurring in the arithmetic logic unit (ALU) (e.g., from loads and stores in the ALU). The DMA controller may be triggered to execute a DMA transaction by the initiation of a merge operation. The DMA controller may then execute such a DMA transaction independently of the arithmetic logic unit (ALU) (e.g., depending only on the availability of read and write access at the valid endpoint of such a DMA transaction).
[0021] Figure 1 is a schematic diagram of a system 100 that performs DMA transactions. The system 100 includes several components that communicate via a bus 101. The components include a host CPU 102, a memory controller 106, RAM 108 (which may form the main system memory), peripheral devices 114, and a DMA controller 112. In this context, “direct memory access” as used herein means a process performed by one or more hardware subsystems and / or circuits to access a specific memory independently of a particular processing unit and / or central processing unit (for example, independently of the host CPU 102). According to one embodiment, the DMA controller 112 may initiate an operation to access (for example, read access or write access) the random access memory (RAM) 108 via the bus 101 independently of the host CPU 102. For example, the DMA controller and / or engine 112 may control and / or execute transactions to transfer data items (also called data values) between the peripheral device 114 and the RAM 108 (through the memory controller 106), independently of the actions of the host CPU 102. Such DMA transactions may be triggered, for example, by signals, conditions, and / or events (e.g., interrupt signals).
[0022] As described above, the processing of data items for automotive applications or other computing applications can be enhanced through a merging engine (CE). Figures 2A and 2B are schematic diagrams of a computing device 200 comprising a direct memory access (DMA) controller 212 (also called a DMA engine 212) and a buffer 216 for implementing one or more embodiments of a CE according to one embodiment. The computing device 200 may be, for example, a system-on-a-chip (SoC), a microchip, a control circuit, or any other computing device, or may form them. In some implementations, the computing device 200 may form a vehicle controller such as an Advanced Driver Assistance System (ADAS) device, a telematics control unit (TCU), an electronic control unit (ECU), a centralized vehicle computer, or any other vehicle controller. As shown in Figures 2A and 2B, the computing device 200 may further include a CPU 202 (also called the host CPU), transport memory 218, line-accessible memory 208, buffer 216, and a pool of compute nodes (CNs) 220 (also called the CN pool 220).
[0023] In one implementation, the CN may comprise a single processing circuit core capable of performing operations to map input operands to output calculation results. In another implementation, the CN may comprise multiple separate processing cores for performing operations to map input operands to output calculation results. In yet another implementation, two or more non-concurrent CNs may be implemented on the same processing circuit core. For example, a processing circuit may implement a first CN to generate an output result (stored, for example, in transport memory 218) that will become input to a second CN implemented on the same processing circuit that will be executed later.
[0024] A CN in the CN pool 220 may have dedicated local memory (e.g., static random access memory (SRAM)) and general-purpose registers to receive operands for operations to be performed and / or to provide results from the execution of operations. In this example, the host CPU 202 may use transport memory 218 to store data items and control the CNs in the CN pool 220 to perform operations on these data items. Transport memory 218 may be physically closer to the CNs and / or operate with lower access latency and may therefore be used as a cache for storing data items. Line-accessible memory 208 may be outside the transport memory 218 and may provide a larger amount of memory space to the transport memory 218, but may be physically farther from the CN pool 220 and may operate with longer access latency. According to one embodiment, transport memory 218 may have one or more synchronization mechanisms to facilitate inter-CN communication between CNs in the CN pool 220 (e.g., for synchronizing communication between CNs having different execution latencies). In certain implementations, line-accessible memory 208 may be isolated from transport memory 218, CPU 202, and buffer 216 by a bus (not shown). According to one embodiment, buffer 216 may be formed in the circuitry for implementing the core circuitry of the DMA controller and / or engine 212 so that buffer 216 is separate from and isolated from the circuitry for forming transport memory 218. Such formation of buffer 216 in the core circuitry of the DMA controller and / or engine 212 may reduce and / or minimize the latency associated with loading data items into and storing data items from buffer 216 during the process of performing DMA operations.
[0025] In one embodiment, the DMA controller and / or engine 212 may be configured to interface with transport memory 218 and line-accessible memory 208. Transport memory 218 and / or line-accessible memory 208 may be cache line-addressable. In other words, line-accessible memory 208 in this example may be cache line-addressable memory. Transport memory 218 and / or line-accessible memory 208 can provide data items (also called data values) that CNs in the CN pool 220 can operate on, operate on, or otherwise process. According to one embodiment, all or part of transport memory 218 may be configured as a cache that can be integrated with commercially available components. It should also be noted that a cache in either transport memory 218 or buffer 216 may mitigate manufacturing defects and / or allow the use of embodiments larger than the expected application size.
[0026] In one implementation, the DMA controller and / or engine 212 may be configured to process (handle) cache line-sized data items (e.g., 64 bytes or 128 bytes), and such data items are locatable and accessible by cache line addresses, even in distributed-collection operations. Such a cache line address for a 64-byte cache line may be represented by a binary notation ending in six zeros. Similarly, a cache line address for a 128-byte cache line may be represented by a binary notation ending in seven zeros. According to one embodiment, the DMA controller and / or engine 212 may be configured to process (handle) word-sized data items, even while the line-accessible memory 208 may remain addressable only in cache lines. To facilitate distributed or collection operations for transferring word-sized data items to and from the line-accessible memory 208, the DMA controller and / or engine 212 may implement a buffer 216 (e.g., located between the line-accessible memory 208 and the transport memory 218). Buffer 216 may be configured to store together bytes from multiple cache lines that are located in the destination word in transport memory 218. The DMA controller and / or engine 212 may also be capable of performing multicast write operations in either direction (e.g., from transport memory 218 to line-accessible memory 208, or from line-accessible memory 208 to transport memory 218). Buffer 216 may be separate from / different from transport memory 218. As described herein, buffer 216 may be formed within the client core to implement the DMA controller and / or engine 212.
[0027] In this context, “transport memory” as used herein means circuitry for facilitating communication of data items between CNs, such as CNs in a pool of CNs 220. In one particular implementation, such transport memory may transport the result of the execution of a first operation in a first CN to become an input operand for a second operation performed in a second CN (e.g., in a compute pipeline). In a particular implementation, transport memory 218 may be configured as a static random access memory (SRAM) device used as access control memory, or as shared memory as cache memory, a word-addressable or SIMD vector-addressable register file, or as circuitry and / or devices specifically structured to function as a first-in, first-out (FIFO) buffer. Such circuitry and / or devices specifically structured to function as a FIFO buffer may have widths that are, to give a few examples, word-wide or single-instruction, multiple-data (SIMD) vector-wide or network-on-chip (NOC) devices coupled between endpoints, or a combination thereof.
[0028] As described above, CNs in the CN pool 220 (e.g., a pool of CNs) can operate on data items read from memory. In this context, “computation node” as used herein means a set of identifiable and distinct computing resources (e.g., hardware and executable instructions) configurable to perform operations to process input values and provide output values. CNs in the CN pool 220 may comprise scalar CNs and / or processing circuit cores to implement arithmetic logic units (ALUs), digital signal processors (DSPs), vector CNs, VLIW engines, or field-programmable gate array (FPGA) cells, or combinations thereof, to give some examples of specific circuit cores that may be used to implement CNs in the CN pool 220. The CN pool 220 can be implemented according to various architectures. For example, the CN pool 220 may comprise one or more CNs implemented in full-featured or simplified forms according to a reduced instruction set computing (RISC) architecture, a composite instruction set computing (CISC) or very long instruction word (VLIW) architecture, or some combination of these types. The CNs in the CN pool 220 may include scalar, SIMD, or combinations of multiple instruction and single data stream (MISD) ALUs.
[0029] According to one embodiment, the functionality of the computing device 200 may include, for example, commercially available functions such as transport rings, cross-networks (clos networks), and shuffle circuits. The CN pool 220 may facilitate, to name a few, multithreading in the CNs, clustering of the CNs, mailboxes, interrupts, functions for synchronization between CNs, atomic operation at locations in the transport memory 218, and additional functions to meet safety and security requirements.
[0030] In certain implementations, the clustering of CNs within the CN pool 220 may be formed at least partially on a resource trade-off and may be permanently defined in the integrated circuit (IC) device. In certain embodiments, depending on how the CN pool 220 is configured, two of the same integrated circuit (IC) devices may implement different clusterings of related CNs. For example, a processor configuration may define the processing of multiple merges with CN clustering, at least partially on the related merges being performed. Each related cluster may process array merges such that, for example, the output of one cluster provides input to one or more other clusters.
[0031] According to one embodiment, the transport memory 218 may form one or more buffers, each containing one or more first-in, first-out (FIFO) buffers. One or more FIFO buffers may comprise a “vertical” FIFO buffer. Such a vertical FIFO buffer may have an endpoint that interfaces with the DMA controller and / or engine 212 (such endpoint may be called an “outside” endpoint) and another endpoint that forms a register capable of providing data items available as operands for CNs in the CN pool 220 (such endpoint may be called an “inside” endpoint). The inside endpoint of the vertical FIFO buffer may be shared by multiple CNs in the CN pool 220. If such an inside endpoint includes, for example, the FIFO out-end (i.e., the outbound end of the FIFO buffer), broadcasting to multiple CNs in the CN pool 220 can be achieved. If such an inside endpoint includes the FIFO in-end (i.e., the inbound end of the FIFO buffer), hardware locks and / or instructions executed on CNs in the CN pool 220 can prevent race conditions. In some implementations, the FIFO buffer formed within the transport memory 218 may also include a “horizontal” FIFO buffer. A horizontal FIFO buffer may have two endpoints, one providing data items as operands for different CNs in the CN pool 220, or the other simply providing data items at locations within the transport memory 218. It should be noted that the FIFO buffer formed within the transport memory 218 can provide a transparent block and deallocation mechanism when input and output speeds differ. If the outer endpoint of the FIFO buffer in the transport memory 218 is a register or operand consumed / processed by several CNs in the CN pool 220, and the CNs are slow to do so, the DMA controller and / or engine 212 may ultimately perform a block mechanism.Conversely, if the DMA controller and / or engine 212 is slow to write to such FIFO buffers in transport memory 218, the CNs in the CN pool 220 may eventually perform a blocking mechanism. If FIFO buffers are implemented between CNs and / or between locations in transport memory 218, a similar transparent block and deallocation mechanism may exist.
[0032] In an IC device, if the transport memory 218 forms a FIFO buffer, the circuitry at the endpoints of the FIFO buffer may be permanent or configurable (e.g., via internal FPGA circuitry). Certain implementations may include segments of vertical and horizontal FIFO buffers formed in the IC device circuitry. Such buffer circuitry may have endpoints that are configurable at runtime from at least one of the following: operands or registers of CNs in the pool of CNs 220, interfaces to buffer 216, locations in the transport memory 218, and endpoints of other FIFO buffers in the transport memory 218. If the endpoint of the FIFO buffer of the transport memory 218 is an operand or register of a CN in the pool of CNs 220, such FIFO buffer endpoints may be shared among multiple CNs in the pool of CNs 220.
[0033] According to one embodiment, CNs in the CN pool 220 may receive operands for computational operations from, for example, registers in a register file, from the FIFO buffer of the transport memory 218, and / or from special constant and parameter registers (which may include shared registers). In some examples, a portion of the register file may be shared among multiple CNs in the CN pool 220. Within the computing device 200, the host CPU 202 may also have access to constant and parameter registers. The host CPU 202 may provide host functions, for example, initiating a merge operation that is performed after completing a configuration task. Such configuration tasks may include, to name a few, defining clusters of CNs in the CN pool 220, configuring inter-CN communication, configuring the width and depth of the FIFO buffer in the transport memory 218, configuring the endpoints of the FIFO buffer in the transport memory 218, defining manager CNs in the CN pool 220 and CN clusters managed by manager CNs in the CN pool 220, setting up the DMA controller and / or engine 212, setting up atomics, setting up communication and synchronization resources, and monitoring the completion of merges.
[0034] In one embodiment, the buffer 216 can facilitate the merging process in several ways. In one such non-limiting embodiment, the input signal stream in the merging may comprise an "indirection stream" (e.g., a set of addresses in the line-accessible memory 208 to be read). In such a case, such an indirection signal stream does not have to be supplied directly to the CN in the pool of CNs 220, but instead may be supplied to the DMA controller and / or engine 212, which may then transport the signal stream of the read data items to the CN. In other words, the latency of access by random read operations (random line reads, or random word reads, or random indirect reads of lines or words, or a combination thereof) can be masked by extracting patterns of random access, constructing access lists of considerable length, and performing the necessary extraction and indirection in the DMA controller and / or engine 212 itself. In other cases, the input signal stream may include a "double indirection stream" in which the signal stream contains addresses, and the data items at these addresses provide further indirection after some processing. In certain implementations, the first level of indirection in a dual indirection stream may be read and used multiple times. In sensor fusion applications, multiple mergers may be configured, and the initial merger may be constructed (for example, within the transport memory 218 itself) using a lookup table (LUT) for the data items of the first level of indirection and the data items of the second level of indirection. However, it should be understood that these examples are not limiting.
[0035] According to one embodiment, some operations within the confluence of signal streams (for example, within f() and read_next() in the pseudocode example above) can be considered as extracting information from data items distributed across CNs in the CN pool 220. An example of such extraction is shown in Table 1 below. [Table 2]
[0036] According to one embodiment, the results may be provided by CNs in a pool of CNs 220 that operate cooperatively using operations such as shifting, shuffling, broadcasting, and multicasting operands between participating CNs, to give a few examples. Such features may exist within a CN, for example, for the word operands of SIMD vector CNs in the pool of CNs 220. The pool of CNs 220 may be configured to have such features for inter-CN communication via a configurable bridge circuit between CNs. Such a circuit may be hardwired or configurable at runtime.
[0037] According to one embodiment, CNs in a pool of CNs 220 (for example, configured in a CN cluster) may be configured for special processing functions. In certain implementations, such special CNs in the pool of CNs 220 may facilitate the management of the application's processing flow. For example, the pool of CNs 220 may include one or more processing CNs 224 and one or more manager CNs 222. A processing CN 224 may perform processing operations on, for example, sensor observations, measurements, and / or other signals. In this example, a manager CN 222 may manage different sets of processing CNs 224. A processing CN 224 may communicate with one or more manager CNs 222. In some examples, a manager CN 222 may provide information to a processing CN 224 based on communication from the processing CN 224. The processing CNs 224 may, as qualified, continue their processing based on the information provided by the manager CNs 222. In some implementations, a manager CN 222 may consist only of a physically separate processing CN 224, or dedicated circuitry formed within a processing CN 224.
[0038] It should be noted that the rates at which different individual signal streams are generated and consumed (e.g., processed) at a merging operation (e.g., from different sources such as different sensors) may not necessarily be equal. The rate at which such signal streams are consumed or generated may be determined by the operational characteristics of a function such as f() (in the pseudocode example above), for example, the rate at which the signal stream of sensor measurements and / or observations for sensor fusion operation can be calculated / generated by function f(). Such sensor fusion operation may involve an inverse sensor model that results in outputting a signal stream longer than the input signal stream. To handle the merging of longer signal streams, for example, the output rate may be matched to the bandwidth / throughput of line-accessible memory 208. Configuring the output of one CN to become the input to another CN may assist, for example, load balancing.
[0039] According to one embodiment, the computing device 200 can enable the deployment of advanced sensor fusion operations to update the particle filter state with minimal power consumption (e.g., in autonomous driving or other automotive applications). In one application, an instance of the computing device 200 may be implemented, for example, as a sequence of pipeline stages. In such an application, the features of the computing device 200 may be configured in use to have different amounts of resources allocated to different pipeline stages. The exchange of data items between the pipeline stages and the line-accessible memory 208 can be synchronized transparently (e.g., without mutexes, spinlocks, etc.) by using a FIFO buffer (e.g., a FIFO buffer formed in the transport memory 218).
[0040] In another embodiment, the computing device 200 may be configured to optimize power, space, and / or performance (e.g., accuracy and / or latency). While the features of the computing device 200 may be adapted to implement a CE, they may also be adapted to other applications, such as those relying on random access to line-accessible memory 208. The features of the computing device 200 may also be implemented in so-called "supercomputers." In the context of supercomputers, the low-power features of the computing device may help overcome power constraints that might hinder the realization of, for example, exascale supercomputers. Furthermore, the circuitry for implementing the computing device 200 may incorporate safety and security features to meet the requirements of an embedded computing device. Due to the small physical size and low power consumption of the computing device 200, its use is not necessarily limited to use as an external accelerator integrated circuit (IC) device, but may be integrated into a subsystem within an automotive-grade system-on-chip (SOC) IC device.
[0041] In one embodiment, for example, the computing device 200 may comprise a specific arrangement and / or configuration of a CN (e.g., a pool of CNs 220), transport memory 218, and / or a DMA controller and / or engine 212. The computing device 200 may be configured to provide a network of CNs and memory elements adapted to a specific type of computation, such as a specific application for processing signal streams (e.g., carrying measured values and / or observations from sensors). In one implementation, such a network of CNs and memory elements may enable simultaneous processing of multiple signal streams with high throughput and low latency. Such a network of CNs and memory elements may be implemented, at least in part, using physical connections and an implementation of an intra-device communication protocol (e.g., AXI) via a FIFO buffer. As noted above, the endpoints of the FIFO buffer may include, for example, an addressable memory location or register for receiving operands (e.g., a general-purpose register of an ALU configured as a compute node as an operand for a compute operation) or the results computed by a CN. The pool of FIFO buffers may comprise, for example, endpoints configurable to be associated with various CNs or memories.
[0042] In one embodiment, certain embodiments disclosed herein relate to so-called vectorized input / output (I / O) operations, including “distributed” and “collective” operations. Such vectorized operations can enable high-throughput transfer of large amounts of data into or out of physical memory (e.g., multiple addressable lines of memory in line-accessible memory 208) using a single request or command, for example, to improve efficiency and convenience. For example, a collection operation may include (or may involve) sequentially reading data from multiple memory locations (e.g., buffers) and writing the read data to a signal stream or a contiguous portion of memory in a single transaction. In one implementation, the DMA controller and / or engine 212 may perform a collection operation to service (process, provide) a collection request (e.g., occurring in an application) that specifies multiple memory locations, not necessarily line-aligned, from which data items are read, and a destination (e.g., a memory address) for storing the read items. On the other hand, distributed operations may include (or may involve) reading data items from a signal stream or contiguous memory and writing the read data items to multiple different memory locations that are not necessarily line-aligned. In one implementation, the DMA controller and / or engine 312 may perform collection operations to service distributed requests (e.g., occurring in an application) that can specify the location (e.g., a contiguous memory address) of the data item to be read. Such distributed requests may also indirectly specify the location to which the read data item will be written (by requesting a read of a specific location that allows determining where the read data item will be written). The DMA controller and / or engine 312 may similarly perform collection operations that are indirectly specified.
[0043] According to one embodiment, the use of DMA transactions to assist in processing data items in a signal stream can be enhanced through the use of distributed and collect operations. In certain implementations, content loaded into a buffer from a collect operation may be used to determine one or more addresses for subsequent collect or distributed operations. Figures 3A and 3E are schematic diagrams of a computing device 300 according to one embodiment, comprising a DMA controller and / or engine 312 that includes or communicates with a buffer 316 to facilitate distributed and / or collect operations. In certain implementations, the computing device 300 may include one or more features of the computing device 200 (Figures 2A and 2B). The DMA controller and / or engine 312 may be configured as a distributed collect multicast DMA engine (SGM-DMA), which, for example, has the ability to address words in an addressable line.
[0044] According to one embodiment, the DMA controller and / or engine 312 may receive an input signal stream and / or supply such an input signal stream as blocks on a virtual channel through a FIFO buffer to an initial cluster of CNs. The output signal stream from the initial cluster of CNs may then be supplied to subsequent downstream clusters of CNs as blocks on a virtual channel or as data through a FIFO buffer. The transfer between clusters of CNs may be, for example, a multicast transfer directed by a particular application. At any stage, some or all of the output signal streams from several CN clusters may be returned to the DMA controller and / or engine 312 (for example, in the process of the DMA performing a distribution or collection operation).
[0045] According to one embodiment, the DMA controller and / or engine 312 may perform certain distributed and / or retrieval operations to transfer data items from one non-contiguous memory block to another by using a series of smaller contiguous block transfers. Here, retrieving such data items from a non-contiguous block of source memory may be performed in a retrieval operation. Similarly, writing data items to a non-contiguous block of destination memory may be performed in a distributed operation. In one implementation, the smallest unit of memory that can be accessed in such source or destination memory may be a single addressable line of values and / or states (e.g., a single cache line or word-in-line accessible memory). For example, the DMA controller and / or engine 312 may communicate with line-accessible memory (LAM) 308, which may be line-accessible.
[0046] According to one embodiment, physical memory (such as LAM308) may include bit cells for defining values and / or states to represent information such as 1 or 0. Such physical memory may further configure the bit cells into words (e.g., 4-byte words over 32 bits or 8-byte words over 64 bits) containing 8-bit integers. Furthermore, such physical memory may define line addresses (e.g., word line addresses) associated with contiguous bits that define "addressable lines" of values and / or states. For example, in response to a read or write request (e.g., originating from a host processor), the memory controller may access a portion of memory in a targeted read or write transaction according to the word line address specified in the request. To service a read request, for example, the memory controller may retrieve the values and / or states of all bytes in the line associated with the line address specified in the read request. Similarly, to service a write request, the memory controller may write the values and / or states of all bytes in the addressable line associated with the line address specified in the write request. A line address can specify a memory location that includes all contiguous bytes of an addressable line; however, such a line address does not specify the location of individual bytes or contiguous bytes less than the entire addressable line, or bytes that span the addressable line, or any other individual sub-parts of such an addressable line. Such sub-parts of an addressable line in any other way are referred to herein as “unaddressable parts.”
[0047] According to one embodiment, an addressable line can define the smallest unit of memory that can be located and / or accessed according to a memory addressing scheme. In a particular exemplary embodiment of Figure 3C, such an addressable line 360 may be made up of smaller memory units such as bits, bytes, or words. In a particular illustrated embodiment, the addressable line 360 is made up of n+1 bytes 3620-362 n This includes. As noted above, certain implementations may target updates to unaddressable portions and / or portions less than (smaller than) the entire addressable line and / or unaddressable portions that span multiple lines, with or without an addressable line. In the illustrated exemplary embodiment, bytes 3622 and 3623 may define an unaddressable portion of addressable line 360, where the entire addressable line may be locatable and / or accessible via a unique address according to the memory addressing scheme, but bytes 3622 and 3623 (less than the entire line) are not addressable according to the memory addressing scheme.
[0048] In one implementation, the LAM308 may have a LAM controller 306 configured to receive requests specifying the address of a line of data items stored in the LAM308. Such a line of data items may contain multiple words or multiple bytes and may be the smallest unit of data items that the LAM308 can retrieve and return to other devices (that the LAM308 can retrieve and return to another device). According to one embodiment, the computing device 300 may use a buffer 316 to enable access to smaller, unaddressable portions of a line (e.g., a word or a group of bytes within a word). According to one embodiment, the circuit forming the buffer 316 may be integrated with the circuit forming the DMA controller and / or engine 312 to enable minimum latency for access to the buffer 316 initiated by the DMA controller and / or engine 312. For example, buffer 316 may be configured as a static random access memory (SRAM) device accessible by the DMA controller and / or engine 312 without initiating a request and / or transaction on the main memory bus (e.g., a bus coupled to the LAM308 or a host computer / processor).
[0049] According to one embodiment, the DMA controller and / or engine 312 may communicate with an initiator 322. The initiator 322 may include devices that achieve a specific state for triggering one or more DMA transactions performed by the DMA controller (e.g., at least partially implemented by circuitry and / or logic). For example, the initiator 322 may include, to name a few examples of devices that can initiate a DMA transaction, an ALU output register, a buffer, or a hardware interrupt handler.
[0050] According to one embodiment, the DMA controller and / or engine 312 may obtain a list of collection requests in response to a signal from the initiator 322. In one particular implementation, the initiator 322 may trigger a DMA transaction in response to an event or condition in the execution of the particle filter process. For example, the particle filter process may identify data items in memory that are expected to be retrieved for processing in a future execution cycle. Once a substantial amount of such data items is identified, a list of collection requests identifying such data items (e.g., as redirected collection requests) may be forwarded to the DMA controller and / or engine 312. In one implementation, such a list of collection requests may be provided to the DMA controller and / or engine 312 in shared memory or network-on-chip (NOC), to name a few examples. Once it is known that such a list is available for processing by the DMA controller and / or engine 312, a process for generating the list (e.g., execution of a computer-readable instruction) may trigger the DMA controller and / or engine 312 via an interrupt or post-message. Such a trigger can start the DMA controller and / or engine 312 to initiate one or more acquisition operations (e.g., redirected acquisition operations).
[0051] In response to a signal from initiator 322, the DMA controller and / or engine 312 may obtain a list of acquisition requests in the form of a linked list. Such a linked list may be locatable in memory (e.g., reconfiguration buffer 316 or line-accessible memory (LAM) 318) according to addresses provided, for example, by initiator 322. According to one embodiment, such a list of acquisition requests may include individual acquisition requests that can be served as standalone acquisition requests, independently of other acquisition requests in the list of acquisition requests. The DMA controller and / or engine 312 may combine addresses in such acquisition requests with a (potentially smaller) list of line read requests executed by a memory controller (e.g., memory controller 106 in Figure 1). Each line read request in such a list of line read requests may indicate a physical memory address (e.g., in LAM 308) that specifies the memory location that the memory controller reads to serve the individual line read request. Such a memory controller may serve such line read requests by loading the requested line into buffer 316. When a requested line read by the memory controller arrives in buffer 316, the DMA controller and / or engine 312 may refer to the original list of collection requests (used to form a list of line read requests) to extract the requested data item from the read line arriving in buffer 316. The DMA controller and / or engine 312 may form a packet from the extracted items to be forwarded to one or more requesting entities (e.g., processes running on the CN of the host CPU 202 and / or CN pool 220). In one embodiment, the data item extracted from the line stored in buffer 316 may include two or more unaddressable portions of the line (e.g., selected bytes and / or fields in an addressable line / word, such as an addressable line in memory containing the data item).
[0052] Figure 3B shows a flowchart of a process 350 for an acquisition operation according to one aspect of the present disclosure. In one embodiment, process 350 may include operations 352, 354, 356, and 358 which can be performed by one or more circuits, such as a DMA controller and / or engine 312 and / or buffer 316. Operation 352 may include processing one or more acquisition requests to determine one or more addressable lines of data items to fetch from memory. For example, operation 352 may map parameters in the received request to addresses in LAM 308. Operation 354 may be implemented by a circuit to load values and / or states into memory and may include loading signals and / or states (e.g., representing data items) of one or more addressable lines stored in memory, such as LAM 308, into buffer 316. As shown in Figure 3B, operation 356 may include parsing one or more unaddressable portions of the line loaded into buffer 316 in operation 354 into portions such as a word or a group of bytes. Operation 358 may include completing the processing of two or more collection requests by returning, for example, the unaddressable portions parsed in operation 356.
[0053] As shown in Figure 3B, operation 358 may include responding to multiple acquisition requests presented to the DMA controller and / or engine 312 in a list of acquisition requests. Loading one or more addressable lines into buffer 316 in operation 354 may occur in response to the list of acquisition requests. Operation 354 may parse two or more unaddressable portions according to the data items specified in the list of acquisition requests. For example, operation 356 may decode parameters in an acquisition request that will be mapped from line addresses to byte offsets to the corresponding unaddressable portions. The parsed portions may then be forwarded to the initiator of the acquisition request (e.g., initiator 322). Here, process 350 may enable serving multiple acquisition requests using access to a single addressable line loaded into buffer 316, for example, if multiple acquisition requests specify parsed portions within the same single addressable line. This eliminates the need for the DMA controller and / or engine 312 to access the same addressable line multiple times for separate collection requests for data items within the same addressable line (e.g., LAM308).
[0054] To serve one or more collect requests, process 350 may collect less than the entire addressable line in memory by loading the addressable line into a buffer and parsing the unaddressable portion provided to the requester. Some collect requests may request the collection of less than the entire addressable line, while one or more received collect requests may request the collection of an entire addressable line and / or multiple lines and / or bytes spanning a line. In the case of a collect request requesting a collection of less than the entire addressable line, the DMA controller and / or engine 312 may execute process 350. According to one embodiment, in the case of a collect request requesting the collection of an entire addressable line, the DMA controller and / or engine 312 may bypass operations 354, 356, and 358 and perform the collect operation without loading the addressable line into buffer 316.
[0055] In another implementation, the DMA controller and / or engine 312 may obtain a list of distributed requests in response to a signal from the initiator 322. The DMA controller and / or engine 312 may obtain such a list of distributed requests in the form of a memory-locatable linked list, according to the addresses provided by the initiator 322. The DMA controller and / or engine 312 may then combine the addresses accessed by such distributed requests with a (potentially smaller) list of line read requests performed by a memory controller (e.g., memory controller 106). For example, distributed requests in a list of distributed requests that refer to data items on the same addressable line in memory may be combined such that only the requested single line read is required (to access the data item in order to serve multiple distributed requests). The requested line read by such a memory controller may be loaded into buffer 316.
[0056] According to one embodiment, the acquired list of distributed requests may indicate a specific unaddressable portion (e.g., individual bytes or fields) of an addressable line to be read and loaded into buffer 316. The unaddressable portion in the addressable line loaded into buffer 316 may then be modified and / or overwritten. When the requested line read by the memory controller arrives in buffer 316, the DMA controller and / or engine 312 can refer to the original list of distributed requests to determine the specific unaddressable portion of the read line that arrived in buffer 316 to be modified and / or overwritten. The DMA controller and / or engine 312 may form a packet from the modified line in buffer 316 and write it back to memory via the memory controller.
[0057] Figure 3D shows a flowchart of process 370 for distributed operation according to one embodiment. Process 370 may include operations 372, 374, 376, and 378. Operation 372 may include, for example, receiving one or more distributed requests from initiator 322. Operation 374 may include loading signals and / or states (e.g., representing data items) of one or more addressable lines in memory, such as LAM 308, into buffer 316. In a particular implementation, the DMA controller and / or engine 312 may determine which addressable lines are fetched from LAM 308 in operation 374 by processing one or more retrieval requests. Operation 376 may include writing values and / or states to at least one unaddressable portion of the lines loaded into buffer 316 in operation 374 in order to at least partially modify one or more addressable lines stored in buffer 316. For example, operation 374 may write values and / or states to two or more unaddressable portions based on multiple distributed requests received in block 372. Operation 378 may complete servicing (processing, providing) such distributed requests by initiating a write operation to write back at least one of the one or more addressable lines modified in operation 374 to memory (e.g., LAM308). For example, operation 378 may write values and / or states to two or more unaddressable portions based on multiple distributed requests received in block 372. Ta A Dress selection available line to LAM308 The process of writing back to the line address may begin.
[0058] In certain implementations, operation 374 may be initiated by a distributed request received in operation 372. Here, one or more addressable lines of values and / or states loaded into buffer 316 in operation 374 may be acquired by a memory controller (e.g., memory controller 106) servicing one or more line read requests. Operation 378 may include initiating the memory controller to perform one or more operations to write one or more modified addressable lines in order to write the modified addressable lines. Here, process 370 may be able to serve multiple distributed requests using access to a single addressable line loaded into buffer 316. Multiple distributed requests received in operation 372 may, for example, specify words, bytes, fields, etc., within the same single addressable line. This eliminates the need for the DMA controller and / or engine 312 to access the same addressable line multiple times for separate distributed requests for data items within the same addressable line (e.g., in LAM 308). Process 370 may further include translating the multiple distributed requests received in operation 372 into a list of line read requests issued to the memory controller (a list of line read requests specifying one or more addressable lines of value and / or state). Such a translation of the multiple distributed requests into a list of line read requests may further include constructing at least one single line read request for an addressable line in memory containing the data items requested by at least two distributed requests received in operation 372.
[0059] To serve one or more distributed requests, process 370 may update less than the entire addressable line in memory by loading the addressable line into buffer 316, updating some of the loaded addressable line, and keeping others unchanged. Some distributed requests may request updates less than the entire addressable line, but one or more received distributed requests may request updates to the entire addressable line, which the DMA simply writes to rather than performing a read-modify-write operation. In the case of distributed requests requesting updates less than the entire addressable line, the DMA controller and / or engine 312 may execute process 370. The DMA controller and / or engine 312 may also be configured to serve distributed requests requesting updates to the entire addressable line by bypassing loading the addressable line into buffer 316. Here, to complete such an update to the entire addressable line, the DMA controller and / or engine 312 may initiate a write operation to the addressable line in LAM 308 without loading the addressable line into buffer 316.
[0060] According to one embodiment, the DMA controller and / or engine 312 may receive multiple distributed requests that collectively request updates to the same overlapping portion of an addressable line in the LAM 308. This can cause a conflict over how to update the overlapping portion to supply the multiple distributed requests. Such multiple distributed requests may be ordered, for example, according to creation time or reception time. According to one embodiment, a conflict over updating a portion of an addressable line by multiple distributed requests may be resolved, for example, according to the most recently created or received distributed request.
[0061] In specific implementations of processes 350 and 370, the unaddressable portion of a line stored in buffer 316 may be a byte, a set of bytes or fields, etc. Although the specific actions in the distributed and collected operations described above are described as occurring in a specific sequence, some actions may be performed concurrently, and some of these actions may be performed concurrently or in a specific sequence, as a matter of engineering choice. Furthermore, physical optimizations such as the number and type of processing cores used to realize the features of the DMA controller and / or engine 312 and the interface engine, the number and type of memory blocks for the associated memory elements and buffer 316, and the number of memory ports and addressability features forming the LAM 308 may be selected as a matter of engineering choice. For example, buffer 316 may or may not be byte-addressable, and the DMA controller and / or engine 312 may comprise, for example, a scalar or a vector engine.
[0062] Figures 4A and 4D are schematic diagrams of a computing device 400 comprising a DMA controller and / or engine 412 that communicates with a buffer 416 and an initiator 422 to facilitate the redirection of DMA transactions, according to one embodiment. In a particular implementation, the computing device 400 may include one or more features of the computing device 200 (Figures 2A and 2B). According to one embodiment, the DMA controller and / or engine 412 may obtain a list of redirected read requests in response to a signal from the initiator 422. According to one embodiment, a read request may include a message and / or signal specifying one or more target memory addresses to be accessed in a read transaction to service the read request. For example, a read request may specify one or more target addresses as wordline addresses of locations in memory containing content to be retrieved in a memory read transaction to service the read request. As used herein, “redirected read request” means a read request that has been transformed or modified so that the content(s) at the original target memory address(s) is modified and / or translated to a different memory address(s). Here, different memory addresses(s) specify locations in memory containing content to be retrieved in a read transaction to service a redirected read request. The DMA controller and / or engine 412 will interpret the addresses specified in such a redirected read request as a word collection request, for example, to store the collected word in buffer 416. The size of the associated word initially read to service the word collection request portion of the redirected read request may depend, at least in part, on how the associated word is interpreted to form the target address for the redirection. In one implementation, the word to be read may include, for example, an address or an index of an array that can be translated to an address.In one implementation, the DMA controller and / or engine 412 may then translate the collected word stored in buffer 416 into an address and merge the address (i.e., the translated collected word) with a redirected read request to form a collect request. The DMA controller and / or engine 412 may then service the formed collect request to send the resulting data items to a destination, as described in the redirected read request acquired in response to a signal from initiator 422.
[0063] Figure 4B is a flowchart of a process 450 for facilitating the redirection of a DMA transaction according to one embodiment. Process 450 may include operations 452, 454, 456, and 458. Operation 454 may include performing a collect operation (e.g., a word collect operation) at least in part on one or more redirected read requests received in operation 452. In this operation, the DMA controller and / or engine 412 may acquire or receive such redirected read requests, for example, in response to a signal from the initiator 422. In certain implementations, the DMA controller and / or engine 412 may interpret one or more redirected read requests as requests for the first collect operation of two collect operations to be performed. In the course of the first collect operation performed in operation 454, “collected” data items may be loaded into buffer 416.
[0064] Operation 456 may include translating one or more collected data items stored in buffer 416 to one or more addresses to specify a subsequent collection operation. For example, operation 456 may include translating one or more words obtained from the execution of a first word collection operation (performed in operation 454) to one or more addresses (also called one or more address values). In one particular implementation, operation 456 may include parsing the values and / or states in buffer 416 (from the collection operations) to determine one or more memory addresses in LAM 408. Operation 456 may further include applying one or more arithmetic operations to the parsed values and / or states to determine one or more memory addresses in LAM 408. For example, operation 456 may apply one or more arithmetic operations to the parsed values and / or states stored in the buffer, which form memory addresses to memory locations in LAM 408. Such formed addresses to memory locations in LAM 408 may form the basis for a subsequent collection operation.
[0065] According to one embodiment, the arithmetic operation applied in operation 456 may be defined according to formula (1) as follows: address=base+x×element_size (1) During the ceremony, The address is the target address (for example, for the collection operation determined in operation 456, or for the distribution operation in operation 476), x is a value obtained from a collection operation (for example, in operation 454 or 474), and base and element_size are parameters provided in a redirected request (for example, a redirected read request received in operation 452, or a redirected write request received in operation 472).
[0066] Operation 458 may include performing a second (e.g., subsequent) collection operation to transfer data items located at one or more determined addresses to a destination. Such a destination may be determined at least in part on one or more redirected read requests. In another implementation, operation 458 may perform two or more collection operations based on one or more addresses obtained in operation 456. According to one embodiment, when performing a collection operation, operation 458 may interpret one or more redirected read requests as two requests for word collection operations.
[0067] According to one embodiment, a write request may include a message and / or signals specifying one or more target memory addresses to be accessed in a memory write transaction to service the write request. For example, a write request may specify one or more target addresses as wordline addresses of locations in memory to be written in a memory write transaction (to service the write request). As used herein, a “redirected write request” means a write request that has been transformed or modified so that the original target memory address(s) is modified and / or translated to a different target memory address(s), where the different target memory address(s) specify locations in memory to be written in a write transaction to service the redirected write request.
[0068] In another specific implementation, the DMA controller and / or engine 412 may, in response to a signal from the initiator 422, obtain a list of redirected write requests. The DMA controller and / or engine 412 may interpret the addresses specified in such redirected write requests as collection requests. Such collection requests may, for example, involve loading collected words into buffer 416. The associated words read out may have a size that depends at least in part on how the associated words are interpreted to form a target address for redirection. One or more such words loaded into buffer 416 may be interpreted to form a target address for redirection. In one implementation, such words loaded into buffer 416 may include, for example, an address or an index of an array that can be translated to an address. In one implementation, the DMA controller and / or engine 412 may then translate the collected words stored in buffer 416 into addresses and merge the addresses with the redirected write requests to form a distributed request. The DMA controller and / or engine 412 may then service the formed distributed request and bring about line updates in accordance with redirected write requests obtained in response to signals from the initiator 422.
[0069] Figure 4C is a flowchart of a process 470 for facilitating the redirection of a DMA transaction according to one embodiment. Operation 472 may include performing a word collection operation based at least in part on one or more redirected write requests received in operation 472. The DMA controller and / or engine 412 may receive such redirected write requests in operation 472, for example, in response to a signal from initiator 422. In certain implementations, the DMA controller and / or engine 412 may interpret one or more redirected write requests as requests for a word collection operation to be performed, followed by a word distribution operation. In the course of the word collection operation performed in operation 474, “collected” data items may be loaded into buffer 416. Operation 476 may include translating one or more collected data items stored in buffer 416 (from the collection operation) to one or more addresses to specify a subsequent word distribution operation. In one particular implementation, operation 476 may include parsing values and / or states in buffer 416 to determine one or more memory addresses in LAM 408. For example, operation 476 may be performed at least in part by circuitry (e.g., a DMA controller and / or engine adapted to parsing values and / or states) that parses values and / or states to determine at least one of the one or more addresses. For example, operation 476 may apply one or more arithmetic operations to the parsed values and / or states stored in the buffer for forming memory addresses in LAM 408. Operation 476 may further include applying one or more arithmetic operations to the parsed values and / or states to determine one or more memory addresses to memory locations in LAM 408.
[0070] Operation 478 may include performing a distribution to write a specific data item to one or more addresses obtained in operation 476, at least in part on one or more redirected read requests. According to one embodiment, operation 476 may apply arithmetic operations to compute the target address of the distributed operation performed in operation 478 according to equation (1). For example, a DMA controller and / or engine 412 may form such a distributed request, at least in part on the content at one or more addresses determined in operation 476. Such a specific data item to be written from such a distributed request may be specified, for example, in one or more redirected write requests received in operation 472 in response to a signal from initiator 422. In one particular implementation, operation 478 may include interpreting the content at the addresses determined in operation 476 as addresses for the distributed operation.
[0071] One particular implementation of a computing device for processing multiple signal streams from multiple related sources is shown by computing device 500 in Figures 5A and 5C. In a particular implementation, computing device 500 may include one or more features of computing device 200 (Figures 2A and 2B). Computing device 500 may comprise a pool 520 of compute nodes (CNs) having related register files and local memory, a pool of block memory and FIFO buffers, and a NOC for interconnecting various components. As noted above, the CNs in the pool 520 may comprise scalar CNs and / or processing circuit cores for implementing ALUs, digital signal processors (DSPs), vector CNs, VLIW engines, or field-programmable gate array (FPGA) cells, or combinations thereof, to give some examples of specific circuit cores that may be used to implement the CNs in the pool 520. In one implementation, a CN may comprise a single processing circuit core capable of performing operations for mapping input operands to output computed results. In another implementation, the CN may comprise multiple separate processing cores for performing operations to map input operands to output computation results. The CNs within the pool of CNs 520 may comprise dedicated local memory (e.g., SRAM) and general-purpose registers for receiving operands for operations to be performed and / or providing results from the execution of operations. In one particular implementation, the CNs within the pool of CNs 520 comprise elements of the pool of memory 508 (e.g., transport memory 218), such as FIFO buffers. and Embodied / in It may be formed from one or more configured processing circuit cores. In one embodiment, the memory pool 508 is C N pool 520 It can provide shared memory resources between CNs within the C N pool 520 This facilitates communication between internal CNs. For example, integrating with CN pool 520 / inThe endpoints of the configured FIFO buffer may be local to the CN or include memory blocks isolated from the CN, and / or the register file of the CN in CN520. The functions of the CNs in the CN pool 520 may, to name a few, be the functions of a host processor, a processing manager, or a signal processor, or a combination thereof. According to one embodiment, C N pool 520 The CN inside is C N pool 520 To facilitate signal communication between internal CNs, atomic, interrupt, and other inter-processor communication (IPC) functions may be supported.
[0072] According to one embodiment, the computing device 500 can receive a signal stream containing data items (e.g., sensor signals, observed and / or measured values, timestamps, metadata, etc.) supplied from an external source such as a sensor and / or memory. The external memory (not shown) may be coupled with a DMA controller and / or engine (e.g., DMA controller and / or engine 312 and / or 412). CNs in the CN pool 520 can also provide a source for a signal stream containing data items. For example, a CN in the CN pool 520 can supply data items in a signal stream that are loaded into a FIFO buffer. A CN in the CN pool 520 may also supply data items in a signal stream by transporting packets of output parameters in an IC device in the NOC as blocks of data items. According to one embodiment, a single logical signal stream may be transmitted in a multicast manner to multiple CNs in the CN pool 520. Alternatively, a single logical signal stream may be segmented into sub-signal streams supplied to a subset of CNs in the CN pool 520. In another embodiment, several CNs in the CN pool 520 may process data items from various signal streams and provide output signal streams as input streams to other CNs in the CN pool 520. The final output of the CNs in the CN pool 520 may include, to name a few, signal streams output from a computing device 500 to a sink such as an actuator, memory, storage, or display device. The processing between CNs in the CN pool 520 may be controlled and / or organized via a combination of interrupts, status flag polling, and periodic checks of work, to name a few.
[0073] Figure 5B is a flowchart of a method 500 (also called process 500) for processing multiple signal streams from multiple related sources according to one embodiment. According to one embodiment, process 500 can combine data items received in signal streams from multiple sources to update the state of a particle filter (for example, to support autonomous driving of a vehicle). Such multiple sources may include, for example, multiple sensor devices (e.g., mounted in a vehicle), such as cameras, speedometers, active sensing devices (e.g., radar / lidar), environmental sensors (e.g., thermometers, light sensors, altimeters, etc.), and microphones. Such signal streams from multiple sources may be received in signal packets from a communication network (e.g., signal packets containing sensor measurements and / or observations acquired from a remote vehicle). In a particular implementation, the data items received in the signal streams from multiple sources may include sensor measurements and / or observations having common attributes that can be identified by process 500. In certain implementations, such common attributes may be identifiable by co-located metadata associated with measurements and / or observations received from multiple sources. For example, such common attributes may be associated with time (e.g., a timestamp indicating the time of measurement and / or observation), space (e.g., a location relative to an origin, such as a point on a vehicle), and source (e.g., any particular physical sensor on a vehicle, or any particular type of sensor among sensors mounted on a vehicle).
[0074] Operation 552 may include associating data items received from multiple data streams based at least partially on attributes common to the data items. Such multiple data streams may be provided as outputs of CN and / or DMA collection or word collection or redirected read / collection operations. For example, operation 552 may sort and / or correlate measured and / or observed values received from different signal streams by time (e.g., according to timestamps) and / or space (e.g., the position of the observed object relative to a reference point) (e.g., "bucketize"). In certain implementations, operation 552 may associate measured and / or observed values received from different sources (e.g., different sensors) that are acquired almost simultaneously with the position of particles defined in the current state of the particle filter. In one particular embodiment, operation 552 may associate measured and / or observed values from different sources according to the locality of a particular object observed and / or measured by such associated measured and / or observed values. In another specific implementation, data items from each of several signal streams may be loaded into a buffer (e.g., buffer 216) associated with a direct memory access DMA controller associated with the signal stream. Operation 552 may then identify at least one of the common attributes associated with the data items loaded into the buffer, based at least partially on the contents of the data items loaded into the buffer.
[0075] Operation 554 may include simultaneously loading the data items associated in operation 552 into one or more registers of the compute node (e.g., general-purpose registers of the ALU or other processing cores forming the compute node) without storing them in line-accessible memory. According to one embodiment, the registers of the compute node may be loaded with data items to be retrieved in the compute node's execution cycle. For example, data items loaded into the compute node's registers in one execution cycle (e.g., at the endpoint of a FIFO buffer) may provide operands for compute operations performed by the compute node in the next execution cycle. In one particular implementation, one or more registers of the compute node may include endpoints of associated FIFO buffers formed by internal memory (e.g., a memory pool 508). Multiple such FIFO buffers having endpoints in the compute node's registers may be synchronized to apply data items from multiple sources (e.g., data items loaded from different sensors) as operands of the compute node having common attributes (e.g., temporal and spatial attributes). Operation 556 may include the execution of a compute node to process the data items loaded concurrently in operation 554 and to process the loaded data items as operands for one or more compute operations (for example, to perform one or more functions such as updating the state of a particle filter). The data items output from the execution of one or more compute operations in operation 556 may form data items for additional signal streams processed by additional compute nodes and / or for storage in memory.
[0076] In one embodiment, operation 552 may be performed by a first compute node that sorts the sensor observations and / or measurements received in multiple signal streams, at least partially based on the associated timestamp and locality of the object observed and / or measured by the sensor observations and / or measurements. The sorted sensor observations and / or measurements can then be loaded into one or more registers of a second compute node in block 554. The execution of the second node in block 556 can then combine the sorted sensor observations and / or measurements.
[0077] In another embodiment, the data items of the additional signal stream as the output of operation 556 may be loaded into one or more registers of a subsequent compute node as operands for one or more additional compute operations. For example, one or more direct memory access transactions may be performed to store the data items of the additional signal stream in external memory, to perform a word distribution operation to write the data items of the additional signal stream, to perform a redirected write operation to write the data items of the additional signal stream, or to provide the data items of the additional signal stream as control signals to one or more actuators, or a combination thereof.
[0078] In this context, “simultaneously load” as used herein means the loading of data items processed by the compute node in the same execution cycle. When data items from synchronized signal streams are simultaneously loaded into the compute node’s registers, such data items may be loaded into the registers in the same execution cycle of the compute node (e.g., to become operands of a compute operation performed in the next execution cycle). When data items from each asynchronous signal stream are simultaneously loaded into the compute node’s registers, such data items may be loaded into the registers in different (e.g., adjacent) execution cycles of the compute node. For example, the execution of the compute node may be paused for one or more execution cycles to allow multiple data items from different asynchronous signal streams to be loaded into the registers as operands of a compute operation in the compute node’s execution cycle. In another embodiment, a transparent block and release mechanism can be applied to the compute node to facilitate the simultaneous loading of data items from asynchronous signal streams into the compute node’s registers.
[0079] According to one embodiment, data items in at least two associated signal streams in operation 552 may include associated sensor observations and / or measurements in operation 554. Data items loaded simultaneously in operation 554 may then include associated sensor observations from multiple signal streams (e.g., from multiple different sources). In a particular implementation, operation 552 may include combining sensor observations and / or measurements from at least two signal streams to provide combined sensor observations and / or measurements. Operation 552 may then sort and / or correlate (e.g., bucket) the combined sensor observations and / or measurements based at least partially on the associated timestamps and locations of the objects observed and / or measured by the sensor observations and / or measurements.
[0080] In certain implementations, operation 554 may be implemented at least partially using process 350 (Figure 3B) and / or process 450 (Figure 4B). For example, a DMA controller and / or engine may perform an acquisition operation as described in process 350 and / or 450, filling queues maintained by FIFO buffers with (e.g., associated measured and / or observed values) and loading registers of a compute node. Such an execution by the DMA controller and / or engine may selectively load data items into one or more registers, at least partially based on indications of common attributes in the content of the data items loaded into the FIFO buffers. In one implementation, multiple FIFO buffers may have endpoints to corresponding registers of a compute node, and queues of different FIFO buffers may be filled with data items from different signal streams (e.g., measured and / or observed values from different sensors). The queues of multiple FIFO buffers can be filled so that data items from different signal streams associated by specific attributes (e.g., temporal and spatial attributes) are simultaneously loaded into the endpoints of the FIFO buffers (e.g., into general-purpose registers of the compute node). According to one embodiment, operations 552 and 554 may be performed by the first compute node in one execution cycle of the second compute node to simultaneously load the associated data items into one or more registers of the second compute node. In subsequent execution cycles of the second compute node, the second compute node can perform compute operations using the associated data items simultaneously loaded in block 554 as operands. In one implementation, the simultaneous loading of data items in block 554 can be facilitated by a FIFO buffer having an endpoint in a register of the second compute node, where the first compute node can fill the queue of the FIFO buffers with associated data items from different signal streams so that each associated data item arrives at the endpoint of the FIFO buffer in the same or adjacent execution cycles.By filling the FIFO buffer queue with related data items in this way, the first compute node no longer needs to store (storage) the related data items in line-accessible memory (e.g., line-accessible memory 208 or RAM 108).
[0081] In another specific implementation, the result of the computation node's execution in operation 556 may provide data items for one or more additional signal streams. In one implementation, such data items for additional signal streams may be loaded into one or more registers of subsequent computation nodes and / or downstream computation nodes. In one example, data items in the computation node's output registers (e.g., loaded from the execution of a computation operation) may be transferred to line-accessible memory (e.g., according to process 470 in Figure 4C) by DMA write transactions, word-distributed DMA transactions, and / or redirect-distributed DMA transactions. Such DMA transactions may, for example, store data items for additional signal streams in line-accessible memory. In another example, such data items provided in the computation node's output registers may be applied as data items for input signal streams to one or more other computation nodes. In one specific implementation, the result from the execution of the first computation node in operation 556 may be loaded (e.g., from memory pool 508) into an output register defined as the first endpoint of the associated FIFO buffer. The associated FIFO buffer may then have a second endpoint defined as an input register for the second compute node to receive the result determined by the execution of the first compute node as an operand.
[0082] According to one embodiment, operations 552 and 554 may be performed by a first CN of the CN pool 520, and operation 556 may be performed by a second CN of the CN pool 520. In operation 554, the first compute node may simultaneously load related data items (e.g., data items containing sensor observations and / or measurements associated at least partially on spatial and temporal attributes, where concurrently loaded data items include associated sensor observations and / or measurements) into one or more registers of the second CN of the CN pool 520. The second CN may then, in operation 556, process the data items concurrently loaded by the first CN as operands to one or more compute operations. In one implementation, the first FIFO buffer may define a first endpoint as a register of the first CN (e.g., an output register of the first CN). The second endpoint of the first FIFO buffer may define a first register (e.g., an input register of the second CN) among one or more registers of the second CN. As can be observed, the first FIFO may, as noted above, allow data items to be loaded simultaneously into one or more registers of the second CN in operation 554 without storing the associated data items in line-accessible memory. In another implementation, the FIFO buffer may define the first endpoint as a register of the third CN in the pool of CNs 520, and the second endpoint as at least the second register of one or more registers of the second CN. Here, both the first and third CNs may, in operation 554, load associated data items simultaneously into one or more registers of the second CN node without storing the associated data items in line-accessible memory.
[0083] According to one embodiment, all or part of computing devices 200, 300 (including features that implement processes 350 and / or 370, such as circuits for forming a DMA controller), 400 (including features that implement processes 450 and / or 470, such as circuits for forming a DMA controller), and / or 500 (including features that implement process 550, for example) may be formed by transistors and / or lower metal interconnects (not shown) in a process (e.g., a front-end-of-line and / or back-end-of-line process), such as a process for forming a complementary metal-oxide-semiconductor (CMOS) circuit, and / or represented in whole or in part therein. However, it should be understood that this is only one example of how circuits may be formed within a device in a front-end-of-line process, and the claimed subject matter is not limited to this.
[0084] It should be noted that the various circuits disclosed herein are described using computer-aided design tools and their behavior, register transfers, logic components, transistors, layout geometry, and / or other characteristics may be represented (or expressed) as data and / or computer-readable instructions embodied in various computer-readable media (e.g., non-temporary storage media). The formats of files and other objects in which such circuit representations may be implemented (e.g., in circuit devices) include, but are not limited to, formats supporting behavioral languages such as C, Verilog, and Very High Speed Integrated Circuit Hardware Description Language (VHDL), formats supporting register-level description languages such as Register Transfer Language (RTL), formats supporting geometry description languages such as Graphic Design System II (GDSII), Graphic Design System III (GDSIII), Graphic Design System IV (GDSIV), Caltech Intermediate Format (CIF), Manufacturing Electron Beam Exposure System (MEBES), and any other suitable formats and languages. The storage medium in which such formatted data and / or instructions may be embodied may include, but are not limited to, various forms of non-volatile storage mediums (e.g., optical storage mediums, magnetic storage mediums, or semiconductor storage mediums) and carrier waves that may be used to transfer such formatted data and / or instructions through wireless, optical, or wired signaling media or any combination thereof. Examples of the transfer of such formatted data and / or instructions by carrier waves may include, but are not limited to, transfers over the Internet and / or other computer networks via one or more electronic communication protocols (e.g., Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Simple Mail Transfer Protocol (SMTP), etc.) (uploads, downloads, emails, etc.).
[0085] When received within a computer system via one or more machine-readable media, such data and / or instruction-based representations of the above-mentioned circuits may be processed by processing entities within the computer system (e.g., one or more processors) in conjunction with the execution of one or more other computer programs, including, but not limited to, netlist generation programs, arrangement and routing programs, in order to generate representations or images of the physical representations of such circuits. Such representations or images may then be used in device manufacturing, for example, by enabling the generation of one or more masks used to form various components of the circuit in a device manufacturing process (e.g., a wafer manufacturing process).
[0086] In the context of this patent application, the term “between” and / or similar terms are understood to include “among” where appropriate for the particular use, and vice versa. Similarly, in the context of this patent application, the terms “compatible with,” “comply with,” and / or similar terms are understood to include substantial compatibility and / or substantial compliance, respectively.
[0087] Unless otherwise specified, in the context of this patent application, the term “or” used to associate a list such as A, B, or C is intended to mean A, B, and C in an inclusive sense, as well as A, B, or C in an exclusive sense. In this understanding, “and” is intended to mean A, B, and C in an inclusive sense, while “and / or” may be used, though not required, with sufficient care to clarify that all of the aforementioned meanings are intended. Furthermore, the terms “one or more” and / or similar terms are used in the singular to describe any feature, structure, property, etc., while “and / or” is also used to describe multiple features, structures, properties, etc., and / or any other combination thereof. Similarly, the terms “based on” and / or similar terms are understood not to necessarily convey an exhaustive list of factors, but rather to allow for the existence of additional factors not explicitly described.
[0088] Algorithmic descriptions and / or symbolic representations are examples of techniques used by those skilled in the art to communicate the nature of their work to others skilled in the art. In the context of this patent application, an algorithm is generally understood to be a self-consistent sequence of actions and / or similar signal processing that leads to a desired result. In the context of this patent application, actions and / or processing involve the physical manipulation of physical quantities. Typically, but not always, such quantities may take the form of electrical and / or magnetic signals and / or states that can be stored, transferred, combined, compared, processed, and / or otherwise manipulated, for example, as electronic signals and / or states that constitute components of various forms of digital content such as signal measurements, text, images, video, and audio.
[0089] For reasons of general use, it has sometimes proven convenient to refer to such physical signals and / or physical states as bits, values, elements, parameters, symbols, characters, terms, samples, observations, weights, numbers, digits, measurements, content, etc. However, it should be understood that all of these terms and / or similar terms should be associated with the appropriate physical quantities and are merely convenient labels. Unless otherwise specified, as is evident from the above explanation, throughout this specification, the terms "process," "computing," and calculation Descriptions using terms such as "calculating," "determining," "establishing," "obtaining," "identifying," "selecting," and "generating" are specific to dedicated computers and / or similar dedicated computing devices and / or network devices. Device It should be understood that this can refer to the actions and / or processes of a dedicated computer and / or similar dedicated computing device and / or network device. Therefore, in the context of this specification, a dedicated computer and / or similar dedicated computing device and / or network device can process, manipulate, and / or convert signals and / or states, typically in the form of physical electronic and / or magnetic quantities, within the memory, registers, and / or other storage devices, processing devices, and / or display devices of the dedicated computer and / or similar dedicated computing device and / or network device. Therefore, in the context of this particular patent application, as stated above, the term “specific device” includes general-purpose computing devices and / or network devices, such as general-purpose computers, that have been programmed to perform specific functions, such as following program software instructions.
[0090] In some situations, the operation of a memory device, such as a change of state from binary 1 to binary 0, or vice versa, may involve transformations such as physical transformations. In certain types of memory devices, such physical transformations may involve the physical transformation of the product to a different state or material. For example, but not limited to, in some types of memory devices, a change of state may involve the accumulation and / or storage of charge, or the release of stored charge. Similarly, in other memory devices, a change of state may involve physical changes such as a change in magnetic orientation. Similarly, physical changes may involve transformations of molecular structure, such as from a crystalline form to an amorphous form, or vice versa. In yet other memory devices, a change of physical state may involve quantum mechanical phenomena such as superposition and entanglement, which may include, for example, qubits. The foregoing is not intended to be an exhaustive list of all examples in which a change of state from binary 1 to binary 0, or vice versa, in a memory device may involve transformations such as physical but non-transient transformations. Rather, the foregoing are intended as illustrative examples.
[0091] Figure 6 shows an exemplary sensor signal acquisition pattern of one embodiment of a vehicle 2200 that can operate in one or more automated driving modes (e.g., fully autonomous mode, semi-autonomous mode, driver assistance mode). As shown, an automated vehicle such as vehicle 2200 may include several sensors to provide measured values and / or observations that are processed according to processes 350, 370, 450, 470, and / or 550, etc. While specific patterns and / or specific numbers of sensors are shown in Figure 6, the subject is not limited in these respects. For example, a system or device such as vehicle 2200 may include any number of sensors in any of a wide range of arrangements and / or configurations. Also, although Figure 6 is shown as a two-dimensional representation, sensors in a vehicle that can operate in one or more automated modes may generate signals and / or signal packets that represent, for example, the state surrounding vehicle 2200 in three-dimensional space. In implementations, one-dimensional and / or two-dimensional sensor measurements can be combined and / or processed in other ways to generate, for example, a three-dimensional representation of the state surrounding vehicle 2200.
[0092] In various implementations, for example, various sensors may be mounted on the vehicle 2200 to capture observations and / or measurements of different parts of the environment surrounding and / or adjacent to the vehicle. In these implementations, the vehicle 2200 may include a number of different sensors capable of detecting input signals such as optical signals, electromagnetic signals, and / or audio signals. Each sensor may have a different observation field of view of the environment surrounding the vehicle 2200. Exemplary fields of view 2210a to 2210h are shown, but naturally, the subject matter is not limited to these points.
[0093] In the implementation, the sensor signals and / or signal packets may be used by the processing system of the vehicle 2200 to autonomously guide the vehicle through the environment, for example, or by at least one processor of the vehicle 2200 to identify objects and / or other environmental conditions in the vicinity of the vehicle 2200. Exemplary objects that may be detected in the environment surrounding a vehicle such as the vehicle 2200 may include other vehicles, trucks, cyclists, pedestrians, animals, rocks, trees, lampposts, guardrails, painted lines, signal lights, buildings, road signs, and several others. object The object may remain stationary, while other objects, such as pedestrians, may move within the environment.
[0094] In one implementation, one or more sensors of an exemplary vehicle 2200 may generate signals and / or signal packets that can represent at least a portion of the environment surrounding and / or adjacent to the vehicle 2200. Other sensors may provide signals and / or signal packets that represent the vehicle 2200's speed, acceleration, orientation, position (e.g., via a Global Navigation Satellite System (GNSS)), etc. As will be described in more detail below, the sensor signals and / or signal packets may be processed through a particle filter, etc., to generate a plurality of particles. Such particles may be used, at least in part, to influence the operation of the vehicle 2200. In the implementation, as the vehicle 2200 moves, for example, through the environment, the sensor signals and / or signal states may be used, at least in part, to update the particle filter, for example, to further influence the vehicle's operation. As will be described in more detail below, one of several sorting operations may be performed on the sensor signals and / or signal packets so that the particle filter, etc., can utilize a relatively wide range of signals and / or signal packets generated by various sensors.
[0095] In this context, “particle” refers to a digital representation of environmental conditions for a specific point in a specific coordinate system at a specific time, at least partially derived from sensor signals and / or signal packets. For example, a specific particle may include an array of parameters describing a specific point in the environment surrounding vehicle 2200 at a specific time. In implementations, “particle filters” or the like can be used to process sensor signals and / or signal packets to generate multiple particles describing environments, such as the environment surrounding vehicle 2200. Naturally, particle filters are merely illustrative types of processing that can be performed on sensor signals and / or signal packets, and the subject is not limited in this respect.
[0096] In implementations, a specific coordinate system may be specified, but the subject is not limited to any particular coordinate system within its scope. In implementations, such a coordinate system may include a three-dimensional parameter space, but other embodiments may specify other numbers of dimensions. In implementations, individual particles may be associated with a specific position in a particular three-dimensional space (e.g., the X, Y, and Z axes).
[0097] Figure 7 shows an exemplary schematic block diagram of an exemplary vehicle 2200. As described above, in the implementation, the vehicle 2200 may include several sensors 2210. Such sensors may include, for example, image acquisition (image capture) (e.g., camera), radar, lidar, and / or ultrasonic sensors, to give some non-limiting examples. In the implementation, the sensors 2210 may generate sensor signals and / or signal packets 2215 that are supplied to and / or acquired by a control system such as a control system 2220.
[0098] In an implementation, the control system 2220 may include, for example, at least one processor, at least one memory device, and / or at least one communication interface. In an implementation, the control system 2220 may include, for example, one or more central processing units (CPUs), neural network processors (NNPs), and / or graphics processing units (GPUs). In an implementation, the control system 2220 may process sensor signals and / or signal packets to generate signals and / or signal packets that may affect the operation of the vehicle 2200. For example, signals and / or signal packets may be generated by the control system 2220, provided to the drive system 2230, and / or acquired by the drive system 2230. In an implementation, the processing of sensor signals and / or signal packets by the control system 2220 may include, for example, a particle filter, but other embodiments may utilize other signal processing algorithms, techniques, approaches, etc., and the subject matter is not limited in this respect. In the implementation, the drive system 2230 may include, for example, devices, mechanisms, and systems for influencing the operation of the vehicle 2200. As described above, additional sensor signals and / or signal packets may be acquired and processed so that the operation of the vehicle 2200 can be updated over time as the vehicle 2200 traverses the environment.
[0099] The preceding description described various aspects of the claimed subject matter. For explanatory purposes, details such as quantities, systems, and / or configurations were described as examples. In other examples, well-known features were omitted and / or simplified so as not to obscure the claimed subject matter. While certain features have been illustrated and / or described herein, many modifications, substitutions, changes, and / or equivalents will be conceivable to those skilled in the art. Therefore, it should be understood that the attached claims are intended to encompass all modifications and / or changes that fall within the scope of the claimed subject matter.
Claims
1. It is a system, The system comprises multiple computing nodes and a transport memory for transporting data and / or values, wherein the transport memory includes a transport memory buffer. Each of the plurality of computing nodes is a circuit adapted to perform computational operations on data, and the plurality of computing nodes includes at least a first computing node and a second computing node. The first computing node of the plurality of computing nodes is Receiving multiple signal streams from multiple sources, wherein each of the multiple signal streams includes a set of data items, and the data items in at least two of the multiple signal streams include sensor observations and / or measurements. From each set of the aforementioned data items, identify the data items that have common attributes, Associating the sensor observations and / or measurements based at least partially on spatial and temporal attributes, The identified data items received from two or more of the plurality of signal streams are loaded simultaneously into one or more registers of the second compute node in the same execution cycle via endpoints configured as one or more registers of the second compute node within the transport memory buffer, wherein the simultaneously loaded data items are associated based on common attributes and include the associated sensor observations and / or measurements. Sort the sensor observations and / or measurements based at least partially on the relevant timestamps and locations of the objects observed and / or measured by the sensor observations and / or measurements, It is configured to do the following: The second computing node described above is The aforementioned simultaneously loaded data items are processed as operands for one or more calculation operations, Combining the sorted sensor observations and / or measured values, It is configured to do the following: system.
2. The transport memory comprises a first transport memory buffer among the transport memory buffers, The first transport memory buffer has a first endpoint and a second endpoint, The plurality of computing nodes are configured to communicate with external memory located outside the system via a bus, The first computing node includes a register that forms the first endpoint of the first transport memory buffer, The first register among the one or more registers of the second computing node forms the second endpoint of the first transport memory buffer. The first computing node is configured to load the data items into one or more registers of the second computing node simultaneously in the same execution cycle, without storing the data items in the external memory. The system according to claim 1.
3. The transport memory comprises a second transport memory buffer among the transport memory buffers, the second transport memory buffer having a first endpoint and a second endpoint, The first endpoint of the second transport memory buffer is formed by a register of the third computing node among the plurality of computing nodes, The second endpoint of the second transport memory buffer is formed by the second register of the one or more registers of the second compute node, The first and third computing nodes are configured to load the data items into one or more registers of the second computing node simultaneously in the same execution cycle, without storing the data items in the external memory. The system according to claim 2.
4. Associating data items from multiple signal streams from multiple sources based at least partially on attributes common to the data items, wherein associating the data items of the multiple signal streams is Execute a direct memory access (DMA) controller to load data items from each of the plurality of signal streams into the buffer associated with the signal stream, Identifying at least one common attribute between the loaded data item and at least one other data item, based at least partially on the content of the data item loaded into the buffer, From the aforementioned plurality of signal streams, the associated data items are simultaneously loaded into one or more registers of the compute node within the same execution cycle via endpoints configured as one or more registers of the compute node, which are transport memory buffers formed in the transport memory for transporting data and / or values from the aforementioned plurality of signal streams. The calculation node is executed to process the simultaneously loaded related data items as operands for one or more calculation operations, Methods that further include the above.
5. The data items in at least two of the plurality of signal streams include sensor observations and / or measured values. The sensor observations and / or measurements in at least two of the plurality of signal streams are associated at least partially with spatial and temporal attributes. The associated data items loaded simultaneously include sensor observations and / or measurements that are associated at least partially based on spatial and temporal attributes. The method according to claim 4.
6. Associating the data items of the multiple signal streams from multiple sources is In the calculation node, the sensor observations and / or measurements are sorted at least partially based on the relevant timestamps and locations of the objects observed and / or measured by the sensor observations and / or measurements. In the aforementioned computing node, the sorted sensor observations and / or measured values are combined, The method according to claim 5, including the method described in claim 5.
7. Associating data items of multiple signal streams from multiple sources, at least partially based on attributes common to the data items, wherein associating the data items of the multiple signal streams is Execute a direct memory access (DMA) controller to load data items from each of the plurality of signal streams into the buffer associated with the signal stream, Identifying at least one common attribute between the loaded data item and at least one other data item, based at least partially on the content of the data item loaded into the buffer, The process involves simultaneously loading associated data items from the aforementioned multiple signal streams into one or more registers of a compute node, Execute the compute node to process the simultaneously loaded related data items as operands for one or more compute operations, It further includes, Executing the aforementioned DMA controller means Loading one or more addressable lines of values and / or states stored in memory into the buffer, Analyze one or more unaddressable portions of at least one of the addressable lines of the loaded value and / or state, Processing one or more collection requests based at least partially on one or more unaddressable portions analyzed, Methods that include...
8. Associating data items of multiple signal streams from multiple sources, at least partially based on attributes common to the data items, wherein associating the data items of the multiple signal streams is Execute a direct memory access (DMA) controller to load data items from each of the plurality of signal streams into the buffer associated with the signal stream, Identifying at least one common attribute between the loaded data item and at least one other data item, based at least partially on the content of the data item loaded into the buffer, The process involves simultaneously loading associated data items from the aforementioned multiple signal streams into one or more registers of a compute node, Execute the compute node to process the simultaneously loaded related data items as operands for one or more compute operations, It further includes, Executing the aforementioned DMA controller means Performing a first word collection operation based at least partially on one or more redirected read requests, Converting one or more words obtained from the execution of the first word collection operation into one or more addresses, Perform a second word collection operation to transfer the data items located at the one or more addresses to a destination determined at least partially based on the one or more redirected read requests, Methods that include...
9. Associating data items of multiple signal streams from multiple sources, at least partially based on attributes common to the data items, wherein associating the data items of the multiple signal streams is Execute a direct memory access (DMA) controller to load data items from each of the plurality of signal streams into the buffer associated with the signal stream, Identifying at least one common attribute between the loaded data item and at least one other data item, based at least partially on the content of the data item loaded into the buffer, The process involves simultaneously loading associated data items from the aforementioned multiple signal streams into one or more registers of a compute node, Execute the compute node to process the simultaneously loaded related data items as operands for one or more compute operations, It further includes, Simultaneously loading the associated data items from the plurality of signal streams into one or more registers of the compute node is: Loading the data items of the plurality of signal streams into buffers associated with the plurality of signal streams, Execute a direct memory access (DMA) controller associated with the plurality of signal streams to selectively load data items into one or more registers, at least partially based on the display of common attributes in the content of the data items loaded into the buffer, Methods that include...
10. It is a system, Multiple sensors for generating multiple associated signal streams, Multiple computing nodes coupled to the aforementioned multiple sensors, A transport memory for transporting data and / or values, The transport memory includes a transport memory buffer, Each of the plurality of computing nodes is a circuit adapted to perform computational operations on data, and the plurality of computing nodes includes at least a first computing node and a second computing node. At least the first computing node among the plurality of computing nodes is The system is configured to load data items generated by two or more of the sensors from two or more of the associated signal streams into one or more registers of the second compute node via endpoints configured as one or more registers of the second compute node within the transport memory buffer, in the same execution cycle, wherein the simultaneously loaded data items are associated at least partially based on attributes common to the simultaneously loaded data items. The second computing node described above is The simultaneously loaded data items are configured to be processed as operands for one or more calculation operations. The system is further configured to update the state of the particle filter based at least partially on the simultaneously loaded data items. system.
Citation Information
Patent Citations
Inter-processor transfer system for message information
JP1989072255A
Inter-processor communication equipment
JP1989240963A
Data management method, information processor, and program
JP2014071495A
System parameter identification device, system parameter identification method, and computer program therefor
JP2017083922A