A system, device, and / or method for providing services for redirected memory access requests by utilizing direct memory access distribution and / or collection operations.
Patent Information
- Application Number
- KR1020257022240
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-01-04
- Filing Date
- 2023-12-19
- Publication Date
- 2026-09-02
- Estimated Expiration
- 2043-12-19
Smart Images

Figure 112025074782425-PCT00002_ABST
Abstract
Description
Technology Field
[0001] The subject matter of the claims disclosed herein relates to processing signals received from streams from multiple data sources. Background Technology
[0002] Autonomous driving and / or automated driving applications and other automotive and robotic applications may rely on the fusion of signals, measurements, and / or observation reports generated by multiple sensors. Processing for these applications may involve the manipulation of arrays of data elements. These applications may be executed and / or implemented by commercially available central processing units (CPUs) and / or graphics processing units (GPUs). These commercially available processing units may be configured to manipulate elements of input arrays to generate elements of output arrays.
[0003] One embodiment disclosed herein relates to a system, wherein the system comprises a memory comprising one or more memory devices; and a Direct Memory Access (DMA) controller coupled to the memory by a bus, and the DMA controller is configured to receive one or more redirected write requests; to execute a collection operation based at least partially on the received one or more redirected write requests to acquire one or more words from the memory; to translate one or more words acquired from the execution of the collection operation into one or more addresses in the memory; and to execute a distribution operation to write a data item to one or more addresses in the memory. In one particular embodiment, the memory is a first memory, and the system further comprises a second memory capable of / acting as a buffer, and the buffer is configured to store values and / or states retrieved from the first memory. The DMA controller may be further configured to parse the values and / or states in the buffer to determine at least one address among one or more addresses in the first memory. In one embodiment, the DMA controller is further configured to determine at least one address by applying one or more arithmetic operations to at least one of the parsed values and / or states.
[0004] In another specific embodiment, the DMA controller is configured to interpret one or more redirected write requests as requests for a collection operation followed by a distributed operation. In another specific embodiment, the DMA controller is further configured to form a request for a distributed operation for the redirected write request based at least partially on interpreting and / or converting data from the first collection operation to one or more addresses in memory for writing data. In another specific embodiment, the system further includes an initiator configured to initiate one or more redirected write requests, and the DMA controller is configured to execute the first collection operation in response to a signal from the initiator. The initiator of one or more redirected write requests may include one or more registers of a computing node or local memory configured to receive sensor measurements and / or observations as data items in a signal stream.
[0005] Another specific embodiment disclosed herein relates to a method in a direct memory access (DMA) controller, the method comprising: executing a first acquisition operation based at least partially on one or more redirected read requests; converting one or more words obtained from the execution of the first acquisition operation into one or more addresses; and executing a second acquisition operation to forward data items located at one or more addresses to a destination determined at least partially based on one or more redirected read requests. In one embodiment, the step of converting one or more words obtained from the execution of the first acquisition operation into one or more addresses further comprises parsing values and / or states in a buffer to determine a memory address. In another embodiment, the method further comprises applying one or more arithmetic operations to the parsed values and / or states to determine a memory address.
[0006] In one specific embodiment, the method further comprises the step of interpreting one or more redirected read requests as a first collection operation request among two collection operation requests. In another specific embodiment, the method further comprises the step of forming a second collection operation request among two collection operation requests based at least partially on one or more addresses determined from the result of the first collection operation. In yet another specific embodiment, the method further comprises the step of executing a first collection operation in response to a signal from an initiator of one or more redirected read requests. In one embodiment, an initiator of one or more redirected read requests includes a process executed by a computing node to execute one or more sensor fusion operations.
[0007] Another specific embodiment disclosed herein relates to a system, wherein the system comprises a memory comprising one or more memory devices; and a Direct Memory Access (DMA) controller coupled to the memory by a bus. The DMA controller is configured to receive one or more redirected read requests; to execute a first collection operation at least partially based on one or more received redirected read requests to acquire one or more words from the memory; to translate one or more words acquired from the execution of the first collection operation to one or more addresses; and to execute a second collection operation to forward data items located at one or more addresses to a destination determined at least partially based on one or more redirected read requests. Brief explanation of the drawing
[0008] The claimed subject matter is specifically identified and clearly claimed in the concluding section of this specification. However, regarding the organization and / or method of operation and the resulting purpose, features, and / or benefits, they are best understood by referring to the following detailed description together with the accompanying drawings. FIG. 1 is a schematic diagram of a computing device according to embodiments. FIGS. 2a and 2b are schematic diagrams of a computing device including an engine including a direct memory access (DMA) controller and / or a buffer, according to one embodiment. FIGS. 3a and 3e are schematic diagrams of a computing device including a DMA controller and / or engine including a buffer for facilitating distributed collection operations, according to one embodiment. FIG. 3c is a schematic diagram illustrating the unaddressable portions of an addressable line according to one embodiment. FIGS. 3b and FIGS. 3d are flowcharts of a process to facilitate distributed operation and collection operation according to one embodiment. FIGS. 4a and 4d are schematic diagrams of a computing device including a DMA controller including a buffer for facilitating the redirection of DMA transactions, according to one embodiment. FIGS. 4b and FIGS. 4c are flowcharts of a process for facilitating the redirection of DMA transactions according to one embodiment. FIGS. 5a and 5c are schematic diagrams of a computing device for facilitating the processing of multiple signal streams from multiple associated sources, according to one embodiment. FIG. 5b is a flowchart of a method for processing multiple data streams from multiple associated sources according to one embodiment. FIG. 6 is a drawing illustrating exemplary sensor signal collection for an exemplary vehicle according to one embodiment. FIG. 7 illustrates an exemplary schematic block diagram of exemplary vehicle features in an autonomous driving / automatic driving application according to one embodiment. In the following detailed description, reference is made to the attached drawings, which form part of the description. In the drawings, the same reference numerals may correspond throughout and / or designate similar parts. It will be recognized that the drawings are not necessarily drawn to actual scale, for example, to simplify and / or clarify the examples. For example, the dimensions of some embodiments may be exaggerated compared to others. It should also be understood that other embodiments may be utilized. Furthermore, structural changes and / or other modifications may be made without departing from the claimed subject matter. Throughout this specification, the reference to "claimed subject matter" refers to the subject matter intended to be covered by one or more claims or parts thereof, and is not intended to refer to a complete set of claims, a specific combination of sets of claims (e.g., method claims, device claims, etc.), or a specific claim. It should also be noted that directions and / or references, such as up, down, upper, lower, etc., are used to facilitate discussion of the drawings and are not intended to limit the application of the claimed subject matter. Therefore, the following detailed description shall not be construed as limiting the claimed subject matter and / or equivalents. Specific details for implementing the invention
[0009] Throughout this specification, references to "one embodiment," "one embodiment," "one embodiment," "one embodiment," etc., mean that specific features, structures, properties, etc., described in relation to a specific embodiment and / or embodiment are included in at least one embodiment and / or embodiment of the claimed subject matter. Accordingly, for example, the appearance of such phrases in multiple places throughout this specification is not intended to refer to the same embodiment and / or embodiment or to any single specific embodiment and / or embodiment. Furthermore, it should be understood that the specific features, structures, properties, etc. described may be combined in various ways in one or more embodiments and / or embodiments, and thus fall within the scope of the intended claims. In general, as is always the case with the specification of a patent application, these and other matters may vary in the specific context of use. In other words, throughout this disclosure, the description and / or use of a specific context provides useful guidance regarding reasonable inferences to be drawn, but likewise, the phrase "in this context" generally without further limitation refers at least to the context of this patent application.
[0010] To facilitate the efficient processing of data items within one or more signal streams, a Direct Memory Access (DMA) controller coupled to memory by a bus may be configured to receive one or more redirected write requests. The DMA controller may execute a collection operation at least partially based on one or more received redirected write requests to acquire one or more words from memory, translate one or more words acquired from the execution of the collection operation to one or more addresses in memory, and execute a distribution operation to write data items to one or more addresses in memory determined by the translation. In another embodiment, the DMA controller may process one or more redirected read requests by executing a first collection operation at least partially based on one or more redirected read requests, translate one or more words acquired from the execution of the first collection operation to one or more addresses, and execute a second collection operation to forward data items located at one or more addresses determined by the translation to a destination determined at least partially based on one or more redirected read requests.
[0011] In some cases, the host central processing unit (CPU) may be part of a computing device within the vehicle and may process data items in memory for automotive applications. For example, in automotive applications such as autonomous / automated driving applications (e.g., fully autonomous, semi-autonomous, driver assistance systems, etc.), signal streams of sensor signals and / or observations may be fused to use a particle filter, for example, to update the state of a particle filter. Such applications may be implemented in systems such as automated machinery (vehicles, trucks, etc.). In this context, the "signal stream" referred to herein means a time-varying progression of a sequence of encoded data items to be transmitted to a receiving device via a signal transmission medium. The encoded data items transmitted in the signal stream (also referred to as data values) may represent attributes indicating conditions and / or events, subject identifiers, timestamps indicating the time of an event, and metadata, merely for the purpose of naming certain attributes that may be represented in the encoded data items transmitted in the signal stream. In a specific embodiment, the signal stream may carry sensor measurements and / or observations along with associated timestamps to indicate the times when these measurements and / or observations are acquired.
[0012] In one aspect of an embodiment, information originating from different sources and reaching each signal stream may be processed as an information "confluence" to determine a calculation result. This information confluence may be processed by associating and / or correlating information items from different sources by specific attributes (e.g., time, space, reliability, reliability, etc.). Then, processing of the information confluence may include performing one or more operations on the items based on one or more attributes to produce a result. In a specific embodiment, data items within the information confluence may be processed by updating one or more states of a particle filter. For example, such a particle filter may implement processing of the confluence of arrays of sensor signals / observations (e.g., those received from signal streams) to update the states of measurement-particles, filter-particles, static-particles, and / or dynamic-particles. In one embodiment, measurements and / or observations in the confluence of arrays of measurements may be generated from different sensors. Nevertheless, such measurements and / or observations generated by different sensors may be associated and / or correlated by time and space. According to one embodiment, associating data items from different sources in the confluence of arrays may be implemented at least partially using radix alignment applied to an array of keys. In a specific embodiment, sensor signals and / or observations may be formatted in arrays to be processed to generate the confluence of arrays. An exemplary procedure for generating such confluence of arrays may be executed according to the following pseudocode:
[0013] / / Number of input streams, S_in[] is k
[0014] / / Number of output arrays, S_out[] is l
[0015] / / Pseudo code that takes in S_in[] and some parameters
[0016] / / and outputs S_out[] in_index[0...(k-1)] = 0
[0017] out_index[0...(l-1)] = 0 parameters[]
[0018] while(1){ count = 0
[0019] value_in[0...(k-1)] = 0
[0020] value_out[0...(l-1)] = 0
[0021] for(i: 0..(k-1)){
[0022] if(is_valid(in_index[i], i, parameters)){ value_in[i] = S_in[i][in_index[i]] count++
[0023] }
[0024] }
[0025] if(count == 0){break}
[0026] (value_out[], s_out_enable[]) = f(value_in[], parameters) for(i: 0..(l-1)){
[0027] if(s_out_enable[i]{ s_out[i][out_index[i]] = value_out[i] out_index[i]++
[0028] }
[0029] }
[0030] / / typically, addition below is saturating in_index[] += read_next(value_in[], parameters)
[0031] / / typically, read_next() returns 0, 1, or N[]
[0032] }
[0033] According to one embodiment, the confluence of a plurality of signal streams may include mapping data items in different input signal streams to data items in one or more output signal streams. For example, the data items in these input signal streams may include sensor measurements and / or observations of a sensor associated with the input signal stream. Accordingly, the confluence of these input signal streams may include mapping sensor measurements and / or observations (e.g., from different / separate sensors associated with the input signal streams) to data items of the output signal stream. These data items of the output signal stream may include values inferred / calculated based on the sensor measurements and / or observations. In the pseudocode example provided above, a plurality of input signal streams may be defined as S_in[] (containing data item value_in[]), and a plurality of output signal streams may be defined as S_out[] (containing data item value_out[]). Here, the expression "(value_out[], S_out_enable[]) = f(value_in[], parameter) for(i: 0..(l-1))" can map the data item value_in[] of the input signal stream S_in[] to the data item value_out[] of the output signal stream S_out[] according to the function f().
[0034] According to one embodiment, a computing device, circuitry, and / or logic may form a "confluence engine" (CE), also referred to as a "confluencer" and / or a "confluence processor" (CP), to process the confluence of data items as discussed above. The confluence of these data items may include, for example, confluences of signal streams and / or sequences of confluences of signal streams, with reduced latency and / or reduced computing resources (e.g., power, memory, etc.). In a specific embodiment, the output of a confluence operation of this CE may provide all or part of the input to a subsequent confluence operation. Characteristics of the confluence, such as functions like k, f(), and read_next() as shown in the pseudocode example above, may be part of the runtime programming of the CE.
[0035] According to one embodiment, the CE may utilize a Direct Memory Access (DMA) subsystem that may include a DMA controller (also referred to as a DMA "engine(s)") configurable to initiate read and write operations between, for example, line-accessible memory and transfer memory. The specific process by which the CE accesses memory may be determined parametrically at compile time and physically at runtime. As such, in a particular embodiment, the events triggering the execution of the DMA controller may not be limited to events occurring in the Arithmetic Logic Unit (ALU) (e.g., from the loads and stores of the ALU). The DMA controller may be triggered to execute DMA transactions by the initiation of a Confluence operation. Subsequently, the DMA controller may execute these DMA transactions independently of the Arithmetic Logic Units (ALUs) (e.g., depending entirely on the availability of read and write access at the valid endpoints of these DMA transactions).
[0036] FIG. 1 is a schematic diagram of a system (100) for executing DMA transactions. The system (100) includes a plurality of components that communicate via a bus (101). The components include a host CPU (102), a memory controller (106), RAM (108) (which may form main system memory), a peripheral device (114), and a DMA controller (112). In this context, "direct memory access" as referred to herein means a process executed by one or more hardware subsystems and / or circuits to access a specific memory independently of a specific processing unit and / or central processing unit (e.g., independently of the host CPU (102)). According to one embodiment, the DMA controller (112) may initiate operations to access (e.g., read access or write access) the random access memory (RAM) (108) via the bus (101) independently of the host CPU (102). For example, the DMA controller and / or engine (112) can control and / or execute transactions to transfer data items (also referred to as data values) between a peripheral device (114) and RAM (108) independently of the operation by the host CPU (102) (via the memory controller (106)). Such DMA transactions may be triggered by, for example, signals, conditions, and / or events (e.g., interrupt signals).
[0037] As discussed above, the processing of data items for automotive applications or other computing applications can be enhanced through a Confluence Engine (CE). FIGS. 2a and 2b are schematic diagrams of a computing device (200) comprising a Direct Memory Access (DMA) controller (212) (also referred to as a DMA engine (212)) and a buffer (216) for implementing one or more aspects of the CE according to one embodiment. The computing device (200) may be, for example, a system-on-chip (SoC), a microchip, a control circuit, or some other computing device, or may form such a device. In some embodiments, the computing device (200) may form a vehicle controller such as an advanced driver assistance system (ADAS) device, a telematics control unit (TCU), an electronic control unit (ECU), a centralized vehicle computer, or some other vehicle controller. As illustrated in FIGS. 2a and 2b, the computing device (200) may further include a CPU (202) (also referred to as a host CPU), a transfer memory (218), a line-accessible memory (208), a buffer (216), and a computing node (CN) pool (220) (also referred to as a CN pool (220)).
[0038] In one embodiment, the CN may include a single processing circuit core capable of executing operations for mapping input operands to output computing results. In another embodiment, the CN may include a plurality of separate processing cores for executing operations for mapping input operands to output computing results. In another embodiment, two or more asynchronous execution CNs may be implemented on the same processing circuit core. For example, the processing circuit may implement a first CN to generate an output result (e.g., stored in a transfer memory (218)) which is an input to a second processing circuit that subsequently executes the CN implemented on the same processing circuit.
[0039] A CN in a pool of CNs (220) may include dedicated local memory (e.g., static random access memory (SRAM)) and general registers for receiving operands for operations to be executed and / or providing results from the execution of operations. In this embodiment, the host CPU (202) may use a transfer memory (218) to store data items and may control the CNs in the CN pool (220) to perform operations on these data items. The transfer memory (218) may be physically closer to and / or may operate with lower access latency and thus may be used as a cache for storing data items. Line accessable memory (208) may be outside the transfer memory (218) and may provide a larger amount of memory space compared to the transfer memory (218), but may be physically further from the CN pool (220) and may operate with longer access latency. According to one embodiment, the transfer memory (218) may include one or more synchronization mechanisms to facilitate communication between CNs and / or between CNs within the CN pool (220) (e.g., for synchronizing communications between CNs having different execution latencies). In a specific embodiment, the line-accessible memory (208) may be separated from the transfer memory (218), CPU (202), and buffer (216) by a bus (not shown). According to one embodiment, the buffer (216) may be formed in a circuit for implementing the core circuit of the DMA controller and / or engine (212) so that the buffer (216) is distinguished and separated from the circuit for forming the transfer memory (218).This formation of the buffer (216) in the core circuit of the DMA controller and / or engine (212) can reduce and / or minimize the latency associated with loading data items into the buffer (216) and storing data items from the buffer (216) during the process of executing DMA operations.
[0040] In one embodiment, the DMA controller and / or engine (212) may be configured to interface with the transfer memory (218) and the line-accessible memory (208). The transfer memory (218) and / or the line-accessible memory (208) may also be cache line-addressable. That is, in this embodiment, the line-accessible memory (208) may be a cache line-addressable memory. The transfer memory (218) and / or the line-accessible memory (208) may provide data items (also referred to as data values) that a CN in the pool of CNs (220) can manipulate, operate on, or otherwise process. According to one embodiment, all or part of the transfer memory (218) may be organized as a cache that can be integrated with commercially available components. Additionally, it should be noted that caches within the transfer memory (218) or buffer (216) may mitigate manufacturing defects and / or enable the use of embodiments larger than the expected application size.
[0041] In one embodiment, the DMA controller and / or engine (212) may be configured to process cache line-sized data items (e.g., 64 bytes or 128 bytes), and even in distributed collection operations, such data items are locable and accessible by cache line addresses. Such cache line addresses for a 64-byte cache line may be represented in six-zero-terminated binary notation. Likewise, cache line addresses for a 128-byte cache line may be represented in seven-zero-terminated binary notation. According to one embodiment, the DMA controller and / or engine (212) may be configured to process word-sized data items, but the line-accessible memory (208) may continue addressing only in cache lines. The transfer memory (218) may be organized as a set of transfer memory buffers, such as a cache, a word-addressable register, or word-first-in-first-out (FIFO) buffers, or a combination thereof. In one embodiment, such a transfer memory buffer within the transfer memory (218) may include a circuit and / or device specifically designed to function as a FIFO buffer. In another embodiment, such a transfer memory buffer within the transfer memory (218) may include a static random access memory (SRAM) device or a network-on-chip (NOC) device specifically configured to function as a FIFO buffer. To facilitate a distribution or collection operation for transferring word-sized data items to and from the line-accessible memory (208), the DMA controller and / or engine (212) may implement a buffer (216) (e.g., located between the line-accessible memory (208) and the transfer memory (218). The buffer (216) may be configured to store together bytes from multiple cache lines to be placed at the destination word within the transfer memory (218).The DMA controller and / or engine (212) may also perform multi-casting write operations in either direction (e.g., from the transfer memory (218) to the line-accessible memory (208), or from the line-accessible memory (208) to the transfer memory (218). The buffer (216) may be distinct from the transfer memory (218) and the memory DMA controller and / or engine (212). As discussed herein, the buffer (216) may be formed in a client core to implement the DMA controller and / or engine (212).
[0042] In this context, the “transfer memory” referred to herein means a circuit that facilitates communication of data items between CNs and / or between CNs, such as CNs in a pool of CNs (220). In one particular embodiment, such a transfer memory may transfer the result from the execution of a first operation in a first CN to become the input operands of a second operation to be executed in a second CN. In a particular embodiment, the transfer memory (218) may be organized as a static random access memory (SRAM) device used as an access control memory, or as a shared memory, word addressable or SIMD vector addressable register file, or a circuit and / or device specifically structured to function as a FIFO buffer as a cache memory. These circuits and / or devices, specifically structured to function as FIFO buffers, may have a word-wise or single-instruction, multiple-data (SIMD)-vector wide or circuit, or network-on-chip (NOC) devices combined between endpoints, or a combination thereof—these are merely examples intended to be given.
[0043] As described above, CNs within the CN pool (220) (e.g., a pool of CNs) can operate on data items read from memory. In this context, the term "computing node" as used herein refers to an identifiable, distinct set of computing resources (e.g., hardware and executable instructions) that is configurable to perform operations for processing input values to provide output values. CNs within the pool of CNs (220) may include an Arithmetic Logic Unit (ALU), a Digital Signal Processor (DSP), a Vector CN, a VLIW engine, or a Field Programmable Gate Array (FPGA) cell, or scalar CNs and / or processing circuit cores for implementing a combination thereof—these are merely intended to present some examples of specific circuit cores that may be used to implement CNs within the pool of CNs (220). The CN pool (220) may be implemented according to various architectures. For example, the CN pool (220) may include one or more CNs implemented according to a RISC (reduced instruction set computing) architecture, a CISC (complex instruction set computing) or VLIW (very long instruction word) architecture, or some combination of these types, in a fully functional or simplified form. The CNs in the CN pool (220) may also include a combination of scalar, SIMD or multiple instruction, and single data stream (MISD) ALUs.
[0044] According to one embodiment, the features of the computing device (200) may include commercially available features such as, for example, transmission rings, cross networks, and shuffle circuits. A pool of CNs (220) may facilitate multithreading in the CNs, clustering of the CNs, features for mailboxes, interrupts, synchronization between CNs, atomic operations in locations within the transmission memory (218), and features added to satisfy safety and security requirements—these are merely a few examples.
[0045] In a specific embodiment, the clustering of CNs within the CN pool (220) may be formed at least partially based on the trade-offs of resources and may be permanently defined in the integrated circuit (IC) device. In a specific embodiment, depending on how the CN pool (220) is configured, two of the same integrated circuit (IC) devices may implement different clusterings of associated CNs. For example, the processor configuration may be defined at least partially based on the associated confluences to be executed for processing multiple confluences using the clustering of CNs. Each associated cluster may process the confluence of arrays, for example, such that the output of one cluster provides an input to one or more other clusters.
[0046] According to one embodiment, the transfer memory (218) may form one or more buffers, including one or more FIFO buffers. One or more FIFO buffers may include "vertical" FIFO buffers. These vertical FIFO buffers may be buffers having an endpoint (which may be referred to as an "external" endpoint) that interfaces with the DMA controller and / or engine (212), and another endpoint (which may be referred to as an "internal" endpoint) that forms registers capable of providing data items available as operands for CNs in the pool of CNs (220). The internal endpoints of the vertical FIFO buffers may be shared by a plurality of CNs in the pool of CNs (220). If such internal endpoints include a FIFO-out end (i.e., an outbound end of the FIFO buffer), for example, a broadcast to a plurality of CNs in the pool of CNs (220) may be achieved. If such internal endpoints include a FIFO-like end (i.e., an inbound end of the FIFO buffer), hardware locks and / or instructions executed on a CN within the pool of CNs (220) can prevent race conditions. In some embodiments, the FIFO buffer formed in the transfer memory (218) may also include a "horizontal" FIFO buffer. A horizontal FIFO buffer may be a buffer having two endpoints that provide data items as operands to different CNs within the pool of CNs (220) or simply provide data items to locations in the transfer memory (218). It should be noted that the FIFO buffers formed in the transfer memory (218) may provide a transparent block and release mechanism for instances with different input and output rates.If the outer end of a FIFO buffer in the transfer memory (218) is a register or operand to be consumed / processed by some CNs in the pool of CNs (220) and the CNs are slow to do so, the DMA controller and / or engine (212) may eventually execute a block mechanism. Conversely, if the DMA controller and / or engine (212) is slow to write to these FIFO buffers in the transfer memory (218), the CNs in the pool of CNs (220) may eventually execute a block mechanism. If FIFO buffers are implemented between CNs and / or between CNs and / or between locations and / or between locations in the transfer memory (218), similar transparent block and release mechanisms may exist.
[0047] In an IC device, where the transfer memory (218) forms a FIFO buffer, the circuitry at the endpoints of the FIFO buffers may be permanent or configurable (e.g., through internal FPGA circuitry). A specific embodiment may include segments of vertical FIFO buffers and horizontal FIFO buffers formed in the IC device circuitry. Such buffer circuitry may have configurable endpoints at runtime from at least one of the operands or registers of CNs in the pool of CNs (220), the interface (216) to the buffer, locations within the transfer memory (218), and endpoints of other FIFO buffers in the transfer memory (218). If the endpoints of the FIFO buffers of the transfer memory (218) are the operands or registers of CNs in the pool of CNs (220), these FIFO buffer endpoints may be shared among multiple CNs in the pool of CNs (220).
[0048] According to one embodiment, a CN in a pool of CNs (220) may receive operands for computing operations, for example, from registers in a register file, from FIFO buffers in a transfer memory (218), and / or from special constant and parameter registers (which may include shared registers). In some cases, a portion of the register file may be shared among multiple CNs in a pool of CNs (220). Within the computing device (200), a host CPU (202) may also access the constant and parameter registers. The host CPU (202) may provide host functions, for example, such as launching confluence operations to be performed after configuration tasks are completed. These configuration tasks may include defining clusters of CNs within a pool of CNs (220), configuring communication between CNs, configuring the width and depth of FIFO buffers in the transfer memory (218), configuring endpoints of FIFO buffers in the transfer memory (218), defining manager CNs within a pool of CNs (220) and defining clusters of CNs within a pool of CNs (220) to be managed by manager CNs, configuring a DMA controller and / or engine (212), configuring atomics, configuring communication and synchronization resources, and monitoring the termination of confluences—these are merely examples intended.
[0049] In one embodiment, the buffer (216) can facilitate the processing of confluences in a number of embodiments. In one non-limiting embodiment, the input signal stream within the confluence may include a "direction stream" (e.g., a collection of addresses in line-accessible memory (208) to be read). In this case, this indirect signal stream may not be supplied directly to a CN in the pool of CNs (220), but instead may be provided to a DMA controller and / or engine (212), while the DMA controller and / or engine (212) may transmit the signal stream of read data items to the CN. That is, the latency of access by random read operations (random read of lines, or random read of words, or random indirect read of lines or words or combinations thereof) may be hidden by extracting a pattern of random accesses, constructing an access list of considerable length, and performing the necessary extraction and indirect read within the DMA controller and / or engine (212) itself. In other cases, the input signal stream may include a "dual-direction stream" in which the signal stream may contain addresses, and data items at these addresses provide additional direction change after some processing. In a specific embodiment, the directionality of the first level of the dual-direction stream may be read and used multiple times. In a sensor fusion application, multiple confluences may be configured, which can be constructed with a lookup table (LUT) for the first level indirect data items and the second level indirect data items (e.g., in the transmission memory (218) itself). However, it should be understood that these examples are not limited.
[0050] According to one embodiment, some processing within the confluence of signal streams (e.g., within f() and read_next() in the pseudocode example above) may be seen as the extraction of information from data items distributed across CNs within a pool of CNs (220). An example of such extraction is provided in Table 1 below.
[0051] [Table 1]
[0052]
[0053] According to one embodiment, results may be provided by CNs in a pool of cooperatively operating CNs (220) using operations such as shift, shuffle, broadcast, and multicast of operands among the participating CNs—these are merely examples intended to be given. These features may exist within the CNs, for example, for the word-operands of SIMD-vector CNs in the pool of CNs (220). The pool of CNs (220) may be configured to have these features for inter-CN communication through a configurable bridge circuit between the CNs. This circuit may be hardwired or configurable at runtime.
[0054] According to one embodiment, CNs within a pool of CNs (220) (e.g., composed of CN clusters) may be configured for specialized processing functions. In a specific embodiment, such specialized CNs within a pool of CNs (220) may facilitate the management of the processing flow of an application. For example, a pool of CNs (220) may include one or more processing CNs (224) and one or more manager CNs (222). Processing CNs (224) may perform processing operations, for example, on sensor observations, measurements, and / or other signals. In this embodiment, manager CNs (222) may manage different sets of processing CNs (224). Processing CNs (224) may communicate with one or more manager CNs (222). In some cases, the manager CN (222) may provide information to the processing CNs (224) based on communication from the processing CNs (224). The processing CNs (224) may continue their processing as verified by the information provided by the manager CN (222). In some embodiments, the manager CNs (222) may include physically separate processing CNs (224), or only specialized circuits formed within the processing CNs (224).
[0055] It should be noted that the rates at which different individual signal streams of the confluence (from different sources, such as different sensors) are generated and consumed (e.g., processed) may not necessarily be the same. The rates at which these signal streams are consumed or generated may be determined by the characteristics of an operation, for example, the function f() (in the pseudocode example above). For example, there is a rate at which the signal stream of the sensor's measurements and / or observations for the sensor fusion operation can be calculated / generated by the function f(). This sensor fusion operation may involve an inverse sensor model that outputs a signal stream longer than the input signal stream. To process the confluence of the longer signal streams, for example, the output rate may be matched to the bandwidth / throughput of the line-accessible memory (208). Configuring the output of one CN as the input to another CN may be helpful, for example, for load balancing.
[0056] According to one embodiment, the computing device (200) can enable the deployment of advanced sensor fusion operations to update particle filter states while consuming very little power (e.g., in autonomous driving or other automotive applications). In one application, instances of the computing device (200) may be implemented, for example, as a sequence of pipeline stages. In such an application, the features of the computing device (200) may be configured so that different amounts of resources are allocated to different pipeline stages when in use. The exchange of data items between and / or between the pipeline stages and the line-accessible memory (208) can be transparently synchronized (e.g., without mutexes, spin-locks, etc.) by using FIFO buffers (e.g., FIFO buffers formed in the transfer memory (218).
[0057] In another embodiment, the computing device (200) may be configured to be optimized for power, space, and / or performance (e.g., accuracy and / or latency). Features of the computing device (200) may be configured to implement CE, but features of the computing device (200) may be adapted to other applications, such as applications that rely on random access to line-accessible memory (208). Features of the computing device (200) may also be implemented in a so-called "supercomputer." In the context of a supercomputer, the low-power features of the computing device may help overcome power constraints that, for example, prevent the realization of an exascale supercomputer. Additionally, the circuitry for implementing the computing device (200) may incorporate safety and security features in a manner that meets the requirements of an embedded computing device. According to the computing device (200) having a small physical size and low power consumption, the use of the computing device (200) is not necessarily limited to being used as an external accelerator integrated circuit (IC) device, but may also be integrated into a subsystem within an automotive-grade system-on-chip (SOC) IC device.
[0058] In one embodiment, for example, the computing device (200) may include a specific array and / or configuration of CNs (e.g., a pool of CNs (220)), a transfer memory (218), and / or a DMA controller and / or engine (212). The computing device (200) may be configured to provide a network of CNs and memory elements adapted to a specific type of computation, such as a specific application for processing signal streams (e.g., returning measurements and / or observations from sensors). In one embodiment, such a network of CNs and memory elements may enable the simultaneous processing of multiple signal streams at high throughput and low latency. Such a network of CNs and memory elements may be implemented, at least in part, using implementations of in-device communication protocols (e.g., AXI) through physical connections and FIFO buffers. As noted above, the endpoints of a FIFO buffer may include, for example, an addressable memory location or register for receiving an operand (e.g., a general register of an ALU configured as a computing node as an operand for a computing operation) or a result calculated by a CN. For example, a pool of FIFO buffers may include endpoints that can be configured to be associated with various CNs or memories.
[0059] In one aspect, the specific embodiments disclosed herein relate to so-called vectorized input / output (I / O) operations, including “distribute” operations and “collect” operations. These vectorized operations may enable high-throughput transfer of large amounts of data to or from physical memory (e.g., multiple addressable memory lines within line-accessible memory (208)) in a single request or command to improve efficiency and convenience, for example. For example, a collect operation may involve sequentially reading data from multiple memory locations (e.g., buffers) and writing the read data to a signal stream or a contiguous portion of memory in a single transaction. In one embodiment, a DMA controller and / or engine (212) may execute a collect operation to service a collect request (e.g., a request originating from an application) that specifies multiple memory locations where data items are to be read, where line alignment is not necessarily required, and specifies a destination (e.g., a memory address) for storing the read items. On the other hand, a distributed operation may involve reading data items from a signal stream or contiguous memory and writing the read data items to a number of different memory locations where line alignment is not necessarily required. In one embodiment, the DMA controller and / or engine (312) may execute a collection operation to service a distributed request (e.g., a request originating from an application), which may specify the locations of the data items to be read (e.g., contiguous memory addresses) and indirectly specify the locations where the read data items will be written (by requiring a read of specific locations that allows the locations where the read data items will be written to be determined). The DMA controller and / or engine (312) may likewise execute indirectly specified collection operations.
[0060] According to one embodiment, using DMA transactions to aid in the processing of data items in signal streams can be enhanced by using distributed and collect operations. In a specific embodiment, contents loaded into a buffer from a collect operation may be used to determine one or more addresses for a subsequent collect or distributed operation. FIGS. 3a and 3e are schematic diagrams of a computing device (300) according to one embodiment, comprising a DMA controller and / or engine (312) that includes or communicates with a buffer (316) to facilitate distributed and / or collect operations. In a specific embodiment, the computing device (300) may include one or more features of the computing device (200) ( FIGS. 2a and 2b). The DMA controller and / or engine (312) may be configured as a distributed collect multicast DMA engine (SGM-DMA), but, for example, has the ability to address words within an addressable line.
[0061] According to one embodiment, the DMA controller and / or engine (312) may receive input signal streams and / or supply these input signal streams as blocks to an initial cluster of CNs via virtual channels through a FIFO buffer. In turn, output signal streams from the initial cluster of CNs may be supplied as blocks via virtual channels or as data via a FIFO buffer to a subsequent downstream cluster of CNs. Transmission between clusters of CNs may be a multicast transmission, for example, as directed by a particular application. At any stage, some or all of the output signal streams from some CN clusters may be returned to the DMA controller and / or engine (312) (for example, during the process in which the DMA executes distribution or collection operations).
[0062] According to one embodiment, the DMA controller and / or engine (312) may perform specific distribution and / or collection operations to transfer data items from one discontinuous memory block to another memory block using a series of smaller, consecutive block transfers. Here, acquiring these data items from a discontinuous block of source memory may be performed in a collection operation. Likewise, writing data items to a discontinuous block of destination memory may be performed in a distribution operation. In one embodiment, the smallest unit of memory that can be accessed in this source memory or destination memory may be a single addressable line of values and / or states (e.g., a single cache line or a word within line-accessible memory). For example, the DMA controller and / or engine (312) may communicate with line-accessible memory (LAM) (308) that can be accessed line by line.
[0063] According to one embodiment, physical memory (e.g., LAM (308)) may include bit cells for defining values and / or states to represent information such as 1 or 0. Such physical memory may further organize bit cells into words containing integer numbers of 8 bit bytes (e.g., 4-byte words exceeding 32 bits or 8-byte words exceeding 64 bits). Additionally, such physical memory may define line addresses (e.g., word line addresses) associated with consecutive bits defining "addressable lines" of values and / or states. For example, in response to a read or write request (e.g., originating from a host processor), a memory controller may access portions of memory in targeted read or write transactions according to the word line address specified in the request. To service a read request, for example, a memory controller may retrieve values and / or states for all bytes of the line associated with the line address specified in the read request. Likewise, to service a write request, the memory controller may write values and / or states for all bytes to an addressable line associated with the line address specified in the write request. A line address may specify a memory location containing all consecutive bytes of an addressable line, but such a line address does not specify the locations of individual sub-parts of an addressable line, such as individual bytes or consecutive bytes, that are smaller than the entire addressable line or the bytes spanning the addressable lines. Such sub-parts of an addressable line that are not are referred to herein as "unaddressable parts."
[0064] According to one embodiment, an addressable line may define a minimum unit of memory that is locable and / or accessible according to a memory locating scheme. In a specific exemplary embodiment of FIG. 3c, such an addressable line (360) may consist of smaller units of memory, such as bits, bytes, or words. In a specific embodiment illustrated, the addressable line (360) is n+1 bytes (3620 to 362 n ...includes ). As noted above, certain specific embodiments may relate to updates of unaddressable parts and / or parts that are less than the whole of addressable lines and / or unaddressable parts spanning multiple lines, with or without addressable lines. In the illustrated exemplary embodiment, bytes (3622 and 3623) may define the unaddressable part of an addressable line (360). Here, the entire addressable line may be locable and / or accessible via a unique address according to a memory addressing scheme, but the bytes (3622 and 3623) (less than the whole) may not be addressable according to a memory addressing scheme.
[0065] In one embodiment, the LAM (308) may have a LAM controller (306) configured to receive a request to specify an address for a line of data items stored in the LAM (308). Such lines of data items may include multiple words or multiple bytes and may be the smallest unit of data items that the LAM (308) can retrieve and return to another device. According to one embodiment, the computing device (300) may use a buffer (316) to enable access to smaller unaddressable portions of the line (e.g., a single or multiple bytes within a word or a group of words). According to one embodiment, the circuit forming the buffer (316) may be integrated with the circuit forming the DMA controller and / or engine (312) to enable minimum latency for accessing the buffer (316) initiated by the DMA controller and / or engine (312). For example, the buffer (316) may be formed as a static random access memory (SRAM) device accessible by a DMA controller and / or engine (312) without initiating requests and / or transactions on the main memory bus (e.g., LAM (308) or a bus coupled to the host computer / processor).
[0066] According to one embodiment, the DMA controller and / or engine (312) may communicate with the initiator (322). The initiator (322) may include a device (e.g., at least partially implemented by circuitry and / or logic) that achieves a specific state for triggering one or more DMA transactions to be executed by the DMA controller. For example, the initiator (322) may include an output register of an ALU, a buffer, or a hardware interrupt handler—these are merely intended to provide a few examples of devices capable of initiating a DMA transaction.
[0067] According to one embodiment, the DMA controller and / or engine (312) may obtain a list of collection requests in response to a signal from the initiator (322). In one particular embodiment, the initiator (322) may trigger a DMA transaction in response to an event or condition in the execution of the particle filter process. For example, the particle filter process may identify data items in memory that are expected to be retrieved for processing in future execution cycles. Once a significant amount of these data items are identified, a list of collection requests (e.g., redirected collection requests) identifying these data items may be forwarded to the DMA controller and / or engine (312). In one embodiment, the list of these collection requests may be provided to the DMA controller and / or engine (312) from shared memory or a network-on-chip (NOC)—these are merely intended to provide a few examples. Once it is known that such a list is available for processing by the DMA controller and / or engine (312), a process for generating the list (e.g., executing computer-readable instructions) may trigger the DMA controller and / or engine (312) via an interrupt or a posted message. Such a trigger may initiate the DMA controller and / or engine (312) to initiate one or more collection operations (e.g., redirected collection operations).
[0068] In response to a signal from the initiator (322), the DMA controller and / or engine (312) may obtain a list of collection requests in the form of a linked list. This linked list may be locable in memory (e.g., a reconfiguration buffer (316) or line accessable memory (LAM) (318)) according to an address provided by the initiator (322). According to one embodiment, this list of collection requests may include individual collection requests that are serviceable as standalone collection requests independently of other collection requests within the list of collection requests. The DMA controller and / or engine (312) may combine the addresses in these collection requests with a (potentially smaller) list of line read requests to be executed by a memory controller (e.g., the memory controller (106) of FIG. 1). Individual line read requests within a list of such line read requests may represent physical memory addresses (e.g., in LAM (308)) that specify the memory location to be read by the memory controller to service the individual line read request. The memory controller may service such line read requests by loading the requested lines into a buffer (316). As the requested lines read by the memory controller reach the buffer (316), the DMA controller and / or engine (312) may refer to the original list of collection requests (used to form the list of line read requests) to extract requested data items from the read lines reaching the buffer (316). The DMA controller and / or engine (312) may form packets from the extracted items to be forwarded to one or more request entities (e.g., host CPU (202) and / or processes running on the CN of the CN pool (220)).In one embodiment, data items extracted from lines stored in the buffer (316) may include two or more unaddressable parts of the lines (e.g., selected bytes and / or fields within an addressable line / word, such as an addressable line in memory containing data items).
[0069] FIG. 3b illustrates a flowchart of a process (350) for a collection operation according to one aspect of the present disclosure. In one embodiment, the process (350) may include operations (352, 354, 356, and 358) that can be performed by one or more circuits such as a DMA controller and / or engine (312) and / or buffer (316). Operation 352 may include processing one or more collection requests to determine one or more addressable lines of data items to be fetched from memory. For example, operation 352 may map parameters in the received requests to addresses in the LAM (308). Operation 354 may be implemented by a circuit for loading values and / or states into memory and may include loading signals and / or states (e.g., representing data items) of one or more addressable lines stored in memory, such as the LAM (308), into the buffer (316). As illustrated in FIG. 3b, operation 356 may include, for example, parsing one or more unaddressable parts of lines loaded into buffer (316) in operation 354 into parts such as single or multiple bytes within a word or group of words. Operation 358 may include, for example, completing the processing of two or more collection requests by returning the unaddressable parts parsed in operation 356.
[0070] As illustrated in FIG. 3b, operation 358 may include the step of responding to a number of collection requests presented to the DMA controller and / or engine (312) in a list of collection requests. In operation 354, loading one or more addressable lines into the buffer (316) may occur in response to the list of collection requests. Operation 354 may parse two or more unaddressable parts according to data items specified in the list of collection requests. For example, operation 356 may decode parameters in the collection requests so as to map them from line addresses to byte offsets to corresponding unaddressable parts. Then, the parsed parts may be forwarded to the initiator of the collection requests (e.g., initiator (322)). Here, the process (350) may enable the service of a number of collection requests, for example, by accessing a single addressable line loaded into the buffer (316) if the number of collection requests specify parsed parts within the same single addressable line. This eliminates the need for the DMA controller and / or engine (312) to access the same addressable line multiple times for separate collection requests for data items within the same addressable line (e.g., within the LAM (308)).
[0071] To service one or more collection requests, the process (350) may collect less than the entire addressable line in memory by loading the addressable line into a buffer and parsing the unaddressable portion to be provided to the requester. While some collection requests may call for a collection of less than the entire addressable line, one or more received collection requests may call for the collection of the entire addressable line and / or multiple lines and / or bytes spanning across the lines. For collection requests that request a collection of less than the entire addressable line, the DMA controller and / or engine (312) may execute the process (350). According to one embodiment, for collection requests that request the collection of the entire addressable line, the DMA controller and / or engine (312) may bypass operations (354, 356 and 358) to execute the collection operation without loading the addressable line into the buffer (316).
[0072] In another embodiment, the DMA controller and / or engine (312) may obtain a list of distributed requests in response to a signal from the initiator (322). The DMA controller and / or engine (312) may obtain this list of distributed requests in the form of a linked list locable in memory according to the address provided by the initiator (322). The DMA controller and / or engine (312) may then combine the addresses to be accessed by such distributed requests with a (potentially smaller) list of line read requests to be executed by a memory controller (e.g., memory controller (106)). For example, distributed requests in a list of distributed requests referencing data items within the same addressable memory line may be combined such that only a single requested line read is required (to access data items to service multiple distributed requests). The requested lines read by this memory controller may be loaded into a buffer (316).
[0073] According to one embodiment, a list of acquired distributed requests may indicate specific unaddressable portions (e.g., individual bytes or fields) of addressable lines to be read and loaded into buffer (316). Subsequently, unaddressable portions within the addressable lines loaded into buffer (316) may be modified and / or overwritten. As requested lines read by the memory controller reach buffer (316), the DMA controller and / or engine (312) may refer to the original list of distributed requests to determine specific unaddressable portions of the read lines reaching buffer (316) to be modified and / or overwritten. The DMA controller and / or engine (312) may form packets from the modified lines within buffer (316) to be written back into memory via the memory controller.
[0074] FIG. 3d illustrates a flowchart of a process (370) for a distributed operation according to one embodiment. The process (370) may include operations (372, 374, 376, and 378). Operation 372 may include, for example, receiving one or more distributed requests from an initiator (322). Operation 374 may include loading signals and / or states (e.g., representing data items) of one or more addressable lines into a buffer (316) in memory such as LAM (308). In a particular embodiment, a DMA controller and / or engine (312) may determine the addressable lines to be fetched from LAM (308) in operation 374 by processing one or more collect requests. Operation 376 may include writing values and / or states to at least one unaddressable portion of the lines loaded into the buffer (316) in operation (374) in order to at least partially modify one or more addressable lines stored in the buffer (316). For example, operation 374 may write values and / or states to two or more unaddressable portions based on a number of distributed requests received in block 372. Operation 378 may complete the service of these distributed requests by initiating write operations to write at least one of the one or more addressable lines modified in operation 374 back to memory (e.g., LAM (308)). For example, operation 378 may initiate operations to write the addressable lines loaded in operation 374 and modified in operation 376 back to the line addresses of LAM (308).
[0075] In a specific embodiment, operation 374 may be initiated by distributed requests received in operation 372. Here, one or more addressable lines of values and / or states loaded into buffer (316) in operation 374 may be obtained by the service of one or more line read requests by a memory controller (e.g., memory controller (106)). Operation 378 may include initiating the memory controller to execute one or more operations for writing one or more modified addressable lines in order to write modified addressable lines. Here, the process (370) may enable multiple distributed requests to be serviced by access to a single addressable line loaded into buffer (316). The multiple distributed requests received in operation 372 may, for example, specify words, bytes, fields, etc. within the same single addressable line. This eliminates the need for the DMA controller and / or engine (312) to access the same addressable line multiple times for separate distributed requests for data items within the same addressable line (e.g., within the LAM (308)). The process (370) may further include converting the multiple distributed requests received in operation 372 into a list of line read requests to be issued to the memory controller (a list of line read requests to specify one or more addressable lines of values and / or states). This conversion of the multiple distributed requests into a list of line read requests may further include the configuration of at least one single line read request for an addressable line in memory containing data items requested by at least two distributed requests received in operation 372.
[0076] To service one or more distributed requests, the process (370) may update less than the whole of an addressable line in memory by loading an addressable line into a buffer (316) and updating a portion of the loaded addressable line while keeping other portions unchanged. While some distributed requests may call for an update of less than the whole of an addressable line, one or more received distributed requests may call for an update of the whole of an addressable line to be written by the DMA rather than executing a read-modify-write operation. For distributed requests requesting an update of less than the whole of an addressable line, the DMA controller and / or engine (312) may execute the process (370). The DMA controller and / or engine (312) may also be configured to service distributed requests requesting an update of the whole of an addressable line by bypassing the loading of the addressable line into the buffer (316). Here, to complete these updates for the entire addressable line, the DMA controller and / or engine (312) may initiate a write operation for the addressable line of the LAM (308) without loading the addressable line into the buffer (316).
[0077] According to one embodiment, the DMA controller and / or engine (312) may receive multiple distributed requests that collectively request updates to the same overlapping portion of an addressable line within the LAM (308). This may cause conflicts regarding how the overlapping portion is updated to service multiple distributed requests. These multiple distributed requests may be ordered, for example, by creation time or reception time. According to one embodiment, conflicts regarding updating a portion of an addressable line by multiple distributed requests may be resolved, for example, by the most recent distributed request created or received.
[0078] In a specific embodiment of the processes (350 and 370), the unaddressable portion of the line stored in the buffer (316) may be bytes, a collection of bytes or fields, etc. Although specific operations in the distribution and collection operations are described as occurring in a specific sequence, specific operations may be executed simultaneously and may be executed simultaneously or in a specific sequence as a matter of engineering choice. Additionally, physical optimizations such as the number and type of processing cores to be used to realize the features of the DMA controller and / or engine (312) and interface engines, the number and type of memory blocks for the associated memory elements and buffer (316), the number of ports of the memories forming the LAM (308), and addressability features may be selected as a matter of engineering choice. For example, the buffer (316) may or may not be byte-addressable, and the DMA controller and / or engine (312) may include, for example, a scalar or vector engine.
[0079] FIGS. 4a and 4d are schematic diagrams of a computing device (400) according to one embodiment, comprising a DMA controller and / or engine (412) communicating with a buffer (416) and an initiator (422) to facilitate the redirection of DMA transactions. In a specific embodiment, the computing device (400) may include one or more features of the computing device (200) ( FIGS. 2a and 2b). According to one embodiment, the DMA controller and / or engine (412) may obtain a list of redirected read requests in response to a signal from the initiator (422). According to one embodiment, a read request may include a message and / or signal specifying one or more target memory addresses to be accessed in a read transaction to service the read request. For example, a read request may specify one or more target addresses as word-line addresses of locations in memory containing content to be retrieved in a memory read transaction to service the read request. The “redirected read request” referred to herein means a read request that has been converted or altered so that the content(s) at the original target memory address(s) are modified and / or converted to different memory address(s). Here, the different memory address(s) designate a location in memory containing the content to be retrieved in a read transaction to service the redirected read request. The DMA controller and / or engine (412) will interpret the addresses specified in these redirected read requests as, for example, word collection requests to store collected words in a buffer (416). The size of the associated word to be read first to service the word collection request portion of the redirected read request may depend, for example, at least partially on how the associated word is interpreted to form the target address for redirection. In one embodiment, the word to be read may include, for example, an address or index of an array that can be converted into an address.In one embodiment, the DMA controller and / or engine (412) may then convert the collected words stored in the buffer (416) into addresses and merge the addresses (i.e., converted collected words) with redirected read requests to form a collection request. Then, the DMA controller and / or engine (412) may service the formed collection request to transmit result data items to a destination as presented in the redirected read requests acquired in response to a signal from the initiator (422).
[0080] FIG. 4b is a flowchart of a process (450) for facilitating the redirection of DMA transactions according to one embodiment. The process (450) may include operations (452, 454, 456, and 458). Operation 454 may include executing a collection operation (e.g., a word collection operation) based at least partially on one or more redirected read requests received in operation 452. In this operation, the DMA controller and / or engine (412) may acquire or receive these redirected read requests in response to, for example, a signal from the initiator (422). In a specific embodiment, the DMA controller and / or engine (412) may interpret one or more redirected read requests as requests for a first collection operation among two collection operations to be executed. In the process of the first collection operation executed in operation 454, "collected" data items can be loaded into the buffer (416).
[0081] Operation 456 may include converting one or more collected data items stored in buffer (416) into one or more addresses to specify a subsequent collection operation. For example, operation 456 may include converting one or more words obtained from the execution of the first word collection operation (executed in operation 454) into one or more addresses (also referred to as one or more address values). In one particular embodiment, operation 456 may include parsing values and / or states within buffer (416) (from the collection operation) to determine one or more memory addresses in LAM (408). Operation 456 may further include applying one or more arithmetic operations to the parsed values and / or states to determine one or more memory addresses in LAM (408). For example, operation 456 may apply one or more arithmetic operations to the parsed values and / or states to be stored in buffer to form memory addresses at memory locations in LAM (408). These formed addresses for memory locations of LAM (408) can form a basis for subsequent collection operations.
[0082] According to one embodiment, the arithmetic operation applied in operation 456 can be defined according to the following equation (1).
[0083] Address = base + x × element_size Equation (1)
[0084] In the above formula, the address is a target address (e.g., for a collection operation determined in operation 456 or for a distribution operation in operation 476), and
[0085] x is a value obtained from the collection operation (e.g., in operation 454 or 474), and
[0086] Criteria and element_size are parameters provided in the redirected request (e.g., the redirected read request received in operation 452 or the redirected write request received in operation 472).
[0087] Operation 458 may include executing a second (e.g., subsequent) collection operation to forward data items located at one or more determined addresses to a destination. Such destination may be determined at least partially based on one or more redirected read requests. In another embodiment, operation 458 may execute two or more collection operations based on one or more addresses obtained in operation 456. According to one embodiment, when executing a collection operation, operation 458 may interpret one or more redirected read requests as two requests for a word collection operation.
[0088] According to one embodiment, a write request may include a message and / or a signal specifying one or more target memory addresses to be accessed in a memory write transaction to service the write request. For example, the write request may specify one or more target addresses as word-line addresses of locations in memory to be written in a memory write transaction (to service the write request). As referred to herein, "redirected write request" means a write request that has been converted or altered such that the original target memory address(s) are modified and / or converted to different target memory address(s). Here, the different target memory address(s) specify locations in memory to be written in a write transaction to service the redirected write request.
[0089] In another specific embodiment, the DMA controller and / or engine (412) may obtain a list of redirected write requests in response to a signal from the initiator (422). The DMA controller and / or engine (412) may interpret the addresses specified in these redirected write requests as collection requests. These collection requests may, for example, be to load collected words into a buffer (416). The associated words to be read may have a size that depends at least partially on, for example, how the associated words are interpreted to form a target address for redirection. One or more of these words loaded into the buffer (416) may be interpreted to form a target address for redirection. In one embodiment, such words loaded into the buffer (416) may include, for example, an address or index of an array that can be converted into an address. In one embodiment, the DMA controller and / or engine (412) may then convert the collected words stored in the buffer (416) into addresses and merge the addresses with redirected write requests to form a distributed request. The DMA controller and / or engine (412) may then service the formed distributed request and, as a result, update the lines according to the redirected write requests obtained in response to a signal from the initiator (422).
[0090] FIG. 4c is a flowchart of a process (470) for facilitating the redirection of DMA transactions according to one embodiment. Operation 472 may include executing a word collection operation based at least partially on one or more redirected write requests received in operation 472. The DMA controller and / or engine (412) may receive these redirected write requests in operation 472, for example, in response to a signal from an initiator (422). In a specific embodiment, the DMA controller and / or engine (412) may interpret one or more redirected write requests as requests for a word collection operation to be executed, followed by the execution of a word distribution operation. During the process of the word collection operation executed in operation 474, "collected" data items may be loaded into a buffer (416). Operation 476 may include converting one or more collected data items stored in buffer (416) (from the collection operation) into one or more addresses to specify a subsequent word distribution operation. In one particular embodiment, operation 476 may include parsing values and / or states within buffer (416) to determine one or more memory addresses in LAM (408). For example, operation 476 may be executed at least partially by a circuitry (e.g., a DMA controller and / or engine adapted to parse values and / or states) for parsing values and / or states in buffer to determine at least one of one or more addresses. For example, operation 476 may apply one or more arithmetic operations to the parsed values and / or states to be stored in buffer to form memory addresses in LAM (408). Operation 476 may further include applying one or more arithmetic operations to parsed values and / or states to determine one or more memory addresses for memory locations within LAM (408).
[0091] Operation 478 may include executing a distribution to write specific data items to one or more addresses obtained in operation 476, based at least partially on one or more redirected read requests. According to one embodiment, operation 476 may apply an arithmetic operation to calculate a target address for the distribution operation to be executed in operation 478 according to Equation (1). For example, the DMA controller and / or engine (412) may form such a distribution request based at least partially on the content at one or more addresses determined in operation 476. These specific data items to be written from such a distribution request may be specified, for example, in one or more redirected write requests received in operation 472 in response to a signal from the initiator (422). In one particular embodiment, operation 478 may include interpreting the content at the addresses determined in operation 476 as addresses for the distribution operation.
[0092] One specific embodiment of a computing device for processing multiple signal streams from multiple associated sources is illustrated in FIGS. 5a and 5c as a computing device (500). In a specific embodiment, the computing device (500) may include one or more features of the computing device (200) ( FIGS. 2a and 2b). The computing device (500) may include a pool of computing nodes (CNs) (520) having associated register files and local memories, a pool of block memories and FIFO buffers, and an NOC interconnecting various components. As noted above, the CNs in the pool of CNs (520) may include scalar CNs and / or processing circuit cores for implementing an ALU, a digital signal processor (DSP), a vector CN, a VLIW engine, or a field programmable gate array (FPGA) cell, or a combination thereof—these are merely intended to provide some examples of specific circuit cores that may be used to implement the CNs in the pool of CNs (520). In one embodiment, a CN may include a single processing circuit core capable of executing operations to map input operands to output computing results. In another embodiment, a CN may include a plurality of separate processing cores for executing operations to map input operands to output computing results. A CN in the pool of CNs (520) may include dedicated local memory (e.g., SRAM) and general registers for receiving operands for operations to be executed and / or providing results from the execution of operations. In one particular embodiment, the CNs in the pool of CNs (520) may be formed from one or more processing circuit cores that are integrated with or configured with elements of the pool of memories (e.g., transfer memory (218)) such as FIFO buffers.In one embodiment, a pool of memories (508) may provide memory resources to be shared among CNs within a pool of CNs (520) and may facilitate communication between CNs and / or between CNs within the pool of CNs (520). For example, endpoints of FIFO buffers integrated into / configured with the pool of CNs (520) may include memory blocks local to or separated from the CNs and / or register files of the CNs within the pool of CNs (520). The functionality of the CNs within the pool of CNs (520) may be that of a host processor, a processing manager or a signal processor, or a combination thereof—these are merely intended to provide a few examples. According to one embodiment, CNs in a pool of CNs (520) may support atomic, interrupt, and other inter-processor communication (IPC) features to facilitate signal communication between CNs and / or between CNs in the pool of CNs (520).
[0093] According to one embodiment, the computing device (500) may receive signal streams containing data items supplied from external sources such as sensors and / or memories (e.g., sensor signals, observations and / or measurements, timestamps, metadata, etc.). External memories (not shown) may be combined with DMA controllers and / or engines (e.g., DMA controllers and / or engines (312 and / or 412)). CNs in a pool of CNs (520) may also provide sources of signal streams containing data items. For example, a CN in a pool of CNs (520) may supply data items in a signal stream to be loaded into a FIFO buffer. A CN in a pool of CNs (520) may also supply data items in a signal stream by transmitting packets of output parameters in an IC device in a NOC as blocks of data items. According to one embodiment, a single logic signal stream may be transmitted in a multicast manner to multiple CNs within a pool of CNs (520). Additionally, a single logic signal stream may be segmented into sub-signal streams to be supplied to a subset of CNs within the pool of CNs (520). In another embodiment, some CNs within the pool of CNs (520) may process data items from various signal streams to provide output signal streams as input streams to other CNs within the pool of CNs (520). The final output of the CNs within the pool of CNs (520) may include signal streams output from the computing device (500) into sinks such as actuators, memories, storages, or display devices—these are merely a few examples intended to be presented.Processing between CNs and / or between CNs in a pool of CNs (520) can be controlled and / or coordinated through a combination of interrupts, polling of status flags, and periodic checks for operations—these are merely examples to show.
[0094] FIG. 5b is a flowchart of a method (500) (also referred to as process (500)) for processing multiple signal streams from multiple associated sources according to one embodiment. According to one embodiment, the process (500) can update the state of a particle filter (e.g., to support the automated operation of a vehicle) by combining data items received from signal streams from multiple sources. These multiple sources may include multiple sensor devices (e.g., placed in a vehicle), such as cameras, speedometers, active sensing devices (e.g., radar / lidar), environmental sensors (e.g., thermometers, light sensors, altimeters, etc.), and microphones—these are merely a few examples intended to be given. These signal streams from multiple sources may be received from a communication network as signal packets (e.g., signal packets containing sensor measurements and / or observations obtained from a remote vehicle). In a specific embodiment, data items received from signal streams from multiple sources may include sensor measurements and / or observations having common attributes identifiable by the process (500). In a specific embodiment, these common attributes may be identified by metadata placed together with the measurements and / or observations received from multiple sources. For example, these common attributes may be related by time (e.g., a timestamp indicating the time of measurement and / or observation), space (e.g., a location relative to an origin such as a point on the vehicle), and source (e.g., a specific physical sensor on the vehicle, or a specific type of sensor among the sensors mounted on the vehicle).
[0095] Operation 552 may include associating data items received from multiple data streams based at least partially on attributes common to the data items. These multiple data streams may be provided as outputs of CN and / or DMA acquisition or word acquisition or redirected read / acquire operations. For example, operation 552 may align and / or correlate (e.g., “bucketize”) measurements and / or observations received from different signal streams according to time (e.g., according to timestamps) and / or space (e.g., the position of the observed object relative to a reference point). In a specific embodiment, operation 552 may associate measurements and / or observations received from different sources (e.g., different sensors) acquired at approximately the same time with the position of a particle defined in the current state of the particle filter. In one specific embodiment, operation 552 may associate measurements and / or observations from different sources according to the locality of the specific object observed and / or measured by such associated measurements and / or observations. In another specific embodiment, data items from each of a plurality of signal streams may be loaded into a buffer (e.g., buffer (216)) associated with a direct memory access DMA controller associated with the signal stream. Operation 552 may then identify at least one of common attributes associated with the data item loaded into the buffer based at least partially on the content of the data item loaded into the buffer.
[0096] Operation 554 may include loading the associated data items from operation 552 into one or more registers of the computing node (e.g., general registers of the ALU or other processing core forming the computing node) simultaneously without storing them in line-accessible memory. According to one embodiment, the registers of the computing node may be loaded with data items to be retrieved during the execution cycles of the computing node. For example, data items loaded into the registers of the computing node in one execution cycle (e.g., at the endpoints of the FIFO buffers) may provide operands for computing operations to be executed by the computing node in the next execution cycle. In one particular embodiment, one or more registers of the computing node may include endpoints of associated FIFO buffers formed by internal memory (e.g., a pool of memories (508)). A plurality of such FIFO buffers having endpoints in the registers of a computing node may be synchronized to apply data items from a plurality of sources (e.g., loaded data items from different sensors) as operands of a computing node having common attributes (e.g., time and space attributes). Operation 556 may include the execution of a computing node to process data items simultaneously loaded in operation 554 in order to process data items loaded as operands for one or more computing operations (e.g., to perform one or more functions such as updating the state of a particle filter). Data items output from the execution of one or more computing operations in operation 556 may form data items for additional signal streams to be processed by additional computing nodes and / or for storage in memory.
[0097] In one embodiment, operation 552 may be performed by a first computing node that aligns sensor observations and / or measurements received in a plurality of signal streams based at least partially on the associated timestamps and locations of objects observed and / or measured by sensor observations and / or measurements. The aligned sensor observations and / or measurements may then be loaded into one or more registers of a second computing node in block 554. The execution of the second node in block 556 may then combine the aligned sensor observations and / or measurements.
[0098] In another embodiment, data items of an additional signal stream as outputs of operation 556 may be loaded into one or more registers of a subsequent computing node as operands for one or more additional computing operations. For example, one or more direct memory access transactions may be executed to store data items of an additional signal stream in external memory, to execute a word distribution operation to write data items of an additional signal stream, to execute a redirected write operation to write data items of an additional signal stream, to provide data items of an additional signal stream as control signals to one or more actuators, or to provide a combination thereof.
[0099] In this context, "concurrent loading" as used herein refers to the loading of data items to be processed by a computing node in the same execution cycle. When data items from synchronized signal streams are loaded concurrently into the registers of a computing node, such data items may be loaded into the registers in the same execution cycle of the computing node (e.g., may be operands of a computing operation executed in the next execution cycle). When data items from each unsynchronized signal stream are loaded concurrently into the registers of a computing node, such data items may be loaded into the registers in different (e.g., adjacent) execution cycles of the computing node. For example, the execution of the computing node may be suspended for one or more execution cycles to allow multiple data items from different unsynchronized signal streams to be loaded into the registers as operands of a computing operation in the execution cycle of the computing node. In other embodiments, transparent block and release mechanisms may be applied to the computing node to facilitate the concurrent loading of data items from unsynchronized signal streams into the registers of the computing node.
[0100] According to one embodiment, data items in at least two signal streams associated in operation 552 may include sensor observations and / or measurements associated in operation 554. Data items loaded simultaneously in operation 554 may then include sensor observations associated from a plurality of signal streams (e.g., from a plurality of different sources). In a specific embodiment, operation 552 may include combining sensor observations and / or measurements in at least two signal streams to provide combined sensor observations and / or measurements. Then, operation 552 may align and / or correlate (e.g., bucketize) the combined sensor observations and / or measurements based at least partially on the associated timestamps and locations of the objects observed and / or measured by the sensor observation and / or measurement.
[0101] In a specific embodiment, operation 554 may be implemented using at least partially process (350, FIG. 3b) and / or process (450, FIG. 4b). For example, a DMA controller and / or engine may load registers of a computing node by executing acquisition operations as described in processes (350 and / or 450) to fill queues maintained by FIFO buffers (e.g., with associated measurements and / or observations). Such execution by the DMA controller and / or engine may selectively load data items into one or more registers based at least partially on indications of common attributes within the contents of the data items loaded into the FIFO buffers. In one embodiment, multiple FIFO buffers may have endpoints in corresponding registers of a computing node where queues of different FIFO buffers are filled with data items from different signal streams (e.g., measurements and / or observations from different sensors). Queues of multiple FIFO buffers may be filled so that data items from different signal streams associated by specific attributes (e.g., time and space attributes) are simultaneously loaded into the endpoints of the FIFO buffers (e.g., general registers of the computing node). According to one embodiment, operations 552 and 554 may be performed by the first computing node to simultaneously load associated data items into one or more registers of the second computing node in one execution cycle of the second computing node. In a subsequent execution cycle of the second computing node, the second computing node may execute a computing operation using the associated data items simultaneously loaded in block 554 as operands. In one embodiment, the simultaneous loading of data items in block 554 may be facilitated by FIFO buffers having endpoints in the registers of the second computing node.Here, the first computing node may fill the queues of FIFO buffers with associated data items from different signal streams so that each associated data item reaches the endpoints of the FIFO buffers in the same or adjacent execution cycle. Filling the queues of FIFO buffers with associated data items in this way eliminates any need for the first computing node to store the associated data items in line-accessible memory (e.g., line-accessible memory (208) or RAM (108)).
[0102] In another specific embodiment, the result of executing the computing node in operation 556 may provide data items for one or more additional signal streams. In one embodiment, these data times for the additional signal streams may be loaded into one or more registers of a subsequent computing node and / or a downstream computing node. In one embodiment, data items in the output registers of the computing node (e.g., loaded from the execution of the computing operation) may be transferred to line-accessible memory by DMA write transactions, word-distributed DMA transactions, and / or redirected distributed DMA transactions (e.g., according to process (470) of FIG. 4c). These DMA transactions may store data items of the additional signal streams, for example, in line-accessible memory. In another embodiment, these data items provided in the output registers of the computing node may be applied as data items of input signal streams to one or more other computing nodes. In one particular embodiment, as a result of executing the first computing node in operation 556, results may be loaded into output registers defined as first endpoints of associated FIFO buffers (e.g., from a pool of memories (508)). However, the associated FIFO buffers may then have second endpoints defined as input registers of the second computing node to receive the results determined by the execution of the first computing node as operands.
[0103] According to one embodiment, operations 552 and 554 may be executed by a first CN of a pool of CNs (520), while operation 556 may be executed by a second CN of a pool of CNs (520). In operation 554, the first computing node may simultaneously load associated data items (e.g., data items including associated sensor observations and / or measurements based at least partially on spatial and temporal attributes, and simultaneously loaded data items including associated sensor observations and / or measurements) into one or more registers of the second CN of the pool of CNs (520). Then, in operation 556, the second CN may process the data items simultaneously loaded by the first CN as operands for one or more computing operations. In one embodiment, the first FIFO buffer may define a first endpoint as a register of the first CN (e.g., an output register of the first CN). The second endpoint of the first FIFO buffer may define a first register (e.g., an input register of the second CN) among one or more registers of the second CN. As can be observed, the first FIFO may allow data items to be loaded simultaneously into one or more registers of the second CN in operation 554 without storing the associated data items in line-accessible memory as noted above. In another embodiment, the FIFO buffer may define the first endpoint as a register of the third CN of the pool of CNs (520) and the second endpoint as at least a second register among one or more registers of the second CN. Here, both the first CN and the third CN may allow data items to be loaded simultaneously into one or more registers of the second CN in operation 554 without storing the associated data items in line-accessible memory.
[0104] According to one embodiment, all or part of a computing device (200, 300) (e.g., including features for implementing processes (350 and / or 370), such as a circuit for forming a DMA controller), a computing device (400) (e.g., including features for implementing processes (450 and / or 470), such as a circuit for forming a DMA controller), and / or a computing device (500) (e.g., including features for implementing process (550)) may be formed wholly or partially by transistors and / or lower metal interconnects (not shown) in processes such as processes for forming a complementary metal oxide semiconductor (CMOS) circuit (e.g., front-end-of-line and / or back-end-of-line processes), merely as an example. However, it should be understood that this is merely an example of how a circuit part can be formed on a device in a front-end-of-line process, and that the claimed subject matter is not limited to this aspect.
[0105] It should be noted that the various circuits disclosed herein may be described using computer-aided design tools and may be represented (or expressed) as data and / or computer-readable instructions implemented on various computer-readable media (e.g., non-transient storage media) in terms of their behavior, register transfers, logic components, transistors, layout geometries, and / or other characteristics. Formats of files and other objects in which such circuit representations may be implemented include, but are not limited to, formats supporting behavioral languages such as C, Verilog, and Very High Speed Integrated Circuit Hardware Description Language (VHDL); formats supporting register-level description languages such as Register Transfer Language (RTL); formats supporting geometric description languages such as Graphic Design System II (GDSII), Graphic Design System III (GDSIII), Graphic Design System IV (GDSIV), Caltech Intermediate Form (CIF), Manufacturing Electron Beam Exposure System (MEBES); and other suitable formats and languages. Storage media in which such formatted data and / or instructions may be implemented include, but are not limited to, various forms of non-volatile storage media (e.g., optical storage media, magnetic storage media, or semiconductor storage media) and carriers that may be used to transmit such formatted data and / or instructions via wireless, optical, or wired signaling media, or any combination thereof. Examples of transmissions of such formatted data and / or instructions via carriers may include, but are not limited to, transmission (upload, download, email, etc.) over the Internet and / or other computer networks via one or more electronic communication protocols (e.g., Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Simple Mail Transfer Protocol (SMTP), etc.).
[0106] When received within a computer system via one or more machine-readable media, these data-based and / or instruction-based representations of the aforementioned circuits may be processed by a processing entity (e.g., one or more processors) within the computer system in conjunction with the execution of one or more other computer programs, including, without limitation, a net-list generation program, a location and root program, etc., to generate a representation or image of a physical manifestation of these circuits. This representation or image may then be used in device manufacturing, for example, by enabling the generation of one or more masks used to form various components of the circuits in a device manufacturing process (e.g., a wafer manufacturing process).
[0107] In the context of this patent application, the terms "between" and / or similar terms are understood to include "among" where appropriate for a particular use, and vice versa. Likewise, in the context of this patent application, the terms "compatibility," "compliance," and / or similar terms are understood to include substantial compatibility and / or substantial compliance, respectively.
[0108] Unless otherwise indicated, in the context of this patent application, the term “or” is intended to mean A, B, and C as used in an inclusive sense herein, as well as A, B, or C as used exclusively, when used to associate a list such as A, B, or C. With this understanding, “and” is used in an inclusive sense and is intended to mean A, B, and C, whereas “and / or” may be used with sufficient care to make it clear that all the foregoing meanings are intended, but such usage is not required. Additionally, the terms “one or more” and / or similar terms are used to describe any feature, structure, characteristic, etc. in the singular form, and “and / or” is also used to describe features, structures, characteristics, etc. in the plural and / or some other combination. Likewise, the terms “based on” and / or similar terms are understood not to convey a comprehensive list of factors, but to allow for the existence of additional factors not explicitly described.
[0109] Algorithm descriptions and / or symbolic representations are examples of techniques used by a person skilled in the art of signal processing and / or related art to convey the substance of their work to another person skilled in the art. In the context of this patent application, an algorithm is generally considered to be a consistent sequence of operations and / or similar signal processing that produces a desired result. In the context of this patent application, operations and / or processing involve the physical manipulation of physical quantities. Typically, but not necessarily, such physical quantities may take the form of electrical and / or magnetic signals and / or states that can be stored, transmitted, combined, compared, processed, and / or otherwise manipulated as electronic signals and / or states constituting components of various forms of digital content, such as signal measurements, text, images, video, audio, etc.
[0110] Sometimes, primarily for reasons of general usage, it has proven convenient to refer to such physical signals and / or physical states as bits, values, elements, parameters, symbols, characters, terms, samples, observations, weights, numbers, digits, measurements, content, and / or similar terms. However, it should be understood that all such terms and / or similar terms will be associated with appropriate physical quantities and are merely convenient labels. Unless otherwise specifically stated, as is evident from the preceding discussion, it is recognized that discussions throughout this specification using terms such as “processing,” “computing,” “calculating,” “determining,” “establishing,” “acquiring,” “identifying,” “selecting,” “generating,” etc., may refer to the operations and / or processes of specific devices, such as special-purpose computers and / or similar special-purpose computing and / or network devices. Accordingly, in the context of this specification, a special purpose computer and / or similar special purpose computing and / or network device may process, manipulate, and / or convert signals and / or states, typically in the form of physical electronic quantities and / or magnetic quantities, within the memory, registers, and / or other storage devices, processing devices, and / or display devices of the special purpose computer and / or similar special purpose computing and / or network device. Accordingly, in the context of this particular patent application, the term “specific device” as referred to includes a general-purpose computing and / or network device, such as a general-purpose computer, when programmed to perform a specific function, such as based on program software instructions.
[0111] In some situations, for example, operations of a memory device, such as a state change from binary 1 to binary 0 or vice versa, may involve transformations such as physical transformations. For certain types of memory devices, such physical transformations may involve physically converting an article into a different state or thing. For example, but without limitation, for some types of memory devices, a change in state may involve the accumulation and / or storage of charge or the release of stored charge. Likewise, in other memory devices, a change in state may involve physical changes such as a change in magnetic orientation. Likewise, a physical change may involve a transformation of molecular structure, such as from a crystalline form to an amorphous form or vice versa. In other memory devices, a change in physical state may involve quantum mechanical phenomena such as superposition, entanglement, etc., which may involve, for example, quantum bits (qubits). The foregoing is not intended to be a comprehensive list of all examples that may include a transition from binary 1 to binary 0 or vice versa in a memory device, such as a transition that is physical but non-transient. Rather, the foregoing is intended as exemplary examples.
[0112] FIG. 6 is an example illustrating an exemplary sensor signal acquisition pattern for an embodiment of a vehicle (2200) capable of operating in one or more automated driving modes (e.g., fully autonomous mode, semi-autonomous mode, driver assistance mode). As illustrated, an automated vehicle such as the vehicle (2200) may include a plurality of sensors to provide measurements and / or observations to be processed according to, for example, processes (350, 370, 450, 470 and / or 550). Although a specific pattern and / or a specific number of sensors are illustrated in FIG. 6, the scope of the claims is not limited to these points. For example, a system or device such as the vehicle (2200) may include any number of sensors in any wide array and / or configuration. Additionally, although FIG. 6 is illustrated in a two-dimensional representation, for example, the sensors of a vehicle capable of operating in one or more automated modes may generate signals and / or signal packets representing conditions surrounding the vehicle (2200) in three-dimensional space. In embodiments, for example, to generate a three-dimensional representation of the conditions surrounding the vehicle (2200), one-dimensional and / or two-dimensional sensor measurements may be combined and / or processed.
[0113] In embodiments, various sensors may be mounted on the vehicle (2200) to capture observations and / or measurements of various parts of the environment surrounding and / or adjacent to the vehicle, for example. In embodiments, the vehicle (2200) may include a number of different sensors capable of detecting incoming signals, such as, for example, optical signals, electromagnetic signals, and / or sound signals. Individual sensors may have different observation areas of the environment surrounding the vehicle (2200). Exemplary views (2210a to 2210h) are illustrated, but of course, the scope of the claim is not limited to this aspect.
[0114] In embodiments, sensor signals and / or signal packets may be utilized, for example by at least one processor of the vehicle (2200), to identify objects and / or other environmental conditions near the vehicle (2200) that can be utilized by the processing system of the vehicle (2200) to autonomously guide the vehicle through the environment. Exemplary objects that may be detected in the environment surrounding a vehicle such as the vehicle (2200) may include other vehicles, trucks, cyclists, pedestrians, animals, rocks, trees, streetlights, guardrails, painted lines, traffic lights, buildings, road signs, etc. Some objects may be stationary, while other objects, such as pedestrians, may move through the environment.
[0115] In one embodiment, one or more sensors of an exemplary vehicle (2200) may generate signals and / or signal packets that may represent at least a portion of the environment surrounding or adjacent to the vehicle (2200). Other sensors may provide signals and / or signal packets that represent the speed, acceleration, orientation, position (e.g., via a Global Navigation Satellite System (GNSS)), etc. of the vehicle (2200). As described more fully below, the sensor signals and / or signal packets may be processed, for example, through a particle filter to generate a plurality of particles. These particles may be utilized to influence the operation of the vehicle (2200) at least partially. In embodiments, for example, as the vehicle (2200) proceeds through the environment, the sensor signals and / or signal states may be utilized to update the particle filter to further influence the operation of the vehicle, for example, at least partially. As discussed more fully below, any of a number of alignment operations may be performed on the sensor signals and / or signal packets in order to enable particle filters, etc. to use a relatively wide range of signals and / or signal packets generated by various sensors.
[0116] In this context, "particle" refers to a digital representation of environmental conditions for a specific point in a specific coordinate system at a specific point in time, derived at least partially from sensor signals and / or signal packets. For example, a specific particle may include an array of parameters describing a specific point within the environment surrounding the vehicle (2200) at a specific point in time. In embodiments, a "particle filter," etc., may be utilized to process sensor signals and / or signal packets to generate a plurality of particles describing an environment such as the environment surrounding the vehicle (2200). Of course, a particle filter is merely an exemplary type of processing that may be performed on sensor signals and / or signal packets, and the scope of the claim is not limited in this respect.
[0117] In embodiments, a specific coordinate system may be specified, but the scope of the claim is not limited to any specific coordinate system. In embodiments, such a coordinate system may include a three-dimensional parameter space, but other embodiments may specify a different number of dimensions. In embodiments, individual particles may belong to a specific location within a specific three-dimensional space (e.g., X, Y, and Z axes).
[0118] FIG. 7 illustrates an exemplary schematic block diagram of an exemplary vehicle (2200). As mentioned, in embodiments, the vehicle (2200) may include a plurality of sensors (2210). These sensors may include, for example, image capture (e.g., camera), radar, lidar, and / or ultrasonic sensors. In embodiments, the sensors (2210) may generate sensor signals and / or signal packets (2215) that may be provided to and / or acquired by a control system such as a control system (2220).
[0119] In embodiments, the control system (2220) may include, for example, at least one processor, at least one memory device, and / or at least one communication interface. In embodiments, the control system (2220) may include, for example, one or more central processing units (CPUs), neural network processors (NNPs), and / or graphics processing units (GPUs). In embodiments, the control system (2220) may process sensor signals and / or signal packets to generate signals and / or signal packets that may affect the operation of the vehicle (2200). For example, signals and / or signal packets may be generated by the control system (2220) and provided to and / or obtained by the driving system (2230). In embodiments, the processing of sensor signals and / or signal packets by the control system (2220) may include, for example, a particle filter, but other embodiments may utilize other signal processing algorithms, techniques, approaches, etc., and the scope of the claim is not limited in this respect. In embodiments, the driving system (2230) may include, for example, a device, mechanism, system, etc. that affects the operation of the vehicle (2200). As mentioned, as the vehicle (2200) traverses the environment, additional sensor signals and / or signal packets may be acquired and processed so that the operation of the vehicle (2200) can be updated over time.
[0120] In the foregoing description, various aspects of the claimed subject matter have been described. For the purposes of explanation, details such as quantity, system, and / or configuration have been provided as examples. In other cases, well-known features have been omitted and / or simplified to avoid obscuring the claimed subject matter. While specific features have been exemplified and / or described herein, many modifications, substitutions, changes, and / or equivalents will come to mind for those skilled in the art. Accordingly, the appended claims should be understood as intended to encompass all modifications and / or changes to the claimed subject matter.
[0121] * Exemplary Examples
[0122] 1. As a system,
[0123] Memory comprising one or more memory devices; and
[0124] The memory includes a Direct Memory Access (DMA) controller coupled to the memory by a bus, and the DMA controller
[0125] To receive one or more redirected record requests;
[0126] To obtain one or more words from the above memory, to execute a collection operation based at least partially on one or more received redirected record requests;
[0127] To convert one or more words obtained from the execution of the above collection operation into one or more addresses in the memory; and
[0128] A system configured to execute distributed operations to write data items to one or more addresses in memory.
[0129] 2. In the first item, the memory is a first memory, and the system further comprises a second memory capable of operating / functioning as a buffer, the buffer is configured to store values and / or states retrieved from the first memory, and the DMA controller is further configured to parse the values and / or states in the buffer to determine at least one address among the one or more addresses in the memory.
[0130] 3. In the second item, the system is further configured such that the DMA controller applies one or more arithmetic operations to at least one of the parsed values and / or states to determine the at least one address.
[0131] 4. In the first item, the system is further configured such that the DMA controller interprets the one or more redirected write requests as requests for the collection operation.
[0132] 5. In the first item, the system is further configured such that the DMA controller forms a request for the distributed operation based at least partially on one or more addresses in the memory.
[0133] 6. A system according to the first item, further comprising an initiator configured to initiate one or more redirected write requests, wherein the DMA controller is configured to execute a first collection operation in response to a signal from the initiator.
[0134] 7. In item 6, the initiator of the one or more redirected record requests comprises one or more registers of a computing node or local memory configured to receive sensor measurements and / or observations as data items within a signal stream.
[0135] 8. As a method in a Direct Memory Access (DMA) controller,
[0136] A step of executing a first collection operation based at least partially on one or more redirected read requests;
[0137] A step of converting one or more words obtained from the execution of the first collection operation into one or more addresses; and
[0138] A method comprising the step of executing a second collection operation to forward data items located at one or more of the above addresses to a destination determined at least partially based on one or more of the above redirected read requests.
[0139] 9. In item 8, the step of converting the one or more words obtained from the execution of the first collection operation into the one or more addresses further comprises the step of parsing values and / or states in a buffer to determine a memory address.
[0140] 10. The method of item 9, further comprising the step of applying one or more arithmetic operations to the parsed values and / or states to determine the memory address.
[0141] 11. A method according to item 8, further comprising the step of interpreting one or more redirected read requests as a first collection operation request among two collection operation requests.
[0142] 12. The method of item 8, further comprising the step of forming a second collection operation request among two collection operation requests based at least partially on one or more addresses.
[0143] 13. A method according to item 8, further comprising the step of executing the first collection operation in response to a signal from the initiator of the one or more redirected read requests.
[0144] 14. In the 13th item, the initiator of one or more redirected read requests comprises a process executed by an ALU to execute one or more sensor fusion operations.
[0145] 15. As a system,
[0146] Memory comprising one or more memory devices; and
[0147] The memory includes a Direct Memory Access (DMA) controller coupled to the memory by a bus, and the DMA controller
[0148] To receive one or more redirected read requests;
[0149] To obtain one or more words from the above memory, to execute a first collection operation based at least partially on one or more received redirected read requests;
[0150] To convert one or more words obtained from the execution of the first collection operation into one or more addresses; and
[0151] A system configured to execute a second collection operation to forward data items located at one or more of the above addresses to a destination determined at least partially based on one or more of the above redirected read requests.
[0152] 16. In item 15, the system further configured such that the DMA controller parses values and / or states in a buffer to determine one or more addresses.
[0153] 17. In item 15, the DMA controller
[0154] To parse values and / or states within the buffer; and
[0155] A system further configured to determine one or more memory addresses by applying one or more arithmetic operations to parsed values and / or states.
[0156] 18. In item 15, the DMA controller
[0157] A system further configured to interpret the above one or more redirected read requests as the first collection operation request among two collection operation requests.
[0158] 19. A system further configured in item 15 to form a second collection operation request among two collection operation requests based at least partially on one or more of the above addresses.
[0159] 20. In item 15, the system further configured such that the DMA controller executes the first acquisition operation in response to a signal from the initiator of the one or more redirected read requests.
Claims
Claim 1 A system comprising: a memory including one or more memory devices; and a direct memory access (DMA) controller coupled to the memory by a bus, wherein the DMA controller is configured to receive one or more redirected write requests; to execute a collection operation based at least partially on the received one or more redirected write requests to acquire one or more words from the memory; to translate one or more words acquired in the execution of the collection operation to one or more addresses in the memory; and to execute a distribution operation to write data items to the one or more addresses in the memory. Claim 2 A system according to claim 1, wherein the memory is a first memory, and the system further comprises a second memory capable of operating / functioning as a buffer, wherein the buffer is configured to store values and / or states retrieved from the first memory, and the DMA controller is further configured to parse the values and / or states in the buffer to determine at least one address among the one or more addresses in the memory. Claim 3 A system according to paragraph 2, wherein the DMA controller is further configured to apply one or more arithmetic operations to at least one of the parsed values and / or states to determine the at least one address. Claim 4 A system according to claim 1, wherein the DMA controller is further configured to interpret the one or more redirected write requests as requests for the collection operation. Claim 5 A system according to claim 1, wherein the DMA controller is further configured to form a request for the distributed operation based at least partially on the one or more addresses in the memory. Claim 6 A system according to claim 1, further comprising an initiator configured to initiate one or more redirected write requests, wherein the DMA controller is configured to execute the collection operation in response to a signal from the initiator. Claim 7 In paragraph 6, the initiator of the one or more redirected record requests comprises one or more registers of a computing node or local memory configured to receive sensor measurements and / or observations as data items within a signal stream. Claim 8 A method executed by a Direct Memory Access (DMA) controller comprising: executing a first collection operation based at least partially on one or more redirected read requests; converting one or more words acquired in the execution of the first collection operation to one or more addresses; and executing a second collection operation to forward data items located at the one or more addresses to a destination determined at least partially based on the one or more redirected read requests. Claim 9 A system comprising: a memory including one or more memory devices; and a Direct Memory Access (DMA) controller coupled to the memory by a bus, wherein the DMA controller is configured to receive one or more redirected read requests; to execute a first collection operation at least partially based on the received one or more redirected read requests to acquire one or more words from the memory; to translate one or more words acquired in the execution of the first collection operation to one or more addresses; and to execute a second collection operation to forward data items located at one or more addresses to a destination determined at least partially based on the one or more redirected read requests. Claim 10 A system according to claim 9, wherein the DMA controller is further configured to parse values and / or states from a buffer and to determine the one or more memory addresses by applying one or more arithmetic operations to the parsed values and / or states.
Citation Information
Patent Citations
Methods for traffic dependent direct memory access optimization and devices thereof
US11855898B1
Circuits, system, and methods for processing multiple data streams
US6055619A
System and method for scatter gather cache processing
US8495301B1