Data adaptation in hardware acceleration subsystems

By using the scheduler and type adapter in the enhanced hardware acceleration subsystem, the problem of low memory data processing efficiency in the prior art is solved, enabling flexible data processing and efficient memory transfer, thereby improving system performance and adaptability.

CN115023686BActive Publication Date: 2026-02-13TEXAS INSTRUMENTS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180012055.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-31
Filing Date
2021-01-04
Publication Date
2026-02-13
Estimated Expiration
2041-01-04

AI Technical Summary

Technical Problem

Existing hardware acceleration subsystems suffer from low access efficiency when processing data from external memory. In particular, the low efficiency of DDR memory due to the fixed block size scheme and the complex row-to-block and block-to-row conversion problems limit the flexibility and performance of the subsystem.

Method used

An enhanced hardware acceleration subsystem is employed, comprising a scheduler, DMA controller, hardware accelerator, and load storage engine. It enables conversion between rows, blocks, and aggregated blocks through data aggregation and type adapters. The scheduler and type adapter coordinate the workflow of the hardware accelerator and DMA controller to achieve flexible processing and efficient transfer of data elements.

Benefits of technology

It improves memory access efficiency, enhances the flexibility and performance of the hardware acceleration subsystem, adapts to the processing needs of different data elements, reduces CPU involvement, and improves the overall efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115023686B_ABST
    Figure CN115023686B_ABST
Patent Text Reader

Abstract

Methods, apparatus, systems, and articles of manufacture are described herein to implement data aggregation and pattern adaptation in a hardware acceleration subsystem. In some examples, a hardware acceleration subsystem (310) includes a first scheduler (382a), a first hardware accelerator (350a) coupled to the first scheduler (382a) to process at least a first data element and a second data element, and a first load-store engine (352a) coupled to the first hardware accelerator (350a), the first load-store engine (352a) configured to communicate with the first scheduler (382a) at a superblock level by sending a completion signal to the first scheduler (382a) in response to determining that a block count equals a first BPR value, and aggregate the first data element and the second data element to generate a first aggregated data element based on the first BPR value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] This description generally relates to hardware acceleration subsystems, and more particularly, to enhanced external memory transfers and pattern adaptation in hardware acceleration subsystems. BACKGROUND

[0002] While central processing units (CPUs) have improved to meet the demands of modern applications, computer performance is still limited by the large amount of data that must be processed by the CPU at the same time. Hardware accelerator subsystems can provide improved performance and / or power consumption by offloading tasks from the central processing unit (CPU) of a computer to hardware components that are specifically designed to perform those tasks. BRIEF DESCRIPTION OF DRAWINGS

[0003] Figure 1 is an example diagram of a block-based processing and storage subsystem to perform processing tasks on macroblocks fetched from external memory.

[0004] Figure 2 is an example diagram of a hardware acceleration subsystem to process data elements fetched from external memory.

[0005] Figure 3 is a block diagram of an example hardware acceleration subsystem to implement data aggregation and pattern adaptation in hardware acceleration.

[0006] Figure 4 is an example diagram illustrating an example data element aggregation to generate aggregated data elements.

[0007] Figure 5 is an example diagram illustrating an example pattern adaptation process implemented by an example pattern adapter to convert data blocks to row data elements.

[0008] Figure 6 is an example user-defined graph illustrating an example multi-consumer / multi-producer hardware acceleration subsystem for image, vision, and / or video processing.

[0009] Figure 7 is an example diagram illustrating an example multi-consumer / multi-producer hardware acceleration scheme.

[0010] Figure 8 illustrates an example multi-producer lens distortion correction (LDC) hardware accelerator to output first data elements on a first channel and second data elements on a second channel.

[0011] Figure 9 is a flowchart of machine-readable instructions of an example hardware acceleration subsystem that can be executed to implement Figure 3 .

[0012] Figure 10 It is structured for execution Figure 9 To implement the instructions Figure 3 A block diagram of the device's instance processor platform.

[0013] Figure 11 It is used to combine software (e.g., with) Figure 9 A block diagram of an instance software distribution platform that distributes the software corresponding to the instance computer-readable instructions to client devices (e.g., consumers (e.g., for licensing, selling and / or using), retailers (e.g., for selling, reselling, licensing and / or sublicensing) and / or original equipment manufacturers (OEMs) (for example, for inclusion in products to be distributed to retailers and / or directly to customers)).

[0014] The figures are not to scale. Instead, the thickness of layers or regions may be enlarged in the figures. While the figures show layers and regions with clean lines and boundaries, some or all of these lines and / or boundaries may be idealized. In reality, boundaries and / or lines may be unobservable, mixed, and / or irregular. Generally, throughout the figures and accompanying written description, the same reference numerals will be used to refer to the same or similar parts. As used herein, unless otherwise stated, the term “above” describes the relationship of two parts relative to the Earth. If the second part has at least one part between the Earth and the first part, then the first part is above the second part. Similarly, as used herein, the first part is “below” the second part when the first part is closer to the Earth than the second part. As mentioned above, the first part may be above or below the second part if it has other parts in between, does not have other parts in between, is in contact with the second part, or is not in direct contact with each other. As used herein, a statement that any part (e.g., layer, film, region, area, or plate) is located on another part in any manner (e.g., positioned on another part, located on another part, disposed on another part, or formed on another part, etc.) indicates that the mentioned part is in contact with the other part, or that the mentioned part is above the other part, wherein one or more intermediate portions are located between the two parts. As used herein, unless otherwise indicated, a connection reference (e.g., attachment, coupling, connection, and engagement) may include intermediate portions between elements referenced by the connection reference and / or relative movement between those elements. Thus, a connection reference does not necessarily imply that two elements are directly connected and / or fixed to each other. As used herein, a statement that any part is in contact with another part is defined as meaning that there is no intermediate portion between the two parts.

[0015] Unless specifically stated otherwise, the use of descriptions such as "first," "second," "third," etc., as used herein do not imply or otherwise connote priority, physical order, a list in a sequence, and / or any other meaning in any manner, but are simply used as labels and / or arbitrary names to identify elements so as to facilitate understanding of the examples described. In some examples, a description "first" can be used to refer to an element in the detailed description, while the same element can be referenced with a different description (e.g., "second" or "third") in the technical solutions. Such descriptions are used only to explicitly identify those elements that can otherwise share the same name. As used herein, "substantially real-time" refers to occurring in a near-instantaneous manner, recognizing that there can be real-world delays in computing time, transmission, etc. Thus, unless otherwise specified, "substantially real-time" refers to actual time + / - 1 second. DETAILED DESCRIPTION

[0016] In some cases, hardware acceleration can be used to reduce latency, increase throughput, reduce power consumption, and enhance parallelization of computing tasks. Commonly used hardware accelerators include graphics processing units (GPUs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), complex programmable logic devices (CPLDs), and system on chips (SoCs).

[0017] Hardware acceleration has various applications across many different fields, including the automotive industry, advanced driver systems (ADAS), manufacturing, high performance computing, robotics, drones, and other industries involving complex high-speed processing, such as hardware-based encryption, computer-generated graphics, artificial intelligence, and digital image processing, the latter of which involves various complex processing operations performed on individual images or video streams, such as lens distortion correction, scaling, warping, dense optical flow, pyramid representation, stereoscopic door effect (SDE), and other processing operations. Many of the computational tasks associated with these operations involve significant processing power, and in some cases, such as processing video streams in real-time, the amount of processing power required to process the image or video stream can place significant stress on a CPU.

[0018] Many hardware accelerators are designed to perform various computational tasks on data fetched from external memory. Many example hardware accelerators are configured to perform processing tasks on data elements in blocks or rows. For example, in image processing where imaging / visual algorithms are typically based on two-dimensional (2D) blocks of images, a hardware accelerator can be configured to process two-dimensional blocks from an image frame rather than processing the entire image frame as a row. Various example hardware accelerators can operate on block sizes of 16x16 bytes, 32x32 bytes, and 64x32 bytes.

[0019] If the hardware accelerator is implemented on a system on a chip (SoC), a direct memory access (DMA) controller can implement direct memory access to fetch data blocks or data lines from external memory and transfer data to local on-chip memory. Many types of external memory, such as double data rate synchronous dynamic random access memory (DDR SDRAM), prefer linear data access based on one-dimensional (ID) lines because a line transfer can not incur a page penalty, which can occur, for example, when two page open / close cycles are needed to access vertically adjacent pixels that fall on different pages (each page open / close cycle has a duration of about 60 ns (e.g., page penalty)).

[0020] Although DDR external memory prefers linear access, the DMA controller can access data from DDR external memory in blocks; however, the data blocks sent by DDR external memory can be fixed-size rectangular blocks with a block height corresponding to the number of lines in the DMA data block request and a fixed block width of 64 bytes. In some cases, the fixed-size rectangular blocks sent by DDR can have a block height managed by the external memory controller and / or a block width that is a function of burst size. Since the hardware accelerator can frequently operate on data lines or data blocks that are smaller than the data blocks sent by DDR external memory, the DMA controller can use only a small portion of the data blocks sent by DDR and discard excess data that must be re-fetched from DDR external memory at a later time for processing. Thus, DDR external memory can send the same data to the DMA controller multiple times before the data is processed by the hardware accelerator. Likewise, the DMA controller can send processed data to DDR external memory multiple times before the processed data is stored in DDR external memory. This redundancy can result in low operational efficiency of DDR external memory.

[0021] For example, if the DMA controller attempts to fetch a 16x16 byte data block (e.g., a data block with a block height of 16 lines and a block width of 16 bytes) from DDR external memory, DDR external memory can return a larger data block, for example, a 16x64 byte data block with a block height of 16 lines and a block width of 64 bytes. The external memory controller can write only 16 bytes in each of the 16 lines and discard the remaining 48 bytes, resulting in 25% DDR memory access efficiency. Likewise, if the DMA controller attempts to write a 16x16 byte data block to DDR external memory, the DMA can effectively consume bandwidth and time to write 16 lines of 64 bytes. Thus, DDR inefficiency can occur when fetching data from and / or writing data to DDR external memory.

[0022] While some hardware accelerators operate on data blocks as described above, other hardware accelerators can operate on data rows. Multiple hardware accelerators can be integrated into a hardware acceleration subsystem to form a hardware acceleration chain, however, as row-to-block and block-to-row conversion become increasingly complex when implemented on hardware accelerators, hardware accelerators in existing subsystems operate on the same type of data element (e.g., blocks or rows). This limitation makes customization of hardware acceleration subsystems very difficult.

[0023] Existing techniques for improving memory access efficiency by hardware accelerators are generally limited to fixed block size schemes with or without software and / or hardware based caching, resulting in inefficient transfers and simple linear hardware acceleration use case chain construction. These existing techniques include hardware acceleration subsystems that fetch fixed size macroblocks from external memory. For example, Figure 1 is an example diagram of a block-based processing and storage subsystem including an image subsystem (ISS) 100 configured using a configuration interconnect 112 configured to fetch fixed size macroblocks from system memory 110 via an ISS data interconnect 138 and store the fixed size macroblocks in local on-chip memory (e.g., a switchable buffer from a set of switchable buffers 134). Lens distortion correction (LDC2) hardware accelerator 128 and / or noise filtering (VTNF) hardware accelerator 130 can perform processing tasks on the macroblocks and send the processed macroblocks to local memory (e.g., a switchable buffer from a set of switchable switchable buffers 134) via a static controller crossbar 132. However, these existing hardware acceleration subsystems can not include hardware that allows the subsystem to adjust the size of the macroblocks as needed by the individual hardware accelerators. Rather, the size of the macroblocks in these subsystems is primarily driven by the input buffer 134, and the output block size is defined by the input block scaling factor, the output image buffer size, and / or the input block local memory size. The buffers in these subsystems are typically switchable buffers without a mechanism for combining blocks. This lack of control and storage capability in existing hardware acceleration subsystems can limit the flexibility of the subsystem to combine blocks into various sizes. While other hardware accelerators can involve a CPU to combine multiple adjacent blocks to create one bounding box, the involvement of the CPU typically results in area and performance cost of the CPU pipeline.

[0024] Figure 2This is an example diagram 200 of a hardware acceleration subsystem 210, which is integrated into a single-chip system (SoC) 220 and configured to retrieve data elements 231, 232 from external memory 230, process data elements 231, 232 to generate processed data elements 236, 238, and write the processed data elements 236, 238 to external memory 230.

[0025] Figure 2 The hardware acceleration subsystem 210 illustrated in the figure includes: a first direct memory access (DMA) controller 240 for facilitating the transfer of data elements 231, 232 from external memory 230 (e.g., from input frame 237 stored in external memory 230) to local memory 260; and a second DMA controller 242 for facilitating the transfer of processed data elements 236, 238 from local memory 260 and / or hardware accelerators 250a, 250b, 250c, 250d to external memory 230. (e.g., the transmission of output frame 239 to external memory 230); four hardware accelerators 250a, 250b, 250c, 250d for performing various processing tasks on data elements 231, 232 to produce intermediate data elements 233, 234 and / or processed data elements 236, 238; local memory 260 for temporarily storing data elements 231, 232 and / or intermediate data elements 233, 234 during processing; and scheduler 280 for coordinating the workflow between hardware accelerators 250a to 250d, local memory 260 and DMA controllers 240, 242.

[0026] exist Figure 2 In the example illustrated herein, hardware accelerators 250a to 250d are configured to consume data elements 231, 232 and / or intermediate data elements 233, 234 as input, perform processing tasks on data elements 231, 232 and / or intermediate data elements 233, 234, and generate processed data elements 236, 238 as output for consumption by another hardware accelerator 250a to 250d, written to local memory 260 and / or written to DDR external memory 230 via DMA controller 242. Figure 2In the example of FIG. 2, the hardware acceleration subsystem 210 is configured to process multiple data elements 231, 232 in parallel, e.g., the first hardware accelerator 250a performs a first processing task on the first data element 231 to produce an intermediate data element 233, while the second hardware accelerator 250b performs a second processing task on the second data element 232 to produce an intermediate data element 234. The scheduler 280 facilitates the workflow of the hardware accelerators 250a-250d, DMA controllers 240, 242, and local memory 260 as the data elements 231, 232 progress along the hardware acceleration pipeline.

[0027] In some examples, an enhanced hardware acceleration subsystem for improved DDR access and implementation of data adaptation for multiple producers and consumers includes a first hardware accelerator to perform a first processing task on a first data element, a scheduler to control workflow of the hardware accelerator and data aggregation, and a load store engine coupled to the first hardware accelerator to aggregate the first data element with a second data element in a local memory. In some examples, the scheduler includes a format adapter to implement conversion between a row, a block, and an aggregated block.

[0028] Figure 3 is a block diagram of an example hardware acceleration subsystem 310 to implement data aggregation and format adaptation. The example hardware acceleration subsystem 310 includes an example DMA controller 340 coupled to an example channel mapper (e.g., an example DMA scheduler 382d), an example first hardware accelerator 350a coupled to an example first scheduler 382a, an example second hardware accelerator 350b coupled to an example second scheduler 382b, an example third hardware accelerator 350c coupled to an example third scheduler 382c, an example local memory 360, and an example master hardware thread scheduler (HTS) 380. In some examples, the example first hardware accelerator 350a includes an example first load store engine 352a, the example second hardware accelerator 350b includes an example second load store engine 352b, and the example third hardware accelerator 350c includes an example third load store engine 352c. In some examples, the example hardware acceleration subsystem 310 includes an example memory mapping register (MMR) controller 392 coupled to the example HTS 380, the example schedulers 382a-382d, the example load store engines 352a-352c, and / or the example DMA controller 340. In some examples, the example MMR controller 392 is software (SW) programmable.

[0029] In Figure 3In the instanced hardware acceleration subsystem 310 illustrated in the diagram, instanced schedulers 382a, 382b, 382c, and 382d respectively include instanced consumer sockets 384a, 384b, 384c, and 384d, which are configured to track input data consumed by the corresponding instanced hardware accelerators 350a to 350c and the corresponding instanced DMA controller 340. Figure 3 In the instanced hardware acceleration subsystem 310 illustrated in the figure, instanced schedulers 382a to 382d respectively include instanced producer sockets 386a, 386b, 386c, and 386d, which are configured to track output data generated by the corresponding instanced hardware accelerators 350a to 350c and the corresponding instanced DMA controller 340. In some instances, instance first scheduler 382a includes instance first producer type adapter 390a coupled to instance first producer socket 386a, instance second scheduler 382b includes instance second producer type adapter 390b coupled to instance second producer socket 386b, instance third scheduler 382c includes instance third producer type adapter 390c coupled to instance third producer socket 386c, and instance DMA scheduler 382d includes instance DMA type adapter 390d coupled to instance DMA producer socket 386d.

[0030] exist Figure 3 In the illustrated instance hardware acceleration subsystem 310, an instance DMA controller 340 facilitates the transfer of data elements (e.g., data blocks) between instance local memory 360 and instance external memory 330 (e.g., DDR external memory or other off-chip memory outside the instance hardware acceleration subsystem 310). In some instances, the instance DMA controller 340 communicates with an external memory controller (e.g., a DDR controller) to transfer data elements between instance local memory 360 and instance external memory 330. In some instances, the instance DMA controller 340 transfers data elements between instance hardware acceleration subsystem 310, instance external memory 330, and / or other components and / or subsystems in the SoC via a common bus.

[0031] In some examples, the example DMA controller 340 communicates with and is coupled to the example DMA schedulers 382d via the example crossbar 370. In some examples, the example DMA schedulers 382d perform scheduling operations similar to the example schedulers 382a-382c corresponding to the example hardware accelerators 350a-350c. In some examples, the example DMA schedulers 382d map DMA channels to the example hardware accelerators 350a-350c. In some examples, when a data transfer is initiated via a DMA channel (e.g., a DMA channel corresponding to hardware accelerator 350a), the example DMA controller 340 communicates a channel start signal. In some examples, when a data transfer is completed via a DMA channel (e.g., a DMA channel corresponding to hardware accelerator 350a), the example DMA controller 340 communicates a channel completion signal.

[0032] In Figure 3 the example hardware accelerator subsystem 310 illustrated in FIG. 3, the example DMA controller 340 fetches data elements from the example external memory 330 for consumption by at least one of the hardware accelerators 350a-350c. In some examples, the data elements are stored contiguously in the example local memory 360. In some examples, in response to instructions from the example schedulers 382a-382d and / or the example HTS 380, the example DMA controller 340 transfers processed data elements to the example external memory 330, e.g., to an output frame.

[0033] In Figure 3 the example hardware accelerator subsystem 310 illustrated in FIG. 3, the example hardware accelerators 350a-350c are configured to perform processing tasks on data elements (e.g., data blocks and / or data lines). In some examples, the example hardware accelerators 350a-350c are configured to perform image processing tasks, e.g., lens distortion correction (LDC), scaling (e.g., MSC), noise filtering (NF), dense optical flow (DOF), stereo door effect (SDE), or any other processing task suitable for image processing.

[0034] In some examples, the example first hardware accelerator 350a operates on data blocks. In some examples, the example first hardware accelerator 350a operates on 16x16B data blocks, 32x32B data blocks, 64x32 data blocks, or any other data block size suitable for performing processing tasks. In some examples, at least the example first scheduler 382a coupled to the example first hardware accelerator 350a includes a plurality of consumer sockets 384a and / or a plurality of example producer sockets 386a that can be connected to the example second scheduler 382b. For example, an ISS hardware accelerator can have 6 outputs (Y12, UV12, U8, UV8, S8, and H3A), and an LDC hardware accelerator can have two outputs (Y, UV) or three outputs (R, G, B).

[0035] In some examples, the example first consumer socket 384a and / or the example first producer socket 386a of the example first hardware accelerator 350a are connected to the example second consumer socket 384b and / or the example second producer socket 386b of the example second hardware accelerator 350b via the example crossbar 370 of the example HTS 380 to form a dataflow chain. In some examples, the dataflow chain is configured by the example MMR controller 392. In some examples, the example MMR controller 392 is software (SW) programmable. In some examples, the example first hardware accelerator 350a is configured to perform a first task on data elements independent of the example second hardware accelerator 350b, e.g., the example hardware accelerators 350a-c are configured to perform processing tasks on data elements in parallel.

[0036] Although Figure 3 Although the example hardware acceleration subsystem 310 includes three example hardware accelerators 350a-c and one example DMA controller 340 for illustrative purposes, the example hardware acceleration subsystem 310 can include any number of example hardware accelerators 350a-c and / or DMA controllers 340. Further, the example hardware acceleration subsystem 310 can include different types of example hardware accelerators 350a-c and / or example hardware accelerators 350a-c that operate on different types of data (e.g., blocks or lines) and / or perform different processing tasks (e.g., LDC, scaling, and noise filtering), allowing a user to customize the example hardware acceleration subsystem 310 for a variety of functions.

[0037] In Figure 3In the example hardware acceleration subsystem 310 illustrated in FIG. 3, the example schedulers 382a-d communicate with the corresponding example hardware accelerators 350a-c and the corresponding example DMA controllers 340 to control the processing workflow of the example hardware accelerators 350a-c and the example DMA controllers 340. In some examples, the example first scheduler 382a controls the workflow of the example first hardware accelerator 350a. In some examples, the example first scheduler 382a sends a start signal (e.g., a Tstart signal) to the example first hardware accelerator 350a to communicate to the example first hardware accelerator 350a to start processing a data element. In some examples, the example first hardware accelerator 350a sends a completion signal (e.g., a Tdone signal) to indicate that the example first hardware accelerator 350a has completed processing a data element. In some examples, in response to receiving the Tdone signal, the example first scheduler 382a instructs the example DMA controller 340 to fetch another data element from the example external memory 330. In some examples, the example first scheduler 382a sends a start signal to the example first hardware accelerator 350a to indicate to the example hardware accelerator 350a that a frame starts processing. In some examples, the example first hardware accelerator 350a sends a frame end signal to the example first scheduler 382a to communicate that a frame ends processing, e.g., that the example first hardware accelerator has completed processing a frame.

[0038] In Figure 3 In the example hardware acceleration subsystem 310 illustrated in FIG. 3, the example schedulers 382a-d include respective example consumer sockets 384a-d to track consumed input data (e.g., data elements fetched from the example local memory 360) and respective example producer sockets 386a-d to track produced output data (e.g., data elements processed by the corresponding example hardware accelerators 350a-c and the corresponding example DMA controllers 340). In some examples, the example first hardware accelerator 350a includes multiple consumer sockets 384a and / or multiple producer sockets 386a. For example, the example first hardware accelerator can include an example first consumer socket 384a and / or an example first producer socket 386a to input / output data on a chroma channel and a second consumer socket and / or an example producer socket to input / output data on a luma channel.

[0039] In some examples, the example consumer sockets 384a-384d include consumer dependencies and the example producer sockets 386a-386d include example producer dependencies. In some examples, the consumer dependencies and the producer dependencies are specific to the corresponding example hardware accelerators 350a-350c and the corresponding example DMA controllers 340. In some examples, the example consumer sockets 384a-384d are configured to generate a signal, e.g., a dec signal, indicating consumption of generated data in response to the corresponding example hardware accelerators 350a-350c and the corresponding example DMA controllers 340 consuming data. In some examples, the example producer sockets 386a-386d are configured to generate a signal, e.g., a pend signal, indicating availability of consumable data in response to the corresponding example hardware accelerators 350a-350c and the corresponding example DMA controllers 340 generating consumable data. In some examples, the example dec signals are routed to the corresponding example producers and the example pend signals are routed to the corresponding example consumers.

[0040] Figure 3 The example schedulers 382a-382c of the example hardware acceleration subsystem 310 illustrated in FIG. 3 include example producer-style adapters 390a, 390b, 390c, 390d coupled to the example producer sockets 386a, 386b, 386c, 386d to perform logical conversions between rows, blocks, and aggregated block formats.

[0041] In some examples, the example schedulers 382a-382d implement aggregation of groups of output data (e.g., a first data element and a second data element when the first data element and the second data element have the same data type). In some examples, the example schedulers 382a-382d implement logical conversions of data elements and / or aggregated data elements between a first data type and a second data type (e.g., row to 2D block and block to 2D row). Thus, the example schedulers 382a-382d implement at least four scenarios, e.g., row to row, row to 2D block, 2D block to row, and 2D block to 2D block.

[0042] Figure 3The instance hardware accelerator subsystem 310 illustrated herein includes instance workload storage engines 352a, 352b, and 352c coupled to corresponding instance hardware accelerators 350a to 350c. Instance workload storage engines 352a, 352b, and 352c are configured to aggregate at least a first data element (e.g., a first data block) and a second data element (e.g., a second data block) in instance local memory 360 to generate aggregated data elements (e.g., superblocks), and / or to partition the aggregated data elements into at least the first data element and the second data element. In some instances, instance first workload storage engine 352a is configured to aggregate the first data element and the second data element in instance local memory 360. In some instances, the instanced first load storage engine 352a horizontally aggregates data elements based on a block-per-row (BPR) value (e.g., CBUF_BPR) programmed into the instanced MMR controller 392 by the user. This allows tuning based on, for example, the instanced hardware accelerators 350a to 350c, output block size, DDR burst size, destination consumption type, and instanced local memory 360. In some instances, the instanced load storage engines 352a to 352c enable a software (SW) programmable circular buffer stored in the instanced local memory 360 for data aggregation based on the block-per-row (BPR). In some instances, the BPR value is determined by software based on available memory in the instanced local memory 360 and / or memory allocated in the instanced local memory 360 for the instanced hardware accelerators 350a to 350c. In some instances, the BPR value is hard-coded into the instanced MMR controller 392.

[0043] Figure 4 This is a diagram illustrating the use of instanced local memory 360 ( Figure 3 In ) through instanced load storage engines 352a to 352c ( Figure 3 The instantiation of data elements 402, 404, 406, and 408 is performed to generate an instantiation schema of aggregated data elements 420a and 420b (e.g., superblocks 420a and 420b). Figure 4 In the example illustrated herein, instance data elements 402, 404, 406, and 408 are stored in local memory in a first configuration 410 (e.g., Figure 3 In instance local memory 360. In some instances, instance data elements 402, 404, 406, 408 are stored in instance local memory 360 with instance first configuration 410 having (for example) a BPR value of 1 (e.g., a block width) and a buffer size of 4 (e.g., CBUF_SIZE=OBH*4). Figure 3) based on the BPR value received from the example MMR controller 392. In some examples, the example first load store engine 352a horizontally aggregates the data elements 402, 404, 406, 408 in the example local memory 360 to generate an example second configuration 420 including two superblocks 420a, 420b, each having a width of two blocks (e.g., BPR = 2), based on the BPR value received from the example MMR controller 392. Figure 4 The example second configuration 420 illustrated in FIG. 4B has a buffer size of 2 (e.g., CBUF SIZE = OBH * 2). The example first load store engine 352a can be configured to horizontally aggregate any suitable number of blocks into example superblocks 420a, 420b having any suitable width as determined by the BPR value from the example MMR controller 392. In some examples, the example first load store engine 352a horizontally aggregates the processed data blocks 402, 404, 406, 408 received from the example first hardware accelerator 350a, writes the processed data blocks 402, 404, 406, 408 to the example local memory 360, and aggregates the data blocks 402, 404, 406, 408 in the example local memory 360 to generate the superblocks 420a, 420b. Figure 4 The example second configuration 420 illustrated in FIG. 4B has a buffer size of 2 (e.g., CBUF SIZE = OBH * 2). The example first load store engine 352a can be configured to horizontally aggregate any suitable number of blocks into example superblocks 420a, 420b having any suitable width as determined by the BPR value from the example MMR controller 392. In some examples, the example first load store engine 352a horizontally aggregates the processed data blocks 402, 404, 406, 408 received from the example first hardware accelerator 350a, writes the processed data blocks 402, 404, 406, 408 to the example local memory 360, and aggregates the data blocks 402, 404, 406, 408 in the example local memory 360 to generate the superblocks 420a, 420b. Figure 4 The example horizontally aggregated superblocks 420a, 420b illustrated in FIG. 4B can enable larger reads / writes between the example local memory 360 Figure 3 ) and the example external memory 330.

[0044] In some examples, the example load store engines 352a-352c are configured to select individual data elements 402, 404, 406, 408 from the corresponding superblocks 420a, 420b. Thus, in some examples, the example load store engines 352a-352c are configured to aggregate individual data elements 402, 404, 406, 408 to produce aggregated data elements 420a, 420b and / or select individual data elements 402, 404, 406, 408 from the corresponding superblocks 420, 420b, depending, for example, on the format of the data that the corresponding example hardware accelerator 350a-350c is configured to operate on and the format of the data transferred from the example external memory 330 to the example local memory 360.

[0045] In some examples, the instance load store engine 352a-c receives processed data elements 402, 404, 406, 408 from the corresponding instance hardware accelerator 350a-c, aggregates the processed data elements 402, 404, 406, 408 to generate aggregated data elements 420a, 420b, and writes the aggregated data elements 420a, 420b to the instance local memory 360. In some examples, the instance load store engine 352a-c receives processed instance aggregated data elements 420a, 420b from the corresponding instance hardware accelerator 350a-c, selects individual processed data elements 402, 404, 406, 408 from the processed aggregated data elements 420a, 420b, and writes the data elements 402, 404, 406, 408 to the instance local memory 360. In some examples, data blocks can be aggregated into rows 430a, 430b (e.g., 2D block-to-row rasterization). In some examples, data blocks can be aggregated into rows 430a, 430b by setting a BPR value as a function of frame width (e.g., BPR = FR_WIDTH / OBW). In some examples, the rasterized data rows 430a, 430b can be transferred to the instance external memory 330 by the instance DMA controller 340 Figure 3 .

[0046] In some examples, the instance first hardware accelerator 350a generates a completion signal (e.g., a Tdone signal) and sends the Tdone signal to the instance first scheduler 382a Figure 3 in response to the instance first hardware accelerator 350a completing processing of the instance data elements 402, 404, 406, 408 Figure 4 . In some examples, in response to receiving the Tdone signal, the instance first scheduler 382a instructs the instance second hardware accelerator 350b to read processed data elements 402, 404, 406, 408 Figure 4 or aggregated data elements 420a, 420b or 430a or 430b. In some examples, the instance second hardware accelerator 350b consumes aggregated data elements with whole row blocks (e.g., aggregated data elements 430a, 430b of Figure 3 In some examples, in response to receiving the Tdone signal, the instance first scheduler 382a Figure 3 instructs the instance DMA controller 340 to write processed data elements 402, 404, 406, 408 Figure 4 or aggregated data elements 420a or 420b or 430a or 430b from the instance external memory 330 .

[0047] In some instances, in response to the instanced first hardware accelerator 350a ( Figure 3 Process data elements 402, 404, 406, and 408. Figure 4 ), instance-based first load storage engine 352a ( Figure 3 The block count is incremented. In this way, the instanced first load storage engine 352a tracks the data generated by the instanced first hardware accelerator 350a. Figure 3 The data elements processed are 402, 404, 406, and 408. Figure 4 The number of (). In some instances, in response to the instanced first load storage engine 352a determining that the block count equals the BPR value, the instanced first load storage engine 352a aggregates the processed data elements 402, 404, 406, 408 in local memory 360. Figure 4 This generates aggregated data elements 420a and 420b. In some instances, in response to the instanced first load storage engine 352a determining that the block count equals the BPR value, the instanced first hardware accelerator 350a sends a Tdone signal to the instanced first scheduler 382a, at which point the instanced first scheduler 382a may instruct the instanced second hardware accelerator 350b or the instanced third hardware accelerator 350c to read the aggregated data elements 420a and 420b from the instanced local memory 360. In some instances, in response to the Tdone signal, the instanced first scheduler 382a instructs the instanced DMA controller 340 to transfer the aggregated data elements 420a, 420b, 430a, or 430b to the instanced external memory 330. Figure 3 ).

[0048] As described above, the instanced first hardware accelerator 350a and / or the instanced first load storage engine 352a may implement counting logic, which includes incrementing the block count in response to the instanced first hardware accelerator 350a processing data elements. In some instances, Figure 3 The instanced hardware acceleration subsystem 310 illustrated in the diagram includes generation mode parameters (e.g., the Tdone_gen_mode parameter) to enable instanced hardware accelerators 350a to 350c to operate at the block level (e.g., at individual data elements 402, 404, 406, 408). Figure 4) or at a superblock level (e.g., at an aggregated data element level) with the corresponding instance-specific scheduler 382a-c. In some instances, the generation mode parameter is MMR programmable and / or based on the BPR value. In some instances, in a first generation mode (e.g., when Tdone_gen_mode = 0), the instance-specific hardware accelerator 350a-c communicates with the instance-specific scheduler 382a-c at a block level, e.g., the instance-specific hardware accelerator 350a-c sends a Tdone signal to the corresponding instance-specific scheduler 382a-c immediately after processing individual data elements 402, 404, 406, 408 ( Figure 4 ). In some instances, in a second generation mode (e.g., when Tdone_gen_mode = 1), the instance-specific hardware accelerator 350a-c communicates with the instance-specific scheduler 382a-c at a superblock level, e.g., the instance-specific hardware accelerator 350a-c sends a Tdone signal to the corresponding instance-specific scheduler 382a-c immediately after processing a superblock based on the BPR value (e.g., when the instance-specific hardware accelerator 350a-c has processed a number of data elements 402, 404, 406, 408 equal to the BPR value). For example, if the BPR value is 2 (e.g., Figure 4 ), and if the instance-specific hardware accelerator 350a-c is processing superblock 420a ( Figure 4 ), is communicating with the corresponding instance-specific scheduler 382a-c (e.g., Tdone_gen_mode = 1), then the instance-specific hardware accelerator 350a-c sends a Tdone signal to the corresponding instance-specific scheduler 382a-c immediately after processing two data elements 402, 404.

[0049] The flexibility of communicating at a block level or at a superblock level prevents the instance-specific scheduler 382a-c and / or the instance-specific HTS 380 from triggering a DMA transfer after the instance-specific first hardware accelerator 350a processes a single data block 402, 404, 406, 408 ( Figure 4 ) of a superblock 420a, 420b. For example, if the instance-specific first load store engine 352a aggregates two data blocks 402, 404 ( Figure 4 ) horizontally into a superblock (e.g., Figure 4If the instance of the first scheduler 382a and / or the instance of the HTS 380 is configured to communicate with the instance of the hardware accelerator 350a at the block level (e.g., first generation mode), then the instance of the first scheduler 382a and / or the instance of the HTS 380 can trigger a DMA transfer after one data block 402 or 404 is processed rather than waiting until both data blocks 402, 404 in the aggregated data block 420a are processed.

[0050] In some instances, the aggregated data elements (e.g., instance of the aggregated data element 430a, 430b of the third configuration 430) have a BPR value equal to a frame width of the input frame processed by the instance of the hardware accelerator subsystem 310 (e.g., BPR = FR_WIDTH / OBW). In scenarios where the frame width of the input frame is not a multiple of the BPR value (e.g., when frame width = 10 blocks and BPR value = 4), the aggregated data element 430a can include an end of row (EOR) trigger mode (e.g., partial_bpr_trigmode) to account for scenarios where the frame width of the input frame is not a multiple of the superblock size. For example, if the instance of the hardware accelerator 350a-c is operating in the EOR trigger mode while the BPR value is 4 and the remaining superblock buffer has two blocks, then the number of blocks in the superblock buffer is 50% of the BPR value and the instance of the hardware accelerator 350a-c will send an EOR trigger to the instance of the corresponding scheduler 382a-c and / or the instance of the HTS 380 after processing the two blocks in the superblock buffer at the end of the row. In some instances, when the instance of the hardware accelerator 350a is operating in the EOR trigger mode (e.g., partial_bpr_trigmode = 1), the instance of the corresponding scheduler 382a-c and / or the instance of the HTS 380 triggers the instance of the DMA controller 340 to transfer the EOR superblock to the instance of the external memory 330 via a separate DMA channel. In some instances, the instance of the first load store engine 352a communicates with the instance of the first scheduler 382a via a partial BPR count (e.g., partial_bpr_count mode) to indicate the remaining block count in the EOR superblock buffer. Figure 3

[0051] ​By combining multiple data blocks horizontally as described herein and enabling instance hardware accelerators 350a to 350c to communicate at the block level and / or superblock level, instance workload storage engines 352a to 352c enable larger reads and writes to instance local memory 360 from / to instance external memory 330, thereby improving DDR efficiency. With larger memory requests (up to frame width), DDR page opening / closing is significantly reduced.

[0052] Figure 5 This is an instance diagram illustrating the instance type adaptation process 500, which is implemented by an instance type adapter to logically convert a 24x32-byte data block 532 into 24 rows of data elements 534. Figure 5 In the instance-type adapter diagram 500, the instance-type consumer socket 584 allows the instance-type first hardware accelerator 350a and / or the instance-type first load storage engine 352a to read from the instance-type local memory 360 and tracks the number of times the instance-type first hardware accelerator 350a and / or the instance-type first load storage engine 352a reads from the instance-type local memory 360. For example, in Figure 5 In the example illustrated, the instance-type adapter 588 can logically convert a 24x32 data block 532 into 24 rows of data elements. In some instances, instance-type hardware accelerators (e.g., Figure 3 The instance of the first hardware accelerator 350a) performs processing tasks on 24 rows of data elements 534.

[0053] In some instances, instance-based schedulers (e.g., Figure 3 The instanced first scheduler 382a) reads the Tdone signal from the instanced first hardware accelerator 350a, and in response, the instanced type adapter 588 logically converts the 24 rows of data elements into 24x32B data blocks, and the instanced producer socket 586 generates processed data blocks as output data for consumption by another hardware accelerator and / or via the instanced DMA controller (e.g., Figure 3 The data is transferred to the instance external memory 330 via the instance DMA controller 340.

[0054] By transforming data elements between rows, blocks, and superblocks, Figure 3 Instance-type producer adapters 390a to 390d and / or Figure 5 The instance-type adapter 588 enables the instance-type hardware acceleration subsystem 310 ( Figure 3) different types of outputs, thereby enabling complex user-defined multi-producer and multi-consumer hardware acceleration schemes for various functions while maintaining improved efficiency of DDR external memory (e.g., example external memory 330) of Figure 3

[0055] Figure 6 is an example user-defined graph 600 illustrating an example multi-consumer / multi-producer hardware acceleration subsystem 610 for image, vision, and / or video processing. Figure 6 The example hardware acceleration subsystem 610 illustrated in includes an example lens distortion correction (LDC) hardware accelerator 650a to perform lens distortion correction on data blocks, an example magnification scaling (MSC) hardware accelerator 650b to perform magnification scaling on data rows, an example noise filtering (NF) hardware accelerator 650c to perform noise filtering on data rows, an example first DMA controller 640 in communication with the example LDC hardware accelerator 650a and the example DDR external memory 630, an example second DMA controller 642 in communication with the example LDC hardware accelerator 650a and the example DDR external memory 630, an example third DMA controller 644 in communication with the example MSC hardware accelerator 650b, the example NF hardware accelerator 650c, and the example DDR external memory 630, and an example fourth DMA controller 646 in communication with the example NF hardware accelerator 650c and the example DDR external memory 630.

[0056] Figure 6 In the example user-defined hardware acceleration subsystem 610 illustrated in Figure 6 , the example LDC hardware accelerator 650a generates multiple outputs including, for example, data blocks 632 consumed by the example first DMA controller 640. In the example of Figure 6 , the example MSC hardware accelerator 650b consumes a set of data rows based on the output of the example LDC hardware accelerator 650a and the example second DMA controller 642 consumes data blocks 634. In the example of Figure 6 , the example MSC hardware accelerator 650b consumes a set of data rows based on the output of the example LDC hardware accelerator 650a, performs magnification scaling operations on the data rows, and generates data row elements 636 consumed by the example NF hardware accelerator 650c and the example third DMA controller 644. In the example ofIn the example of FIG. 6C, the example NF hardware accelerator 650c consumes the data line element 636, performs a noise filtering operation on the data line element 636, and generates a data line element 638 that is consumed by the example fourth DMA controller 646. In the example of FIG. 6C, the example first DMA controller 640, the example second DMA controller 642, the example third DMA controller 644, and the example fourth DMA controller 646 are configured to write the respective data elements 632, 634, 636, 638 to the example DDR external memory 630. Figure 6

[0057] Figure 7 FIG. 7 is an example diagram illustrating an example multi-consumer / multi-producer hardware acceleration scheme. Figure 7 The example hardware acceleration subsystem 710 illustrated in FIG. 7 includes an example DDR external memory 730, an example DMA controller 740, an example LDC hardware accelerator 750a, an example MSC / NF hardware accelerator 750b, and an example HTS 780. In the example hardware acceleration subsystem 710, the example LDC hardware accelerator 750a is configured to perform a lens distortion correction operation on data to generate a data element 732. In the example hardware acceleration subsystem 710, the example MSC / NF hardware accelerator 750b is configured to consume the data element 732, perform scaling and noise filtering operations on the data element 732, and generate a data element 734. In the example hardware acceleration subsystem 710, the example DMA controller 740 is configured to consume the data elements 732, 734 generated by the example LDC hardware accelerator 750a and the example MSC hardware accelerator 750b, respectively, and write the data elements 732, 734 to the example DDR external memory 730. Figure 7

[0058] The aggregation requirements can differ based on the consumer of the data, given local memory availability. For example, the MSC hardware accelerator 650b ( Figure 6 ) can require that a full group of line data be available, while the DMA CH writeout can fit aggregating a few blocks to save DDR bandwidth. To enable different aggregations, each output channel can be programmed in the LSE 352a to include different BPR values.

[0059] Figure 8 FIG. 8 illustrates an example multi-producer LDC hardware accelerator 880 to output a first aggregated data element 832 on a first channel (e.g., a chroma channel) and a second aggregated data element 834 on a second channel (e.g., a luma channel). In the example of FIG. 8, the example LDC hardware accelerator 880 is configured to perform a lens distortion correction operation on a data element 832, 834 to generate a data element 832, 834 that is consumed by the example DMA controller 840. Figure 8 ​​In the example illustrated in FIG. 8, the first channel is associated with a first BPR value (e.g., a first BPR value of 4) and the second channel is associated with a second BPR value (e.g., a BPR value equal to the frame width). In some examples, the first BPR value and / or the second BPR value is based on a number of blocks in the frame row, a number of pixels in the frame row, and / or a number of bytes in the frame row. In Figure 8 In the example illustrated in FIG. 8, the first aggregated data element 832 is output to an external DDR (e.g., the example external memory 330) and the second data element 834 is output to a second hardware accelerator (e.g., an MSC / NF hardware accelerator). Thus, the examples described herein enable separate asymmetric data element aggregation (e.g., on separate data channels). Figure 3

[0060] Although an example manner of implementing the hardware acceleration subsystem 310 of Figure 9 is illustrated in FIG. 8, one or more of the elements, processes and / or devices illustrated in FIG. 8 can be combined, divided, re-arranged, omitted, eliminated and / or implemented in any other way. Further, the example hardware accelerators 350a-350c, the example schedulers 382a-382d, the example load store engines 352a-352c, the example producer style adapters 390a-390d, the example DMA controller 340, the example local memory 360, and / or, more generally, the example hardware acceleration subsystem of FIG. 8 can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, any of the example hardware accelerators 350a-350c, the example schedulers 382a-382d, the example load store engines 352a-352c, the example producer style adapters 390a-390d, the example DMA controller 340, the example local memory 360, and / or, more generally, the example hardware acceleration subsystem of FIG. 8 could be implemented by one or more digital Figure 3 Figure 9 Figure 3 Figure 3 ​​​​Any of the example hardware acceleration subsystems 310 can be implemented by one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, graphics processing units (GPUs), digital signal processors (DSPs), application specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field programmable logic devices (FPLDs). When reading devices or systems claims, of the present patent to cover a purely software and / or firmware Figure 9 elements, processes, and / or devices illustrated in Figure 3 The example hardware acceleration subsystems 310 can include one or more elements, processes, and / or devices, and / or can include more than one of any or all of the illustrated elements, processes, and devices. As used herein, the phrase “in communication,” including variations thereof, encompasses direct communication and / or indirect communication through one or more intermediary components and does not require direct physical (e.g., wired) communication and / or continuous communication, but additionally includes selective, periodic, scheduled, and / or one-time communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events.

[0061] In Figure 9 a flow diagram representing example hardware logic, machine readable instructions, hardware-implemented state machines, and / or any combination thereof for implementing a hardware acceleration subsystem 310 of Figure 3 The machine readable instructions can be one or more executable programs or portions of an executable program for execution by a computer processor and / or processor circuitry, such as the processor 1012 shown in the example processor platform 1000 described below in connection with Figure 10 Figure 9 ​The flowcharts illustrated in the middle describe example procedures, but many other methods of implementing the example hardware acceleration subsystem 310 can alternatively be used. For example, the order of execution of the blocks can be changed, and / or some of the blocks described can be changed, eliminated, or combined. Additionally or alternatively, any one or all of the blocks can be implemented by one or more hardware circuits structured to perform the corresponding operations without performing software or firmware. The processor circuitry can be distributed in different network locations and / or local to one or more devices (e.g., multi-core processors in a single machine, multiple processors distributed across a server rack, etc.).

[0062] The machine-readable instructions described herein can be stored in one or more of a compressed format, an encrypted format, a segmented format, a compiled format, an executable format, a packaged format, etc. The machine-readable instructions, as described herein, can be stored as data or data structures that can be used to create, manufacture, and / or produce machine-executable instructions (e.g., portions of instructions, code, representations of code, etc.). For example, the machine-readable instructions can be segmented and stored on one or more storage devices and / or computing devices (e.g., servers) located at the same or different locations of a network or collection of networks (e.g., in the cloud, in edge devices, etc.). The machine-readable instructions can require one or more of installation, modification, adaptation, updating, combining, supplementing, configuring, decryption, decompression, unpacking, distribution, reassignment, compilation, etc. in order to make them directly readable, interpretable, and / or executable by a computing device and / or other machine. For example, the machine-readable instructions can be stored in multiple portions that are individually compressed, encrypted, and stored on separate computing devices, where the portions, when decrypted, decompressed, and combined, form a set of executable instructions that implement one or more functions that together can form a program (e.g., a program described herein).

[0063] In another example, the machine-readable instructions can be stored in a state in which they can be read by processor circuitry, but require the addition of a library (e.g., a dynamic link library (DLL)), a software development kit (SDK), an application programming interface (API), etc. in order to execute the instructions on a particular computing device or other device. In another example, the machine-readable instructions can require configuration (e.g., stored settings, data input, recorded network addresses, etc.) before the machine-readable instructions and / or corresponding program can be executed in whole or in part. Thus, a machine-readable medium, as used herein, can include machine-readable instructions and / or programs regardless of the particular format or state of the machine-readable instructions and / or programs when stored or otherwise at rest or in transit.

[0064] The machine-readable instructions described herein can be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, the machine-readable instructions can be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, HyperText Markup Language (HTML), Structured Query Language (SQL), Swift, etc.

[0065] As mentioned above, Figure 9 Example processes of the present disclosure can be implemented using executable instructions stored on non-transitory computer and / or machine-readable media (such as hard drives, flash memory, read only memory, optical discs, digital versatile discs, cache memory, random access memory, and / or any other storage devices or storage disks) on which information is stored for any duration (e.g., for extended periods of time, permanently, for brief instances, for temporarily buffering, and / or for caching information). As used herein, the term non-transitory computer-readable medium expressly excludes propagating signals and excludes transmission media.

[0066] “Including” and “comprising” (and any form of these terms, e.g., include, includes, comprises, comprising, includes, having, have, etc.) are open-ended terms that are used to describe various embodiments. Thus, whenever a claim employs any form of “include” or “comprise” (e.g., comprises, includes, comprising, including, having, etc.) as a preamble, it is to be understood that additional elements, terms, etc. can be present in the corresponding claim, in addition to those explicitly recited. As used herein, when the phrase “at least” is used as the transition term in, for example, a claim, it is to be understood that the additional element(s) need not be present, even if the phrases “comprising” or “including” are used in the same or in a different claim. The term “and / or” when used in, for example, the form “A, B, and / or C” runs the gamut from meaning A, B and C, to meaning A, B, or C, to meaning A, or B, to meaning C, to meaning A and B but not C, to meaning A and C but not B, to meaning B and C but not A, to meaning not A, not B, and not C, and the like. As used herein in context with describing structures, compositions, items, objects, and / or things, the phrase “at least one of A and B” is intended to refer to embodiments in which: (1) at least one A is present, (2) at least one B is present, and (3) at least one A and at least one B are present. Similarly, as used herein in context with describing structures, compositions, items, objects, and / or things, the phrase “at least one of A or B” is intended to refer to embodiments in which: (1) at least one A is present, (2) at least one B is present, and (3) at least one A and at least one B are present. As used herein in context with describing performance or execution of processes, instructions, actions, activities, and / or steps, the phrase “at least one of A and B” is intended to refer to embodiments in which: (1) at least one A is performed, (2) at least one B is performed, and (3) at least one A and at least one B are performed. Similarly, as used herein in context with describing performance or execution of processes, instructions, actions, activities, and / or steps, the phrase “at least one of A or B” is intended to refer to embodiments in which: (1) at least one A is performed, (2) at least one B is performed, and (3) at least one A and at least one B are performed.

[0067] As used herein, singular references (e.g., “a,” “an,” “first,” “second,” and the like) are not to be construed as excluding the plural. The term “one” or “a” entity as used herein means one or more than one of the entity. The terms “a” (or “an”), “one or more” and “at least one” can be used interchangeably herein. Furthermore, although individually listed, a plurality of means, elements or method actions can be implemented by, e.g., a single unit or processor. Additionally, although individual features can be included in different examples or claims, these can also be combined, and the inclusion in different examples or claims does not imply that a combination of features is not feasible and / or advantageous.

[0068] Figure 9 is a flowchart representative of a machine readable instruction that can be executed to implement an example hardware acceleration subsystem 310 to implement data aggregation and pattern adaptation. Figure 3

[0069] At block 902, a hardware accelerator (e.g., a lens distortion correction hardware accelerator) processes a data chunk. For example, the example first hardware accelerator 350a Figure 3 may process the data chunk 402 (e.g., from the first configuration 410 of the Figure 4 ).

[0070] At block 904, the hardware accelerator writes the processed data chunk to a local memory. For example, the example first hardware accelerator 350a can write the processed data chunk 402 Figure 4 to the example local memory 360 Figure 3 .

[0071] At block 906, a load store engine coupled to the hardware accelerator determines whether the hardware accelerator is communicating at a chunk level or a superchunk level. For example, the example first load store engine 352a can determine whether the example first hardware accelerator 350a is communicating at a chunk level (e.g., Tdone_gen_mode = 0) or a superchunk level (e.g., Tdone_gen_mode = 1).

[0072] If the load store engine determines that the hardware accelerator is communicating at a chunk level (block 906), the machine readable instructions 900 proceed to block 914, in which the hardware accelerator sends a completion signal to a corresponding scheduler. For example, if the example first load store engine 352a determines that the example first hardware accelerator 350a is communicating at a chunk level (e.g., Tdone_gen_mode = 0), the example first hardware accelerator 350a sends a completion signal (e.g., a Tdone signal) to the example first scheduler 382a. The process ends.​

[0073] In some examples, the scheduler triggers the second hardware accelerator (e.g., a scaling hardware accelerator) to read the processed data block from the local memory or triggers the DMA controller to write the processed data block to the example external memory 330 Figure 3 ). For example, the example first scheduler 382a can trigger the example second hardware accelerator 350b to read the processed data block 402 from the example local memory 360 or trigger the example DMA controller 340 to write the processed data block 402 to the example external memory 330 Figure 3

[0074] If the load store engine determines that the hardware accelerator communicates at the superblock level (e.g., Tdone_gen_mode = 1) (block 906), the hardware accelerator increments the block count (block 908). For example, if the example first load store engine 352a determines that the example first hardware accelerator 350a communicates at the superblock level (e.g., Tdone_gen_mode = 1) (block 906), the example first hardware accelerator 350a can increment the block count by 1.

[0075] In some examples, the load store engine determines whether the block count is equal to the BPR value. If the load store engine determines that the block count is not equal to the BPR value (e.g., the block count is less than the BPR value) (block 910), the machine-readable instructions return to block 902 and the hardware accelerator processes another data block (block 902). For example, if the BPR value is 2 (e.g., BPR = 2) and the block count is 1 (e.g., the hardware accelerator has processed one data block 402), the example load store engine 352a can determine that the block count is not equal to the BPR value (block 910) and the example machine-readable instructions 900 return to block 902, where the example first hardware accelerator 350a processes another data block 404 (e.g., from the first configuration 410 of Figure 4 .

[0076] If the load store engine determines that the block count is equal to the BPR value (block 910), the load store engine aggregates the data blocks in the local memory to generate an aggregated data block based on the BPR value (block 912). For example, if the BPR value is 2 (e.g., BPR = 2) and the block count is 2 (e.g., the example first hardware accelerator 350a has processed two data blocks 402, 404), the example first load store engine 352a aggregates the data blocks 402, 404 and generates an aggregated data block (e.g., a superblock) 420a.

[0077] ​At block 914, the hardware accelerator sends a completion signal to the corresponding scheduler. For example, the example first hardware accelerator 350a can send a completion signal (e.g., a Tdone signal) to the example first scheduler 382a.

[0078] At block 916, in response to the completion signal, the scheduler triggers the second hardware accelerator to read the aggregated data block from the local memory or triggers the DMA controller to write the aggregated data element to the example external memory 330 Figure 3 . For example, in response to the Tdone signal, the example first scheduler 382a can trigger the example second hardware accelerator 350b to read the superblock 420a from the example local memory 360 or trigger the example DMA controller 340 to write the superblock 420a to the example external memory 330 Figure 3 . The process ends.

[0079] Although the example first load store engine 352a aggregates the data blocks 402, 404 (block 912) at the superblock level when communicating in Figure 9 some examples, the example first load store engine 352a aggregates the data blocks 402, 404 when communicating at the block level. Thus, in some examples, the example first load store engine 352a aggregates the data blocks 402, 404 regardless of whether the example first hardware accelerator is communicating at the block level or the superblock level. In some examples, the example first load store engine 352a aggregates the data blocks 402, 404 such that the example first load store engine 352a writes the data blocks as a single block to the example local memory 360 (e.g., the address of the data element 404 is contiguous with the data element 402).

[0080] Figure 10 is a block diagram of an example processor platform 1000 that is structured to execute the instructions of Figure 9 to implement the device of Figure 3 . For example, the processor platform 1000 can be a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a mobile phone, a smart phone, a tablet computer (e.g., iPad TM ), a personal digital assistant (PDA), an Internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a game console, a personal video recorder, a set-top box, a headset, or other wearable device, or any other type of computing device.

[0081] The processor platform 1000 of the illustrated example includes the example HWA subsystem 310 described in connection with Figure 3 .

[0082] The processor platform 1000 of the illustrated example includes a processor 1012. The processor 1012 of the illustrated example is hardware. For example, the processor 1012 can be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer.

[0083] The processor 1012 of the illustrated example includes a local memory 1013 (e.g., cache). The processor 1012 of the illustrated example is in communication with a main memory including a volatile memory 1014 and a non-volatile memory 1016 via a bus 1018. The volatile memory 1014 can be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS® dynamic random access memory (RDRAM®), and / or any other type of random access memory device. The non-volatile memory 1016 can be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 1014, 1016 is controlled by a memory controller.

[0084] The processor platform 1000 of the illustrated example also includes an interface circuit 1020. The interface circuit 1020 can be implemented by any type of interface standard, such as an Ethernet interface, a universal serial bus (USB), a Bluetooth® interface, a near field communication (NFC) interface, and / or a PCI express interface.

[0085] In the illustrated example, one or more input devices 1022 are connected to the interface circuit 1020. The input device(s) 1022 permit(s) a user to enter data and / or commands into the processor 1012. For example, the input device(s) can be implemented by an audio sensor, a microphone, a camera (still or video), a keyboard, a button, a mouse, a touchscreen, a trackpad, a trackball, isopoint, and / or a voice recognition system.

[0086] One or more output devices 1024 are also connected to the interface circuit 1020 of the illustrated example. The output device(s) 1024 can be implemented, for example, by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-place switching (IPS) display, a touchscreen, etc.), a tactile output device, a printer and / or a speaker. Thus, the interface circuit 1020 of the illustrated example, in some cases, includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.

[0087] The interface circuit 1020 illustrated in the diagram also includes communication devices (e.g., transmitter, receiver, transceiver, modem, residential gateway, wireless access point), and / or a network interface for facilitating data exchange with external machines (e.g., any type of computing device) via network 1026. For example, communication may be via Ethernet connection, digital subscriber line (DSL) connection, telephone line connection, coaxial cable system, satellite system, field wireless system, cellular telephone system, etc.

[0088] The processor platform 1000 illustrated in the diagram also includes one or more mass storage devices 1028 for storing software and / or data. Examples of such mass storage devices 1028 include floppy disk drives, hard disk drives, optical disk drives, Blu-ray disc drives, redundant array of independent disks (RAID) systems, and digital multifunction optical disc (DVD) drives.

[0089] Figure 9 The machine-executable instructions 1032 may be stored in a mass storage device 1028, in volatile memory 1014, in non-volatile memory 1016, and / or on a removable non-transitory computer-readable storage medium (e.g., CD or DVD).

[0090] exist Figure 11 The block diagram is used to illustrate the software (e.g., Figure 9 The instanced computer-readable instructions (1032) are distributed to an instanced software distribution platform 1105 for a third party. The instanced software distribution platform 1105 may be implemented by any computer server, data facility, cloud service, etc., capable of storing software and transmitting it to other computing devices. The third party may be a customer of an entity that owns and / or operates the software distribution platform. For example, the entity owning and / or operating the software distribution platform may be the software (e.g., Figure 9 The developer, seller, and / or licensor of the instance computer-readable instructions 1032. Third parties may be consumers, users, retailers, OEMs, etc., who purchase and / or license the software for use and / or resell and / or sublicense. In the illustrated example, the software distribution platform 1205 includes one or more servers and one or more storage devices. The storage devices store the computer-readable instructions 1032, which can be used with... Figure 9The instance computer-readable instruction 1032 corresponds to that described above. One or more servers of the instance software distribution platform 1105 communicate with network 1110, which may correspond to the Internet and / or any one or more of the instance network 1026 described above. In some instances, as part of a business transaction, one or more servers respond to a request to transfer software to a requesting party. Payment for the delivery, sale, and / or licensing of the software may be handled by one or more servers of the software distribution platform and / or via a third-party payment entity. The server enables the purchaser and / or licensor to download the computer-readable instruction 1032 from the software distribution platform 1105. For example, it may be compatible with... Figure 9 The software corresponding to the instanced computer-readable instruction 1032 is downloaded to the instanced processor platform 1000, which will execute the computer-readable instruction 1032 to implement... Figure 3 The equipment. In some instances, one or more servers of the software distribution platform 1105 periodically provide, deliver, and / or force software updates (e.g., Figure 9 The instance computer-readable instructions 1032) ensure that software is distributed and applied for improvements, patches, updates, etc., at the end-user device.

[0091] Based on the foregoing, it will be understood that an exemplary system, method, and apparatus for implementing data aggregation and type adaptation in a hardware acceleration subsystem have been described. The described methods, apparatus, and articles improve the efficiency of using computing devices by improving the efficiency of external memory and implementing user-defined multi-producer and multi-consumer hardware acceleration schemes. The described methods, apparatus, and articles therefore represent one or more improvements to the functionality of a computer.

[0092] The examples described herein include a single-chip system (SoC) comprising: a first scheduler; a first hardware accelerator coupled to the first scheduler to process at least a first data element and a second data element; and a first load storage engine coupled to the first hardware accelerator, the first load storage engine being configured to: communicate with the first scheduler at a superblock level by sending a completion signal to the first scheduler in response to determining that a block count equals a first BPR value; and aggregate the first data element and the second data element based on the first BPR value to generate a first aggregated data element.

[0093] In some instances, the first load storage engine increments the block count in response to the first hardware accelerator processing the first data element and in response to the first hardware accelerator processing the second data element.

[0094] In some instances, the first scheduler, in response to receiving the completion signal from the first hardware accelerator, instructs the DMA controller to store the first aggregated data element to external memory.

[0095] In some instances, the first scheduler instructs the second hardware accelerator to read the first aggregated data element in response to receiving the completion signal from the first hardware accelerator.

[0096] In some instances, the first load storage engine is configured to communicate with the first scheduler at the block level by sending a completion signal to the first scheduler in response to the first hardware accelerator processing the first data block.

[0097] In some instances, the first BPR value is associated with the first data channel.

[0098] In some instances, the SoC includes a software (SW) programmable memory-mapped register (MMR) coupled to the first scheduler, the MMR being used to provide at least the first BPR value to the first load storage engine.

[0099] In some instances, the first load storage engine is configured to aggregate at least a third data element with a fourth data element based on a second BPR value to produce a second aggregated data element.

[0100] In some instances, the second BPR value is associated with a second data channel.

[0101] In some instances, the first load storage engine enables a software (SW) programmable circular buffer in local memory for data aggregation based at least on the first BPR value.

[0102] In some instances, the first scheduler includes a first consumer socket for tracking input data consumed by the first hardware accelerator and a first producer socket for tracking output data generated by the hardware accelerator.

[0103] In some instances, the first scheduler includes a first producer-type adapter coupled to the first producer socket.

[0104] The examples described herein include a method comprising: processing a first data element and a second data element by a first hardware accelerator; sending a completion signal to a first scheduler by a first load storage engine in response to determining that a block count equals a first BPR value; and aggregating the first data element and the second data element by the first load storage engine based on the first BPR value to generate a first aggregated data element.

[0105] In some instances, the method further includes: incrementing the block count by the first load storage engine in response to the first hardware accelerator processing the first data element; and incrementing the block count by the first load storage engine in response to the first hardware accelerator processing the second data element.

[0106] In some instances, the method further includes: in response to receiving the completion signal from the first hardware accelerator, the first scheduler instructs the DMA controller to store the first aggregated data element to external memory.

[0107] In some instances, the method further includes: in response to receiving the completion signal from the first hardware accelerator, the first scheduler instructs a second hardware accelerator to read the first aggregated data element.

[0108] In some instances, the first BPR value is associated with the first data channel.

[0109] In some instances, the method further includes: the first load storage engine aggregating at least a third data element with a fourth data element based on a second BPR value to generate a second aggregated data element.

[0110] In some instances, the second BPR value is associated with a second data channel.

[0111] The examples described herein include a non-transitory computer-readable medium comprising computer-readable instructions that, when executed, cause at least one processor to perform at least the following operations: processing a first data element and a second data element by a first hardware accelerator; sending a completion signal to a first scheduler by a first load storage engine in response to determining that a block count equals a first BPR value; and aggregating the first data element and the second data element by the first load storage engine based on the first BPR value to generate a first aggregated data element.

[0112] In some instances, the computer-readable instructions further cause the at least one processor to perform at least the following operations: incrementing the block count by the first load storage engine in response to the first hardware accelerator processing the first data element, and incrementing the block count by the first load storage engine in response to the first hardware accelerator processing the second data element.

[0113] In some instances, the computer-readable instructions further cause the at least one processor to perform at least the following operations: in response to receiving the completion signal from the first hardware accelerator, the first scheduler instructs the DMA controller to store the first aggregated data element to external memory.

[0114] In some instances, the computer-readable instructions further cause the at least one processor to perform at least the following operations: in response to receiving the completion signal from the first hardware accelerator, the first scheduler instructs the second hardware accelerator to read the first aggregated data element.

[0115] In some instances, the computer-readable instructions further cause the at least one processor to perform at least the following operations: in response to the first hardware accelerator processing the first data block, the first hardware accelerator sends a completion signal to the first scheduler.

[0116] In some instances, the first BPR value is associated with the first data channel.

[0117] In some instances, the computer-readable instructions further cause the at least one processor to perform at least the following operation: the first load storage engine aggregates at least a third data element with a fourth data element based on a second BPR value to generate a second aggregated data element.

[0118] In some instances, the second BPR value is associated with a second data channel.

[0119] The examples described herein include an apparatus comprising: means for processing a first data element and a second data element; means for sending a completion signal to a first scheduler in response to determining that a block count equals a first BPR value; and means for aggregating the first data element and the second data element based on the first BPR value to generate a first aggregated data element.

[0120] In some instances, the device further includes components for incrementing the block count in response to the first hardware accelerator processing the first data element and components for incrementing the block count in response to the first hardware accelerator processing the second data element.

[0121] In some instances, the device further includes a component for instructing a DMA controller to store the first aggregated data element to external memory in response to receiving the completion signal from the first hardware accelerator.

[0122] In some instances, the device further includes a component for instructing a second hardware accelerator to read the first aggregated data element in response to receiving the completion signal from the first hardware accelerator.

[0123] In some instances, the device further includes a component for sending a completion signal to the first scheduler in response to the first hardware accelerator processing the first data block.

[0124] In some instances, the first BPR value is associated with the first data channel.

[0125] In some instances, the device further includes a component for aggregating at least a third data element with a fourth data element based on a second BPR value to generate a second aggregated data element.

[0126] In some instances, the second BPR value is associated with a second data channel.

[0127] While specific exemplary methods, apparatuses, and articles of manufacture have been described herein, the scope of this patent is not limited thereto. Rather, this patent fairly covers all methods, apparatuses, and articles of manufacture that fall within the scope of the claims of this patent.

[0128] The appended claims are hereby incorporated by reference into this detailed description, wherein each claim is an independent embodiment of the present description.

Claims

1. A single-chip system (SoC) comprising: a first scheduler; a first hardware accelerator coupled to the first scheduler to process at least a first data element and a second data element; and a first load-store engine coupled to the first hardware accelerator, the first load-store engine configured to: communicate with the first scheduler at a superblock level by sending a completion signal to the first scheduler in response to determining that a block count of the first data element and the second data element equals a first BPR value; and aggregate the first data element and the second data element based on the first BPR value to generate a first aggregated data element.

2. The SoC of claim 1, wherein the first load-store engine: increments the block count in response to the first hardware accelerator processing the first data element; and increments the block count in response to the first hardware accelerator processing the second data element.

3. The SoC of claim 1, wherein the first scheduler instructs a DMA controller to store the first aggregated data element to an external memory in response to receiving the completion signal from the first hardware accelerator.

4. The SoC of claim 1, wherein the first scheduler instructs a second hardware accelerator to read the first aggregated data element in response to receiving the completion signal from the first hardware accelerator.

5. The SoC of claim 1, wherein the first load-store engine is configured to communicate with the first scheduler at a block level by sending a completion signal to the first scheduler in response to the first hardware accelerator processing the first data element.

6. The SoC of claim 1, wherein the first BPR value is associated with a first data channel.

7. The SoC of claim 1, including a software (SW) programmable memory mapped register (MMR) coupled to the first scheduler, the MMR to provide at least the first BPR value to the first load-store engine.

8. The SoC of claim 1, wherein the first load-store engine is configured to aggregate at least a third data element and a fourth data element based on a second BPR value to generate a second aggregated data element.

9. The SoC of claim 8, wherein the second BPR value is associated with a second data channel.

10. The SoC of claim 1, wherein the first load-store engine enables software (SW) programmable circular buffer storage in a local memory for data aggregation based at least on the first BPR value.

11. The SoC of claim 1, wherein the first scheduler includes a first consumer socket to track input data consumed by the first hardware accelerator and a first producer socket to track output data produced by the hardware accelerator.

12. The SoC of claim 11, wherein the first scheduler includes a first producer-style adapter coupled to the first producer socket. ​ 13. A data processing method comprising: processing, by a first hardware accelerator, a first data element and a second data element; sending, by a first load store engine to a first scheduler, a completion signal in response to determining that a block count of the first data element and the second data element equals a first BPR value; and aggregating, by the first load store engine, the first data element and the second data element to generate a first aggregated data element based on the first BPR value.

14. The data processing method of claim 13, further comprising: incrementing, by the first load store engine, the block count in response to the first hardware accelerator processing the first data element; and incrementing, by the first load store engine, the block count in response to the first hardware accelerator processing the second data element.

15. The data processing method of claim 13, further comprising instructing, by the first scheduler, a DMA controller to store the first aggregated data element to an external memory in response to receiving the completion signal from the first hardware accelerator.

16. The data processing method of claim 13, further comprising instructing, by the first scheduler, a second hardware accelerator to read the first aggregated data element in response to receiving the completion signal from the first hardware accelerator.

17. The data processing method of claim 13, wherein the first BPR value is associated with a first data channel.

18. The data processing method of claim 13, further comprising aggregating, by the first load store engine, at least a third data element and a fourth data element to produce a second aggregated data element based on a second BPR value.

19. The data processing method of claim 18, wherein the second BPR value is associated with a second data channel.

20. A non-transitory computer-readable medium comprising computer-readable instructions that, when executed, cause at least one processor to at least: process, by a first hardware accelerator, a first data element and a second data element; send, by a first load store engine to a first scheduler, a completion signal in response to determining that a block count of the first data element and the second data element equals a first BPR value; and aggregate, by the first load store engine, the first data element and the second data element to generate a first aggregated data element based on the first BPR value.

Citation Information

Patent Citations

  • Vehicle brake system

    US20160159224A1

  • Load store unit with replay mechanism

    WO2004111839A1