Data adaptation in hardware acceleration subsystem
By introducing a scheduler and a load storage engine, the data transfer and format adaptation in the hardware acceleration subsystem are optimized, solving the problem of low memory access efficiency and achieving more efficient data processing and parallel computing.
Patent Information
- Application Number
- CN202610103518.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-31
- Filing Date
- 2021-01-04
- Publication Date
- 2026-03-06
AI Technical Summary
Existing hardware acceleration subsystems suffer from low memory access efficiency when processing data, especially in DDR external memory. The complexity of data block and row conversion makes system customization difficult, and CPU involvement increases cost and latency.
By introducing a scheduler and a load storage engine, data element aggregation and format adaptation are achieved. Coordination between the scheduler, hardware accelerators, and memory is implemented. Combined with a DMA controller, data transfer is optimized, flexible block size adjustment is supported, and memory access efficiency is improved.
It improves the memory access efficiency of the hardware acceleration subsystem, reduces CPU involvement, lowers latency and power consumption, and enhances the parallelization capability of computing tasks.
Smart Images

Figure CN121614421A_ABST
Abstract
Description
[0001] Information related to divisional application
[0002] This application is a divisional application of the invention patent application filed on January 4, 2021, with application number 202180012055.2 and invention title "Data Adaptation in Hardware Acceleration Subsystem". Technical Field
[0003] This description generally relates to hardware acceleration subsystems, and more specifically, to enhanced external memory transfer and type adaptation within hardware acceleration subsystems. Background Technology
[0004] Although central processing units (CPUs) have been improved to meet the demands of modern applications, computer performance is still limited by the large amounts of data that must be processed by the CPU simultaneously. Hardware accelerator subsystems can provide improved performance and / or power consumption by offloading tasks from the computer's CPU to dedicated hardware components that perform those tasks. Summary of the Invention
[0005] In some embodiments disclosed herein, a circuit arrangement is provided. The circuit arrangement includes: a memory; a first hardware accelerator coupled to the memory and configured to: process a set of data to generate a set of data elements; and cause the set of data elements to be stored in the memory; a second hardware accelerator coupled to the memory; and a load storage circuit coupled between the first hardware accelerator and the memory, the load storage circuit being configured to: determine whether the number of elements in the set of data elements satisfies an aggregation threshold; and, based on the number of elements satisfying the aggregation threshold, cause the second hardware accelerator to process the set of data elements.
[0006] In some embodiments of this disclosure, a circuit arrangement is provided. The circuit arrangement includes: a first memory; a direct memory access (DMA) circuit coupled to the first memory and configured to be coupled to a second memory; a hardware accelerator coupled to the first memory and configured to: perform an operation on a set of data to produce a set of data elements; and cause the set of data elements to be stored in the first memory; and a load storage circuit coupled between the hardware accelerator and the first memory and configured to: determine whether the number of elements in the set of data elements satisfies an aggregation threshold; and, based on the number of elements satisfying the aggregation threshold, cause the DMA circuit to store the set of data elements in the second memory.
[0007] In some embodiments disclosed herein, a method is provided. The method includes: processing a first set of data and a second set of data using a first processing circuit to generate a first set of data elements and a second set of data elements associated with a first channel and a second channel, respectively; storing the first set of data elements and the second set of data elements in a memory; determining whether the number of elements in the first set of data elements satisfies a first aggregation threshold associated with the first channel; determining whether the number of elements in the second set of data elements satisfies a second aggregation threshold associated with the second channel; processing the first set of data elements using a second processing circuit based on the fact that the number of elements in the first set of data elements satisfies the first aggregation threshold; and processing the second set of data elements using the second processing circuit based on the fact that the number of elements in the second set of data elements satisfies the second aggregation threshold.
[0008] In some embodiments of this disclosure, a system is provided. The system includes: a plurality of schedulers, each scheduler associated with a type adapter; a plurality of hardware accelerators, each coupled to the plurality of schedulers; and a memory coupled to the plurality of hardware accelerators; wherein a first hardware accelerator among the plurality of hardware accelerators is configured to: read data in a first format from the memory; and convert the data from the first format to a second format using the type adapter corresponding to the first hardware accelerator; process the data in the second format to form processed data in the second format; and cause the processed data in the second format to be stored in the memory; wherein a second hardware accelerator among the plurality of hardware accelerators is configured to determine whether conditions are met regarding the processed data in the second format to determine whether to further process the processed data in the second format.
[0009] In some embodiments of this disclosure, a system is provided. The system includes: a plurality of schedulers, each scheduler associated with a type adapter; a plurality of hardware accelerators, each coupled to the plurality of schedulers; a plurality of load storage engines, each associated with the plurality of hardware accelerators; a memory coupled to the plurality of load storage engines; and direct memory access (DMA) circuitry coupled to the memory. Attached Figure Description
[0010] Figure 1 It is an instance diagram of a block-based processing and storage subsystem used to perform processing tasks on macroblocks fetched from external memory.
[0011] Figure 2 It is an instance diagram of a hardware acceleration subsystem used to process data elements retrieved from external memory.
[0012] Figure 3This is a block diagram of an instanced hardware acceleration subsystem used to implement data aggregation and type adaptation in hardware acceleration.
[0013] Figure 4 It is a diagrammatic illustration of an instance diagram used to generate an instance of aggregated data elements.
[0014] Figure 5 It is an instance diagram illustrating the instance type adaptation process, which is implemented by the instance type adapter to convert data blocks into row data elements.
[0015] Figure 6 This is an instance of a user-defined diagram illustrating an instance of a multi-consumer / multi-producer hardware acceleration subsystem for image, vision, and / or video processing.
[0016] Figure 7 It is an example diagram illustrating an instance of a multi-consumer / multi-producer hardware acceleration solution.
[0017] Figure 8 The diagram illustrates an example of a multi-producer lens distortion correction (LDC) hardware accelerator used to output a first data element on a first channel and a second data element on a second channel.
[0018] Figure 9 It indicates that it can be implemented through execution. Figure 3 A flowchart of machine-readable instructions for an instance of a hardware acceleration subsystem.
[0019] Figure 10 It is structured for execution Figure 9 To implement the instructions Figure 3 A block diagram of the instance processor platform of the device.
[0020] Figure 11 It is used to combine software (e.g., with) Figure 9 A block diagram of an instance software distribution platform that distributes the software corresponding to the instance computer-readable instructions to client devices (e.g., consumers (e.g., for licensing, selling and / or using), retailers (e.g., for selling, reselling, licensing and / or sublicensing) and / or original equipment manufacturers (OEMs) (for example, for inclusion in products to be distributed to retailers and / or directly to customers)).
[0021] The figures are not to scale. Instead, the thickness of layers or regions may be enlarged in the figures. While the figures show layers and regions with clean lines and boundaries, some or all of these lines and / or boundaries may be idealized. In reality, boundaries and / or lines may be unobservable, mixed, and / or irregular. Generally, throughout the figures and accompanying written description, the same reference numerals will be used to refer to the same or similar parts. As used herein, unless otherwise stated, the term “above” describes the relationship of two parts relative to the Earth. If the second part has at least one part between the Earth and the first part, then the first part is above the second part. Similarly, as used herein, the first part is “below” the second part when the first part is closer to the Earth than the second part. As mentioned above, the first part may be above or below the second part if it has other parts in between, does not have other parts in between, is in contact with the second part, or is not in direct contact with each other. As used herein, a statement that any part (e.g., layer, film, region, area, or plate) is located on another part in any manner (e.g., positioned on another part, located on another part, disposed on another part, or formed on another part, etc.) indicates that the mentioned part is in contact with the other part, or that the mentioned part is above the other part, wherein one or more intermediate portions are located between the two parts. As used herein, unless otherwise indicated, a connection reference (e.g., attachment, coupling, connection, and engagement) may include intermediate portions between elements referenced by the connection reference and / or relative movement between those elements. Thus, a connection reference does not necessarily imply that two elements are directly connected and / or fixed to each other. As used herein, a statement that any part is in contact with another part is defined as meaning that there is no intermediate portion between the two parts.
[0022] Unless otherwise specifically stated, descriptive terms such as “first,” “second,” and “third” are used herein without implying or otherwise indicating any meaning of priority, physical order, arrangement in a list, and / or any sorting, but only as labels and / or arbitrary names to distinguish elements for the purpose of understanding the described instances. In some instances, the descriptive term “first” may be used to refer to an element in a detailed description, while the same element may be referenced in a technical solution using different descriptive terms (e.g., “second” or “third”). Such descriptive terms are only used to explicitly identify those elements that may (for example) otherwise share the same name. As used herein, “substantially real-time” means occurring in a near-instantaneous manner, recognizing that real-world delays may exist in computation time, transmission, etc. Therefore, unless otherwise specified, “substantially real-time” means actual time + / - 1 second. Detailed Implementation
[0023] In some scenarios, hardware acceleration can be used to reduce latency, increase throughput, reduce power consumption, and enhance the parallelization of computational tasks. Commonly used hardware accelerators include graphics processing units (GPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), complex programmable logic devices (CPLDs), and system-on-a-chip (SoCs).
[0024] Hardware acceleration has a wide range of applications across many different fields, including the automotive industry, advanced driver-assistance systems (ADAS), manufacturing, high-performance computing, robotics, drones, and other industries involving complex, high-speed processing, such as hardware-based encryption, computer-generated graphics, artificial intelligence, and digital image processing. The latter involves various complex processing operations performed on a single image or video stream, such as lens distortion correction, scaling, transformation, noise filtering, dense optical flow, pyramid representation, stereoscopic screen door effects (SDE), and other processing operations. Many of the computational tasks associated with these operations involve significant amounts of processing power, and in some cases, such as real-time processing of video streams, the amount of processing power required to process images or video streams can put significant stress on the CPU.
[0025] Many hardware accelerators are designed to perform various computational tasks on data retrieved from external memory. Many instance-based hardware accelerators are configured to perform processing tasks on data elements presented as blocks or rows. For example, in image processing where imaging / vision algorithms are typically based on two-dimensional (2D) blocks, a hardware accelerator can be configured to process 2D blocks from an image frame rather than processing the entire image frame as rows. Various instance-based hardware accelerators can operate on block sizes of 16x16 bytes, 32x32 bytes, and 64x32 bytes.
[0026] If a hardware accelerator is implemented on a single-chip system (SoC), a direct memory access (DMA) controller can implement direct memory access to fetch blocks or rows of data from external memory and transfer the data to local on-chip memory. Many types of external memory (e.g., dual data rate synchronous dynamic random access memory (DDR SDRAM)) prefer linear data access based on one-dimensional (1D) rows because row transfers do not incur page penalties, which can occur (for example,) when two page open / close cycles are required to access vertically adjacent pixels falling on different pages (each page open / close cycle has a duration of approximately 60 ns (e.g., page penalty)).
[0027] Although DDR external memory prefers linear access, the DMA controller can access data from DDR external memory in block form; however, the data blocks sent from DDR external memory can be fixed-size rectangular blocks with a block height corresponding to the number of rows in the DMA data block request and a fixed block width of 64 bytes. In some cases, the fixed-size rectangular blocks sent from DDR can have a block height managed by the external memory controller and / or a block width that is a function of the burst size. Because the hardware accelerator can operate on data rows or data blocks smaller than those sent from DDR external memory frequently, the DMA controller can use only a small portion of the data blocks sent from DDR and discard excess data, which in many cases must be retrieved from DDR external memory at a later time for processing. Therefore, DDR external memory can send the same data to the DMA controller multiple times before the data is processed by the hardware accelerator. Similarly, the DMA controller can send processed data to DDR external memory multiple times before the processed data is stored in DDR external memory. This redundancy leads to low operating efficiency of DDR external memory.
[0028] For example, if a DMA controller attempts to fetch a 16x16 byte data block from DDR external memory (e.g., a data block with a block height of 16 rows and a block width of 16 bytes), the DDR external memory can return a larger data block, for example, a 16x64 byte data block with a block height of 16 rows and a block width of 64 bytes. The external memory controller can then write only 16 bytes from each of the 16 rows and discard the remaining 48 bytes, resulting in a 25% reduction in DDR memory access efficiency. Similarly, if the DMA controller attempts to write a 16x16 byte data block to DDR external memory, the DMA can effectively consume the bandwidth and time required to write 16 rows of 64 bytes. Therefore, DDR inefficiency can occur when fetching data from and / or writing data to DDR external memory.
[0029] While some hardware accelerators operate on data blocks as described above, others can operate on data rows. Multiple hardware accelerators can be integrated into a hardware acceleration subsystem to form a hardware acceleration chain; however, as row-to-block and block-to-row transformations become increasingly complex when implemented on hardware accelerators, existing subsystems often operate on the same type of data elements (e.g., blocks or rows). This limitation makes customizing hardware acceleration subsystems very difficult.
[0030] Existing techniques for improving memory access efficiency via hardware accelerators are typically limited to fixed-block-size schemes with or without software- and / or hardware-based caches, resulting in inefficient transfers and a simple linear chain of hardware-accelerated use cases. These existing techniques include hardware-accelerated subsystems that fetch fixed-size macroblocks from external memory. For example, Figure 1 This is an example diagram of a block-based processing and storage subsystem that includes an Image Subsystem (ISS) 100 configured using a configuration interconnect 112, which is configured to fetch fixed-size macroblocks from system memory 110 via an ISS data interconnect 138 and store the fixed-size macroblocks in local on-chip memory (e.g., from a set of switchable buffers 134). A Lens Distortion Correction (LDC2) hardware accelerator 128 and / or a Noise Filtering (VTNF) hardware accelerator 130 can perform processing tasks on the macroblocks and send the processed macroblocks to local memory (e.g., from a set of switchable buffers 134) via a static controller crossbar switch 132. However, these existing hardware acceleration subsystems may not include hardware that allows the subsystem to adjust the macroblock size according to the needs of individual hardware accelerators. Instead, the macroblock size in these subsystems is primarily driven by the input buffer 134, and the output block size is defined by the input block scaling factor, the output image buffer size, and / or the input block local memory size. The buffers in these subsystems are typically switchable buffers without a mechanism for combining blocks. This lack of control and storage in existing hardware-accelerated subsystems limits the flexibility of combining blocks into various sizes. While other hardware accelerators may involve a CPU to combine multiple adjacent blocks to create a bounding box, CPU involvement typically incurs area and performance costs for the CPU pipeline.
[0031] Figure 2 This is an example diagram 200 of a hardware acceleration subsystem 210, which is integrated into a single-chip system (SoC) 220 and configured to retrieve data elements 231, 232 from external memory 230, process data elements 231, 232 to generate processed data elements 236, 238, and write the processed data elements 236, 238 to external memory 230.
[0032] Figure 2The hardware acceleration subsystem 210 illustrated in the figure includes: a first direct memory access (DMA) controller 240 for facilitating the transfer of data elements 231, 232 from external memory 230 (e.g., from input frame 237 stored in external memory 230) to local memory 260; and a second DMA controller 242 for facilitating the transfer of processed data elements 236, 238 from local memory 260 and / or hardware accelerators 250a, 250b, 250c, 250d to external memory 230. (e.g., the transmission of output frame 239 to external memory 230); four hardware accelerators 250a, 250b, 250c, 250d for performing various processing tasks on data elements 231, 232 to produce intermediate data elements 233, 234 and / or processed data elements 236, 238; local memory 260 for temporarily storing data elements 231, 232 and / or intermediate data elements 233, 234 during processing; and scheduler 280 for coordinating the workflow between hardware accelerators 250a to 250d, local memory 260 and DMA controllers 240, 242.
[0033] exist Figure 2 In the example illustrated herein, hardware accelerators 250a to 250d are configured to consume data elements 231, 232 and / or intermediate data elements 233, 234 as input, perform processing tasks on data elements 231, 232 and / or intermediate data elements 233, 234, and generate processed data elements 236, 238 as output for consumption by another hardware accelerator 250a to 250d, written to local memory 260 and / or written to DDR external memory 230 via DMA controller 242. Figure 2 In one example, the hardware acceleration subsystem 210 is configured to process multiple data elements 231, 232 in parallel. For instance, a first hardware accelerator 250a performs a first processing task on the first data element 231 to produce an intermediate data element 233, while a second hardware accelerator 250b performs a second processing task on the second data element 232 to produce an intermediate data element 234. The scheduler 280 facilitates the workflow of hardware accelerators 250a to 250d, DMA controllers 240, 242, and local memory 260 as data elements 231, 232 travel along the hardware acceleration pipeline.
[0034] In some instances, an enhanced hardware acceleration subsystem for improving DDR access and enabling data adaptation among multiple producers and consumers includes a first hardware accelerator for performing a first processing task on a first data element, a scheduler for controlling the workflow of the hardware accelerator and data aggregation, and a load-balanced storage engine coupled to the first hardware accelerator to aggregate the first data element with a second data element in local memory. In some instances, the scheduler includes a type adapter for implementing conversions between rows, blocks, and aggregated blocks.
[0035] Figure 3 This is a block diagram of an instanced hardware acceleration subsystem 310 used to implement data aggregation and type adaptation. The instanced hardware acceleration subsystem 310 includes an instanced DMA controller 340 coupled to an instanced channel mapper (e.g., an instanced DMA scheduler 382d), an instanced first hardware accelerator 350a coupled to an instanced first scheduler 382a, an instanced second hardware accelerator 350b coupled to an instanced second scheduler 382b, an instanced third hardware accelerator 350c coupled to an instanced third scheduler 382c, an instanced local memory 360, and an instanced main hardware thread scheduler (HTS) 380. In some instances, the instanced first hardware accelerator 350a includes an instanced first load storage engine 352a, the instanced second hardware accelerator 350b includes an instanced second load storage engine 352b, and the instanced third hardware accelerator 350c includes an instanced third load storage engine 352c. In some instances, the instance hardware acceleration subsystem 310 includes an instance memory-mapped register (MMR) controller 392, instance schedulers 382a to 382d, instance workload storage engines 352a to 352c, and / or an instance DMA controller 340 coupled to the instance HTS 380. In some instances, the instance MMR controller 392 is software (SW) programmable.
[0036] exist Figure 3 In the instanced hardware acceleration subsystem 310 illustrated in the diagram, instanced schedulers 382a, 382b, 382c, and 382d respectively include instanced consumer sockets 384a, 384b, 384c, and 384d, which are configured to track input data consumed by the corresponding instanced hardware accelerators 350a to 350c and the corresponding instanced DMA controller 340. Figure 3In the instanced hardware acceleration subsystem 310 illustrated in the figure, instanced schedulers 382a to 382d respectively include instanced producer sockets 386a, 386b, 386c, and 386d, which are configured to track output data generated by the corresponding instanced hardware accelerators 350a to 350c and the corresponding instanced DMA controller 340. In some instances, instance first scheduler 382a includes instance first producer type adapter 390a coupled to instance first producer socket 386a, instance second scheduler 382b includes instance second producer type adapter 390b coupled to instance second producer socket 386b, instance third scheduler 382c includes instance third producer type adapter 390c coupled to instance third producer socket 386c, and instance DMA scheduler 382d includes instance DMA type adapter 390d coupled to instance DMA producer socket 386d.
[0037] exist Figure 3 In the illustrated instance hardware acceleration subsystem 310, an instance DMA controller 340 facilitates the transfer of data elements (e.g., data blocks) between instance local memory 360 and instance external memory 330 (e.g., DDR external memory or other off-chip memory outside the instance hardware acceleration subsystem 310). In some instances, the instance DMA controller 340 communicates with an external memory controller (e.g., a DDR controller) to transfer data elements between instance local memory 360 and instance external memory 330. In some instances, the instance DMA controller 340 transfers data elements between instance hardware acceleration subsystem 310, instance external memory 330, and / or other components and / or subsystems in the SoC via a common bus.
[0038] In some instances, instance DMA controller 340 communicates with instance schedulers 382a to 382c via instance crossbar switch 370 and is coupled to instance DMA scheduler 382d. In some instances, instance DMA scheduler 382d performs scheduling operations similar to those of instance schedulers 382a to 382c corresponding to instance hardware accelerators 350a to 350c. In some instances, instance DMA scheduler 382d maps DMA channels to instance hardware accelerators 350a to 350c. In some instances, instance DMA controller 340 transmits a channel start signal when a data transfer is initiated via a DMA channel (e.g., the DMA channel corresponding to hardware accelerator 350a). In some instances, instance DMA controller 340 transmits a channel completion signal when a data transfer is completed via a DMA channel (e.g., the DMA channel corresponding to hardware accelerator 350a).
[0039] exist Figure 3 In the illustrated instance hardware accelerator subsystem 310, the instance DMA controller 340 retrieves data elements from instance external memory 330 for consumption by at least one of hardware accelerators 350a to 350c. In some instances, the data elements are stored contiguously in instance local memory 360. In some instances, in response to instructions from instance schedulers 382a to 382d and / or instance HTS 380, the instance DMA controller 340 transfers processed data elements to instance external memory 330, for example, to an output frame.
[0040] exist Figure 3 In the exemplary hardware accelerator subsystem 310 illustrated in the diagram, exemplary hardware accelerators 350a to 350c are configured to perform processing tasks on data elements (e.g., data blocks and / or data rows). In some instances, exemplary hardware accelerators 350a to 350c are configured to perform image processing tasks, such as lens distortion correction (LDC), scaling (e.g., MSC), noise filtering (NF), dense optical flow (DOF), stereo screen door effect (SDE), or any other processing task suitable for image processing.
[0041] In some instances, the instanced first hardware accelerator 350a operates on data blocks. In some instances, the instanced first hardware accelerator 350a operates on 16x16B data blocks, 32x32B data blocks, 64x32 data blocks, or any other data block size suitable for performing processing tasks. In some instances, at least an instanced first scheduler 382a coupled to the instanced first hardware accelerator 350a includes multiple consumer sockets 384a and / or multiple instanced producer sockets 386a that can be connected to an instanced second scheduler 382b. For example, the ISS hardware accelerator may have six outputs (Y12, UV12, U8, UV8, S8, and H3A), and the LDC hardware accelerator may have two outputs (Y, UV) or three outputs (R, G, B).
[0042] In some instances, the instance first consumer socket 384a and / or the instance first producer socket 386a of instance first hardware accelerator 350a are connected via the instance cross switch 370 of instance HTS 380 to the instance second consumer socket 384b and / or the instance second producer socket 386b of instance second hardware accelerator 350b to form a data stream chain. In some instances, the data stream chain is configured by instance MMR controller 392. In some instances, instance MMR controller 392 is software (SW) programmable. In some instances, instance first hardware accelerator 350a is configured to perform a first task on data elements independently of instance second hardware accelerator 350b; for example, instance hardware accelerators 350a to 350c are configured to perform processing tasks on data elements in parallel.
[0043] although Figure 3 For illustrative purposes, the instance hardware acceleration subsystem 310 includes three instance hardware accelerators 350a to 350c and one instance DMA controller 340, but the instance hardware acceleration subsystem 310 may include any number of instance hardware accelerators 350a to 350c and / or DMA controller 340. Furthermore, the instance hardware acceleration subsystem 310 may include different types of instance hardware accelerators 350a to 350c and / or instance hardware accelerators 350a to 350c that operate on different types of data (e.g., blocks or rows) and / or perform different processing tasks (e.g., LDC, scaling, and noise filtering), thereby allowing the user to customize the instance hardware acceleration subsystem 310 for various functions.
[0044] exist Figure 3In the illustrated instance hardware acceleration subsystem 310, instance schedulers 382a to 382d communicate with corresponding instance hardware accelerators 350a to 350c and corresponding instance DMA controllers 340 to control the processing workflow of instance hardware accelerators 350a to 350c and instance DMA controllers 340. In some instances, instance first scheduler 382a controls the workflow of instance first hardware accelerator 350a. In some instances, instance first scheduler 382a sends a start signal (e.g., a Tstart signal) to instance first hardware accelerator 350a to communicate with instance first hardware accelerator 350a to begin processing data elements. In some instances, instance first hardware accelerator 350a sends a completion signal (e.g., a Tdone signal) to indicate that instance first hardware accelerator 350a has finished processing data elements. In some instances, in response to receiving a Tdone signal, instance-level first scheduler 382a instructs instance-level DMA controller 340 to fetch another data element from instance-level external memory 330. In some instances, instance-level first scheduler 382a sends a start signal to instance-level first hardware accelerator 350a to indicate to instance-level hardware accelerator 350a that frame processing has begun. In some instances, instance-level first hardware accelerator 350a sends a frame end signal to instance-level first scheduler 382a to communicate the end of frame processing, for example, to communicate that instance-level first hardware accelerator has finished processing the frame.
[0045] exist Figure 3 In the instanced hardware acceleration subsystem 310 illustrated in the figure, instanced schedulers 382a to 382d include corresponding instanced consumer sockets 384a to 384d for tracking consumed input data (e.g., data elements retrieved from instanced local memory 360), and corresponding instanced producer sockets 386a to 386d for tracking generated output data (e.g., data elements processed by corresponding instanced hardware accelerators 350a to 350c and corresponding instanced DMA controller 340). In some instances, the instanced first hardware accelerator 350a includes multiple consumer sockets 384a and / or multiple producer sockets 386a. For example, an instance of a first hardware accelerator may include an instance of a first consumer socket 384a and / or an instance of a first producer socket 386a for inputting / outputting data on a chroma channel, and a second consumer socket and / or an instance of a producer socket for inputting / outputting data on a luminance channel.
[0046] In some instances, instance consumer sockets 384a to 384d include consumer dependencies and instance producer sockets 386a to 386d include instance producer dependencies. In some instances, the consumer and producer dependencies are specific to the corresponding instance hardware accelerators 350a to 350c and the corresponding instance DMA controller 340. In some instances, instance consumer sockets 384a to 384d are configured to generate signals indicating the consumption of generated data, such as a dec signal, in response to data consumption by the corresponding instance hardware accelerators 350a to 350c and the corresponding instance DMA controller 340. In some instances, instance producer sockets 386a to 386d are configured to generate signals indicating the availability of consumable data, such as a pend signal, in response to the generation of consumable data by the corresponding instance hardware accelerators 350a to 350c and the corresponding instance DMA controller 340. In some instances, the instance-specific dec signal is routed to the corresponding instance-specific producer and the instance-specific pend signal is routed to the corresponding instance-specific consumer.
[0047] Figure 3 The instance schedulers 382a to 382c of the instance hardware acceleration subsystem 310 illustrated in the figure include instance producer type adapters 390a, 390b, 390c, and 390d coupled to instance producer sockets 386a, 386b, 386c, and 386d to perform logical conversions between line, block, and aggregated block formats.
[0048] In some instances, instance schedulers 382a to 382d implement aggregation (e.g., row-to-row, block-to-block) of multiple sets of output data (e.g., first data element and second data element when the first data element and the second data element have the same data type). In some instances, instance schedulers 382a to 382d implement logical transformations of data elements and / or aggregated data elements between first and second data types (e.g., row-to-2D block and block-to-2D row). Therefore, instance schedulers 382a to 382d implement at least four scenarios, such as row-to-row, row-to-2D block, 2D block-to-row, and 2D block-to-2D block.
[0049] Figure 3The instance hardware accelerator subsystem 310 illustrated herein includes instance workload storage engines 352a, 352b, and 352c coupled to corresponding instance hardware accelerators 350a to 350c. Instance workload storage engines 352a, 352b, and 352c are configured to aggregate at least a first data element (e.g., a first data block) and a second data element (e.g., a second data block) in instance local memory 360 to generate aggregated data elements (e.g., superblocks), and / or to partition the aggregated data elements into at least the first data element and the second data element. In some instances, instance first workload storage engine 352a is configured to aggregate the first data element and the second data element in instance local memory 360. In some instances, the instanced first load storage engine 352a horizontally aggregates data elements based on a block-per-row (BPR) value (e.g., CBUF_BPR) programmed into the instanced MMR controller 392 by the user. This allows tuning based on, for example, the instanced hardware accelerators 350a to 350c, output block size, DDR burst size, destination consumption type, and instanced local memory 360. In some instances, the instanced load storage engines 352a to 352c enable a software (SW) programmable circular buffer stored in the instanced local memory 360 for data aggregation based on the block-per-row (BPR). In some instances, the BPR value is determined by software based on available memory in the instanced local memory 360 and / or memory allocated in the instanced local memory 360 for the instanced hardware accelerators 350a to 350c. In some instances, the BPR value is hard-coded into the instanced MMR controller 392.
[0050] Figure 4 This is a diagram illustrating the use of instanced local memory 360 ( Figure 3 In ) through instanced load storage engines 352a to 352c ( Figure 3 The instantiation of data elements 402, 404, 406, and 408 is performed to generate an instantiation schema of aggregated data elements 420a and 420b (e.g., superblocks 420a and 420b). Figure 4 In the example illustrated herein, instance data elements 402, 404, 406, and 408 are stored in local memory in a first configuration 410 (e.g., Figure 3 In instance local memory 360. In some instances, instance data elements 402, 404, 406, 408 are stored in instance local memory 360 with instance first configuration 410 having (for example) a BPR value of 1 (e.g., a block width) and a buffer size of 4 (e.g., CBUF_SIZE=OBH*4). Figure 3In some instances, aggregation is performed horizontally based on the BPR value received from the instanced MMR controller 392. Figure 4 The instance data elements 402, 404, 406, and 408. For example, the instance first load storage engine 352a may receive a BPR value of 2 from the instance MMR controller 392 and horizontally aggregate the data elements 402, 404, 406, and 408 in the instance local memory 360 to generate an instance second configuration 420 containing two superblocks 420a and 420b, each superblock 420a and 420b having a width of two blocks (e.g., BPR=2). Figure 4 The instanced second configuration 420 illustrated in the figure has a buffer size of 2 (e.g., CBUF_SIZE = OBH * 2). The instanced first load storage engine 352a can be configured to horizontally aggregate any suitable number of blocks into instanced superblocks 420a, 420b having any suitable width as determined by the BPR value from the instanced MMR controller 392. In some instances, the instanced first load storage engine 352a horizontally aggregates processed data blocks 402, 404, 406, 408 received from the instanced first hardware accelerator 350a, writes the processed data blocks 402, 404, 406, 408 to the instanced local memory 360, and aggregates the data blocks 402, 404, 406, 408 in the instanced local memory 360 to generate superblocks 420a, 420b. Figure 4 The instanced local memory 360, illustrated in the diagram, can be implemented via horizontally aggregated superblocks 420a and 420b. Figure 3 Larger read / write operations between the instanced external memory 330 and the instanced external memory 330.
[0051] In some instances, instance-specific load storage engines 352a to 352c are configured to select individual data elements 402, 404, 406, and 408 from corresponding superblocks 420a and 420b. Therefore, in some instances, instance-specific load storage engines 352a to 352c are configured to aggregate individual data elements 402, 404, 406, and 408 to produce aggregated data elements 420a and 420b and / or select individual data elements 402, 404, 406, and 408 from corresponding superblocks 420 and 420b, which (for example) depends on the data format to which the corresponding instance-specific hardware accelerators 350a to 350c are configured to operate and the format of the data transferred from instance-specific external memory 330 to instance-specific local memory 360.
[0052] In some instances, instance-based workload storage engines 352a to 352c receive processed data elements 402, 404, 406, and 408 from corresponding instance-based hardware accelerators 350a to 350c, aggregate the processed data elements 402, 404, 406, and 408 to generate aggregated data elements 420a and 420b, and write the aggregated data elements 420a and 420b to instance-based local memory 360. In some instances, instance-based workload storage engines 352a to 352c receive processed instance-based aggregated data elements 420a and 420b from corresponding instance-based hardware accelerators 350a to 350c, select individual processed data elements 402, 404, 406, and 408 from the processed aggregated data elements 420a and 420b, and write data elements 402, 404, 406, and 408 to instance-based local memory 360. In some instances, data blocks can be aggregated into rows 430a and 430b (e.g., 2D block-to-line rasterization). In some instances, data blocks can be aggregated into rows 430a and 430b by setting the BPR value as a function of the frame width (e.g., BPR = FR_WIDTH / OBW). In some instances, this can be achieved via the instanced DMA controller 340 ( Figure 3 The rasterized data rows 430a and 430b are transferred to the instanced external memory 330.
[0053] In some instances, the instance-first hardware accelerator 350a completes instance-first data elements 402, 404, 406, 408 in response to the instance-first hardware accelerator 350a. Figure 4 The processing generates a completion signal (e.g., a Tdone signal) and sends the Tdone signal to the instanced first scheduler 382a. Figure 3 In some instances, in response to receiving the Tdone signal, the instanced first scheduler 382a instructs the instanced second hardware accelerator 350b to read processed data elements 402, 404, 406, 408. Figure 4 ) or aggregated data elements 420a, 420b, or 430a or 430b. In some instances, the instanced second hardware accelerator 350b consumes aggregated data elements having entire row blocks (e.g., Figure 4 (Aggregated data elements 430a, 430b). In some instances, in response to receiving the Tdone signal, instanced first scheduler 382a ( Figure 3 ) instructs the instance DMA controller 340 to write from instance external memory 330 ( Figure 3 The processed data elements 402, 404, 406, and 408 () Figure 4 (or aggregated data elements 420a or 420b or 430a or 430b.)
[0054] In some instances, in response to the instanced first hardware accelerator 350a ( Figure 3 Process data elements 402, 404, 406, and 408. Figure 4 ), instance-based first load storage engine 352a ( Figure 3 The block count is incremented. In this way, the instanced first load storage engine 352a tracks the data generated by the instanced first hardware accelerator 350a. Figure 3 The data elements processed are 402, 404, 406, and 408. Figure 4 The number of (). In some instances, in response to the instanced first load storage engine 352a determining that the block count equals the BPR value, the instanced first load storage engine 352a aggregates the processed data elements 402, 404, 406, 408 in local memory 360. Figure 4 This generates aggregated data elements 420a and 420b. In some instances, in response to the instanced first load storage engine 352a determining that the block count equals the BPR value, the instanced first hardware accelerator 350a sends a Tdone signal to the instanced first scheduler 382a, at which point the instanced first scheduler 382a may instruct the instanced second hardware accelerator 350b or the instanced third hardware accelerator 350c to read the aggregated data elements 420a and 420b from the instanced local memory 360. In some instances, in response to the Tdone signal, the instanced first scheduler 382a instructs the instanced DMA controller 340 to transfer the aggregated data elements 420a, 420b, 430a, or 430b to the instanced external memory 330. Figure 3 ).
[0055] As described above, the instanced first hardware accelerator 350a and / or the instanced first load storage engine 352a may implement counting logic, which includes incrementing the block count in response to the instanced first hardware accelerator 350a processing data elements. In some instances, Figure 3 The instanced hardware acceleration subsystem 310 illustrated in the diagram includes generation mode parameters (e.g., the Tdone_gen_mode parameter) to enable instanced hardware accelerators 350a to 350c to operate at the block level (e.g., at individual data elements 402, 404, 406, 408). Figure 4The instance hardware accelerators 350a to 350c communicate with the corresponding instance schedulers 382a to 382c at either the block level or the superblock level (e.g., at the aggregated data element level). In some instances, the generation mode parameters are MMR programmable and / or based on BPR values. In some instances, in the first generation mode (e.g., when Tdone_gen_mode = 0), the instance hardware accelerators 350a to 350c communicate with the instance schedulers 382a to 382c at the block level, for example, when processing individual data elements 402, 404, 406, 408 (…). Figure 4 Immediately after processing the superblock, the instance hardware accelerators 350a to 350c send a Tdone signal to the corresponding instance schedulers 382a to 382c. In some instances, in the second generation mode (e.g., when Tdone_gen_mode = 1), the instance hardware accelerators 350a to 350c communicate with the instance schedulers 382a to 382c at the superblock level. For example, after processing the superblock based on the BPR value (e.g., when the instance hardware accelerators 350a to 350c have processed a certain number of data elements 402, 404, 406, 408 equal to the BPR value), the instance hardware accelerators 350a to 350c immediately send a Tdone signal to the corresponding instance schedulers 382a to 382c. For example, if the BPR value is 2 (e.g., ... Figure 4 Superblock 420a), and if instance hardware accelerators 350a to 350c are processing superblock 420a ( Figure 4 If the instance hardware accelerators 350a to 350c are communicating with the corresponding instance schedulers 382a to 382c (e.g., Tdone_gen_mode = 1), then the instance hardware accelerators 350a to 350c will immediately send the Tdone signal to the corresponding instance schedulers 382a to 382c after processing the two data elements 402 and 404.
[0056] The flexibility of communicating at the block level or the superblock level prevents the instanced schedulers 382a to 382c and / or the instanced HTS 380 from processing individual data blocks 402, 404, 406, and 408 of the superblocks 420a and 420b on the instanced first hardware accelerator 350a. Figure 4 This triggers a DMA transfer. For example, if instance-based first load storage engine 352a transfers two data blocks 402 and 404 based on a BPR value of 2... Figure 4 ) Horizontally aggregated into superblocks (e.g., Figure 4If an instance superblock 420a is provided, and the corresponding instance hardware accelerator 350a communicates with the instance first scheduler 382a and / or the instance HTS 380 at the block level (e.g., in a first generation mode), then the instance first scheduler 382a and / or the instance HTS 380 can trigger a DMA transfer after a data block 402 or 404 is processed, instead of waiting until both data blocks 402 and 404 in the aggregated data block 420a are processed.
[0057] In some instances, aggregated data elements (e.g., instance-specific aggregated data elements 430a, 430b of the third configuration 430) have an equal value to those generated by... Figure 3 The instance hardware accelerator subsystem 310 processes the BPR value of the frame width of the input frame (e.g., BPR = FR_WIDTH / OBW). In scenarios where the frame width of the input frame is not a multiple of the BPR value (e.g., when the frame width = 10 blocks and the BPR value = 4), the aggregated data element 430a may include an end-of-line (EOR) trigger mode (e.g., partial_bpr_trigmode) to account for scenarios where the frame width of the input frame is not a multiple of the superblock size. For example, if instance hardware accelerators 350a to 350c operate in EOR trigger mode, and the BPR value is 4 and the remaining superblock buffer has two blocks, then the number of blocks in the superblock buffer is 50% of the BPR value, and instance hardware accelerators 350a to 350c will send an EOR trigger to the corresponding instance schedulers 382a to 382c and / or instance HTS 380 after processing the two blocks in the superblock buffer at the end of the line. In some instances, when the instance hardware accelerator 350a operates in EOR-triggered mode (e.g., partial_bpr_trigmode = 1), the instance corresponding schedulers 382a to 382c and / or the instance HTS 380 trigger the instance DMA controller 340 to transfer the EOR superblock to the instance external memory 330 via a separate DMA channel. In some instances, the instance first load storage engine 352a communicates with the instance first scheduler 382a via partial BPR counting (e.g., partial_bpr_count mode) to indicate the remaining block count in the EOR superblock buffer.
[0058] By combining multiple data blocks horizontally as described herein and enabling instance hardware accelerators 350a to 350c to communicate at the block level and / or superblock level, instance workload storage engines 352a to 352c enable larger reads and writes to instance local memory 360 from / to instance external memory 330, thereby improving DDR efficiency. With larger memory requests (up to frame width), DDR page opening / closing is significantly reduced.
[0059] Figure 5 This is an instance diagram illustrating the instance type adaptation process 500, which is implemented by an instance type adapter to logically convert a 24x32-byte data block 532 into 24 rows of data elements 534. Figure 5 In the instance-type adapter diagram 500, the instance-type consumer socket 584 allows the instance-type first hardware accelerator 350a and / or the instance-type first load storage engine 352a to read from the instance-type local memory 360 and tracks the number of times the instance-type first hardware accelerator 350a and / or the instance-type first load storage engine 352a reads from the instance-type local memory 360. For example, in Figure 5 In the example illustrated, the instance-type adapter 588 can logically convert a 24x32 data block 532 into 24 rows of data elements. In some instances, instance-type hardware accelerators (e.g., Figure 3 The instance of the first hardware accelerator 350a) performs processing tasks on 24 rows of data elements 534.
[0060] In some instances, instance-based schedulers (e.g., Figure 3 The instanced first scheduler 382a) reads the Tdone signal from the instanced first hardware accelerator 350a, and in response, the instanced type adapter 588 logically converts the 24 rows of data elements into 24x32B data blocks, and the instanced producer socket 586 generates processed data blocks as output data for consumption by another hardware accelerator and / or via the instanced DMA controller (e.g., Figure 3 The data is transferred to the instance external memory 330 via the instance DMA controller 340.
[0061] By transforming data elements between rows, blocks, and superblocks, Figure 3 Instance-type producer adapters 390a to 390d and / or Figure 5 The instance-type adapter 588 enables the instance-type hardware acceleration subsystem 310 ( Figure 3This enables the use of corresponding instance-specific hardware accelerators 350a to 350c and / or corresponding instance-specific DMA controllers 340 to process data elements, thereby implementing complex user-defined multi-producer and multi-consumer hardware acceleration schemes for various functions, while maintaining DDR external memory (e.g., Figure 3 Improved efficiency of the instance external memory 330.
[0062] Figure 6 This is an instance user-defined diagram 600 illustrating an instance of a multi-consumer / multi-producer hardware acceleration subsystem 610 for image, vision, and / or video processing. Figure 6 The exemplary hardware acceleration subsystem 610 illustrated herein includes: an exemplary lens distortion correction (LDC) hardware accelerator 650a for performing lens distortion correction on data blocks; an exemplary scaling (MSC) hardware accelerator 650b for performing scaling on data lines; an exemplary noise filtering (NF) hardware accelerator 650c for performing noise filtering on data lines; an exemplary first DMA controller 640 communicating with the exemplary LDC hardware accelerator 650a and the exemplary DDR external memory 630; an exemplary second DMA controller 642 communicating with the exemplary LDC hardware accelerator 650a and the exemplary DDR external memory 630; an exemplary third DMA controller 644 communicating with the exemplary MSC hardware accelerator 650b, the exemplary NF hardware accelerator 650c, and the exemplary DDR external memory 630; and an exemplary fourth DMA controller 646 communicating with the exemplary NF hardware accelerator 650c and the exemplary DDR external memory 630.
[0063] exist Figure 6 In the exemplary user-defined hardware acceleration subsystem 610 illustrated in the diagram, the exemplary LDC hardware accelerator 650a generates multiple outputs including (for example) data blocks 632 consumed by the exemplary first DMA controller 640. Figure 6 In one instance, the instance MSC hardware accelerator 650b consumes a set of data rows based on the output of the instance LDC hardware accelerator 650a, and the instance second DMA controller 642 consumes data block 634. Figure 6 In one instance, the instance MSC hardware accelerator 650b consumes a set of data lines based on the output of the instance LDC hardware accelerator 650a, performs a scaling operation on the data lines, and generates data line elements 636 consumed by the instance NF hardware accelerator 650c and the instance third DMA controller 644. Figure 6In this instance, the instanced NF hardware accelerator 650c consumes data row element 636, performs noise filtering on data row element 636, and generates data row element 638 consumed by the instanced fourth DMA controller 646. Figure 6 In the example, the instance first DMA controller 640, instance second DMA controller 642, instance third DMA controller 644 and instance fourth DMA controller 646 are configured to write the corresponding data elements 632, 634, 636 and 638 to the instance DDR external memory 630.
[0064] Figure 7 This is an example diagram 700 illustrating an example of a multi-consumer / multi-producer hardware acceleration scheme. Figure 7 The illustrated instance hardware acceleration subsystem 710 includes an instance DDR external memory 730, an instance DMA controller 740, an instance LDC hardware accelerator 750a, an instance MSC / NF hardware accelerator 750b, and an instance HTS 780. Figure 7 In the instance hardware acceleration subsystem 710, the instance LDC hardware accelerator 750a is configured to perform lens distortion correction operations on the data to generate data element 732. In the instance hardware acceleration subsystem 710, the instance MSC / NF hardware accelerator 750b is configured to consume data element 732, perform scaling and noise filtering operations on data element 732, and generate data element 734. In the instance hardware acceleration subsystem 710, the instance DMA controller 740 is configured to consume data elements 732 and 734 generated by the instance LDC hardware accelerator 750a and the instance MSC hardware accelerator 750b, respectively, and write data elements 732 and 734 to the instance DDR external memory 730.
[0065] Considering local storage availability, aggregation requirements can vary based on the data consumer. For example, the MSC Hardware Accelerator 650b ( Figure 6 It may require all rows of data to be available, while DMA CH writes can be adapted to aggregate several blocks to save DDR bandwidth. To achieve different aggregations, each output channel can be programmed in LSE 352a to include different BPR values.
[0066] Figure 8 The diagram illustrates an instance of a multi-producer LDC hardware accelerator 880, which is used to output a first aggregated data element 832 on a first channel (e.g., a chroma channel) and a second aggregated data element 834 on a second channel (e.g., a luminance channel). Figure 8In the illustrated example, the first channel is associated with a first BPR value (e.g., first BPR value 4) and the second channel is associated with a second BPR value (e.g., a BPR value equal to the frame width). In some instances, the first BPR value and / or the second BPR value are based on the number of blocks in a frame line, the number of pixels in a frame line, and / or the number of bytes in a frame line. Figure 8 In the example illustrated in the diagram, the first aggregated data element 832 is output to an external DDR (e.g., Figure 3 The instance external memory 330) and the second data element 834 are output to a second hardware accelerator (e.g., an MSC / NF hardware accelerator). Thus, the instance described herein implements the aggregation of individual asymmetric data elements (e.g., on a separate data channel).
[0067] Despite Figure 9 Chinese illustrations explain implementation Figure 3 The hardware acceleration subsystem 310 is an instance of the hardware acceleration subsystem 310, but it can be combined, partitioned, rearranged, omitted, eliminated, and / or implemented in any other way. Figure 9 One or more of the elements, processes, and / or devices illustrated in the diagram. Additionally, instance hardware accelerators 350a to 350c, instance schedulers 382a to 382d, instance load storage engines 352a to 352c, instance producer-type adapters 390a to 390d, instance DMA controller 340, instance local memory 360, and / or more generally... Figure 3 The instanced hardware acceleration subsystem can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, instanced hardware accelerators 350a to 350c, instanced schedulers 382a to 382d, instanced load storage engines 352a to 352c, instanced producer-type adapters 390a to 390d, instanced DMA controllers 340, instanced local memory 360, and / or more generally... Figure 3Any of the instance hardware acceleration subsystems 310 may be implemented by one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, graphics processing units (GPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field-programmable logic devices (FPLDs). When the device or system claims of this patent are read to cover purely software and / or firmware implementations, at least one of the instance hardware accelerators 350a to 350c, instance schedulers 382a to 382d, instance load storage engines 352a to 352c, instance producer-type adapters 390a to 390d, instance DMA controllers 340, and instance local memory 360 is hereby explicitly defined as including a non-transitory computer-readable storage device or storage disk, such as a memory, digital versatile optical disc (DVD), optical disc (CD), Blu-ray disc, etc., including software and / or firmware. Furthermore, in addition to or replacing Figure 9 The elements, processes, and / or apparatus illustrated herein. Figure 3 The instance hardware acceleration subsystem 310 may include one or more elements, processes, and / or devices, and / or may include more than one of any or all of the elements, processes, and devices illustrated herein. As used herein, the phrase “to communicate” includes variations thereof, encompassing direct communication and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or continuous communication, but further includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-off events.
[0068] exist Figure 9 The middle section shows the implementation. Figure 3 The hardware acceleration subsystem 310 includes an instance of hardware logic, machine-readable instructions, a hardware-implemented state machine, and / or any combination thereof, as shown in the flowchart. The machine-readable instructions may be provided by a computer processor and / or processor circuitry (e.g., in conjunction with the following). Figure 10 The processor 1012 shown in the described exemplary processor platform 1000 executes one or more executable programs or portions thereof. The program may be embodied in software stored on a non-transitory computer-readable storage medium (e.g., CD-ROM, floppy disk, hard disk drive, DVD, Blu-ray disc, or memory associated with the processor 1012), but the entire program and / or portions thereof may alternatively be executed by a device other than the processor 1012 and / or embodied in firmware or dedicated hardware. Furthermore, although references... Figure 9The flowcharts illustrated herein describe an exemplary program, but many other methods of implementing the exemplary hardware acceleration subsystem 310 may be used alternatively. For example, the execution order of the boxes may be changed and / or some boxes in the described boxes may be changed, eliminated, or combined. Additionally or alternatively, any or all of the boxes may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuit systems, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) structured to perform the corresponding operations without executing software or firmware. The processor circuit system may be distributed across different network locations and / or local to one or more devices (e.g., a multi-core processor in a single machine, multiple processors distributed across a server rack, etc.).
[0069] The machine-readable instructions described herein may be stored in one or more of the following formats: compressed format, encrypted format, segmented format, compiled format, executable format, and packaged format. As described herein, machine-readable instructions may be stored as data or data structures (e.g., portions of instructions, code, representations of code, etc.) that can be used to create, manufacture, and / or produce machine-executable instructions. For example, machine-readable instructions may be segmented and stored on one or more storage devices and / or computing devices (e.g., servers) located at the same or different locations within a network or network set (e.g., in the cloud, at an edge device, etc.). Machine-readable instructions may require one or more of the following to be installed, modified, adapted, updated, combined, supplemented, configured, decrypted, decompressed, unpacked, distributed, reassigned, compiled, etc., so that they can be directly read, interpreted, and / or executed by computing devices and / or other machines. For example, machine-readable instructions may be stored in multiple portions that are individually compressed, encrypted, and stored on separate computing devices, wherein the portions, when decrypted, decompressed, and combined, form a set of executable instructions that implement one or more functions that together form a program (such as the program described herein).
[0070] In another instance, machine-readable instructions may be stored in a state in which they can be read by the processor's circuitry, but libraries (e.g., dynamic link libraries (DLLs)), software development kits (SDKs), application programming interfaces (APIs), etc., need to be added to execute the instructions on a specific computing device or other device. In yet another instance, machine-readable instructions (e.g., storage settings, data input, recorded network addresses, etc.) may need to be configured before they can be executed in whole or in part. Therefore, machine-readable media as used herein may contain machine-readable instructions and / or programs, regardless of their specific format or state at storage or otherwise at rest or in transit.
[0071] The machine-readable instructions described in this article can be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, machine-readable instructions can be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, Hypertext Markup Language (HTML), Structured Query Language (SQL), Swift, etc.
[0072] As mentioned above, Figure 9 An instance of such a process may be performed using executable instructions (e.g., computer and / or machine-readable instructions) stored on a non-transitory computer and / or machine-readable medium (e.g., hard disk drive, flash memory, read-only memory, optical disk, digital versatile optical disk, cache memory, random access memory, and / or any other storage device or storage disk) where information is stored for any duration (e.g., extended time period, permanent, short duration, temporary buffer, and / or for caching information). As used herein, the term non-transitory computer-readable medium is explicitly defined as comprising any type of computer-readable storage device and / or storage disk, excluding propagation signals and transmission media.
[0073] The terms “including” and “comprising” (and all their forms and tenses) are used herein as open-ended terms. Therefore, whenever a claim uses any form of “including” or “comprising” (e.g., includes, includes, comprising, having, etc.) in the preamble or in any kind of claim statement, additional elements, terms, etc., may be present without exceeding the scope of the corresponding claim or statement. As used herein, the phrase “at least” is open-ended when used as a transitional term in the preamble of a claim, in the same way that the terms “including” and “comprising” are open-ended. The term “and / or”, when used (for example) in the form of, for example, A, B, and / or C, refers to any combination or subset of A, B, and C, such as (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, and (7) A and B and C. As used herein in the context of describing structures, components, items, objects, and / or things, the phrase “at least one of A and B” means an implementation comprising any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein in the context of describing structures, components, items, objects, and / or things, the phrase “at least one of A or B” means an implementation comprising any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. As used herein in the context of describing the execution or implementation of processes, instructions, actions, activities, and / or steps, the phrase “at least one of A and B” means an implementation comprising any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein in the context of describing the execution or implementation of processes, instructions, actions, activities and / or steps, the phrase “at least one of A or B” means an implementation comprising any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B.
[0074] As used herein, singular references (e.g., "a (a, an)", "first", "second", etc.) do not exclude plurals. The term "a (a or an)" as used herein refers to one or more of the entities mentioned. The terms "a (a)" (or "a (an)"), "one or more", and "at least one" are used interchangeably herein. Furthermore, although individually listed, multiple components, elements, or method actions may be implemented by, for example, a single unit or processor. Additionally, while individual features may be included in different instances or claims, these features may be combined, and their inclusion in different instances or claims does not imply that such combinations are impractical and / or disadvantageous.
[0075] Figure 9 This is a flowchart representing machine-readable instructions that can be executed to perform... Figure 3 The instanced hardware acceleration subsystem 310 is used to achieve data aggregation and type adaptation.
[0076] At box 902, a hardware accelerator (e.g., a lens distortion correction hardware accelerator) processes the data block. For example, instance-first hardware accelerator 350a ( Figure 3 It can process data blocks 402 (e.g., from...). Figure 4 The first configuration 410).
[0077] At box 904, the hardware accelerator writes the processed data block to local memory. For example, instance-first hardware accelerator 350a can write the processed data block 402 ( Figure 4 Write to instance local memory 360 ( Figure 3 ).
[0078] At box 906, the load storage engine coupled to the hardware accelerator determines whether the hardware accelerator communicates at the block level or the superblock level. For example, instanced first load storage engine 352a can determine whether instanced first hardware accelerator 350a communicates at the block level (e.g., Tdone_gen_mode = 0) or at the superblock level (e.g., Tdone_gen_mode = 1).
[0079] If the load storage engine determines that the hardware accelerator is communicating at the block level (box 906), then machine-readable instruction 900 proceeds to box 914, where the hardware accelerator sends a completion signal to the corresponding scheduler. For example, if instance-first load storage engine 352a determines that instance-first hardware accelerator 350a is communicating at the block level (e.g., Tdone_gen_mode = 0), then instance-first hardware accelerator 350a sends a completion signal (e.g., a Tdone signal) to instance-first scheduler 382a. The program ends.
[0080] In some instances, the scheduler triggers a second hardware accelerator (e.g., scales the hardware accelerator) to read processed data blocks from local memory or triggers the DMA controller to write processed data blocks to instanced external memory 330. Figure 3 For example, instance-first scheduler 382a may trigger instance-second hardware accelerator 350b to read processed data block 402 from instance-local memory 360 or trigger instance-DMA controller 340 to write processed data block 402 to instance-external memory 330. Figure 3 ).
[0081] If the load storage engine determines that the hardware accelerator communicates at the superblock level (e.g., Tdone_gen_mode = 1) (box 906), then the hardware accelerator increments the block count (box 908). For example, if instance-first load storage engine 352a determines that instance-first hardware accelerator 350a communicates at the superblock level (e.g., Tdone_gen_mode = 1) (box 906), then instance-first hardware accelerator 350a may increment the block count by 1.
[0082] In some instances, the load storage engine determines whether the block count is equal to the BPR value. If the load storage engine determines that the block count is not equal to the BPR value (e.g., the block count is less than the BPR value) (box 910), then a machine-readable instruction returns to box 902 and the hardware accelerator processes another data block (box 902). For example, if the BPR value is 2 (e.g., BPR = 2) and the block count is 1 (e.g., the hardware accelerator has processed one data block 402), then the instance load storage engine 352a may determine that the block count is not equal to the BPR value (box 910) and the instance machine-readable instruction 900 returns to box 902, where the instance first hardware accelerator 350a processes another data block 404 (e.g., from...). Figure 4 The first configuration 410).
[0083] If the load storage engine determines that the block count equals the BPR value (box 910), then the load storage engine aggregates data blocks in local memory based on the BPR value to generate an aggregated data block (box 912). For example, if the BPR value is 2 (e.g., BPR = 2) and the block count is 2 (e.g., the instanced first hardware accelerator 350a has processed two data blocks 402, 404), then the instanced first load storage engine 352a aggregates data blocks 402, 404 and generates an aggregated data block (e.g., a superblock) 420a.
[0084] At box 914, the hardware accelerator sends a completion signal to the corresponding scheduler. For example, instance first hardware accelerator 350a may send a completion signal (e.g., a Tdone signal) to instance first scheduler 382a.
[0085] At box 916, in response to a completion signal, the scheduler triggers a second hardware accelerator to read aggregated data blocks from local memory or triggers a DMA controller to write aggregated data elements to instanced external memory 330. Figure 3 For example, in response to the Tdone signal, instance-first scheduler 382a may trigger instance-second hardware accelerator 350b to read superblock 420a from instance-local memory 360 or trigger instance-DMA controller 340 to write superblock 420a to instance-external memory 330. Figure 3 The program has ended.
[0086] Despite the instance-first load storage engine 352a in Figure 9 When communicating at the superblock level, data blocks 402 and 404 (box 912) are aggregated. However, in some instances, the instanced first workload storage engine 352a aggregates data blocks 402 and 404 when communicating at the block level. Therefore, in some instances, the instanced first workload storage engine 352a aggregates data blocks 402 and 404 regardless of whether the instanced first hardware accelerator communicates at the block level or the superblock level. In some instances, the instanced first workload storage engine 352a aggregates data blocks 402 and 404 such that the instanced first workload storage engine 352a writes the data blocks as single blocks to the instanced local memory 360 (e.g., the address of data element 404 is contiguous with that of data element 402).
[0087] Figure 10 This is a block diagram of an instance processor platform 1000, which is structured to execute... Figure 9 To implement the instructions Figure 3 The processor platform 1000 can be used as a server, personal computer, workstation, self-learning machine (e.g., neural network), mobile device (e.g., mobile phone, smartphone), tablet computer (e.g., iPad). TM Personal digital assistants (PDAs), Internet devices, DVD players, CD players, digital video recorders, Blu-ray players, game consoles, personal video recorders, set-top boxes, headphones or other wearable devices, or any other type of computing device.
[0088] The illustrated example processor platform 1000 includes a combination of Figure 3 The described instance of HWA subsystem 310.
[0089] The illustrated example processor platform 1000 includes a processor 1012. The illustrated example processor 1012 is hardware. For example, the processor 1012 may be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. The hardware processor may be a semiconductor-based (e.g., silicon-based) device.
[0090] The processor 1012 illustrated in the example includes local memory 1013 (e.g., cache memory). The processor 1012 of the illustrated example communicates via bus 1018 with main memory, which includes volatile memory 1014 and non-volatile memory 1016. Volatile memory 1014 may be implemented using synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS® dynamic random access memory (RDRAM®), and / or any other type of random access memory device. Non-volatile memory 1016 may be implemented using flash memory and / or any other desired type of memory device. Access to main memory 1014, 1016 is controlled by a memory controller.
[0091] The processor platform 1000 illustrated in the diagram also includes interface circuitry 1020. Interface circuitry 1020 can be implemented using any type of interface standard, such as Ethernet interface, Universal Serial Bus (USB), Bluetooth® interface, Near Field Communication (NFC) interface, and / or PCI High Speed interface.
[0092] In the illustrated example, one or more input devices 1022 are connected to interface circuitry 1020. Input devices 1022 allow users to input data and / or commands into processor 1012. For example, input devices may be implemented as audio sensors, microphones, cameras (still or video), keyboards, buttons, mice, touchscreens, tracking pads, trackballs, isopoint devices, and / or voice recognition systems.
[0093] One or more output devices 1024 are also connected to the interface circuit 1020 of the illustrated example. For example, the output device 1024 may be implemented by a display device (e.g., light-emitting diode (LED), organic light-emitting diode (OLED), liquid crystal display (LCD), cathode ray tube display (CRT), in-place switching (IPS) display, touch screen, etc.), a haptic output device, a printer, and / or a speaker. Therefore, the interface circuit 1020 of the illustrated example may include a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0094] The interface circuit 1020 illustrated in the diagram also includes communication devices (e.g., transmitter, receiver, transceiver, modem, residential gateway, wireless access point), and / or a network interface for facilitating data exchange with external machines (e.g., any type of computing device) via network 1026. For example, communication may be via Ethernet connection, digital subscriber line (DSL) connection, telephone line connection, coaxial cable system, satellite system, field wireless system, cellular telephone system, etc.
[0095] The processor platform 1000 illustrated in the diagram also includes one or more mass storage devices 1028 for storing software and / or data. Examples of such mass storage devices 1028 include floppy disk drives, hard disk drives, optical disk drives, Blu-ray disc drives, redundant array of independent disks (RAID) systems, and digital multifunction optical disc (DVD) drives.
[0096] Figure 9 The machine-executable instructions 1032 may be stored in a mass storage device 1028, in volatile memory 1014, in non-volatile memory 1016, and / or on a removable non-transitory computer-readable storage medium (e.g., CD or DVD).
[0097] exist Figure 11 The block diagram is used to illustrate the software (e.g., Figure 9 The instanced computer-readable instructions (1032) are distributed to an instanced software distribution platform 1105 for a third party. The instanced software distribution platform 1105 may be implemented by any computer server, data facility, cloud service, etc., capable of storing software and transmitting it to other computing devices. The third party may be a customer of an entity that owns and / or operates the software distribution platform. For example, the entity owning and / or operating the software distribution platform may be the software (e.g., Figure 9 The developer, seller, and / or licensor of the instance computer-readable instructions 1032. Third parties may be consumers, users, retailers, OEMs, etc., who purchase and / or license the software for use and / or resell and / or sublicense. In the illustrated example, the software distribution platform 1205 includes one or more servers and one or more storage devices. The storage devices store the computer-readable instructions 1032, which can be used with... Figure 9The instance computer-readable instruction 1032 corresponds to that described above. One or more servers of the instance software distribution platform 1105 communicate with network 1110, which may correspond to the Internet and / or any one or more of the instance network 1026 described above. In some instances, as part of a business transaction, one or more servers respond to a request to transfer software to a requesting party. Payment for the delivery, sale, and / or licensing of the software may be handled by one or more servers of the software distribution platform and / or via a third-party payment entity. The server enables the purchaser and / or licensor to download the computer-readable instruction 1032 from the software distribution platform 1105. For example, it may be compatible with... Figure 9 The software corresponding to the instanced computer-readable instruction 1032 is downloaded to the instanced processor platform 1000, which will execute the computer-readable instruction 1032 to implement... Figure 3 The equipment. In some instances, one or more servers of the software distribution platform 1105 periodically provide, deliver, and / or force software updates (e.g., Figure 9 The instance computer-readable instructions 1032) ensure that software is distributed and applied for improvements, patches, updates, etc., at the end-user device.
[0098] Based on the foregoing, it will be understood that an exemplary system, method, and apparatus for implementing data aggregation and type adaptation in a hardware acceleration subsystem have been described. The described methods, apparatus, and articles improve the efficiency of using computing devices by improving the efficiency of external memory and implementing user-defined multi-producer and multi-consumer hardware acceleration schemes. The described methods, apparatus, and articles therefore represent one or more improvements to the functionality of a computer.
[0099] The examples described herein include a single-chip system (SoC) comprising: a first scheduler; a first hardware accelerator coupled to the first scheduler to process at least a first data element and a second data element; and a first load storage engine coupled to the first hardware accelerator, the first load storage engine being configured to: communicate with the first scheduler at a superblock level by sending a completion signal to the first scheduler in response to determining that a block count equals a first BPR value; and aggregate the first data element and the second data element based on the first BPR value to generate a first aggregated data element.
[0100] In some instances, the first load storage engine increments the block count in response to the first hardware accelerator processing the first data element and in response to the first hardware accelerator processing the second data element.
[0101] In some instances, the first scheduler, in response to receiving the completion signal from the first hardware accelerator, instructs the DMA controller to store the first aggregated data element to external memory.
[0102] In some instances, the first scheduler, in response to receiving the completion signal from the first hardware accelerator, instructs the second hardware accelerator to read the first aggregated data element.
[0103] In some instances, the first load storage engine is configured to communicate with the first scheduler at the block level by sending a completion signal to the first scheduler in response to the first hardware accelerator processing the first data block.
[0104] In some instances, the first BPR value is associated with the first data channel.
[0105] In some instances, the SoC includes a software (SW) programmable memory-mapped register (MMR) coupled to the first scheduler, the MMR being used to provide at least the first BPR value to the first load storage engine.
[0106] In some instances, the first load storage engine is configured to aggregate at least a third data element with a fourth data element based on a second BPR value to produce a second aggregated data element.
[0107] In some instances, the second BPR value is associated with a second data channel.
[0108] In some instances, the first load storage engine enables a software (SW) programmable circular buffer in local memory for data aggregation based at least on the first BPR value.
[0109] In some instances, the first scheduler includes a first consumer socket for tracking input data consumed by the first hardware accelerator and a first producer socket for tracking output data generated by the hardware accelerator.
[0110] In some instances, the first scheduler includes a first producer-type adapter coupled to the first producer socket.
[0111] The examples described herein include a method comprising: processing a first data element and a second data element by a first hardware accelerator; sending a completion signal to a first scheduler by a first load storage engine in response to determining that a block count equals a first BPR value; and aggregating the first data element and the second data element by the first load storage engine based on the first BPR value to generate a first aggregated data element.
[0112] In some instances, the method further includes: incrementing the block count by the first load storage engine in response to the first hardware accelerator processing the first data element; and incrementing the block count by the first load storage engine in response to the first hardware accelerator processing the second data element.
[0113] In some instances, the method further includes: in response to receiving the completion signal from the first hardware accelerator, the first scheduler instructs the DMA controller to store the first aggregated data element to external memory.
[0114] In some instances, the method further includes: in response to receiving the completion signal from the first hardware accelerator, the first scheduler instructs a second hardware accelerator to read the first aggregated data element.
[0115] In some instances, the first BPR value is associated with the first data channel.
[0116] In some instances, the method further includes: the first load storage engine aggregating at least a third data element with a fourth data element based on a second BPR value to generate a second aggregated data element.
[0117] In some instances, the second BPR value is associated with a second data channel.
[0118] The examples described herein include a non-transitory computer-readable medium comprising computer-readable instructions that, when executed, cause at least one processor to perform at least the following operations: processing a first data element and a second data element by a first hardware accelerator; sending a completion signal to a first scheduler by a first load storage engine in response to determining that a block count equals a first BPR value; and aggregating the first data element and the second data element by the first load storage engine based on the first BPR value to generate a first aggregated data element.
[0119] In some instances, the computer-readable instructions further cause the at least one processor to perform at least the following operations: incrementing the block count by the first load storage engine in response to the first hardware accelerator processing the first data element, and incrementing the block count by the first load storage engine in response to the first hardware accelerator processing the second data element.
[0120] In some instances, the computer-readable instructions further cause the at least one processor to perform at least the following operations: in response to receiving the completion signal from the first hardware accelerator, the first scheduler instructs the DMA controller to store the first aggregated data element to external memory.
[0121] In some instances, the computer-readable instructions further cause the at least one processor to perform at least the following operations: in response to receiving the completion signal from the first hardware accelerator, the first scheduler instructs the second hardware accelerator to read the first aggregated data element.
[0122] In some instances, the computer-readable instructions further cause the at least one processor to perform at least the following operations: in response to the first hardware accelerator processing the first data block, the first hardware accelerator sends a completion signal to the first scheduler.
[0123] In some instances, the first BPR value is associated with the first data channel.
[0124] In some instances, the computer-readable instructions further cause the at least one processor to perform at least the following operation: the first load storage engine aggregates at least a third data element with a fourth data element based on a second BPR value to generate a second aggregated data element.
[0125] In some instances, the second BPR value is associated with a second data channel.
[0126] The examples described herein include an apparatus comprising: means for processing a first data element and a second data element; means for sending a completion signal to a first scheduler in response to determining that a block count equals a first BPR value; and means for aggregating the first data element and the second data element based on the first BPR value to generate a first aggregated data element.
[0127] In some instances, the device further includes components for incrementing the block count in response to the first hardware accelerator processing the first data element and components for incrementing the block count in response to the first hardware accelerator processing the second data element.
[0128] In some instances, the device further includes a component for instructing a DMA controller to store the first aggregated data element to external memory in response to receiving the completion signal from the first hardware accelerator.
[0129] In some instances, the device further includes a component for instructing a second hardware accelerator to read the first aggregated data element in response to receiving the completion signal from the first hardware accelerator.
[0130] In some instances, the device further includes a component for sending a completion signal to the first scheduler in response to the first hardware accelerator processing the first data block.
[0131] In some instances, the first BPR value is associated with the first data channel.
[0132] In some instances, the device further includes a component for aggregating at least a third data element with a fourth data element based on a second BPR value to generate a second aggregated data element.
[0133] In some instances, the second BPR value is associated with a second data channel.
[0134] While specific exemplary methods, apparatuses, and articles of manufacture have been described herein, the scope of this patent is not limited thereto. Rather, this patent fairly covers all methods, apparatuses, and articles of manufacture that fall within the scope of the claims of this patent.
[0135] The appended claims are hereby incorporated by reference into this detailed description, wherein each claim is an independent embodiment of the present description.
Claims
1. A circuit device comprising: a memory; a first hardware accelerator coupled to the memory and configured to: process a set of data to produce a set of data elements; and cause the set of data elements to be stored in the memory; a second hardware accelerator coupled to the memory; a load store circuit coupled between the first hardware accelerator and the memory, and the load store circuit is configured to: determine whether a number of elements in the set of data elements satisfies an aggregation threshold; and based on the number of elements satisfying the aggregation threshold, cause the second hardware accelerator to process the set of data elements.
2. The circuit device of claim 1, further comprising a scheduler circuit coupled to the first hardware accelerator and the second hardware accelerator, wherein the load store circuit is configured to cause the second hardware accelerator to process the set of data elements by causing a completion signal to be provided to the scheduler circuit.
3. The circuit device of claim 1, wherein the load store circuit is configured to aggregate the set of data elements into an aggregated data element based on the number of elements satisfying the aggregation threshold.
4. The circuit device of claim 3, wherein the load store circuit is configured to cause the set of data elements to be stored in the memory as the aggregated data element.
5. The circuit device of claim 1, further comprising a memory map register configured to store the aggregation threshold.
6. The circuit device of claim 1, wherein the set of data is a set of image data, and each data element in the set of data elements is a two-dimensional block of image data.
7. The circuit device of claim 1, wherein: the set of elements is a first set of elements, and the set of elements is associated with a first channel; the aggregation threshold is a first aggregation threshold, and the aggregation threshold is associated with the first channel; the first hardware accelerator is configured to: produce a second set of data elements associated with a second channel; and cause the second set of data elements to be stored in the memory; the load store circuit is further configured to: determine whether a number of elements in the second set of data elements satisfies a second aggregation threshold associated with the second channel and different from the first aggregation threshold; and based on the number of elements in the second set of data elements satisfying the second aggregation threshold, cause the second hardware accelerator to process the second set of data elements.
8. The circuit device of claim 7, wherein the first channel is a chroma channel, and the second channel is a luma channel.
9. The circuit device of claim 1, wherein each data element in the set of data elements has a size selected from the group of: 16x16 bytes, 32x32 bytes, and 64x32 bytes.
10. A circuit device comprising: a first memory; a direct memory access (DMA) circuit coupled to the first memory and configured to be coupled to a second memory; a hardware accelerator coupled to the first memory and configured to: perform an operation on a set of data to produce a set of data elements; and cause the set of data elements to be stored in the first memory; a load store circuit coupled between the hardware accelerator and the first memory and configured to: determine whether a number of elements in the set of data elements satisfies an aggregation threshold; and based on the number of elements satisfying the aggregation threshold, cause the DMA circuit to cause the set of data elements to be stored in the second memory.
11. The circuit device of claim 10, further comprising a scheduler circuit coupled to the DMA circuit and the hardware accelerator, wherein the load store circuit is configured to cause the DMA circuit to cause the set of data elements to be stored in the second memory by causing a completion signal to be provided to the scheduler circuit.
12. The circuit device of claim 10, wherein the load store circuit is configured to aggregate the set of data elements into an aggregated data element based on the number of elements satisfying the aggregation threshold.
13. The circuit device of claim 12, wherein the load store circuit is configured to cause the set of data elements to be stored in the first memory as the aggregated data element.
14. A method comprising: processing a first set of data and a second set of data using first processing circuitry to produce a first set of data elements and a second set of data elements associated with a first channel and a second channel, respectively; storing the first set of data elements and the second set of data elements in a memory; determining whether a number of elements of the first set of data elements satisfies a first aggregation threshold associated with the first channel; determining whether a number of elements of the second set of data elements satisfies a second aggregation threshold associated with the second channel; based on the number of elements of the first set of data elements satisfying the first aggregation threshold, processing the first set of data elements using second processing circuitry; and based on the number of elements of the second set of data elements satisfying the second aggregation threshold, processing the second set of data elements using the second processing circuitry.
15. The method of claim 14, further comprising providing a completion signal to a scheduler circuit based on the number of elements of the set of data elements satisfying the aggregation threshold, wherein processing the set of data elements using the second processing circuitry is further based on the completion signal.
16. The method of claim 14, further comprising aggregating the set of data elements into an aggregated data element in the memory based on the number of elements of the set of data elements satisfying the aggregation threshold.
17. The method of claim 14, wherein the set of data is a set of image data, and each data element in the set of data elements is a two-dimensional block of image data.
18. The method of claim 14, wherein the first aggregation threshold is different than the second aggregation threshold.
19. The method of claim 18, wherein the first channel is a chroma channel and the second channel is a luma channel.
20. The method of claim 14, wherein each data element in the set of data elements has a size selected from the group of: 16x16 bytes, 32x32 bytes, and 64x32 bytes.
21. A system comprising: a plurality of schedulers, each scheduler associated with a type adapter; a plurality of hardware accelerators respectively coupled to the plurality of schedulers; a memory coupled to the plurality of hardware accelerators; wherein a first hardware accelerator of the plurality of hardware accelerators is configured to: read data in a first format from the memory; and convert the data from the first format to a second format using the type adapter corresponding to the first hardware accelerator; process the data in the second format to form processed data in the second format; and cause the processed data in the second format to be stored in the memory; wherein a second hardware accelerator of the plurality of hardware accelerators is configured to determine whether a condition is satisfied with respect to the processed data in the second format to determine whether to further process the processed data in the second format.
22. The system of claim 21, wherein, the second hardware accelerator is further configured to, when it is determined that the condition is satisfied with respect to the processed data in the second format: convert the processed data in the second format to data in a third format using the type adapter corresponding to the second hardware accelerator; process the data in the third format to form processed data in the third format; and cause the processed data in the third format to be stored in the memory.
23. The system of claim 21, further comprising: a load store circuit coupled to the first hardware accelerator, the second hardware accelerator, and the memory, wherein the load store circuit is configured to determine whether the condition is satisfied. To determine whether the condition is satisfied, the load store circuit is configured to:
24. The system of claim 23, wherein, determine whether a number of elements in the processed data in the second format satisfies an aggregation threshold; and based on the number of elements satisfying the aggregation threshold, cause the second hardware accelerator to read the processed data in the second format.
25. The system of claim 21, further comprising a direct memory access (DMA) controller coupled to the memory, and a channel mapper coupled to the DMA controller and the plurality of schedulers.
26. The system of claim 21, further comprising a plurality of load store circuits respectively associated with the plurality of hardware accelerators.
27. The system of claim 21, further comprising a memory mapped register configured to store the condition.
28. The system of claim 21, wherein a size of each of the elements in the processed data in the second format is one of 16x16 bytes, 32x32 bytes, and 64x32 bytes.
29. A system comprising: a plurality of dispatchers, each dispatcher associated with a type adapter; a plurality of hardware accelerators respectively coupled to the plurality of dispatchers; a plurality of load store engines respectively associated with the plurality of hardware accelerators; a memory coupled to the plurality of load store engines; and direct memory access (DMA) circuitry coupled to the memory.
30. The system of claim 29, wherein each dispatcher of the plurality of dispatchers comprises at least one consumer socket and at least one producer socket.
31. The system of claim 29, further comprising a memory-mapped register.
32. The system of claim 29, further comprising an interface coupled to the plurality of dispatchers.
33. The system of claim 29, wherein each type adapter is configured to convert data from a block format to a row format, or each type adapter is configured to convert data from a row format to a block format.
34. The system of claim 29, wherein each load store engine is configured to move data from the corresponding hardware accelerator to the memory when an element of the data satisfies an aggregation threshold.