Hardware coalescing unit for optimized execution of memory transactions on an integrated circuit
The hardware coalescing mechanism addresses the underutilization of memory bandwidth in special-purpose processors by combining transaction descriptors, thereby enhancing the efficiency and performance of memory transactions for machine-learning computations.
Patent Information
- Application Number
- PCT/US2024/059773
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-19
AI Technical Summary
Existing designs of IP blocks for special-purpose processors result in underutilization of available memory bandwidth, leading to increased execution latency and degraded performance for memory transactions in machine-learning computations.
A hardware coalescing mechanism is implemented to optimize memory transactions by combining multiple transaction descriptors and saturating the available memory transaction bandwidth, thereby executing memory operations more efficiently.
The hardware coalescing mechanism maximizes memory transaction bandwidth and accelerates the execution of machine-learning workloads, such as image warping operations, by reducing the number of memory transactions and improving memory utilization.
Smart Images

Figure US2024059773_19062025_PF_FP_ABST
Abstract
Description
HARDWARE COALESCING UNIT FOR OPTIMIZED EXECUTION OF MEMORY TRANSACTIONS ON AN INTEGRATED CIRCUITBACKGROUND
[0001] This specification generally relates to memory transactions for machine-learning computations.
[0002] Modem computing systems usually incorporate a wide variety of compute processing units that each offer different computing capabilities and trade-offs. Efficient execution of a given compute job often involves parsing computations into meaningful subtasks or workloads that are mapped to available processors or processor cores of a computing system. The computations may be parsed and mapped based on suitability criteria, such as processor capability7, performance, and power. Generally, this overall process of allocating portions of a compute to appropriate processor resources is referred to as heterogeneous compute.
[0003] At least one processor core of the computing system can be an Intellectual Property block (“IP block”) that executes a respective portion of a computational operation for different multimedia use cases. An example workload can involve using a specialpurpose processor of an IP block to process image data captured by a camera of the mobile device. Processing the image data often requires executing multiple memory transactions to route data to and from memory resources of the IP block. The workload and associated memory7transactions can be issued and / or executed using a host that is external to the memory resources. Existing designs of IP blocks for special-purpose processors result in underutilization of the available memory bandwidth, which increases execution latency and degrades performance.SUMMARY
[0004] This specification describes a hardware coalescing mechanism for optimizing memory transactions that facilitate machine-learning (“ML”) computations performed using a special-purpose integrated circuit of a system-on-chip (“SoC”). The SoC may be integrated in a consumer electronic device, such as a smartphone or tablet with two or more digital cameras for capturing digital images.
[0005] Relative to existing approaches, the hardware coalescing mechanism disclosed in this document can be implemented with certain data communication techniques to accelerate memory operations for particular types of ML workloads. In some cases, a type of MLworkload can be an example warping operation used in computer vision applications. For context, image warping is the process of digitally manipulating an image (e.g., a digital image), for example, such that one or more shapes portrayed in a first image can be modified and / or “morphed” to generate a second image. Example warping operations may be used for correcting image distortion as well as for creative purposes (e.g., morphing).
[0006] Performing these image processing workloads often requires executing multiple memory transactions to route data to and from memory resources of an IP block or specialpurpose integrated circuit. In this context, the disclosed hardware coalescing mechanism is used, and / or configured, to maximize an available memory transaction bandwidth at the IP block and accelerate execution of the particular types of ML workloads, such as the example warping operation described above. The hardware coalescing mechanism can be implemented as a discrete storage & processor device, as a combination of data storage and data processing resources, or both.
[0007] To perform an example workload, a set of memory' operations are executed through a host interface block (“HIB”) of a special-purpose processor or IP block. The host interface block is used to fetch or read image data from a memory' (e.g., dynamic randomaccess memory (DRAM)) coupled to the SoC as yvell as to write an output of the ML workload to the memory'. The host interface block generates read and write transactions that have a particular size, such as a 3B size, 8B size, or 12B size, where “B” is one byte (or 8 bits). The read / write transactions can be Advanced extensible Interface (AXI) transactions generated based on the AXI on-chip communication bus protocol.
[0008] The special-purpose processor and host interface block use the hardyvare coalescing mechanism and associated data processing techniques to optimize execution of memory transactions at the IP block. For example, the host interface block can have a threshold memory bandyvidth that is a multiple of the particular size (e.g., 32B). For example, the threshold memory bandyvidth can be a 4x, 8x, or 1 Ox multiple of the particular transaction size. The hardyvare coalescing mechanism can optimize memory operations of the host interface block by combining multiple transaction descriptors to saturate available memory transaction bandwidth of the host interface block. For example, the hardware coalescing mechanism is used to execute a single, combined read / rite transaction in a data size that satisfies or substantially exhausts the threshold memory' transaction bandyvidth of the host interface block.
[0009] Other implementations of this and other aspects include corresponding systems, apparatus, and computer programs, configured to perform the actions of the methods.encoded on computer storage devices. A system of one or more computers can be so configured by virtue of software, firmware, hardware, or a combination of them installed on the system that in operation causes the system to perform the actions. One or more computer programs can be so configured by virtue of having instructions that, when executed by a data processing apparatus, cause the apparatus to perform the actions.
[0010] The details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Fig. 1 is a block diagram of an example computing system with at least one SoC.
[0012] Fig. 2 shows an example memory transaction pipeline.
[0013] Fig. 3 show s example data structures of a hardw are coalescing unit in a host interface block.
[0014] Fig. 4 shows examples stages of the memory transaction pipeline of Fig. 2.
[0015] Figs. 5A-5E show example uses cases of coalescing memory transaction information using the hardware coalescing unit of Fig. 3.
[0016] Fig. 6 show s an example response buffer of the host interface block of Fig. 1 and Fig. 2.
[0017] Fig. 7 is a first example process for executing memory’ transactions for a ML workload.
[0018] Fig. 8 is a second example process for executing memory transactions for a ML workload.
[0019] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION
[0020] Fig. 1 is a block diagram of an example computing system 100 that includes a system-on-chip 102 (“SoC 102”). The SoC 102 includes a central processing unit 104 (“CPU 104”), a memory controller 105. a shared memory 106 (“memory 106”), a resource manager 108, and an IP / circuit block 110. In some implementations, system 100 can include multiple SoCs and any descriptions for the SoC 102 will apply’ equally to each of the multiple SoCs that may be included at system 100.
[0021] The CPU 104 can be a general-purpose CPU (e.g.. a single or multi-coreCPU). The CPU 104 generates one or more indicators, such as an app-launch indicator or a function call that is triggered in response to executing or launching an application at a user device. For example, the application can be a camera application that uses an imaging sensor to generate image data or a gaming application that requires substantial memory7and graphics processing resources to render graphical content of the game. The CPU 104 also generates one or more application values, such as pixel values or frame rate. The application values may be associated with a function call, may be descriptive of an event that occurs during execution of the application, or both.
[0022] The memory 106 is a system memory, shared memory , or both. In the example of Fig. 1, memory 106 is depicted external to circuit block 110. However, memory 106 can include portions of memory that are: i) specific to circuit block 110, ii) external to circuit block 110, or iii) both. The memory7106 can be random access memory of the SoC 102, such as static random-access memory (SRAM), dynamic random-access memory7(DRAM), a synchronous DRAM (SDRAM), or double data rate (DDR) SDRAM.
[0023] In some implementations, aspects of memory 106 are configured as a shared scratchpad memory that supports parallel access of its memory resources by two or more processors of the circuit 110. The memory7106 can also include various other ty pes of memory, such as high bandwidth memory (HBM), narrow memory (e.g., for storing 8-bit values), wide memory (e.g., for storing 16-bit or 32-bit values), etc.
[0024] The resource manager 1 8 is implemented in hardware, software, or both. Aspects of the resource manager 108 can be also implemented as firmware of the SoC 102 or firmware of a device of the SoC 102, such as a DRAM memory device or the CPU 104. The resource manager 108 is a processor-in-memory (PiM) resource manager (“PiM resource manager 108’’) that includes control logic implemented in hardware, software, or both. For example, the PiM resource manager 108 can include resources such as flip-flops, registers, buffers, etc. that are implemented in hardware and control logic (e.g., programmed code) that is implemented in software.
[0025] The circuit block 110 generally includes individual IP devices such as processors, processor cores, or special-purpose processing devices. For example, the circuit block 1 10 can include an image signal processor (ISP) 112, a tensor processing unit (TPU) 114, a digital signal processor (DSP) 116, and a graphics processing unit (GPU) 118. The circuit block 110 is referred to alternatively as an IP block 110, where the IP block can include one or more proprietary hardware elements. For example, each of the ISP 112, TPU 114, DSP 116, andGPU 118 can be a respective proprietary' IP block (or IP device) of a particular entity or device manufacturer.
[0026] One or more aspects of the PiM resource manager 108 can be implemented as a software routine (or module) of the CPU 104, which uses one or more hardware resources of the CPU 104, such as registers, buffers, etc. The CPU 104 can be configured as an instruction and vector data processing engine that processes data obtained from a system memory of the SoC 102. such as memory 106. In some implementations, each processor (e.g., ISP 1 12, DSP 116, TPU 114, GPU 118) of the SoC 102 includes multiple cores and the CPU 104 and / or the PiM resource manager 108 can generate control signaling to manage and distribute memory intensive compute operations to a memory device 125 (e.g., DRAM) to minimize the processing load at each core of the processors. The control signaling is routed at system 100 using an example bus 122 of the SoC 102. The control signaling can include commands, requests, data, instructions, or a combination of these.
[0027] The PiM resource manager 108 cooperates with the CPU 104 and memory' controller 105 to dynamically control and manage one or more compute-in-memory (CiM) operations. In some implementations, the CIM operations are executed at the SoC 102 in support of a heterogeneous compute operation between two or more processing units that are included among the IP block 110, the CPU 104, or both. More specifically, the PiM resource manager 108 is configured to generate control signaling and use one or more discrete signal values of the control signaling to manage and execute debug operations and access control operations at the memory device 125. The debug and access control operations are executed in support of CiM and PiM operations performed at the memory device 125.
[0028] The system 100 includes an example memory device 125. The memory device 125 can include multiple memory dies. For example, the memory device 125 can include N memory die, where A is an integer greater than 1. The memory device 125 can be a dynamic random-access memory (DRAM) or Double Data Rate (DDR) synchronous DRAM (SDRAM). The memory' device 125 is configured to perform or support various ty pes of PiM operations, CiM operations, and memory-near-computing operations (“MnC operations’7). The memory device 125 performs or supports these operations using its multiple PiM compute elements, which are described below with reference to Figs. 2-4.
[0029] The TPU 114 includes an example host interface block 120 that determines and manages indices associated with memory transactions for requesting data from and providing data to memory resources that are external to the TPU 114. For example, the host interface block 120 is used to fetch or read image data from a memory (e.g., dynamic random-accessmemory' (DRAM)) of the system 100 as well as write an output of the ML workload to the memory. For example, the memory can be a shared memory or shared DRAM, such as memory device 125.
[0030] The SoC 102 cooperates with the memory device 125 to perform computations across one or more memory' die of the memory' device 125. The computations can be for operations or workloads that involve one or more of the processors at IP block 110. Additionally, the computations can be for a heterogenous operation that spans multiple processors of IP block 110, multiple IP blocks 110, or both. In at least one example the memory device 125 may be external to the SoC 102, whereas in another example the memory device 125 may be internal to the SoC 102.
[0031] In the example of Fig. 1, system 100 and the SoC 102 is an integrated circuit of an example user / client device 130, consumer electronic device, or mobile device, where each of these devices can include items such as a smartphone 130a, tablet 130b, or laptop 130c, or even a wearable device or autonomous vehicle. The devices 130 may also include other items such as an eNotebook. Netbook, mobile computer, or any device capable of executing machine-learning models for image processing applications. In some implementations, the system 100 and the SoC 102 are integrated circuits of a desktop computer, network server, or related cloud-based asset.
[0032] Fig. 2 shows an example memory transaction pipeline 200. The memory transaction pipeline 200 includes a host 205, a host controller 210, and a corresponding hardware coalescing unit 230 of the host interface block 120. The host 205 can be an IP block, a special-purpose processor, a hardware ML accelerator, or an application-specific integrated circuit (ASIC) of the SoC 102. In the example of Fig. 2, the host 205 is represented as the TPU 114 and the host controller 210 is represented by an example higher- level controller of the TPU 114.
[0033] The host interface block 120 is configured to send data (e.g., via a write request) to the SoC bus 122, for example, by popping data such as inputs and parameters that are stored among compute tiles of the TPU 114. Relatedly. the host interface block 120 is configured to send requests (e.g., read requests) to the SoC bus 122 to obtain data from the memory device 125, memory 106, or other shared / system memory of the SoC 102. In some implementations, the host controller 210 executes programmed instructions to analyze a data stream associated with the received weights and inputs.
[0034] In this example, inputs and parameters relate to a neural network layer of an artificial neural network implemented at an integrated hardware circuit such as the TPU 114.More specifically, the parameters can be a set of weights for a neural network layer, whereas the inputs can be layer inputs or activations for processing through the neural network layer in accordance with the set of weights for the layer. In some implementations, the host interface block 120 receives data representing inputs / activations and parameters via a receiving port of a local buffer of the host interface block 120.
[0035] The hardware coalescing unit 230 includes various resources for implementing coalescing operations on a group of indices that include {x. y} pixel coordinates of image pixel data. For example, the hardware coalescing unit 230 can include memory resources such as buffers, registers (e.g., control status registers (CSR)), SRAM for capturing / storing sets of indices and maintaining data structures such as tables and arrays used for coalescing operations against the indices.
[0036] The hardware coalescing unit 230 can include control and / or data processing logic for making determinations about which {x, y) pixel indices will be coalesced to generate a combined descriptor, including how many {x, y} pixel indices, and controlling, executing, or managing sorting operations to identify candidate {x, y} pixel indices for generating a combined memory transaction descriptor / request 235. The combined memory transaction 235 is executed at system 100 using the memory device 125 and, for example, the memon controller 105 and / or control signaling generated by the host interface block 120.
[0037] The host 205 obtains the data 240 from the memory device 125. memory 106, or other shared / system memory of the SoC 102. For example, the host 205 obtains the data 240 in response to triggering execution of the combined memory transaction 235 at the system 100. In some implementations, the control logic associated with the hardware coalescing unit 230 resides in the host controller 210, or the host interface block 120, which can generate control and / or data signals to trigger a coalesce operation at the hardware coalescing unit 230.
[0038] The control logic of the hardware coalescing unit 230 is configured to coalesce or combine memon' transaction descriptors based on the threshold (or available) memory transaction bandwidth of the host interface block 120. In some implementations, the hardware coalescing unit 230 is configured to coalesce as many indices as possible without exceeding a threshold memory transaction size, such as 32B or 64B. The hardware coalescing unit 230 can be configured to coalesce as many indices as possible but will stop / pause coalescing if coalescing a next index will result in a memon' transaction size that exceeds a threshold memory transaction size (e.g., 256B).
[0039] The control logic of the hardware coalescing unit 230 is configured to determine whether a row value of a first pixel index and a row value of a second, different pixel indexare from the same row of pixels in an image. If the control logic of the hardware coalescing unit 230 determines that the row value of the first pixel index and the row value of the second, different pixel index are from the same row of pixels in an image, then the hardware coalescing unit 230 can coalesce the two indices. The hardware coalescing unit 230 can determine whether to coalesce the two indices from the same row based on a distance between the two indices along that row. In some implementations, pixels along a given row of an image maybe stored at memory locations / cells of memory array 250, such that a layout of the pixels in the memory array 250 is along contiguous sequence of address locations / cells in a row or bank of the memory array 250.
[0040] The threshold memory transaction size / bandwidth can be defined based on a memory interface width of the host interface block 120. In general, the threshold memory transaction size / bandwidth indicates a maximum size or amount of data that can be transmitted to / from the host 205, via the host interface block 120, in a single clock cycle. For clarity7, the threshold memory7transaction bandwidth can be the same as the threshold memory transaction size, and thus may be referred to alternatively as the threshold memory transaction size.
[0041] The system 100 can be configured to use 3B (or 3 bytes) to store one pixel. In some implementations, for a Red, Green, Blue (RGB) image, the system 100 can be configured to use 3B (or 3 bytes) to store one pixel, where IB is allocated for storing each R, G, B color component of the pixel. The system 100 can also uses 4B to store one pixel. The system 100 can also be configured to use more or fewer bytes to store one pixel, e.g., based on design preference. In some implementations, the image pixel data for a source image is stored in memory7device 125 before being routed for processing by a processor of IP block 110, such as the TPU 114 or GPU 118.
[0042] The pixels of an image can identified be using {x, y} dimensional coordinates. Dimensional coordinates for image pixels are referred to alternatively in this document using formats, such as: {x, y} pixel coordinates; {x, y} pixel index; (x, y); and (y, x). In each example format, “x” denotes a row dimension, whereas “y” denotes a column dimension.
[0043] Fig. 3 shows example data structures of the hardware coalescing unit 230. The data structures include a combined memory transaction tracking table 302, a warp index tracking table 304, a combined memory7transaction identification (ID) list 306.
[0044] As described above, in addition to its control logic, the hardware coalescing unit 230 can include memory resources such as buffers, registers (e.g., control status registers (CSR)), SRAM for capturing / storing sets of indices and maintaining data structures such astables and arrays used for coalescing operations against the indices. Each of the combined memory transaction tracking table 302, warp index tracking table 304, and combined memory transaction identification (ID) list 306 can be instantiated, updated, managed, or otherwise controlled using the control logic and memon resources of the hardware coalescing unit 230. Additionally, each of tables 302, 304 and ID list 306 is described in detail below.
[0045] Fig. 4 shows example stages of the memory transaction pipeline of Fig. 2. More specifically, Fig. 4 shows example steps or stages of an operation 400 that is performed at system 100 to generate combined memory transactions. In some implementations, the operation 400 is performed using the memory transaction pipeline of Fig. 2
[0046] Using the example 3B for storing one pixel, as described above, the host interface block 120 is configured to issue requests that trigger execution of memory transactions to obtain (or fetch) data, e.g., from the memory device 125, for one or more pixels of a source image. Prior approaches for fetching multiple pixels from a source image stored in the memory7device 125, required a host to issue multiple discrete 3B memory' transactions (e.g., one memory fetch transaction of 3B for each pixel). However, these prior approaches would severely underutilize available memory bandwidth of an interface block of the host, which increases execution latency for memory transactions and degrades performance.
[0047] The host interface block 120 can have a threshold or available memory' transaction bandwidth 32B, 64B, or 256B. Rather than underutilize the threshold memory transaction bandwidth, for example, by executing single 3B (or 4B) memory transactions, the system 100 uses the hardware coalescing unit 230 to coalesce or combine multiple requests so as to optimize memory7operations of the host interface block 120. For example, the memory7operations are optimized by combining multiple transaction descriptors to execute read / write transactions using a single combined request, e.g.. 32B request or even 64B or 256B request. Thus, this technique of combining multiple transaction descriptors allows for issuing / generating a single request that satisfies or substantially exhausts the threshold memory' transaction bandwidth of the host interface block 120.
[0048] As described above, the hardware coalescing unit 230 is implemented with certain data communication techniques to accelerate memory operations for particular types of ML workload. In some cases, a type of ML workload can be an example warping operation used in computer vision applications. For context, image warping is the process of digitally manipulating an image (e.g., a digital image), for example, such that one or more shapes portrayed in a first image can be modified and / or "morphed" to generate a second image. Example warping operations may be used for correcting image distortion as well as forcreative purposes (e.g., morphing). In some implementations, a warping operation described herein can represent an example image beautification algorithm.
[0049] In this context, an example “image warp’’ defines a mapping from input / source image to result / destination image, where the mapping is determined from a mapping function. For example, the system 100 can define or determine a mapping of image pixel data from a source image to a destination image based on a mapping function. The system 100 also generates multiple indices based on the mapping function. The multiple indices can be a group of indices where each index that forms the multiple indices is an {x, y} pixel coordinate of the image pixel data that was mapped to the destination image. Stated another way, each index in the group of indices is an {x, y } pixel coordinate of the source image that is mapped to a corresponding {i. j} pixel coordinate of the destination image.
[0050] The mapping function is derived from, or corresponds to, a machine-learning algorithm that is used to identify pixel data for one or more regions of the source image that require modification (e.g., beautification, de-bluring, enhancing) to generate the destination image. For instance, the mapping function can be defined based on the expression: (x’, y’) = fix. y). where (x’, y’) represent pixel coordinates in the destination image and (x. y) represent pixel coordinates in the source image. In some implementations, the mapping function represents a warping field (warp_field) that is defined based on the multiple indices, where the multiple indices are captured in a data structure (e g., a table, list, or array) stored in a shared memory of a special-purpose processor, such as the TPU 114. In some implementations, a dimensionality of the data structure is the same as a dimensionality of the source image.
[0051] In some implementations, the group of indices is stored using a list structure or array in the local shared memory’ of TPU 114 that is read by the host controller 205, the hardware coalescing unit 230, or both (402). For example, the host controller 205 is configured to read the indices from the shared memory of the TPU 114 and generate / create entries in a warp_index_tracking_table 304 for all (or some) of the indices that are read from the shared memory (404).
[0052] For each {x, y} pixel coordinate in the group of indices: the system 100 can determine or identify' a subset of indices as candidate indices for generating a combined descriptor for a memory transaction. In some implementations, to determine or identify the subset of indices, the controller 205 (or the host interface block 120) determines a measure of locality among each {x, y} pixel coordinate in the group of indices based on a sorting operation. As an example, determining the measure of locality can include performing asorting operation using at least one of a respective x-position / index or y-position / index of each {x, y} pixel coordinate of the group of indices.
[0053] For example, the controller 205 can sort warp indices in a group on a y-index of an {x, y] pixel index (406). The controller 205 can then sort warp indices in the group on a x-index of an {x, y} pixel index, within every y-index (408). The host controller 205 can identify neighboring or adjacent pixels along a particular row or column of the source image based on the sorting operation performed on each {x, y} pixel coordinate of the warp indices (i.e., the group of indices). In some implementations, the host controller 205 identifies neighboring pixels of each {x, y} pixel coordinate in the group of indices based on the measure of locality and determines a subset of indices for a given {x, y} pixel coordinate based on the neighboring pixels of that given {x, y} pixel coordinate.
[0054] A property of the warp field is that, in practice, there is a measure of locality in its access patterns. For example, (17, 20), (18, 20), (19, 20) are pixel coordinates of adjacent pixels in row 20 of a source image, where each of these pixel coordinates uses a (y, x) format, such that “20” denotes a row along which each pixel resides. Thus, for a contiguous layout in memory device 125. these pixels would be stored next to each other in a particular region of memory array 250 or, more generally, memory device 125. In some implementations, the warp field / indices are stored in a multi-ported 8-bank memory' with 32B width per bank. The host interface block 120 is configured to read 256B per cycle from this memory or 64 indices ({x, y} pixel indexes) per cycle (256 / 4), assuming each index pair (x, y) takes 4B to store in memory. This multi-ported 8-bank memory represents the shared memory of the TPU 1 14 (or host), described above.
[0055] The system 100 can perform a sorting operation to scan through the sorted index list to generate combined memory transactions (410). Leveraging the hardware coalescing unit 230, the system 100 can combine the smaller discrete 3B or 4B memory accesses into a larger combined memory' access. Combining the memory access in this manner can reduce pow er consumption, improve performance of an integrated circuit that includes hardware coalescing unit 230, and improve the utilization of available memory bandwidth at the host interface block 120.
[0056] Figs. 5A-5E show example uses cases of coalescing memory transaction information using the hardw are coalescing unit of Fig. 3. In the example of Figs. 5A-5E, the system 100 generates combined memory' transactions based on an example sorted index list (20, 14), (20. 17), (20, 18). (20, 30), (21, 19) generated at step 408 of Fig. 4, where the {y, x} represent pixel coordinates of a pixel in an image.
[0057] In the example of Fig. 5 A, the system 100 initiates its memory operations using the first index in the sorted index list (20, 14). The hardware coalescing unit 230 is used to pop a combined memory transaction ID '‘CD1” from the free combined memory transaction identification (ID) list 306. The hardware coalescing unit 230 can also add the “CD1” transaction ID to the warp index tracking table 304. Since this is the first index in the combined memory transaction, the entry under the “offset bytes” column for the first index in the combined memory transaction can be set to “0.” Additionally, the hardware coalescing unit 230 can add entries to the combined memory transaction tracking table 302, for example, by causing the table to include: an entry of 14 under the “start_x” column (x for the index (20, 14)): an entry of 15 under the “end_x” column; and an entry of 20 under the “y” column (y for the index (20, 14)). The hardware coalescing unit 230 sets at least the 14 and 15 entries since the operation involves fetching a 2 * 2 slice of pixel data.
[0058] In the example of Fig. 5B, the hardware coalescing unit 230 can move to the next index (20, 17) and check the combined memoiy transaction tracking table to determine whether this index can be coalesced with an existing larger memory’ transaction. The hardware coalescing unit 230 can get a hit (e.g.. set a flag) for the combined memory transaction ID “CD1” since the y index (20) of the second index (20, 14) matches the y index (20) of the first index (20, 17). The hardware coalescing unit 230 also determines that coalescing these indices will not make the combined memory transaction larger than 32B. Thus, in the warp index tracking table 304, for a row corresponding to (20. 17), the hardware coalescing unit 230 can update the combined memory’ transaction ID to “CD1 .” Also, the entry’ under the” offset bytes” column can be updated to 9, reflecting that pixel data for this index starts at the 9thbyte in the combined memory transaction.
[0059] In addition, the entry under the end x column for “CD1” in the combined memory transaction tracking table can be updated to 18, and entry under the number of outstanding indices column can be incremented to 2 (reflecting that this combined memory’ transaction is fetching pixel data for two warp indices).
[0060] In the example of Fig. 5C, the system 100 moves to the third index (20, 18) and performs similar steps to those described above with reference to the second index (20, 17).
[0061] In the example of Fig. 5D, the hardware coalescing unit 230 moves or transitions to index (20, 30) and performs a check of the combined memory’ transaction tracking table 302. The hardware coalescing unit 230 can determine that the y index (20) of the fourth index (20, 30) matches the y index (20) of the third index (20, 18). However, the hardware coalescing unit 230 also determines that combining this index with the combined memorytransaction “CD1” will result in a memory transaction size of the combined transaction 'G D I " that exceeds a threshold memory transaction size, such as 32B or 64B. Consequently, the hardware coalescing unit 230 can pop or generate a new combined memory transaction ID ' CD2 " for (20, 30) pixel index from the free combined memory transaction identification (ID) list and update the relevant entries in the combined memory7transaction tracking table 302 and the warp index tracking table 304 as similarly described above.
[0062] Since the sorted index list is sorted according to both y and x indices, all indices the system 100 may encounter thereafter will have greater indices (y >= 20 and x>=30). Consequently, the combined memory' transaction “CD1” can be issued to memory' as no more coalescing will be possible for the index (20, 18).
[0063] In the example of Fig. 5E, the hardware coalescing unit 230 transitions to index (21, 19) in the sorted index list and check the combined memory transaction tracking table. Since the y index does not match the y index of any existing combined memory transaction, the hardware coalescing unit 230 can pop (or generate) a new combined memory transaction ID “CD2” from the free combined memory transaction identification (ID) list and update the relevant entries in the combined memory transaction tracking table 302 and the warp index tracking table 304 as similarly described above.
[0064] This process described above can be repeated for every' set of indices that are read from the shared memory. The combined memory transactions will be issued to memory device 125 once the size becomes 32B or the hardware coalescing unit 230 determines that it is no longer possible to coalesce (e.g., in the case of CD1 in the example above). In some implementations, the hardware coalescing unit 230 issues the first combined memory' transaction in the combined memory transaction tracking table 302 (that has not yet been issued) in the cycle where no other combined memory7transaction was issued.
[0065] Fig. 6 shows an example response buffer 602 of the host interface block 120. When each combined memory7transaction is issued, the host interface block 120 can coordinate with the host controller 205 to reserve a corresponding space / entry in the response buffer 602 to receive the response data obtained from the memory7device 125 following execution of the combined transaction at the system 100.
[0066] In some implementations, the response buffer 602 is implemented as a multiported register file. The host interface block 120 uses the “offset bytes" and “combined memory transaction ID” information from the warp index tracking table 304 to provide response data for each index through the multi-port design of the register file. For example,the multi-port design is configured to enable multiple reads from the response buffer 602 and multiple writes to the warp ID tracking table 304 in a single cycle.
[0067] In some implementations, once the response data has been written to warp index tracking table, the response data can be forwarded to the CPU 104, the circuit block 110, or both in the system 100 by popping entries from response in the same order. In particular, the response data can be provided to the processing unit in the same order that it was requested, since entries in warp index tracking table are in the requested order and they can be popped in order.
[0068] In some implementations, once an entry is popped from the warp index tracking table, the entry under the “number of outstanding indices'’ for the corresponding combined memory transaction ID can be decremented in the combined mem memory transaction table. For example, as the system 100 pops entries corresponding to indices (20, 18) and (21, 19), the entries under the number of outstanding indices for combined memory transaction IDs CD1 and CD3 can be decremented to 2 and 0, respectively.
[0069] In some implementations, once the “number of outstanding indices” for a combined memory transaction reaches 0, the corresponding “combined memory transaction Id” can be reclaimed and put back into the free combined memory transaction ID list. Continuing with the example above, the combined memory transaction ID CD3 can be added back to the free combined memory transaction ID list.
[0070] Fig. 7 is an example process 700 for executing memory transactions for an ML workload. Process 700 can be implemented or executed using system 100 and the hardware coalescing unit 130 described above. Hence, descriptions of process 700 may reference the above-mentioned computing resources of system 100. In some examples, the steps or actions of process 700 are enabled by programmed software instructions, firmware instructions, or both. Each type of instruction may be stored in a non-transitory machine-readable storage device and is executable by one or more of the processors or other resources described in this document.
[0071] In some implementations, the steps of process 700 are performed at a hardware integrated circuit to generate a machine-learning (ML) output, including an output for a neural network layer of a neural network that implements one or more ML models. For example, the output can be a portion of a computation for a ML task or inference workload to generate an image processing or image recognition output. As indicated above, a portion of the integrated circuit can include a special-purpose integrated circuit, a neural networkprocessor, or hardware ML accelerator configured to accelerate computations for generating different types of data processing outputs.
[0072] Referring again to process 700, the system 100 determines a mapping of image pixel data from a source image to a destination image based on a mapping function (702). For example, the mapping function is based on a machine-learning algorithm. The system 100 generates, based on the mapping function, multiple indices that include {x, y} pixel coordinates of the image pixel data that are mapped to the destination image (704).
[0073] For each {x, y} pixel coordinate among the multiple indices, the system 100 determines a subset of indices for generating a combined descriptor for a memory transaction (706). The system 100 performs a coalescing operation using the hardware coalescing unit that coalesces memory transaction information for obtaining pixel data for each {x, y} pixel coordinate in a first subset of indices (708). The system 100 generates a combined descriptor for a single memory transaction that combines respective transactions for obtaining pixel data for each {x, y} pixel coordinate in the first subset of indices (710).
[0074] Fig. 8 is a flow diagram representing another example process 800 for executing memory transactions using resources of the SoC of Fig. 1. Much like process 700. process 800 can be implemented or executed using system 100 and the hardware coalescing unit 130 described above.
[0075] The steps or actions of process 800 are performed using programmed software instructions, firmware instructions, or both. One or more of the instructions may be stored in a non-transitory machine-readable storage device and are executable by one or more of the processors or other resources described in this document. In some examples, process 800 provides a method for accessing memory' banks of a memory device, such as a DRAM or other shared memory. The memory device can have a memory array that includes memorycells arranged along a first dimension (e.g., a row dimension) and a second, different dimension (e.g., a column dimension).
[0076] Referring again to process 800, the system 100 determines, obtains, or otherwise receives multiple memory addresses that each identify a respective memory- cell in the memory array (802). For example, the host interface block can receive a request from a host device of the SoC 102. In some implementations, the host device is a hardware accelerator or special-purpose processor of the IP block 110, such as TPU 114 or GPU 118. The system 100 generates a sorted list of memory addresses by sorting the multiple memory addresses in an order of the respective memory cells along the first dimension (804).
[0077] The system 100 generates one or more combined memory transactions based on the sorted list of memory addresses (806). Each combined memory transaction accesses memory cells identified respectively by two or more of the multiple memory addresses. More specifically, each of the combined memory transactions is executed at system 100 to access pixel data stored in a memory location / cell that is identified by a memory address among the sorted list of memory addresses.
[0078] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory program carrier for execution by, or to control the operation of, data processing apparatus.
[0079] Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0080] The term '‘computing system” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0081] A computer program (which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0082] A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0083] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as. special purpose logic circuitry, e.g., an FPGA (field programmable gate array), an ASIC (application specific integrated circuit), or a GPGPU (General purpose graphics processing unit).
[0084] Computers suitable for the execution of a computer program include, by way of example, can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. Some elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0085] Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory', media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0086] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., LCD (liquid crystal display) monitor, for displaying information to the user and akeyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.
[0087] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (‘'LAN”) and a wide area network (“WAN”), e.g., the Internet.
[0088] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0089] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0090] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0091] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.
[0092] Aspects of the present disclosure may be as set out in the following numbered clauses:Clause 1. A computer-implemented method performed using a hardware coalescing unit of an integrated circuit; determining, based on a mapping function, a mapping of image pixel data from a source image to a destination image; generating, based on the mapping function, a plurality7of indices comprising {x, y } pixel coordinates of the image pixel data that are mapped to the destination image; for each {x, y} pixel coordinate among the plurality of indices: determining a subset of indices for generating a combined descriptor for a memory transaction; performing, using the hardware coalescing unit, a coalescing operation that coalesces memory transaction information for obtaining pixel data for each {x, y} pixel coordinate in a first subset of indices; and generating a combined descriptor for a single memory transaction that combines respective transactions for obtaining pixel data for each {x, y} pixel coordinate in the first subset of indices coalesced by the hardware coalescing unit.Clause 2. The method of clause 1, wherein determining a subset of indices comprises:determining a measure of locality' among each {x, y} pixel coordinate of the plurality of indices based on sorting operation; identifying neighboring pixels of each {x, y} pixel coordinate in the plurality of indices based on the measure of locality; and determining a subset of indices for a given {x, y} pixel coordinate based on the neighboring pixels of that given {x, y} pixel coordinate.Clause 3. The method of clause 2, wherein determining the measure of locality comprises: performing a sorting operation using at least one of a respective x-position or y- position of each {x, y} pixel coordinate of the plurality of indices; and identifying neighboring or adjacent pixels along a particular row or column of the source image based on the sorting operation performed on each {x, y} pixel coordinate of the plurality of indices.Clause 4. The method of any one of clauses 1-3, wherein each index in the plurality of indices is an {x, y} pixel coordinate of the source image that is mapped to a corresponding {i, j} pixel coordinate of the destination image.Clause 5. The method of any one of clauses 1-4, wherein the mapping function is derived from a machine-learning algorithm that is used to identify pixel data for one or more regions of the source image that require modification to generate the destination image.Clause 6. The method of any one of clauses 1-5, wherein determining the mapping comprises: determining that an {x, y} pixel coordinate in the source image corresponds to an {i, j } pixel coordinate in the destination image.Clause 7. The method of any one of clauses 1-6, wherein the mapping function is defined based on the expression: (x’. y?) = f(x. y). where (x’, y’) represent pixel coordinates in the destination image and (x, y) represent pixel coordinates in the source image.Clause 8. The method of any one of clauses 1-7, wherein: the mapping function represents a warping field defined based on the plurality of indices; andthe plurality of indices are captured in a data structure stored in a shared memory of a special-purpose processor.Clause 9. The method of clause 8, wherein a dimensionality of the data structure that captures the plurality of indices is the same as a dimensionality of the source image.Clause 10. The method of any one of clauses 1-9. wherein the destination image represents a morphed or modified version of the source image.Clause 11. A system comprising: a hardware coalescing unit, a processing device, and a non-transitory machine- readable storage device for storing instructions that are executable by the processing device to cause performance of operations comprising: determining, based on a mapping function, a mapping of image pixel data from a source image to a destination image: generating, based on the mapping function, a plurality of indices comprising {x, y} pixel coordinates of the image pixel data that are mapped to the destination image; for each {x, y} pixel coordinate among the plurality7of indices: determining a subset of indices for generating a combined descriptor for a memory transaction; performing, using the hardware coalescing unit, a coalescing operation that coalesces memory transaction information for obtaining pixel data for each {x, y} pixel coordinate in a first subset of indices; and generating a combined descriptor for a single memory transaction that combines respective transactions for obtaining pixel data for each {x, y} pixel coordinate in the first subset of indices coalesced by the hardware coalescing unit.Clause 12. The system of clause 11, wherein determining a subset of indices comprises: determining a measure of locality among each {x, y} pixel coordinate of the plurality of indices based on sorting operation; identifying neighboring pixels of each {x, y} pixel coordinate in the plurality of indices based on the measure of locality; and determining a subset of indices for a given {x, y} pixel coordinate based on the neighboring pixels of that given {x, y} pixel coordinate.Clause 13. The system of clause 12, wherein determining the measure of locality comprises: performing a sorting operation using at least one of a respective x-position or y- position of each {x, y} pixel coordinate of the plurality of indices; and identifying neighboring or adjacent pixels along a particular row or column of the source image based on the sorting operation performed on each {x, y} pixel coordinate of the plurality of indices.Clause 14. The system of any one of clauses 11-13, wherein each index in the plurality of indices is an {x, y} pixel coordinate of the source image that is mapped to a corresponding {i, j } pixel coordinate of the destination image.Clause 15. The system of any one of clauses 11-14, wherein the mapping function is derived from a machine-learning algorithm that is used to identify' pixel data for one or more regions of the source image that require modification to generate the destination image.Clause 16. The system of any one of clauses 11-14. wherein determining the mapping comprises: determining that an {x, y} pixel coordinate in the source image corresponds to an {i, j} pixel coordinate in the destination image.Clause 17. The system of any one of clauses 1 1 -16, wherein the mapping function is defined based on the expression: (x’, y’) = f(x, y), where (x’, y ’) represent pixel coordinates in the destination image and (x, y) represent pixel coordinates in the source image.Clause 18. The system of any one of clauses 1 1-17, wherein: the mapping function represents a warping field defined based on the plurality of indices; and the plurality of indices are captured in a data structure stored in a shared memory of a special-purpose processor.Clause 19. The system of clause 18, wherein: a dimensionality of the data structure that captures the plurality of indices is the same as a dimensionality of the source image; andthe destination image represents a morphed or modified version of the source image.Clause 20. A non-transitory machine-readable storage device storing instructions that are executable by a processing device to cause performance of operations involving a hardware coalescing unit, the operations comprising: determining, based on a mapping function, a mapping of image pixel data from a source image to a destination image; generating, based on the mapping function, a plurality of indices comprising {x, y} pixel coordinates of the image pixel data that are mapped to the destination image; for each {x, y} pixel coordinate among the plurality of indices: determining a subset of indices for generating a combined descriptor for a memory transaction; performing, using the hardware coalescing unit, a coalescing operation that coalesces memory transaction information for obtaining pixel data for each {x, y} pixel coordinate in a first subset of indices; and generating a combined descriptor for a single memory transaction that combines respective transactions for obtaining pixel data for each {x, y} pixel coordinate in the first subset of indices coalesced by the hardware coalescing unit.
Claims
What is claimed is:
1. A computer-implemented method performed using a hardware coalescing unit of an integrated circuit; determining, based on a mapping function, a mapping of image pixel data from a source image to a destination image; generating, based on the mapping function, a plurality of indices comprising {x, y} pixel coordinates of the image pixel data that are mapped to the destination image; for each {x, y} pixel coordinate among the plurality of indices: determining a subset of indices for generating a combined descriptor for a memory transaction; performing, using the hardware coalescing unit, a coalescing operation that coalesces memory transaction information for obtaining pixel data for each {x, y} pixel coordinate in a first subset of indices; and generating a combined descriptor for a single memory transaction that combines respective transactions for obtaining pixel data for each {x, y} pixel coordinate in the first subset of indices coalesced by the hardware coalescing unit.
2. The method of claim 1 , wherein determining a subset of indices comprises: determining a measure of locality among each {x, y} pixel coordinate of the plurality of indices based on sorting operation; identifying neighboring pixels of each {x, y} pixel coordinate in the plurality of indices based on the measure of locality; and determining a subset of indices for a given {x, y} pixel coordinate based on the neighboring pixels of that given {x, y} pixel coordinate.
3. The method of claim 2, wherein determining the measure of locality comprises: performing a sorting operation using at least one of a respective x-position or y- position of each {x, y} pixel coordinate of the plurality of indices; and identifying neighboring or adjacent pixels along a particular row or column of the source image based on the sorting operation performed on each {x. y} pixel coordinate of the plurality of indices.
4. The method of any one of claims 1-3, wherein each index in the plurality of indices is an {x, y} pixel coordinate of the source image that is mapped to a corresponding {i, j} pixel coordinate of the destination image.
5. The method of any one of claims 1-4, wherein the mapping function is derived from a machine-learning algorithm that is used to identify pixel data for one or more regions of the source image that require modification to generate the destination image.
6. The method of any one of claims 1-5, wherein determining the mapping comprises: determining that an {x, y} pixel coordinate in the source image corresponds to an {i, j } pixel coordinate in the destination image.
7. The method of any one of claims 1-6, wherein the mapping function is defined based on the expression: (x’, y’) = fix, y), where (x’, y’) represent pixel coordinates in the destination image and (x, y) represent pixel coordinates in the source image.
8. The method of any one of claims 1-7, wherein: the mapping function represents a warping field defined based on the plurality of indices; and the plurality of indices are captured in a data structure stored in a shared memory of a special-purpose processor.
9. The method of claim 8, wherein a dimensionality of the data structure that captures the plurality of indices is the same as a dimensionality of the source image.
10. The method of any one of claims 1-9, wherein the destination image represents a morphed or modified version of the source image.
11. A system comprising: a hardware coalescing unit, a processing device, and a non-transitory machine- readable storage device for storing instructions that are executable by the processing device to cause performance of operations comprising: determining, based on a mapping function, a mapping of image pixel data from a source image to a destination image;generating, based on the mapping function, a plurality of indices comprising {x, y} pixel coordinates of the image pixel data that are mapped to the destination image; for each {x, y} pixel coordinate among the plurality of indices: determining a subset of indices for generating a combined descriptor for a memory transaction; performing, using the hardware coalescing unit, a coalescing operation that coalesces memory transaction information for obtaining pixel data for each {x. y} pixel coordinate in a first subset of indices; and generating a combined descriptor for a single memoi ' transaction that combines respective transactions for obtaining pixel data for each {x, y} pixel coordinate in the first subset of indices coalesced by the hardware coalescing unit.
12. The system of claim 11, wherein determining a subset of indices comprises: determining a measure of locality among each {x, y} pixel coordinate of the plurality of indices based on sorting operation; identifying neighboring pixels of each {x. y} pixel coordinate in the plurality of indices based on the measure of locality; and determining a subset of indices for a given {x, y} pixel coordinate based on the neighboring pixels of that given {x, y} pixel coordinate.
13. The system of claim 12, wherein determining the measure of locality comprises: performing a sorting operation using at least one of a respective x-position or y- position of each {x, y} pixel coordinate of the plurality of indices; and identifying neighboring or adjacent pixels along a particular row or column of the source image based on the sorting operation performed on each {x, y} pixel coordinate of the plurality of indices.
14. The system of any one of claims 11-13, wherein each index in the plurality of indices is an {x, y} pixel coordinate of the source image that is mapped to a corresponding {i, j} pixel coordinate of the destination image.
15. The system of any one of claims 11-14, wherein the mapping function is derived from a machine-learning algorithm that is used to identify pixel data for one or more regions of the source image that require modification to generate the destination image.
16. The system of any one of claims 11-15, wherein determining the mapping comprises: determining that an {x. y} pixel coordinate in the source image corresponds to an {i, j} pixel coordinate in the destination image.
17. The system of any one of claims 11-16, wherein the mapping function is defined based on the expression: (x’. y’) = fix. y), where (x‘, y') represent pixel coordinates in the destination image and (x, y) represent pixel coordinates in the source image.
18. The system of any one of claims 11-17, wherein: the mapping function represents a warping field defined based on the plurality' of indices; and the plurality of indices are captured in a data structure stored in a shared memory of a special-purpose processor.
19. The system of claim 18, wherein: a dimensionality of the data structure that captures the plurality of indices is the same as a dimensionality of the source image; and the destination image represents a morphed or modified version of the source image.
20. A non-transitory machine-readable storage device storing instructions that are executable by a processing device to cause performance of operations involving a hardware coalescing unit, the operations comprising: determining, based on a mapping function, a mapping of image pixel data from a source image to a destination image: generating, based on the mapping function, a plurality of indices comprising {x, y} pixel coordinates of the image pixel data that are mapped to the destination image; for each {x, y} pixel coordinate among the plurality7of indices: determining a subset of indices for generating a combined descriptor for a memory transaction; performing, using the hardware coalescing unit, a coalescing operation that coalesces memory transaction information for obtaining pixel data for each {x, y} pixel coordinate in a first subset of indices; andgenerating a combined descriptor for a single memory transaction that combines respective transactions for obtaining pixel data for each {x, y} pixel coordinate in the first subset of indices coalesced by the hardware coalescing unit.