Programmable Accelerator for Irregular Operations with Data Dependence
The programmable accelerator addresses inefficiencies in existing accelerators by using a tile-based processor with XPUs and scatter-gather engines to efficiently handle data-dependent and memory-bound operations, improving performance and energy efficiency in machine learning pipelines.
Patent Information
- Application Number
- JP2023570416
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-11-07
- Filing Date
- 2022-11-09
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2042-11-09
AI Technical Summary
Existing accelerators are highly specialized and inefficient for data-dependent, irregular, and memory-bound operations due to unpredictable computational loads and memory access patterns, leading to performance bottlenecks and limited scalability in machine learning pipelines.
A programmable accelerator with a tile-based processor and cross-lane processing units (XPUs) that can dynamically handle data-dependent and memory-bound operations, utilizing scatter-gather engines and cooperative prefetching to manage irregular memory access and computational demands.
The accelerator enhances performance and energy efficiency by reducing instruction fetch bandwidth and handling various types of operations without physical redesign, enabling scalable training of large machine learning models with unpredictable complexity.
Smart Images

Figure 0007702502000001 
Figure 0007702502000002 
Figure 0007702502000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is a continuation of U.S. patent application Ser. No. 17 / 981,617, filed Nov. 7, 2022, which claims the benefit of the filing dates of U.S. Provisional Patent Application No. 63 / 357,281, filed Jun. 30, 2022; U.S. Provisional Patent Application No. 63 / 322,285, filed Mar. 22, 2022; U.S. Provisional Patent Application No. 63 / 281,960, filed Nov. 22, 2021; and U.S. Provisional Patent Application No. 63 / 279,262, filed Nov. 15, 2021, the disclosures of which are incorporated herein by reference. This application also relates to U.S. patent application Ser. No. 17 / 972,681, filed Oct. 25, 2022; U.S. patent application Ser. No. 17 / 972,663, filed Oct. 25, 2022; and U.S. patent application Ser. No. 17 / 722,782, filed Apr. 18, 2022, the disclosures of which are incorporated herein by reference.
Background Art
[0002] Background Hardware acceleration uses computer hardware to more efficiently execute certain types of operations. Exemplary types of operations that can be accelerated include linear algebra operations, such as matrix - to - matrix multiplication or matrix - to - vector multiplication. A device or processor constructed to execute the operations that hardware accelerates is sometimes referred to as an accelerator.
[0003] An accelerator is designed and fabricated to accelerate only a small portion of the desired operations. During the design and fabrication process of an accelerator, assumptions are made regarding the nature of the operations to be accelerated, such as the size and type of the inputs the accelerator receives, the regularity with which the accelerator receives the inputs, or the computational requirements for executing the operations. As a result, an accelerator often becomes highly specialized and may only be able to accelerate a small class of predetermined operations and, if at all, may not be able to efficiently execute other operations.
[0004] Operations outside this class include data-dependent operations for which the computational load on the accelerator cannot be determined before the operation is executed. Multiple instances of this type of accelerated operation can vary depending on various factors, and in a given accelerator design, it may be inefficient to accelerate at least some of these instances. Other types of operations that are difficult to accelerate include memory-bound operations with low computational intensity and limited data reuse. Yet another type of operation that is difficult to accelerate includes irregular operations that can be characterized by random memory access, complex code patterns, and diverse uses of parallel execution of multiple sub-operations performed simultaneously.
[0005] In practice, processing pipelines, such as those for training or deploying a machine learning model, require performing a wide variety of types of operations. Incorporating an accelerator into the pipeline to accelerate only some types of operations and relying on the device to execute other types of operations without hardware acceleration imposes unacceptable delays and memory bandwidth stress on the links and interconnections between the accelerator and non-accelerator, resulting in limited overall performance. Designing and fabricating an accelerator to cover all types of operations is almost always impossible or infeasible. Data-dependent operations do not contribute to acceleration, and the logistical effort to accelerate other types of operations may not be worth the investment in designing, fabricating, and deploying the corresponding accelerator. SUMMARY OF THE INVENTION
[0006] Summary Aspects of the present disclosure provide an accelerator capable of accelerating data-dependent operations, irregular operations, and / or memory-bound operations. The accelerators described herein include a programmable engine for efficiently performing on-chip calculations that are dynamic, irregular, and / or memory-bound, along with a coprocessor configured to accelerate predictable operations with respect to the computational load and behavior on the coprocessor during design and fabrication.
[0007] Dynamic operations are operations in which the calculations to be performed are data-dependent or input-dependent, meaning that the inputs are not known prior to executing the operation. The irregularity of an operation can be due to random memory accesses, complex code patterns, and the various amounts of computational resources and parallelism required to execute the various instances of the operation for various input data. Memory-bound operations are often operations with low arithmetic intensity, e.g., few operations are performed per unit of data transferred during acceleration of the operation, and data reuse is limited.
[0008] The accelerators described herein can coordinate and distribute cross-chip data scatter and gather operations for various sizes of data to scale acceleration on a host device or data center implementing multiple accelerators and other processors. As described herein, the accelerators utilize the building blocks of an architecture that can be configured to accommodate the acceleration of various types of data-dependent operations, irregular operations, and / or memory-bound operations without requiring physical redesign or modification to the hardware circuitry implementing the accelerator itself.
[0009] Aspects of the present disclosure provide an accelerator that can accelerate the computation of neural network layers that exhibit sparsity, for example, in an embedded format. Sparsity calculations refer to calculations in which the fractional part of the computed data value (e.g., input value, output value, or intermediate value) is zero. The fractional part can vary, for example, between 0.1% and 50%. Aspects of the present disclosure provide acceleration for embedded training and processing as part of a machine learning processing pipeline.
[0010] Aspects of the present disclosure provide a processor. The processor includes a plurality of tiles, each of the plurality of tiles including a vector core and a slice of shared software-controlled scratchpad memory. The processor further includes a scalar core configured to dispatch tasks to the plurality of tiles. The processor also includes a memory coupled to the plurality of tiles and the scalar core.
[0011] In one example, each tile is configured to perform independent computations. In another example, the vector core in each of the plurality of tiles includes a plurality of single instruction, multiple data (SIMD) processing lanes. In yet another example, multiple tiles of the plurality of tiles issue memory requests in parallel to main memory.
[0012] In yet another example, the vector core in each of the plurality of tiles is configured to generate a data-dependent address stream to any level of the memory hierarchy. In yet another example, each data-dependent address stream corresponds to a sequence of addresses, and the length and specific values of the addresses in the sequence are data-dependent and are only known at runtime. In yet another example, the vector core in each of the plurality of tiles is configured to represent the data-dependent address stream while decoupling the high-performance service of the data-dependent address stream to the microarchitecture. In yet another example, the microarchitecture includes a scatter-gather engine for the high-performance service of the data-dependent address stream. In yet another example, the data-dependent address stream includes multiple addressing modes, a runtime-configurable transfer size, and indirect memory access with atomic arithmetic updates.
[0013] In yet another example, the vector core in each of the plurality of tiles includes a circular buffer instruction that enables the transfer and access of a dynamically sized data stream over a statically sized region of memory. In yet another example, the processor further includes a microarchitecture configured to track the runtime buffer size of the dynamically sized data stream. In yet another example, the vector core in each of the plurality of tiles is configured to provide a runtime configuration and access area of the tile-local scratchpad memory as a sequential circular first-in-first-out (FIFO) access without excluding unordered access to the same area of the tile-local scratchpad memory. In yet another example, the sequential circular FIFO access enables a dynamically sized data-dependent address stream over the statically sized region of the tile-local scratchpad memory in association with the microarchitecture.
[0014] In yet another example, each tile includes a scatter-gather engine configured to manage the issuance, fetch, tracking, and ordering of data streams. In yet another example, each scatter-gather engine is further configured to maintain at least 256 outstanding read requests in-flight per tile. In yet another example, each scatter-gather engine is further configured to track and update buffer occupancy to manage flow control.
[0015] In yet another example, a subset of the plurality of tiles each further includes a prefetcher unit configured to cooperatively prefetch data stream instructions. In yet another example, the processor further includes a cross-lane processing unit configured to accelerate at least one of an irregular control flow sequence or an intra-vector dependent operation. In yet another example, each tile is configured to support scatter from off-chip memory to its scratchpad memory and gather from its scratchpad memory to off-chip memory.
[0016] In yet another example, a subset of the plurality of tiles are grouped based on a logically configurable vector width. In yet another example, the logically configurable vector width includes a logical SIMD width.
[0017] In yet another example, the processor is part of a machine learning accelerator configured to execute neural network layers exhibiting semantic sparsity. In yet another example, the neural network layer includes an embedded neural network or a graph neural network. In yet another example, the processor is connected to several other processors via a network configured to perform distributed scatter-gather and computation required by neural network layer computations that are dynamic, irregular, and memory-bound. BRIEF DESCRIPTION OF THE DRAWINGS
[0018]
Figure 1A
Figure 1B
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Mode for Carrying Out the Invention
[0019] Detailed Description Overview Aspects of the present disclosure provide an accelerator capable of accelerating data-dependent operations, irregular operations, and / or memory-bound operations. An accelerator configured to accelerate data-dependent operations on sparse inputs may be referred to herein as a sparse accelerator. The sparse accelerator can be configured to efficiently and performantly execute neural network layers exhibiting semantic sparsity, such as embedded neural networks or graph neural networks (GNNs). A graph neural network may correspond to a neural network for processing data that can be represented as a graph.
[0020] The sparse accelerator can be organized as a tile-based processor. The sparse accelerator may include a tile sequencer used to dispatch tasks to tiles. Each tile of the sparse accelerator has a vector core extended with a cross-lane processing unit (XPU) and a slice of a shared software-controlled scratchpad memory. Both the tile and the sequencer can be connected to a high-bandwidth main memory.
[0021] The organization and structure of the sparse accelerators improve data-dependent operations, irregular operations, and / or memory-bound operations through various combinations of the distinct features described herein. The tiles of the accelerators may include one or more of the following features to more efficiently utilize valid computations while data-dependent operations, irregular operations, and / or memory-bound operations are being performed. The cross-lane processing units (XPUs) within each tile accelerate a common irregular control flow sequence and / or intra-vector dependent operations. Each XPU enables the provision of a custom high-speed data path for common intra-vector dependent operations.
[0022] Other features implementable on the accelerators described herein may include cooperative prefetching and stream instructions and ordering. Cooperative prefetching enables performance and energy efficiency improvements by reducing the instruction fetch bandwidth requirements between tiles. The stream instructions described herein enable high-bandwidth data-dependent scatter-gather. Each vector processing unit of the accelerator may generate data-dependent addresses to any level of the memory hierarchy at high bandwidth.
[0023] Tile-local scatter and gather may enable computation without interruption in the presence of data irregularities. Each accelerator tile can substantially support high-performance scatter and gather to its local memory, exposed by the instruction set architecture (ISA), along with available instructions for indirect vector loads, stores, and store adds.
[0024] The software-controlled tile grouping of the implemented XPU can provide a flexible amortization of control overhead. The software can flexibly group tiles to exhibit a logically configurable vector width, e.g., a logical SIMD width. This can help amortize control overhead, such as eliminating redundant instruction fetches between tiles within the same group or enabling coalesced memory accesses between tiles. As a result, better bandwidth efficiency can be achieved. For example, high-bandwidth memory (HBM) can have different bandwidth efficiencies when the access granularity is different. Data-dependent operations (also referred to as "input-dependent operations") are operations whose amount of computational work for executing the operation is not known in advance and depends on the nature of the data. Data-dependent operations can be represented as data-dependent address streams, which correspond to sequences of addresses where the length of the addresses and specific values within the sequence are known at runtime. Computational work can be measured, for example, by the number of operations or processing cycles required to execute a data-dependent operation. Exemplary data-dependent operations include operations for vector sorting, operations for counting duplicate values within a vector, and operations for processing the shape or size of vectors of various lengths. Data-dependent operations are irregular because, at least, there are differences in the random memory access patterns for executing the same type of operation on different inputs. As a result, data-dependent operations are difficult to optimize as opposed to other types of operations where the computational work does not change based on the nature of the input data, such as its shape or degree or sparsity.
[0025] Data-dependent operations include operations performed on sparse data. The sparsity of a data structure is a measure of the ratio of non-empty elements to empty elements. Depending on the data structure, an empty element may be zero, a reserved word indicating that there is no value for the element, or a value that is so small that it is considered to contribute little to the operations performed on the data structure as input. If there are more empty elements than non-empty elements, the data structure is sparse. Some data structures can be more or less sparse than other data structures.
[0026] Exemplary data-dependent operations include generating embeddings for input training examples. The embedding can be a vector or some other data structure mapped from an input having a higher dimension than the embedding. Embedding generation can be performed as part of a workload processed according to a pipeline.
[0027] The XPU of each tile of the sparse accelerator may also perform other data-dependent operations such as vector scatter or gather operations, segment sums, etc., and / or partition sparse data structures such as tensors. The XPU described herein can be a complementary processing unit to other components of a processor or connected components, such as a vector processing unit constructed according to the SIMD parallel processing paradigm. One or more XPUs can be connected in each processor core of a larger processor that may itself include other components for accelerating the performance of a particular workload such as training a neural network.
[0028] Furthermore, the XPU is not limited to executing a specific type of data-dependent operation, and thus, a processor can be designed to include an XPU to complement other types of processing units for multiple different pipelines. Since the XPU can be configured for each workload, the physical footprint of the XPU is reduced compared to other approaches where special circuitry is physically fabricated on the processor as a complementary unit for sparse data calculations. The functionality of the XPU can also be extended by using an instruction set or by extending to an existing instruction set of the host processor, which can further improve the adaptability of various data-dependent operations in response to changes in pipeline data reception. Instructions can be provided as signals to components of the XPU that serve to translate the instructions to configure the individual processing cells and crossbars of the XPU. The XPU can be configured using a program compiled by a corresponding compiler for the hardware circuit implementing the XPU.
[0029] Aspects of the present disclosure provide at least the following technical advantages. The accelerators described herein provide scalable distributed training of large machine learning models that require performing data-dependent operations where the operands are generally not known until runtime.
[0030] The accelerators enable a flexible arithmetic configuration of tiles implementing the XPU within the accelerators to achieve various forms of parallelism, such as task, data, pipeline, and model parallelism, alone and in combination with other processors. In task-level parallelization, each tile may execute independent computations (tasks) in parallel. For memory-level parallelization, the tiles may issue memory requests in parallel to high bandwidth to absorb the available memory bandwidth. This enables the sparse accelerators to handle various amounts of work in separate tasks.
[0031] The accelerators described herein can be dedicated coprocessors alongside other accelerators or general-purpose processors not configured to accelerate data-dependent operations. The accelerators described herein can achieve performance gains over approaches that rely on host memory to fetch data for performing data-dependent operations, such as accelerating sparse or other data-dependent computations, e.g., when increasing processing speed.
[0032] The accelerators described herein can function as a class of architecture relative to coprocessors configured to perform high-density regular computations. Implementing the programmable sparse accelerators described herein communicatively coupled to a separate processor can provide a flexible means for accelerating computations with compile-time unpredictable complexity. The accelerators described herein can target memory-bound operations involving complex data movement, rearrangement, and summarization. This type of operation can include scatter-gather operations, filtering, sorting, uniquing, etc. The accelerators can be targetable according to a compiler stack, e.g., according to a compiler configured to convert source code into instructions executable by the accelerator and its coprocessor.
[0033] The accelerators described in this specification can accelerate embedding generation. Embedding generation can generally be characterized as involving irregular memory accesses with low arithmetic intensity. This is due to the need to perform on-chip table lookups of embedding maps and tables representing variable-width vector calculations, at least based on the sparse and irregular nature of the input vectors for embedding generation. Embedding is a mapping from discrete objects, e.g., other data structures such as vectors or values, to a vector of numerical values such as real numbers. Embeddings generally have lower dimensions and complexity than their corresponding pre-embedding objects. For example, the embedding of one or more words in English can be a vector of real numbers. Embeddings can be used to represent and identify potentially significant features of pre-embedded inputs. For example, these embeddings can be compared by measuring the distance between embeddings of different words to quantify the similarity between two pre-embedded inputs.
[0034] The cooperative prefetching described in this specification enables performance and energy efficiency improvements by reducing the instruction fetch bandwidth requirements between tiles. The tiled architecture of the accelerator additionally provides a configurable subset of tiles that operate in multiple data models for a single program, while exposing a multiple data programming model for multiple programs.
[0035] The stream instructions described herein enable high-bandwidth data-dependent scatter-gather. Each vector processing unit of the accelerator can generate data-dependent addresses to any level of the memory hierarchy with high bandwidth. The stream instructions described herein provide constructs that are structurally visible, e.g., via the ISA, to enable the representation of data-dependent access patterns in software. These access patterns include indirect memory access including multiple addressing modes, configurable transfer sizes, and atomic operations. The stream instructions enable software to specify a "start" address, "size", and "sequence pattern" for memory access. All of this is done while each tile of the accelerator implements a separate core for accessing data and responding to requests using a scatter-gather engine.
[0036] Data can flow through various components of the processor and other off-chip components and through a data stream that is data values corresponding to the addresses in the address stream. The data stream may have a stream descriptor that provides metadata characterizing the stream. As part of the metadata, the data stream may have a stream identifier. The stream identifier can be used to identify on which execution thread of the processor the stream is currently being processed. Multiple streams, e.g., 8, 16, or 32 streams, may be active and flowing through the processor simultaneously.
[0037] The circular buffer instructions described in this specification may enable a variable-sized dynamic data stream. The circular buffer instructions are constructs that are structurally (e.g., by the ISA) visible and that allow software to fetch and operate on a data stream of statically unknown size without explicitly allocating a buffer at compile time. Thus, circular buffer instructions enable the transfer and access of a dynamic-sized data stream over a static-sized region of memory. The scatter-gather engine tracks the runtime buffer sizes of various data streams and manages flow control. Since the circular buffer instructions also allow access to the underlying memory via standard loads and stores, they provide a structural first-in, first-out (FIFO) abstraction for high-speed common cases without excluding software from non-FIFO access patterns. In another example, the vector cores in each of a plurality of tiles are configured to provide a runtime configuration and access area of tile-local scratchpad memory as an ordered circular first-in, first-out (FIFO) access without excluding unordered access to the same area of tile-local scratchpad memory.
[0038] Stream instructions and circular buffer instructions enable general-case FIFO abstraction without preventing software from making "random" accesses to memory when necessary. Software can fetch the buffer as a FIFO, but can access data multiple times within the fetched buffer / window by reusing it as necessary. This is in contrast to the approach in software that requires multiple pops to reuse data and pushes those pops into the queue. Such control balances between software issuing prefetches and the scatter-gather engine that fetches and executes flow control in granularity. Software obtains higher-bit rights regarding approximately how often to issue a stream, and the scatter-gather engine implementing the stream ordering and instructions described herein accommodates fine-grained changes such as latency, buffer occupancy, etc.
[0039] A microarchitecture, e.g., the scatter-gather engine described herein, enables high-performance and irregular memory accesses. The scatter-gather engine is implemented per tile and manages the address stream defined by software, e.g., the issuance, fetch, tracking, and ordering of stream instructions and circular buffer instructions. The engine can maintain several read memory requests, e.g., 256 requests, per tile, e.g., on-chip. The scatter-gather engine tracks and updates buffer occupancy to rate-control address requests. Software can structurally inspect buffer occupancy without explicitly managing latency variations when processing individual address requests.
[0040] The combination of stream instructions, circular buffer instructions, and a scatter-gather engine enables separate access execution and can effectively increase long memory access latency, including irregular memory access streams. This provides software with the flexibility to generate data-dependent memories and the option to schedule them at a "coarse" granularity, where the scatter-gather engine is configured to flow-control and respond to requests.
[0041] Aspects of the present disclosure provide an accelerator configured to execute multiple threads of dynamic small vector tasks across a set of compute tiles. The dynamic tasks include tasks with data-dependent control and memory access. The small vector tasks can include tasks executed on relatively small vectors (e.g., 8-element vectors). These tiles are managed by a single task management core referred to as a sequencer. The ratio of the accelerator's compute to bandwidth is adjusted to accommodate irregular and sparse accesses and computations of large data sets stored in off-processor high-bandwidth memory.
[0042] The accelerator may include a plurality of tiles, each including its own access core and execution core. The access core may be a scalar unit used to decouple data movement from computation. The execution core may be another scalar unit attached to a vector unit in a plurality of SIMD lanes, configured to process data fetched by the corresponding access core. Each tile may also include its own XPU for performing cross-lane reduction, shuffle, sort, prefix sum, etc. The accelerator may also include a sequencer, which may be a scalar unit for task management across tiles and for communicating with other cores. The accelerator may also include a scratchpad memory, e.g., 8 megabytes of shared memory, although in various examples the size of the shared memory may vary. As described herein, the accelerator may implement a streaming memory access interface to keep transactions pending with respect to off-chip memory. The accelerator may also include a high-bandwidth crossbar for connecting the tiles to the shared scratchpad memory and to each other.
[0043] Exemplary System FIG. 1A is a block diagram of a hardware circuit 101 for accelerating data-dependent operations, according to aspects of the present disclosure. The hardware circuit 101 may include a sparse accelerator 103, a coprocessor 104, a high-bandwidth memory 107, and an on-chip interconnect 108. The sparse accelerator 103 may include one or more tiles 102A-102F. Each tile implements its own vector processing unit (VPU) and includes its own cross-lane processing units (XPUs) 101A-101F. The sparse accelerator 103 may include a tile sequencer 106 configured to coordinate input and output data across tiles 102A-120F.
[0044] Tiles 102A to 120F can be interconnected according to a variety of topologies, for example, as a multi-dimensional ring or torus. The interconnection can include, for example, a crossbar that will be described in more detail with reference to FIG. 1B. The crossbar or interconnection for the processor can receive and send data every clock cycle, for example, between tiles 102A to 120F, on-chip memory, and off-chip memory.
[0045] Tile sequencer 106 is a component of sparse accelerator 103 and is configured to receive and distribute instructions for executing operations on tiles 102A to 120F while adjusting them. This adjustment can be managed at least in part by utilizing the distribution relationship of the processing components on sparse accelerator 103, for example, to utilize different types of data or instruction parallelism. Tile sequencer 106 will be described in more detail with reference to FIG. 4.
[0046] Sparse accelerator 103 is configured to execute data-dependent operations using tiles 102A to 120F. As shown and described in more detail with reference to FIG. 3A, each tile can implement several data processing lanes for streaming data through a vector processing unit (VPU) and a cross-lane processing unit (XPU). The tile can fetch streamed data from on-chip memory 105, which can be any of a variety of memory devices including main memory, cache, or persistent storage, such as solid-state storage or hard disk storage. The streamed data can also be fetched from coprocessor 104, from high-bandwidth memory 107 that serves one or both of coprocessors 103 and 104, and / or from another data source connected to hardware circuit 101 via on-chip interconnection 108.
[0047] The on-chip memory 105 can be a scratchpad memory physically distributed across each of the tiles 102A - 120F. The scratchpad memory, in contrast to a hardware cache, can be programmatically managed and, for example, can store data according to software instructions. The on-chip memory 105 can be globally addressable via various interfaces such as direct memory access and / or a stream interface.
[0048] The coprocessor 104 can be any core such as a CPU, a high-density core, etc. For example, the coprocessor 104 can be configured to accelerate specific operations such as matrix-matrix multiplication, matrix-vector multiplication, etc. Examples of operations can include dense matrix calculations. In this case, most of the elements within the multiplied matrices (e.g., in some examples, more than 50 percent) have non-zero values. The complexity of the calculation can be approximated as a function of the dimensions of the multiplied matrices. In some examples, the coprocessor 104 is on a different device than the rest of the hardware circuit 101 and communicates data to the hardware circuit via the on-chip interconnect 108. The on-chip interconnect 108 can be a data bus or any form of interconnect according to any of various communication standards, e.g., PCIe. The on-chip interconnect can also implement a core memory network. The core memory network can be an on-chip network connecting the coprocessor 104 and the sparse accelerator 103.
[0049] Exemplary features of the sparse accelerator 103 are directed towards improving the calculation of sparse operations (e.g., operations on operands or inputs that generally have more zero-valued elements than non-zero-valued elements). This type of feature can include executing sparse operations using the programmable XPU101A - 101F, cooperative memory prefetching, and / or a combination of instruction stream or instruction ordering, as described in more detail herein.
[0050] Examples are provided in the context of sparse computing. In some examples, it should be understood that the sparse accelerator 103 can be used to accelerate other types of operations generally associated with the acceleration of machine learning model processing, including linear algebra operations such as vector / matrix multiplication, calculation of activation function outputs, pooling of layer outputs, normalization of layer outputs, etc. The coprocessor 104 and the sparse accelerator 103 implemented as part of the common hardware circuit 101 can facilitate the distribution of various tasks suitable for either of these two, although not limited to these two. The acceleration of non-sparse operations can include dense calculations, for example, calculations on non-sparse inputs. Examples of high-density calculations can include linear access or strided access to arrays. As other examples, dense matrix multiplication, fully connected layers, and convolutional layers of deep neural networks are included.
[0051] As an example of an input to the hardware circuit 101, there can be data constructed as a tensor. For example, the tensor can represent the input data and / or model parameter values of a machine learning model to be executed using the hardware circuit 101. A tensor is a data structure that generalizes various other common data structure types of different dimensions. A tensor can include zero or more elements that can be one or more various data types such as integers, floating-point values, boolean values, etc. Within each data type, the data type can be parameterized according to a specific level of precision, for example, 8-bit, 16-bit, or 32-bit integer values or floating-point values. The dimension of a tensor is referred to as its "rank". A rank-zero tensor is a single element also referred to as a scalar. A rank-one tensor is also referred to as a vector. A rank-two tensor is also referred to as a matrix. Vectors and matrices can also be referred to as having different ranks. For example, a rank-two vector corresponds to a matrix. A non-zero rank tensor can be described as a set of tensors of one rank lower. For example, a vector or rank one is a set of scalar values, and a rank-two matrix is a set of rank-one vectors.
[0052] Hardware circuit 101 may at least partially implement a processing pipeline for training a neural network. The pipeline may include generating an embedding for an input training example. Feature tensors for various input training examples will have various degrees of sparsity that affect the amount of computational work required to generate the corresponding embeddings. Sparsity accelerator 103 may be configured to receive a tensor of feature values representing a training input example and generate an embedding as a tensor of lower rank than the feature tensor.
[0053] To generate the embedding, sparsity accelerator 103 is configured to implement various data-dependent operations for efficient sparse data computation on XPU 101A - 101F etc. or more generally on a VPU. These operations include sorting or summing sparse vectors, operations for summarizing the content of input vectors, and operations for converting a sparse matrix from one sparse matrix storage format to another.
[0054] Instead of a physical predetermined circuit for accelerating the execution of data-dependent operations, the VPU including XPU 101A - 101F can be configured, for example programmed, to execute a wide variety of data-dependent operations. Sparsity accelerator 103 enables general support for processing sparse data while allowing complementary coprocessor 104 to execute other types of operations.
[0055] The hardware circuit 101 can be any of a variety of types of processing units, such as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC), such as a tensor processing unit (TPU). The hardware circuit 101 can be implemented on a computing device that can itself be part of a system of one or more devices.
[0056] FIG. 1B is a block diagram of an exemplary data path implemented as part of a hardware circuit according to aspects of the present disclosure. The scratch pad memory 162B and the task instruction memory 160B can form part of the on-chip memory 105.
[0057] Various data paths 150B, 152B, 154B, and 156B are shown. These data paths may or may not physically share a DMA path 150B that includes circuit interconnections between the tile sequencer 106 and tiles 102A - 120F for direct memory access (DMA) of the scratch pad memory 162B. The instruction data path 152B includes circuit interconnections between tiles 102A - 120F, the task instruction memory 160B, the scratch pad memory 162B, and the memory transport interface 163B. The scratch pad (spmem) data path 154B indicates a potential path for data from the task instruction memory 160B to tiles 102A - 120F. The control data path between the tile sequencer 106 and tiles 102A - 120F indicates an exemplary path for control signals generated by the tile sequencer 106. When received by tiles 102A - 120F, the control signals can cause the tiles to perform one or more primitive or composite operations, such as reading data, writing data, and / or processing data, according to one or more functions specified by the control signals.
[0058] The task instruction memory 160B is shared by tiles 102A to 120F. The task instruction memory 160B holds programs executable by the tile access core and the tile execution core. In some examples, the task instruction memory 160B is 468 bits wide and 16,000 words deep and can be organized, for example, as 2,000-word banks including error correction codes. Multiple banks enable multiple reads from or multiple writes to the task instruction memory 160B per block cycle. The DMA descriptor for the task instruction memory 160B may include a length field that is a multiple of 64 bytes. Instruction bundles may be zero-padded from the most significant bit up to the 512-bit boundary when stored in high-bandwidth memory. The sparse accelerator 103 may drop the padded bits before writing to the task instruction memory 160B. Each DMA descriptor can transfer one task instruction memory bundle.
[0059] The tile sequencer 106 can be a scalar core that mainly serves to dispatch tasks to tiles and / or initiate DMA transfers. For example, the tile sequencer can receive instructions as bundles. The instructions within a bundle can be executed in parallel across tiles 102A - 120F and can update the architectural state simultaneously. When a bundle is executed, the bundle receives scalar or vector issue before the instructions within the bundle are executed. The scalar or vector issue can be held over one or more cycles due to various conditions documented per scalar instruction for hold - scalar - issue conditions or per vector instruction for hold - vector - issue conditions. The ISA for the sparse accelerator 103 can define some global scalar - hold - issue conditions or global vector - hold - issue conditions that can hold a bundle unconditionally. During scalar or vector issue, several actions may be performed. The predicate of the bundle is evaluated and updated. The values of the registers to be used can be recorded. Branches within the instructions can be executed and the registers are updated.
[0060] Banks of the task instruction memory can implement one or more of the following interfaces. Each bank can include a prefetch request and prefetch response broadcast bus. The prefetch request and response bus architecture is tailored to the SPMD (single program, multiple data) arithmetic mode. The read response may be broadcast to all tiles on the prefetch response broadcast bus. Bundle broadcast, if possible, enables the hardware to replicate requests originating from another tile, reducing the overall bandwidth requirement.
[0061] FIG. 2 is a block diagram of an exemplary environment 200 for implementing the hardware circuit 101. The hardware circuit 101 may be implemented on a device having one or more processors at one or more locations, such as within the server computing device 215. The user computing device 212 and the server computing device 215 may be communicatively coupled to one or more storage devices 230 via the network 260. The storage device 230 may be a combination of volatile memory and non-volatile memory and may be in the same physical location or a different physical location from the computing devices 212, 215. For example, the storage device 230 may include any type of non-transitory computer-readable medium capable of storing information, such as a hard drive, solid state drive, tape drive, optical storage, memory card, ROM, RAM, DVD, CD-ROM, writable memory, and read-only memory.
[0062] The server computing device 215 may include one or more processors 213 and a memory 214. The memory 214 can store information accessible by the processor 213, including instructions 221 executable by the processor 213. The memory 214 may also include data 223 that can be retrieved, manipulated, or stored by the processor 213. The memory 214 may be a type of non-transitory computer-readable medium capable of storing information accessible by the processor 213, such as volatile memory and non-volatile memory. The processor 213 may include one or more central processing units (CPUs), graphics processing units (GPUs), field programmable gate arrays (FPGAs), and / or application specific integrated circuits (ASICs), such as tensor processing units (TPUs). The processor 213 may include a coprocessor and a sparse accelerator implemented as part of the hardware circuit, as described herein with reference to FIGS. 1A-1B.
[0063] Instruction 221 may include one or more instructions. When executed by processor 213, the instructions cause one or more processors to perform the operations defined by the instructions. Instruction 221 may be stored in an object code format for direct processing by processor 213, or may be stored in other formats including interpretable scripts or sets of independent source code modules that are interpreted on demand or pre-compiled. Instruction 221 may include instructions for configuring a stream transfer in accordance with aspects of the present disclosure. Server computing device 215 and / or user computing device 212 may implement a compiler or other program for generating instructions and transmitting them to hardware circuit 101 as control signals for configuring tiles of the circuit.
[0064] Data 223 can be fetched, stored, or modified by processor 213 in accordance with instruction 221. Data 223 can be stored in a computer register, as a table having a plurality of various fields and records, or in a relational or non-relational database as a JSON, YAML, proto, or XML document. Data 223 can also be formatted in a computer-readable format including, but not limited to, binary values, ASCII, or Unicode. Further, data 223 may include sufficient information to identify related information, such as numbers, descriptive text, proprietary codes, pointers, references, or information used by functions for computing related data, regarding data stored in other memories including other network locations.
[0065] The user computing device 212 may also be configured with one or more processors 216, a memory 217, instructions 218, and data 219, similar to the server computing device 215. The user computing device 212 may also include a user output 226 and a user input 224. The user input 224 may include any suitable mechanism or technology for receiving input from a user, such as a keyboard, a mouse, a mechanical actuator, a software actuator, a touch screen, a microphone, and a sensor.
[0066] The server computing device 215 may be configured to send data to the user computing device 212, and the user computing device 212 may be configured to display at least a portion of the received data on a display implemented as part of the user output 226. The user output 226 may also be used to display an interface between the user computing device 212 and the server computing device 215. The user output 226 may alternatively or additionally include one or more speakers, transducers, or other audio outputs, a tactile interface, or other tactile feedback that provides non-visual information and non-audible information to the platform user of the user computing device 212.
[0067] FIG. 2 shows the processors 213, 216 and memories 214, 217 as being within the computing devices 215, 212. However, the components described herein that include the processors 213, 216 and memories 214, 217 may include multiple processors and memories that are operable at different physical locations rather than within the same computing device. For example, some of the instructions 221, 218 and data 223, 219 can be stored on a removable SD card or the like within a read-only computer chip. Some or all of the instructions and data can be stored at a location physically remote from the processors 213, 216 but accessible to the processors 213, 216. Similarly, the processors 213, 216 may include a set of processors capable of performing simultaneous and / or sequential operations. Each of the computing devices 215, 212 may include one or more internal clocks that can provide timing information for use in measuring the time for operations and programs executed by the computing devices 215, 212.
[0068] The server computing device 215 may be configured to receive requests for processing data from the user computing device 212. For example, the environment 200 may be part of a computing platform configured to provide various services to users through various user interfaces and / or APIs that expose platform services. One or more services may be a set of machine learning frameworks or tools for generating neural networks or other machine learning models according to specified tasks and training data. The user computing device 212 may receive and transmit data specifying the type of workload or complex operation to be performed by the XPU of the sparse accelerator 103. The user computing device 212 may directly send instructions to the hardware circuit 101 or, as described herein, cause the server computing device 215 to generate instructions and transmit them as control signals to the hardware circuit 101.
[0069] Devices 212 and 215 may be capable of direct and indirect communication via network 260. Devices 215 and 212 can set up listening sockets that can accept initial connections for sending and receiving information. Network 260 itself can include various configurations and protocols, including the Internet, the World Wide Web, intranets, virtual private networks, wide area networks, local networks, and private networks that use communication protocols specific to one or more companies. Network 260 can support various short-range and long-range connections. The short-range and long-range connections can occur over various bandwidths such as 2.402 GHz to 2.480 GHz generally associated with the Bluetooth® standard, 2.4 GHz and 5 GHz generally associated with the Wi-Fi® communication protocol, or using various communication standards such as the LTE® standard for wireless broadband communication. Network 260 can also support wired connections between devices 212 and 215, including, additionally or alternatively, via various types of Ethernet® connections.
[0070] Although a single server computing device 215 and user computing device 212 are shown in FIG. 2, it should be understood that aspects of the present disclosure may be implemented according to a wide variety of configurations and amounts of computing devices, in a paradigm for sequential or parallel processing, or via a distributed network of multiple devices. In some implementations, aspects of the present disclosure may be implemented on a single device and any combination thereof.
[0071] FIG. 3A is a block diagram of an exemplary tile 102. The XPU 101 is coupled to a crosslane controller 310. The crosslane controller 310 provides a separate thread of control that enables crosslane instructions on the XPU 101. As described herein, the XPU 101 can receive a first instruction, for example, through one or more control signals, and the first instruction can be converted into one or more second instructions and third instructions for performing a composite operation specified by the first instruction, and can be provided to the processing cells and crossbar of the XPU 101, respectively. Instructions for the XPU 101 can be conveyed through control signals, and the processing cells and crossbar of the XPU 101 are configured to interpret and perform corresponding primitive operations. As an exemplary instruction, there can be an opcode of an instruction set architecture (ISA).
[0072] Tile 102 can receive data from the on-chip interconnect 108 and further from the on-chip memory 105, as described with reference to FIG. 1A. The XPU can also receive instructions from the instruction interface 324, for example, from the tile sequencer 106, through the scalar core 312 or the scalar core 320. The scatter-gather engine 322 of tile 102 can receive incoming data and control which data is passed to the memory 306 via the memory scheduler 314. In some examples, instead of the scatter-gather engine, the scatter-gather engine 322 may be referred to as a read-write engine 322.
[0073] The memory scheduler 314 can adjust how data is accessed and retrieved from the memory 306. The memory 306 is dedicated to tile 102 and cannot be accessed by other components connected to tile 102 such as other tiles. The arbiter 304 is configured to manage which of the vector processing units (VPUs) 302A - 302H accesses the memory 306, for example, every clock cycle. Tile 102 can maintain a task queue 308 for the tasks to be executed by tile 102, and the tasks are sent to the scatter - gather engine 322 via the scalar core 320. Tile 102 can also maintain registers for tile synchronization flags 318 and / or memory flags 316 to synchronize tile 102 with other tiles of the hardware circuit and memory 306 respectively.
[0074] The VPUs 302A - 302H are connected to the XPU101 via data processing lanes indicated by solid lines between the XPU101 and the VPUs 302A - 302H. The dashed lines between the XPU101 and the VPUs 302A - 302H represent control signals that can be received by control cells within the XPU101 to configure the XPU101 to perform complex operations corresponding to the received control signals. The vector processing units are configured for efficient operations on input vectors. The length of the vectors processed by tile 102 at one time can depend on the number or width of the VPUs implemented by the tile. For example, the eight VPUs 302A - 302H have a width of eight. The VPUs 302A - 302H can process data along the same data processing lane. The VPUs 302A - 302H can be configured to perform scalar operations on the elements of the received vectors from the memory 306. The VPUs 302A - 302H can receive data from the XPU101, and the XPU101 can process data across the data processing lane, not just along the lanes executed by each of the VPUs 302A - 302H as described herein.
[0075] Figure 3B is a block diagram of another exemplary tile 392 that implements the XPU 101 for stream transfer. Tile 392 can receive data from the on-chip interconnect 108, as well as from the on-chip memory 105, as described with reference to FIG. 1. The XPU 101 can also receive instructions from an instruction interface 324, for example, from a tile sequencer 106. The scatter-gather engine 322 of tile 392 can receive incoming data and control which data is passed to the memory 306.
[0076] This exemplary tile 392 is based on a decoupled access / execution architecture, and a program (and associated instruction sequence) can be separated into two streams. The first stream can be an access stream for fetching operands and storing results. The second stream can be an execution stream for consuming operands, performing calculations, and generating results. These streams are executed on two separate cores, a tile access core (TAC) 332 and a tile execution core (TEC) 330, which form part of the decoupled access / execution architecture. The TAC 332 is a scalar unit used to decouple data movement from data calculation, for example, decoupling the fetch of data from memory and the processing of the fetched memory. The TEC 330 is a scalar unit attached to a VPU with a plurality of SIMD lanes (e.g., eight SIMD lanes) for processing vectors.
[0077] The tile access core 332 is based on the scalar core complex 320 and can serve to prefetch operands to be executed from the memory 306 or from a high-bandwidth memory outside the tile 392. The tile execution core 330 is based on the scalar complex core 312, includes the XPU 101 and the VPU 302, and can serve to execute computational operations on the operands prefetched to generate results. The VPU 302 is connected to the memory 306 via the load store unit 328. The load store unit 328 is configured to perform gather operations and scatter operations on the data passing through the tile 303. The gather-scatter operations can be performed granularly, for example, at a 4-byte granularity. The load store unit 328 can implement a load queue and a store queue for managing bank conflicts. The LSU 328 can provide load / store access to a subset of the scratchpad memory.
[0078] The TAC 332 and the TEC 330 have independent instruction streams and together form a producer-consumer pair. The tile 102 can maintain a task queue 308 for the tasks to be executed by the tile 102, and the tasks are sent to the TAC 332 and the TEC 330.
[0079] The task queue 308 can include one or more queues for push instructions and popping instructions. For example, the task queue 308 can have a first queue (e.g., a first-in-first-out (FIFO) queue) for pushing values from the TAC 332 to the TEC 330. The TEC 330 can pop and process the enqueued values. As another example, the task queue 308 can include an additional queue for connecting the TAC 332 and the TEC 330 in the reverse direction.
[0080] TAC332 and TEC330 communicate with each other via a tile-local scratchpad memory (SpMEM) such as memory 306. TAC332 and TEC330 can also communicate through scalar memory 334, instruction buffer 326, and tile synchronization flag 318. Memory 306 can be used by TAC332 and TEC330 to exchange data and can be used as a software-managed circular buffer to pass data between TAC332 and TEC330 in a first-in first-out order. Tile synchronization flag 318 can be used as a counting semaphore between TAC332 and TEC330. For example, if a circular first-in first-out order is used between two cores, the producer core increments the synchronization flag 318 by the number of bytes after each push and stops when the count reaches the maximum size in the first-in first-out order. Similarly, the consumer decrements the synchronization flag 318 after each pop and stops when there is no data in the buffer. Since the amount of prefetched data can be dynamic, a done bit is used to indicate the end of the stream.
[0081] In some examples, tile synchronization flag 318 can include a set of 32 synchronization flag registers. Each register can store a "done" bit and an "enable_public_access" bit in addition to 32 bits of data. The tile synchronization flag 318 registers may be implemented as a monolithic flip-flop array, allowing simultaneous access from all sources. If there is an address conflict during writing, the access priority between sources can be specified. For example, scalar miscellaneous instructions may have absolute priority. In some examples, only one of the scalar miscellaneous instructions can be issued. Read-modify-write operations are pipelined, and thus, back-to-back read-modify-write operations on any synchronization flag can be supported.
[0082] DMA updates, stream updates, and remote writes may be combined on a single (external) interface using round-robin arbitration. For example, the external interface from the tile to an external source may have a separate access path to a synchronization flag, e.g., a data path as illustrated and described with reference to FIG. 1B.
[0083] Each bank of the task instruction memory 160B can perform arbitration on a per-cycle basis among the tile request data stored in the bank. These banks can select a winner from the TAC332 and TEC330 within the sparse accelerator 103 that obtains access to the bank using a distributed arbitration scheme. The arbitration can ensure that the required bandwidth is equally divided among the requested tiles. The requests can be, for example, to prefetch data by each TAC for each tile. For example, the winning prefetch request is given the highest priority for access to any of the accessed banks.
[0084] The access from the sparse accelerator 103 via the control status register can be next highest in priority. This is only delayed by a prefetch read request. The control status register (CSR) can be implemented as part of the processor for storing additional information results of the machine instructions executed by the sparse accelerator 103. The sparse accelerator 103 can maintain an indirect CSR timeout status bit to determine whether the completion of the CSR access is blocked. After bank access and CSR access, the DMA read or write can be prioritized last. DMA operations, e.g., reads or writes, can be delayed when accessing a busy bank. In some examples, the priorities of how access operations are arbitrated and resolved on the sparse accelerator 103 can be different.
[0085] Separating data access from data execution has at least the following advantages. Since TAC332 can perform address calculations and prefetch the necessary data, long (e.g., 600-cycle) memory latencies are tolerated more effectively. The architecture described above also has a high tolerance for control latencies. Dynamic dependencies make it difficult to resolve loop conditions, which can prevent effective software development. TAC332 can execute ahead of time to resolve these dependencies when the conditions can be determined outside of TEC330, providing a way to hide control latencies.
[0086] Figure 4 is a block diagram of a tile sequencer 400 according to an aspect of the present disclosure. The sequencer memory 410 (sequencer memory: Simem) can have a width of a VLIW instruction bundle and a depth of about 8000 bundles, although the width and number of bundles of the sequencer memory 410 can vary for each implementation example. The sequencer memory 410 may be read or written through direct memory access or through indirect access. This can be done regardless of whether the sequencer 400 is currently executing. The sequencer memory DMA memory descriptor can have a length field that is a multiple of 32 bytes. Each bundle can be zero-padded from the most significant bit to the 256-bit boundary when stored in high-bandwidth memory. The padded bits have RAZ / WI (read-as-zero and write-ignored). Each memory descriptor can transfer one instruction bundle.
[0087] The sequencer memory 410 can be organized as two banks having consecutive addresses between the banks. Each bank can perform one read or one write per clock cycle. Through interleaving, a continuous instruction sequence can consume only half of the bandwidth of each bank. Loops of unusually small size can consume only three-quarters of the bandwidth of a given bank, leaving sufficient bandwidth remaining to advance access in the forward direction.
[0088] Each sequencer memory bank within the sequencer memory 410 can follow the following interface. The sequencer memory 410 can receive a read request from the instruction data path (e.g., the instruction data path illustrated and described with reference to FIG. 1B). The banks of the sequencer memory 410 can also receive control status register accesses via the indirect register interface, in addition to DMA writes and reads. Reads from the instruction data path can have the highest priority for access to each bank within the sequencer memory 410. The priorities of DMA reads, writes, and CSR accesses can be equal and can receive least recently used (LRU) arbitration for banks that are not busy performing an instruction fetch.
[0089] The sequencer 400 can fetch its instructions from the sequencer memory 410. The sequencer 400 executes the control threads of the program. This involves generating descriptors that are dispatched by a descriptor dispatch unit 413 (for task and stream descriptors) and a DMA unit 414 (for DMA descriptors), respectively. The task descriptors are provided to the tile FIFO 416, which are then executed by their respective tiles. The DMA descriptors can be passed to other components of the hardware circuit 101 and other off-chip components. The sequencer communicates with other cores within the system and coordinates tasks across the tiles of the accelerator.
[0090] The current state of the sequencer 400 can be determined by reading the corresponding status register. Other registers related to the fetch and execution of instruction bundles are the program counter and the branch state. The tile sequencer 400 can issue DMA descriptors that can be adjusted and throttled by hardware. The hardware can leave a predetermined number of DMAs issued by the sequencer, which have synchronization flags stored between the synchronization flags 418, unprocessed. Adjusting and throttling can be made possible or impossible as needed.
[0091] The synchronization flags can be displayed in two or more types. The synchronization flags 418 can be displayed in the tile sequencer 400. Other synchronization flags can be stored in the TAC and TEC of each tile. All of the synchronization flags can be arranged in a single address space accessible to DMA operations, atomic remote set / add instructions, and atomic tile set / add instructions in total. There are several interfaces that can be implemented between the synchronization flags within the tile and the synchronization flags within the sequencer 400. The DMA operation can atomically add a value to the synchronization flag during execution and can set the "Done" bit upon completion. The stream operation can atomically add a value to the synchronization flag during execution and can set the "Done" bit upon completion. The remote write instruction generates a single-word control write that can atomically set or add a value to the synchronization flag. These can be from the atomic remote set / add instruction used to update the remote synchronization flag. These can also be from the atomic tile set / add instruction used to update the synchronization flag within the sparse accelerator. Another implementable interface is the write interface, in which case atomic updates to the synchronization flag are performed, optionally including changing the "Done" bit, by the set synchronization flag and the synchronization flag add instruction. Another implementable interface is the read interface.
[0092] The synchronization flag 418 can be organized in memory as several banks. Each entry may include a "Done" bit and an "enable_public_access" bit in addition to several bits of data, e.g., 32-bit data. The banks can perform cycle-by-cycle arbitration separately for the read and write ports, e.g., to prioritize scalar miscellaneous instructions over DMA, stream, and remote write updates. The synchronization flag registers in the TAC and TEC may also store several bits, a "Done" bit, and an "enable_public_access" bit.
[0093] The synchronization flag in the TAC or TEX can be implemented as a monolithic flip-flop array that allows simultaneous access from all sources. Similar to the sequencer synchronization flag, the tile synchronization flag can be managed according to a priority scheme to avoid write conflicts.
[0094] FIG. 5 is a block diagram of an exemplary scratchpad memory 500 having memory across multiple tiles of a sparse accelerator according to aspects of the present disclosure. Each tile may include a portion of a scratchpad memory 502 referred to as a tile scratchpad memory or TileSpmem. The tile scratchpad memory 502 may be accessed by the tile via a load / store interface implemented as one or more circuits. The tile scratchpad memory 502 can also be used as a local buffer for each tile to move data in and out of the tile using stream instruction transfers.
[0095] As shown in FIG. 5, each tile scratchpad memory 502 may include a number of banks, for example, 32 banks each labeled as bank 0 to 31. Each bank can hold a certain amount of data, for example, 16 kilobytes. Each bank can hold a number of words, for example, 4-byte words and 7-bit error correction codes. In some examples, a bank can hold 4096 words. It should be understood that the scratchpad memory 500 can be implemented with a different number of tile scratchpad memories 502 each holding different sized banks. It should also be understood that each bank can be implemented in various examples to store different numbers of words of different sizes.
[0096] Words and banks within the tile scratchpad memory can be accessed through 17-bit instructions and stream addresses, although the exact size and format of the address may vary for each implementation example. Each bank can implement one or more of the following interfaces. Tile scratchpad memory 502 enqueues access requests to a per-bank load queue (not shown) in response to received instructions, and loads or stores and adds data within the banks and words specified by the received address or address range. To store data, tile scratchpad memory 502 enqueues access requests to the per-bank load queue from received store or store-add instructions. Each bank can also receive read accesses from one or more external sources, such as DMA or stream requests for reading / writing data. Exemplary instructions to tile scratchpad memory 502 include vector load instructions and vector store instructions. These and other instructions can be specified in the ISA for hardware circuit 101 and can define instructions that cause hardware circuit 101 to perform specific predetermined operations in response to receipt of the instructions. Vector load instructions can be issued to the bank load queue. Similarly, vector store instructions may be issued to the bank store queue. A vector store-add instruction can refer to a type of instruction that causes hardware circuit 101 to perform a read-modify-write operation within a target range of addresses in one or more banks of tile scratchpad memory 502.
[0097] To handle load and store prioritization on the tile scratchpad memory banks, the store access at the head of the per-bank store queue may have the highest priority for accessing the write port of its respective bank (not shown in FIG. 5). The load access at the head of the per-bank load queue may have the highest priority for accessing the read port of the bank. If the load is for the same address as a store that was enqueued in the queue and the store was from an instruction bundle issued before the load, the bank may instead read the data from the store queue to load the data.
[0098] Accesses from sources external to the tile hosting the bank may undergo multi-level arbitration before the access is presented to the bank. For example, writes or appends from different sources may first undergo least-recently-used (LRU) arbitration among those enqueued in the per-bank write queue of the target bank. The write queue may be, for example, 4 entries deep. Similarly, read access requests from various sources may also undergo LRU arbitration. The winning read request may then undergo a second level of arbitration with read requests resulting from stream append accesses and load accesses to read the bank. Stream append accesses perform read-modify-write operations, which are also enqueued in the per-bank write queue. If the head of the write queue for the bank is a stream append access, its read request can undergo round-robin arbitration with the winning external read request. The winning read request has the second-highest priority behind load accesses for accessing the read port of the bank. The write request at the head of the per-bank write queue may have the second-highest priority behind store accesses for accessing the write port of the bank. In the absence of bank contention, the bank can maintain a throughput of at least one stream append operation per cycle.
[0099] Each bank of the tile scratchpad memory 502 includes ports for reading from and writing to the bank (not shown in FIG. 5). These ports can generate warnings for detecting depletion. External sources of read requests and write requests may be insufficient when accessing a bank that is continuously accessed by a load, store, or store append. The probability of depletion can be reduced by specifying the maximum access of the total number of banks within the tile scratchpad memory 502. For example, only access to a maximum of 8 out of 32 banks is allowed in a given clock cycle.
[0100] Warnings generated due to depletion may be based on a predetermined threshold. For example, the predetermined threshold may be the number of consecutive clock cycles corresponding to the period during which the external source does not access the bank for a read request or a write request.
[0101] If a read port depletion is detected by a bank within the tile scratchpad memory 502, one or more actions can be performed by the scatter / gather engine. Issuing a hold can be performed when the bundle has a load instruction, a store instruction, or a store append instruction. The load queue and store queue for each bank can be drained normally. Issuing a hold can continue until a predetermined maximum number of read requests are handled by the scatter / gather engine or until the scatter / gather engine read queue becomes empty.
[0102] When a write port depletion is detected by a bank within the tile scratchpad memory 502, the following sequence is executed. Issue a hold when the instruction bundle has a store instruction or a store append instruction. The store queue for each bank is drained normally. Issuing can continue to be held until the maximum threshold of the issued requests being held is handled by the scatter / gather engine. The threshold and cycle count may vary for each implementation example.
[0103] In some examples, the synchronization flag memory 412 can be organized as four banks (not shown). Each entry can have 32-bit data. Each bank within the synchronization flag memory 412 can perform per-cycle arbitration separately for read ports and write ports among the following sources. Scalar miscellaneous instructions have absolute priority. In some examples, only one of the scalar miscellaneous instructions can be issued in any given cycle. The read-modify-write operation may be pipelined, and thus, back-to-back read-modify-write operations to any location can be supported. After a scalar miscellaneous instruction, DMA updates, stream updates, and remote writes may be combined onto a single (external) interface via round-robin arbitration and may then have the next highest priority. Access via the control status register from the host to the hardware circuit 101 has the lowest priority.
[0104] Access requests to the banks can be from multiple sources. Core memory network reads are read accesses sent on the core memory network that can be originated by DMA or stream requests. An access request can be a core memory network write for a write access from the core memory network caused by a DMA or stream request. An access request can be from a stream address, for example, an indirect address read from a local tile. An access request can be from a read or write of an internal stream.
[0105] FIG. 6 is an exemplary block diagram of a scalar core complex 600 of a tile sequencer according to aspects of the present disclosure. The scalar core complex can be a 32-bit scalar VLIW core. The scalar execution pipeline 601 can be one or more circuits configured to execute scalar instructions and miscellaneous instructions such as fetch, decode, and execute. The scalar core complex 600 executes a pre-fetched instruction bundle. During execution, one bundle is fetched per block as addressed by the next PC register. Each bundle can include two scalar instructions and one miscellaneous instruction that is decoded and executed simultaneously.
[0106] The scalar execution pipeline 601 can include a 32-bit computing unit that includes a register 604 and an ALU 606. The computing unit provides address calculations, loop control, and scalar operands used to construct descriptors. The complex 600 has a memory 608 accessible via a load / store interface. The memory 608 is used by the complex 601 to store intermediate scalar values during program execution. The core type determines the depth of the memory 608, for example, depending on whether it is part of a tile sequencer or a tile's TAC or TEC.
[0107] The complex 600 can construct a descriptor 609 and use a descriptor scratchpad register 610. When an instruction uses the descriptor scratchpad register 610, at issue, the complex 600 fetches N × 32-bit words starting from the specified descriptor scratch register address within the instruction. The value of N depends on the particular descriptor type. The descriptor is then enqueued into the descriptor issue FIFO 612.
[0108] Cross-lane processing unit (XPU) An exemplary implementation of an XPU will be described with reference to the following description and the accompanying figures. Distributed embedded training is difficult to accelerate because effective acceleration relies on provisioned computing, available on-chip main memory bandwidth (HBM), and inter-chip communication bandwidth. These computations are dynamic and irregular, making it difficult to utilize them. Additionally, the performance bottlenecks vary greatly from model to model, and the problem space has been evolving rapidly with new algorithms, optimizers, model architectures, etc. Flexibility in representing new algorithms and optimizing for various performance bottlenecks is essential.
[0109] Embeddings can be fed into a machine learning model, such as a neural network, for performing machine learning tasks, such as natural language processing tasks, or other tasks that may benefit from using embeddings. An embedding layer is a layer of a neural network trained to generate embeddings from inputs. After generation, the embeddings can then be processed downstream, for example, by a later layer of the neural network implementing the embedding layer. Models with embedding layers pose unique computational challenges, for example, due to low computational density, high stress on memory bandwidth, and large memory footprints. Furthermore, there can be wide differences in performance bottlenecks from model to model for accelerating these types of models.
[0110] An embedding function or map can be implemented, for example, as a lookup table or as a sparse vector-dense matrix multiplication. For example, an embedding function can be implemented as a matrix that, when multiplied by an input vector, generates a corresponding embedding for the input vector. For example, the input vector may be a bit vector representing the presence or absence of various natural language words within an input sentence to a machine learning model. Since the bit vector can include elements for a large vocabulary of potential natural languages that can form the input sentence, this bit vector will generally be a sparse vector having fewer elements than, for example, 50 percent of the elements within a vector having zero values. The output vector obtained by multiplying the input vector by the embedding function matrix is an embedding representing the input sentence.
[0111] The embedding function can be of a very large table, for example, hundreds of gigabytes in size. As a result, the embedding function cannot fit into the main memory of a single accelerator or processor, and thus, embedding generation is distributed across multiple nodes, each with one or more accelerators. The embedding table may be partitioned among multiple devices, for example, among multiple accelerators within a pod of a data center.
[0112] Such distribution makes the processing of these embedding layers implementing these embedding functions complex. Aspects of the present disclosure provide for scattering, gathering, uniquifying, and accelerating various operations for summing various input values in order to facilitate the generation of embeddings for individual or batch input samples. A large embedding table can be partitioned among multiple accelerators. For ease of explanation, the following description will focus on a single accelerator, e.g., hardware circuit 101. Further, while the examples provided herein describe embeddings for natural language processing tasks such as machine translation, it should be understood that aspects of the present disclosure can provide acceleration for any type of machine learning model that at least partially relies on embedding generation for performing each machine learning task. Other examples include recommendation systems such as content recommendation systems found in domains such as multimedia recommendations, search result ranking, and advertising.
[0113] Examples regarding embedding generation are provided, but it should be understood that the same primitive operations described herein can be assembled in other ways to address other sparse problems such as sparse matrix multiplication. This flexibility enables the acceleration of various sparse problems. Sparse computations can be employed in other problem spaces such as scientific computing and / or graph analysis in addition to machine learning and deep neural networks.
[0114] On the forward pass of a neural network processed according to one or more accelerators as described herein, the input can be a batch of one or more input samples. The input samples can be processed by one or more accelerators that execute operations of one or more embedding layers of the neural network. The output of the embedding layer is a batch of one or more embeddings, one for each sample in the input batch. It should be noted that the input samples share a common length, e.g., the number of potential features resulting from each input sample, but the input samples can have some empty or zero-valued feature values.
[0115] In the forward pass, the input batch is partitioned across a plurality of different accelerators, e.g., across one or more host devices.
[0116] The input batch can be represented as two vectors, a vector of values and a vector of indices. The vector of values corresponds to the value for each identifier within the input samples of the batch. The index can point to the position of the value for each identifier within the tensor representing the input batch. The input batch is partitioned such that portions of the input batch are sent to separate accelerators. When an accelerator receives the partitioned input batch, it can “uniquify” the input to remove duplicate identifiers across the input batches. Uniquifying the input means removing multiple instances of the same identifier. One reason for such removal is to conserve bandwidth by reducing network usage between chips and avoiding redundant accesses to the same identifier. Uniquification avoids redundant lookups on the embedding table. After uniquification, the uniquified input batch is distributed across multiple devices each with a portion of the embedding table to generate an output embedding. The generated embeddings can be gathered from the various devices after scatter and returned to another device that requests the embedding.
[0117] During training, the gradients representing the error rate between the ground truth embedding and the predicted embedding can similarly be scattered to the devices that store the partitions of the embedding table to update each respective partition.
[0118] Aspects of the present disclosure are directed to an XPU for performing data-dependent operations across multiple data processing lanes of a processor. Instead of implementing operation-specific circuits physically fabricated for each data-dependent operation, the XPU can be configured to perform various operations in response to input signals that make up individual operations executed by processing cells and crossbars arranged as a stack-type network in the XPU. The XPU operates across values of multiple SIMD data processing lanes. The XPU can be implemented as part of a coprocessor that complements a second coprocessor configured for SIMD parallel processing. The coprocessor implementing the XPU can be configured to perform data-dependent operations.
[0119] Aspects of the present disclosure provide an XPU for initially processing sparse data before passing the data to a downstream coprocessor within a processing pipeline, enabling a broader workload to perform more efficient computations than were previously possible without an XPU. Since the XPU can process a variety of data-dependent operations, the processing pipeline and corresponding processor can be designed without the constraint of predefining input data for processing on existing SIMD architectures. Without the XPU, existing SIMD architectures cannot efficiently accelerate data-dependent operations such as generating embeddings from a sparse set of features into a machine learning model.
[0120] Exemplary data-dependent operations include generating embeddings for input training examples. The embeddings can be vectors or some other data structure mapped from inputs having dimensions higher than the embeddings. Embedding generation can be performed as part of a workload processed according to a pipeline. As other examples, the XPU may perform vector scatter or gather operations, segment sums, and / or partition sparse feature tensors. The XPU described herein can be a processing unit complementary to other components of a processor, such as a vector processing unit constructed according to a SIMD parallel processing paradigm, or connected components. One or more XPU's can be connected in each processor core of a larger processor that may itself include other components for accelerating the performance of a particular workload, such as training a neural network.
[0121] Furthermore, the XPU is not limited to performing a particular type of data-dependent operation, and thus, a processor can be designed to include an XPU to complement other types of processing units for multiple different pipelines. Since the XPU can be configured per workload, the physical footprint of the XPU is reduced compared to other approaches where special circuitry is physically fabricated on the processor as a complementary unit for sparse data computations. The functionality of the XPU can also be extended by using an instruction set or by extending the host processor to an existing instruction set, further improving the adaptability of various data-dependent operations in response to changing receipt of pipeline data. Instructions can be provided as signals to components of the XPU that serve to translate instructions for configuring the individual processing cells and crossbars of the XPU. The XPU can be configured using a program compiled by a corresponding compiler for the hardware circuit implementing the XPU.
[0122] The XPU includes a network of individual processing cells, each of which processes data passing through one or more data processing lanes through crossbar connections between the processing cells. Each data processing lane may include one or more registers for temporarily storing data during processing. Each processing cell is configured to perform one or more primitive operations on a plurality of sets of operands. The first set of operands is provided as an input from the data processing lanes of the processor shared by the processing cells. The second set of operands is provided from a crossbar configured to coordinate data transmission across a plurality of data processing lanes of the XPU.
[0123] The XPU can be divided into a plurality of pipeline stages, each stage including a crossbar, one or more processing cells, and a control cell corresponding to each processing cell. The number of stages may vary, for example, based on the composite operations configured for the XPU to execute for the current workload.
[0124] The XPU executes a composite operation by performing a plurality of primitive operations across pipeline stages of a stacked network of processing elements and crossbars. A composite operation is an operation performed by the XPU on an input to generate an output. A primitive operation is an operation configured to be executed by an individual processing cell of the XPU, and when executed by the XPU, causes the XPU to execute a composite operation. Executing a composite operation may require executing other composite operations. For example, to perform a vector sort, the XPU may execute a prefix addition, i.e., another operation composed of a plurality of primitive operations. Exemplary primitive operations include comparison of input data, arithmetic operations, or operations for bypassing. The XPU executes a composite operation by configuring each of a plurality of individual processing cells and crossbars arranged according to one of a plurality of pipeline stages for the XPU.
[0125] The primitive operations executed at each stage of the XPU can be defined by a program and can vary for each workload. The primitive operations configured to be executed by a processing cell are determined by one or more control signals or instructions received by the respective control cell for the processing cell. The exact primitive operations executed by a processing cell can depend, for example, on the composite operation that the XPU is currently configured to execute. In other examples, the processing cells in various lanes or various stages of the XPU can be configured to always execute one or more predetermined primitive operations. After the XPU generates an output, the output can be passed along multiple data processing lanes to another processing unit or memory unit of the processor implementing the XPU.
[0126] Two exemplary composite operations that the XPU can execute are vector sort and vector duplicate count. Vector sort is an in-place stable sort of (key, value) tuples of an input vector sorted by a key. Vector duplicate count returns a running count of the duplicates of the values of the (key, value) tuples of an input vector. The XPU is configured to execute both vector sort and duplicate count according to the same configuration of processing cells and crossbars as described herein. By using the same configuration, the XPU can execute both composite operations more efficiently because, at least, there is no need to reconfigure the XPU while executing vector sort and vector duplicate count for a given input vector. Other composite operations configured to be executed by the XPU include parallel prefix sum, vector split, vector histogram, vector compact, vector replace, vector reduce, vector shift insert, vector gather, vector scatter, etc. Executing a vector duplicate count makes it possible to identify the presence of duplicate values, which can be used to uniquify vector inputs and avoid redundant processing.
[0127] Aspects of the present disclosure can provide the following technical advantages. Hardware circuits implementing XPUs can provide more flexible and programmable hardware for embedded class workloads, and other data-dependent operations that cannot be efficiently parallelized. The XPU provides an acceleration path for different classes of data-dependent operations for each workload without the need for the XPU to be fixed to perform only specific operations efficiently. By providing programmable units as described herein, the implementing hardware circuits can robustly adapt to the requirements of various workloads and complement parallelizable data-independent SIMD operations, which otherwise may be inefficient or ineffective for workloads requiring data-dependent operations.
[0128] Hardware circuits such as application-specific integrated circuits can also be designed with various amounts of XPUs to further scale, condition, and distribute the workload. The XPUs described herein also enable the efficient execution of multiple operations using the same configuration, further reducing processing time and configuration time. For example, the XPU can be configured to perform both vector sorting and vector duplicate counting instead of separate configurations of the XPU and / or separate instances of dedicated circuits to accelerate those operations.
[0129] FIG. 7 is a block diagram of an exemplary XPU700. The XPU700 includes processing cells 701-709, crossbars 703-711, and a control cell 750 (represented by the hatched blocks in the block diagram of FIG. 7). Data progresses from bottom to top along data processing lanes 700A-700H, starting at stage 1 and ending at stage 6. Stage 1 includes processing cell 701 and crossbar 702. Stage 2 includes processing cell 703 and crossbar 704. Stage 3 includes processing cell 705 and crossbar 706. Stage 4 includes processing cell 707 and crossbar 708. Stage 5 includes processing cell 709 and crossbar 711. Stage 6 includes processing cell 711 and crossbar 712. In another example, the XPU may include more or fewer stages. The XPU may also include crossbar 799.
[0130] For the sake of explanation, a more previous stage is considered to be "upstream" with respect to a more subsequent stage, and a more subsequent stage is considered to be "downstream" with respect to a more previous stage. For example, stage 1 is upstream of stage 5, and stage 4 is downstream of stage 3.
[0131] The crossbar at each stage of the XPU can be any type of circuit configured to replace various input values from each lane to various other processing lanes according to the current configuration of the crossbar. The crossbar can receive one or more control signals from the control cell for each processing cell in the same stage as the crossbar. The crossbar is configured to replace the input values from each processing cell within the same stage according to a fixed pattern. That pattern depends on the complex operation that the XPU is currently configured to execute and does not necessarily cause the crossbar to replace all processing cell outputs. In other words, some processing cell outputs can bypass the crossbar and proceed to the next stage along the same processing lane.
[0132] To configure a processing cell, each processing cell of the XPU700 has respective control cells 750 configured to receive one or more control signals along each processing lane in which the processing cell resides. The processing cell is configured with circuitry to perform a variety of primitive operations and execute those operations according to received control signals or instructions, as will be described in more detail with reference to FIG. 7. The control cell receives instructions along the data processing lane as one or more signals that can be translated by the control cell, for example, to determine which primitive operation its corresponding processing cell is to perform. The control cell can transfer a control signal to the processing cell or transfer a generated control signal configured to be received by the processing cell to process the received instruction or signal to enable or disable execution of a specified primitive operation.
[0133] The processing cell can also be configured to bypass input data received from its respective processing lane for the processing cell. When received, the input is passed to the crossbar from the processing cell at the same stage without being modified and bypassed. Input received from the crossbar of the previous stage by the bypass processing cell can be associated with zero or ignored. The actual behavior of the bypass processing cell can depend on the pipeline stage in which the cell is located and / or the processing lane in which the processing cell is located. FIG. 7 shows an exemplary processing cell configured to perform comparison, arithmetic, and / or bypass primitive operations.
[0134] The XPU700 may be configured to receive instructions defined as part of an instruction set architecture or an extension of an instruction set architecture that a processor implementing the XPU700 is adapted and executed to execute. These instructions can specify various composite operations and / or primitive operations that the XPU and individual processing cells are each configured to execute as corresponding operations. The control cell 750 is configured to receive data representing instructions defined as part of the instruction set or extension, and / or to convert the instructions into control signals for configuring the corresponding processing cells. For example, the control cell 750 can receive signals as the operation codes of an instruction set corresponding to the processor or hardware circuit implementing the XPU700, that is, the code words for the operations configured to be executed by the XPU. When the XPU700 receives instructions for executing composite operations such as vector sorting or vector duplicate counting, the XPU700 can configure each processing cell to execute the respective primitive operations for causing the XPU to execute the instructed composite operation.
[0135] The operations executed by the XPU can be synchronized by clock cycles. For example, the operations executed by the processing cells for each stage can be executed in one or more cycles. For example, the operations for each stage can be executed in a single cycle. The various composite operations executed by the XPU can require various amounts of clock cycles to execute. For example, vector sorting can be executed by the XPU in 6 cycles, by vector prefix sum in 4 cycles, and by vector compact in 2 cycles.
[0136] As will be described in more detail with respect to FIG. 7, the processing cell can be configured to execute arithmetic operations such as addition between various types of operands including floating point values and signed or unsigned integers. Arithmetic operations such as addition can form part of the composite operations executed by the XPU for the scan operation.
[0137] Exemplary instructions include instructions to reset the XPU and to fetch information regarding clock synchronization and primitive operations performed by the XPU. Other instructions include instructions to fetch one or more of both operands, mask values, and / or segment markers from each processing lane. These instructions may include instructions to access data structures stored by the XPU, along with control information specifying each of a wide variety of composite operations supported by the XPU. In a further example, these instructions may include instructions to cause the XPU to push data into various registers, latches, or flip-flops to determine whether the above is valid. The pushed data may include, for example, values being processed as part of the execution of a composite operation, and / or mask values.
[0138] The configured XPU700 is said to implement a processing network for performing a particular composite operation. For example, XPU700 may include 48 XPU cells configurable as follows. That is, 18 cells are configurable for arithmetic operations, 38 cells are configurable for comparing input values (one cell may be configured for both arithmetic and comparison operations), and 10 cells are configured to bypass the input. In response to new instructions, XPU700 can reconfigure itself with a new processing network to perform a different composite operation.
[0139] The XPU can be configured to operate in a variety of arithmetic modes that can be specified as various instructions in an instruction set or extension. The variety of arithmetic modes can include various complex operations for sorting duplicates, counting, scanning data for partitioning, and / or identifying unique values within the data input to the XPU. Further, these instructions can include operands that specify the type of comparison or arithmetic operation to be performed, e.g., comparison of unsigned integers or floating-point addition for sorting or scanning, and the like. Other operands for instructions to perform complex operations include specifying from which processing lane the output of the complex operation should be sent out from the XPU700. Other received operands can include, for example, segment markers for performing complex operations on segments of the input data in order to sort each of a plurality of segments of data received by the XPU700 across data processing lanes.
[0140] When performing vector sorting and / or vector duplicate counting, the XPU700 is configured to include an odd / even merge network and a value shuffle network. The network configuration includes one or more stages of the XPU, and each cell and crossbar of each stage is configured to perform one or more primitive operations.
[0141] The XPU700 can include register files 760A and 760B. The register files 760A and 760B can be coupled to the data processing lanes 700A - 700H between different stages and can be used to store and retrieve data. For example, some data can be stored in the register file 760B after the processing cell 707 in stage 4, while the data output by the XPU700 is stored in the register file 760A.
[0142] Stream Instructions and Ordering Aspects of the present disclosure provide a hardware or software interface for asynchronous data movement between off-core memory and core-local memory, where the movement between memories is referred to as "stream transfer". Stream transfer may include a stream descriptor that enables software to represent common data movement patterns as seen in sparse workloads. The data may be referred to as a stream or data stream. Stream transfer can be initiated by a stream instruction. The stream instruction can encode information necessary for the execution of the stream transfer. Each stream may have an associated stream identifier (ID) indicated by a data synchronization flag ("sync flag") associated with the stream instruction. Stream instructions issued by cores having the same stream ID can form at least partially a single stream.
[0143] A stream descriptor is an internal data structure that can represent information necessary for the execution of a stream transfer. For example, the information may include a source address, a destination address, a stream operation code, and control information such as a linear or circular buffer.
[0144] Stream transfer may only move data to or from core-local memory. Additionally, a core-local synchronization flag may be used to track the progress of a stream. The synchronization flag can track the partial progress of a stream transfer. For example, depending on whether the flag is cleared or set according to a predetermined configuration, the synchronization flag tracks reads from core-local memory when core-local memory is the source, and writes to core-local memory when core-local memory is the destination. The progress of reads and writes to off-core memory may not be tracked, but a scalar fence instruction can be used to enable the selection of which memory accesses a barrier to ensure that outstanding writes to off-core memory are committed.
[0145] Stream transfer may involve indirect scatter-gather memory access in either off-core memory or core-local memory. The address to the source or destination of the access can be stored at a different memory location relative to the address first read. As an example, the indirect address can be sourced from a register file with masking support or sourced from memory. Indirect scatter-gather memory access may further include various addressing modes such as row address or word address. Stream transfer may include support for direct ScatterAdd / GatherAdd mode for memory words. Memory words can be updated atomically. As an example, data types of 32-bit floating point, 32-bit integer, 16-bit floating point, and 16-bit integer can be supported.
[0146] Stream transfer may include support for a circular buffer within the source or destination buffer, thereby simplifying the buffer allocation problem for software. This is because the buffer size is not known during software compilation.
[0147] Furthermore, a stream ordering model is disclosed herein in general terms. Synchronization primitives for these transfers allow the data transfers to be processed in order while the actual data transfers are out of order. Discrete stream instructions issued by cores with the same stream ID form a single stream. The hardware guarantees the ordering for transfers within a single stream, which can span multiple stream instructions.
[0148] Stream instructions belonging to a stream are processed in order. In the case of indirect stream instructions, the offset list is ordered. For example, offset elements within the offset list are processed in order. Writes are issued to the destination memory in order and can be committed out of order by the destination memory. Reads are issued to the source memory in order and can be served out of order by the source memory.
[0149] The synchronization flag is updated to indicate a monotonic progress for the stream. When the core-local memory is the source, the synchronization flag tracks reads to the core-local memory. The synchronization flag value of N when the core-local memory is the source of the read indicates that the first N chunks of data are overwritable within the core-local memory, where a chunk of data is a predetermined unit of size for measuring data within the core-local memory. When the core-local memory is the destination, writes to the core-local memory are tracked by the synchronization flag. The synchronization flag value N in an example where the core-local memory is the destination indicates that subsequent reads to the first N chunks of data within the core-local memory will return the requested data.
[0150] A stream can terminate when the data for requests preceding the last stream descriptor and for the request including the last stream descriptor are fully committed to memory. As an example, when the core-local memory is the source, the stream can terminate when all reads are complete. When the core-local memory is the destination, the stream can terminate when all writes are committed.
[0151] Aspects of the present disclosure enable software to more efficiently represent common data movement patterns, specifically those seen in sparse workloads. Aspects of the present disclosure can also provide an effective solution to complexity while hiding long memory access latency while maintaining an in-order core computational core and software programming model.
[0152] Stream transfer enables tile 102 and tile sequencer 106 to move data between tile-local memory such as memory 306 or scalar memory 334 and off-tile memory such as memory 105 or high-bandwidth memory 107. Tile-local memory is an example of core-local memory as physically local memory to sparse accelerator 103 such as memory 306, scalar memory 334, etc. Since memory physically remote from TEC 330 or TAC 332 may include memory 105 and / or high-bandwidth memory 107, off-tile memory is an example of off-core memory. Each stream has an associated stream ID indicated by a synchronization flag 318 associated with the stream instruction. Discrete stream instructions having the same stream ID form a single stream having the shared stream ID.
[0153] Data can be moved to or from the tile local memory by stream transfer. The tile local synchronization flag 318 is used to track the progress of the stream. The synchronization flag 318 tracks the partial progress of the stream transferred to and from various memories via the tile 102. For example, the synchronization flag 318 tracks a read operation (or "read") from the tile local memory if the tile local memory is the source, or the synchronization flag 318 tracks a write operation (or "write") to the tile local memory if the tile local memory is the destination. The progress of reads and writes to the off-tile memory need not be tracked. To ensure that all outstanding writes to the off-tile memory are committed, a scalar fence instruction can be used to enable selection of which memory accesses the barrier. The scatter-gather engine 322 tracks the status of stream transfers issued for each specific memory and communicates this status to the scalar core 320. When a scalar fence is issued for a barrier on a specific memory, the scatter-gather engine 322 waits for the status to indicate that all outstanding stream transfers targeting that memory (read or write) are fully committed. When that condition is met, the fence wait state is released on the scalar core 320.
[0154] Stream transfer can support efficient scatter-gather operations using a strided stream for accessing off-tile memory and an indirect stream for accessing off-tile memory from tile-local memory or a register file. Whether it is a strided stream or an indirect stream can be based on the software access pattern. If the software desires to access every Nth element within a tensor, a strided stream is preferred, but an indirect stream can still function. However, if the software desires to access a random set of elements within a tensor, an indirect stream should be used. Stream transfer can also support circular buffer semantics on tile-local memory.
[0155] Stream transfer supports the following data movements, where the granularity and alignment of the data movement depend on the pair of memory source and destination. Data can be transferred from memory 306 to on-chip memory 105 and from on-chip memory 105 to memory 306. Data can also be transferred from memory 306 to high-bandwidth off-chip memory 107 and from off-chip memory 107 to memory 306. Data can further be transferred from scalar memory 334 to on-chip memory 105 and from on-chip memory 105 to scalar memory 334. As an example, the minimum granularity, source alignment, and destination alignment can be 4 bytes. As another example, 32-byte accesses can be used to support 4-byte accesses to off-chip memory 107. As yet another example, 32-byte alignment and 128-byte minimum length can ensure the performance for streams to or from off-chip memory 107.
[0156] Each tile within the processor can implement its own scatter-gather engine (SGE) to coordinate the movement of data from the tile to the scratchpad memory and / or the movement of data between various memories. These various memories can include scratchpad memory, off-tile memory including high-bandwidth memory, and on-tile memory. The SGE can support multiple outstanding stream requests that may originate from one or both of the TEC and TAC of the tile implementing the SGE. Requests to read data can be processed as gather operations by the SGE, and requests to write data can be processed as scatter operations. The SGE can also be implemented in the tile sequencer to handle reads and writes to data streams between the sequencer and the tile and / or memory.
[0157] FIG. 8 is an exemplary functional diagram of a scatter-gather engine 1500. The SGE 1500 can support a varying number of execution threads, e.g., eight threads. The threads can be selected based on a stream identifier for an incoming request and the availability of an address generator thread and stream type, e.g., high-bandwidth memory and / or scratchpad memory. In some examples, a stream request with an in-machine stream identifier in an address generator thread needs to be mapped to the same thread.
[0158] The SGE 1500 can ensure the ordering for a particular type of request (e.g., a gather request or a scatter request) across multiple streams that belong to the same stream and target the same external interface (e.g., the processor's crossbar or other interconnect) originating from the same tile core. These requests can be identified as belonging to the same stream if they have the same stream identifier. Synchronization flags managed on the tile can be used to track the ordering between requests as described herein.
[0159] The SGE1500 receives scatter / gather requests, expands the requests, moves data on the processor interconnect to a remote Spmem slice or to a high-bandwidth memory in the core memory network CMN, sends a synchronization flag update to the tile synchronization flag memory of one of the cores, and updates this according to the progress of the transaction.
[0160] The SGE1500 also services DMA requests received from the CMN interface, performs writes to and reads from the tile Spmem slice, and processes reads and writes created by the SGE of a remote tile targeting the Spmem slice local to this tile.
[0161] The stream request can be one of several different types. The stream request can include a scatter / gather request from the core of a tile. These requests can be processed by a scatter / gather engine that can itself be implemented on the tile and / or tile sequencer. One type of stream request is a linear request, in which case the SGE can expand the request into a plurality of smaller requests based on the length of the stream request. Another type of stream request is a strided request in which the SGE expands the request into a plurality of requests based on the stride and length. Another type of stream request is an indirect request in which the SGE expands a list of addresses of the same length and they are expanded into separate requests. Another type of stream request is an indirect request in which the SGE receives a list of addresses.
[0162] The SGE1500 can implement several stages, for example, a descriptor dispatch stage 1500A, an address generator stage 1500B, and a data transfer stage 1500C. Each stage will be described in turn.
[0163] SGE can interface directly with descriptor generators within each tile core, such as TAC and TEC. The descriptor generator is configured to enqueue the stream descriptors generated by the core. The stream descriptors are sendable to the descriptor generator, where metadata related to the descriptors is enqueued into the descriptor issue FIFO and the actual descriptors are written into the descriptor RAM. The content of the FIFO may include a pointer within the descriptor RAM where the descriptor resides, the memory type, and the stream identifier attached to the corresponding stream. If valid stream descriptor metadata is available at the head of the TAC or TEC FIFO, the following sequence may occur.
[0164] First, SGE may perform a resource check on the metadata of each core. Next, assuming that both cores have the required resources as determined by the resource check, SGE can perform the longest unused arbitration to be selected between these two cores. If only one of the cores has the resources required, this will be the next core to receive service. The stream identifier attached to the metadata is looked up in the stream identifier in the address generator map and the associated counter is incremented, or a new entry is created in the map to select the address generator thread at the address generator stage which is the destination to send the descriptor. Then, the metadata is placed into the descriptor metadata queue associated with the selected thread. The resource check can be performed in the descriptor FIFO associated with the descriptor metadata queue. The longest unused arbitration selects one of the descriptor metadata queues. In this case, the winning entry is popped and its metadata is transferred to the descriptor RAM request FIFO. The request is popped and sent to the descriptor RAM, and the data response stored in the descriptor FIFO is transferred to the address generator stage.
[0165] The stream identifier-address generator map is used to calculate the address generator thread within address generator stage 1500B that should transmit the descriptor metadata, and to ensure that the ordering between subsequent stream requests belonging to the same stream is maintained. This structure can hold an active stream identifier, the address generator thread to which these identifiers are mapped, a bit indicating that the mapped stream identifier to the thread identifier entry is valid, and a count of the number of descriptors belonging to this stream in the pipeline from the queue of the stream FIFO.
[0166] The above structure can be sized to hold various maximum stream identifiers, for example 16. Each time a stream descriptor is issued, along with the stream identifier that is not currently active, the available space within the descriptor metadata queue, and / or the available space within the stream identifier-address generator map, the SGE1500 can perform one or more actions. The SGE selects the next available thread, stores this value in the stream identifier-address generator map, and increments the counter associated with this thread by one. All subsequent results for the same stream will be sent to the same address generator thread here and increment the counter.
[0167] If there is no available space in the descriptor metadata queue for this request for that address generator thread, or if there is no space to allocate a new entry in the address generator map for a new stream identifier (if this is a new stream identifier), or if the counter within the map is almost full, the FIFO will not be dequeued until space becomes available within all necessary resources. This check is performed prior to arbitration between the two core FIFOs.
[0168] Among the address generator threads in the Address Generator Stage, 4 are associated with the CMN interface and the other 2 are associated with the SC data crossbar. The CMN interface address generator threads expand requests targeting HBM, while the SC data crossbar interface address generator threads target remote Spmem or tile SpmemN. To determine the next available thread to store along with the stream ID, the stream interface metadata determines which address generator thread to select from, and LRU arbitration is used across the relevant threads that also have space in the corresponding descriptor metadata queue.
[0169] When a stream request is expanded by an address generator thread within the address generator stage and the synchronization flag tracking structure within the data transfer stage is updated, the counter within the stream identifier in the address generator map can be decremented. This decrement update to this counter from the data transfer stage originates from the following four parts of the data transfer pipeline. Namely, stream scatter to the core memory network (xl), stream gather to the core memory network (xl), stream scatter to the processor data crossbar (x2), and stream gather to the processor data crossbar (x2). This is done by the last request to be expanded into a descriptor by the address generator thread.
[0170] When the counter associated with the stream identifier in the stream identifier in the address generator map reaches zero, this map entry can be invalidated, and any subsequent requests to the same stream can be remapped to any of the available address generator threads for that interface (CMN / data crossbar), but this will ultimately reach the same synchronization flag tracking structure that maintains the order between them.
[0171] For example, if the descriptor is in the aircraft (up to the data transfer stage), the stream ID is first mapped to thread_iD#2 over a time window. The mapped entry will be invalidated if no new descriptors are created with the same stream ID for some time (the associated counter reaches zero). After the mapped entry is invalidated, if a new descriptor is created with the same stream ID, the new map created can have the stream ID mapping thread_id #1, but this should not be a problem as there is already space allocated in the synchronization flag tracking structure within the data transfer stage that serves to maintain ordering. Note that this is assuming that thread_id#1 and thread_id#2 belong to the same stream interface type (CMN / data crossbar) and are tracking updates to the same memory (scatter (scatter) vs gather (gather)).
[0172] Since all information from the stream FIFO is only used in part of the initial stage of the descriptor dispatch stage 1500A, it does not need to be enqueued in the descriptor metadata queue. Only the address for the descriptor ram is required in the latter part of the stage to issue a read to the descriptor ram. All other metadata being carried in the FIFO is only used to map the descriptor to a specific address generator. After this point, this metadata can be discarded. The descriptor metadata queue only carries pointers to the descriptor ram.
[0173] At the output of the descriptor metadata queue, there is LRU arbitration across all address generator threads to obtain access to the descriptor RAM in order to maintain fairness.
[0174] Referring to the address generator stage 1500B, descriptors from the descriptor FIFO are popped by the stream descriptor manager logic within the address generator stage and then passed to the address expansion state machine for further expansion into requests to remote memory. Each sub-block of the address generator will be described in more detail below.
[0175] The stream descriptor manager logic is a state machine that expands the address list of an indirect stream before passing the descriptor to the address expansion state machine. Although this specification describes the logic in the form of a state machine, it should be noted that in an actual implementation example, the states are not explicitly labeled in this way. This is done to simplify the code structure. However, the actual functions executed by the logic do not change and are the same as those described in this specification. This performs the following tasks in each state.
[0176] Idle state: In this state, the logic pops from the descriptor data FIFO and checks the off_tile_stream_type field of the descriptor. In the expanded address list state, this state is reached only when, for example, there are more than 4x addresses to be expanded in the address list of an indirect stream.
[0177] The address expansion state machine is used to expand each stream descriptor presented to it into one or more read / write requests to the core memory network interface (directed to HBM) or the data crossbar interface (directed to remote Spmem).
[0178] Note that the address expansion state machine must indicate to the data transfer stage the stream scatter or gather that is the last request expanded from each stream descriptor. This information is passed from the stream descriptor manager logic to the address expansion state machine as part of the descriptor metadata by indicating the last entry in the address expansion input FIFO associated with a particular stream descriptor. This information is required by the data transfer stage to track the "committed" and "retired" states of the descriptor for fence instructions.
[0179] In the data transfer stage 1500C, the SGE1500 addresses the format settings of the scatter and gather of the output stream to meet the interface requirements of the CMN and the data crossbar. It manages the DMA access to Spmem and the synchronization flag tracking for the scatter and gather of the unprocessed streams. This pipeline stage also caters for the incoming read and write accesses from the remote tiles.
[0180] The SGE1500 increments a transaction counter to track the status of the fence instruction based on the source core ID and the target remote memory. For fences, for example, there could be six counters per core type [total 12] - Spmem (retired write, committed write, retired read), HBM (retired write, committed write, retired read). Each transaction that wins arbitration to the data transfer stage will increment both the retired and committed counters for the associated memory and core type.
[0181] For the fence descriptor counter present there, if this is the last transaction associated with the descriptor being expanded, SGE1500 sends a decrement to the descriptor dispatch stage. This can also be done for the first transfer associated with the descriptor, but since the "last transfer" status is already available, no additional information needs to be tracked if the decrement is done based on the last transfer. This also ensures that an incorrect status is not provided to the core complex. The decrement is performed at the granularity for the next interface transfer.
[0182] If this is the last transaction associated with the descriptor being expanded, SGE1500 sends a decrement to the descriptor dispatch stage regarding the synchronization flag id to the address generator map.
[0183] Since the committed state of the descriptor regarding the stream gatherer is the same as the retired state of the descriptor, the stream gatherer needs to be tracked by only one type of status counter. In the case of the stream scatterer, two counters will be updated at different points in the flow. Also, since the granularity of the update in the pipeline that updates the synchronization flag may be different from the granularity at which the request is tracked up to the remote memory, the logic needs to maintain the state in these two separate parts of the data transfer stage.
[0184] The data transfer stage 1500C maintains a counter for each source core for each remote memory type that is incremented by each instance of the synchronization flag tracking logic. The counter is incremented when the synchronization flag message is enqueued to the message router interface by the synchronization flag tracker. The synchronization flag update sent to the message router associates the remote memory and the source core with the transaction to which the update belongs and is sent back to the SGE when the synchronization flag update is complete to decrement these counters in the data transfer stage.
[0185] Each synchronization flag tracker also maintains a status as to whether there is an ongoing transaction being tracked thereby (per source core per memory type).
[0186] As long as there is a synchronization flag tracker that is "in progress" for the stream scatterer to a particular interface, all descriptors for that memory type are not yet "committed" or "retired". Note that if the "retired" counter for the stream scatterer for a particular memory type is 0 but the associated synchronization flag tracker still has an "in progress" transaction, the fence status for that memory will still indicate that not all are "committed" or "retired".
[0187] In the case of a stream gatherer, even if all "retired" counters are decremented, as long as there is an "in progress" transaction within the synchronization flag tracker, the current status for the descriptor dispatch stage cannot be set to "committed" or "retired". When the synchronization flag tracker for a particular memory reports that nothing is being tracked there, the status can be updated to the descriptor dispatch stage.
[0188] The descriptor tracking logic in data transfer stage 1500C sets the ongoing fence status to descriptor dispatch stage 1500A.
[0189] As described herein, the SGE1500 can be part of a tile sequencer and can be implemented in each tile. The scatter-gather engine within the tile sequencer will be a parameterized version of the SGE subsystem described above, although with some differences detailed in this paragraph. The "local" memory in the case of the tile sequencer is always shared memory and has lower read / write bandwidth requirements compared to scratchpad memory.
[0190] A stream descriptor is a data structure that can represent all the information for the scatter-gather engine 322 to perform a stream transfer. A stream instruction can fully encode the fields for a stream descriptor. The following are examples of fields for a stream descriptor.
[0191] In the case of a stream operation code, a gather stream reads off-tile memory and stores data in or adds data to tile-local memory. A scatter stream reads from tile-local memory and stores data in or adds data to off-tile memory. The off-tile memory and tile-local memory are determined by fields such as the off-tile memory type and tile-local memory type, respectively.
[0192] Additional variations of stream instructions can support both floating-point and signed integer addition operations. It is possible to support gather (scatter) of signed integer addition and gather of variations of floating-point addition for tile-local memory. It is possible to support scatter signed integer addition and scatter variations of floating-point addition for off-tile memory and tile-local memory. If an improper combination is detected, a program error can be presented by the engine.
[0193] The tile local stream type indicates the address pattern used to access tile local memory. For example, a linear stream facilitates access to several consecutive words starting from a tile local start offset. The number of consecutive words can have a 4-byte length. As another example, a circular buffer stream enables software to construct a logical circular buffer within tile local memory. In this exemplary access pattern, the base field, size field, and offset field within the circular buffer metadata are used to generate addresses for several words. The number of words can have a 4-byte length. If the effective length of the granularity is greater than the size field within the circular buffer metadata, a program error may be presented.
[0194] The off-tile stream type indicates the address pattern used for access to off-tile memory. A linear stream facilitates access to several consecutive positions starting from an off-tile start offset. The actual word size depends on the off-tile memory type. A strided stream facilitates converting a strided access pattern into a multi-dimensional array stored in the off-tile memory type. The stream transfer can support a single level of stride. An indirect stream enables a random scatter-gather access pattern to a table. Here, an indirect offset list is used, and each entry in the list accesses data of the same length.
[0195] The source of the indirect offset list can be tile local memory or a register file. When the source is tile local memory, the indirect offset field has a start offset to the tile local memory where several offsets are stored. When the source is a register file, the indirect offset field has a register file and indicates several lanes containing valid offsets. These offsets are used to perform a scatter operation or a gather operation as indicated by the stream operation code.
[0196] The core type indicates the type of core within tile 102 that generated the stream descriptor, such as tile executor core 330 or tile access core 332. The synchronization flag core type indicates the core type of synchronization flag 318 that tracks the progress of the stream. Encoding can be the same as the core type that enables the start of the stream by the tile access core tracked by the tile executor core and enables the start of the stream by the tile executor core tracked by the tile access core.
[0197] The synchronization flag ID indicates an offset within the target synchronization flag memory. The synchronization flag ID can also be used as a stream ID and can ensure ordering, as further described below. The set done bit indicates that the current descriptor is the last within the stream. The done bit is set after all data for the current descriptor and preceding descriptors within the stream have been fully committed to the tile local memory.
[0198] The synchronization flag count type indicates the type of count that synchronization flag 318 is tracking, whether it be the number of words or the number of descriptors. In either case, synchronization flag 318 tracks the monotonic step - by - step progress of the stream at different granularities.
[0199] The tile local memory type indicates the type of tile local memory involved in the stream transfer and can include scalar memory 334 or a local bank of memory 306.
[0200] The tile local start offset field is used when the tile local stream type is linear. It indicates an aligned start offset word, such as a 4 - byte word, within the tile local memory accessed by this transfer. The actual access type depends on the stream operation code.
[0201] The tile local stride encodes the stride size and number of bytes accessed per stride, which is used to access the tile local memory selected by the tile local memory type. The length, which can be, for example, 4 bytes, does not have to be a multiple of the number of bytes accessed per stride. The last request for a strided access will access the remaining transfer words, which can be less than the length per stride. The stride calculation can be the same for both linear and circular buffer stream types.
[0202] When the tile local stream type is a circular buffer, the circular buffer metadata field is used. The size of the circular buffer can be a multiple of the granularity of the off-tile memory type and the offset can be aligned. When the circular buffer wraps around, the request is split into multiple requests and an error can be presented if the resulting requests are not multiples of the granularity of the off-tile memory type. An error can also be presented if the total length of the stream transfer is greater than the size of the circular buffer.
[0203] The off-tile memory type indicates the type of off-tile memory involved in the transfer. This includes on-chip memory 105 and high-bandwidth memory 107. A high-bandwidth memory view that allows access with 4-byte granularity and 4-byte alignment can also be used. If sequencer 106 is the initiator of the stream transfer, this field does not have to have an encoded high-bandwidth memory 107.
[0204] The tile ID field can be used to select the tile ID for memory slicing. The off-tile start offset includes the start offset word within the off-time memory 105 indicated by the associated off-tile memory type. The unit of the offset can be equal to the value indicated in the offset alignment column within the off-tile memory type. For example, in the case of the high-bandwidth memory 107, an offset value of 1 would be translated to a byte address of 32. If the off-tile stream type is indirect, this field can function as the base address that is added to the offset read from the indirect offset list before accessing the memory.
[0205] An indirect offset can be used if the off-tile stream type is indirect. If the source is tile-local memory, the indirect offset provides the word start offset within the tile-local memory that stores the indirect offset list. If the source is a register file, the indirect offset provides the file register index that is the source of the indirect offset list. The register file can be read at the time of issuing the stream instruction.
[0206] If the off-tile stream type is indirect, the indirect list size can be used. If the source is tile-local memory, the number of elements within the offset list is stored in the tile-local memory. If the source is a register file, the number of lanes that contain valid offsets is stored. The completion of the transfer is maintained in order with the remaining descriptors within the stream.
[0207] If the off-tile stream type is indirect and indicates the type of the offset stored in the offset list, the indirect list type can be used. This can include word offsets and row offsets. The indirect list stride is used if the off-tile stream type is indirect and indicates the distance between two address words within the offset list stored in the tile-local memory. This can be a signed integer.
[0208] The indirect filter field is used when the off-tile stream type is indirect. When this field is set, the indirect memory addresses that match the indirect filter value are filtered out. The indirect filter value indicates the value of the elements in the indirect access list that need to be filtered out. This value is of the type indicated by the indirect list type. When the indirect filter field is set for an indirect stream and / or when the value of an element in the indirect offset list matches this field, filtering can be enabled. The off-tile access and tile-local access corresponding to the filtered elements are dropped, but the tile-local buffer will still be advanced by the size of the filtered access.
[0209] Lengths such as 4 bytes or a multiple of 512 bytes, for example, indicate the total number of words accessed by the stream. When the off-tile stream type is linear or strided, this field indicates the total number of words accessed by the stream. When the off-tile stream type is indirect, this field indicates the number of words accessed from each address in the indirect offset list. A program error may be presented if the actual value of this field is not a multiple of the granularity of the off-tile memory type. A program error may also be presented if the generated address exceeds the boundary of the off-tile memory 105.
[0210] The stride size field indicates the stride size in units of the granularity of the off-tile memory type. This can be a signed integer. For example, the length per stride such as a length per stride that is a multiple of 4 bytes or 512 bytes indicates the number of words accessed per stride. This is a signed field but should contain non-negative values. The length need not be a multiple of this field. This field should be a multiple of the granularity of the off-tile memory type selected by this stream descriptor. The last request for a strided access will access the remaining transfer words which can be less than the length per stride. If the length per stride is 0, negative, or not a multiple of the off-tile memory access granularity, a program error may be presented. A program error may also be presented if the generated address exceeds the boundary of the off-tile memory 105.
[0211] The trace field indicates whether to trace the stream transfer. Tracing may include logging information about the actions taken during the stream transfer as part of debugging.
[0212] FIG. 9 is a flowchart of an exemplary process 1600 for expanding a stream descriptor into a configured off-tile stream request or a tile-local stream request. The exemplary process 1600 can be executed on a system of one or more processors at one or more locations. For example, the hardware circuit 101 can execute the process 1600 as described above.
[0213] As shown in block 1610, the process includes a process of receiving the size of the off-tile memory, such as receiving a 4-byte size. Further, the process includes a process of receiving the maximum chunk size of a stream request targeting the off-tile memory type. As shown in block 1620, the process further includes a process of converting an indirect offset read at an offset such as a 4-byte offset from a file register or tile-local memory to the off-tile memory based on the indirect list type.
[0214] As shown in block 1630, the process also includes a process of generating a strided request and / or an indirect request. In the case of a strided request, the process may include a process of partially expanding a strided stream descriptor into a set of requests each accessing consecutive addresses within the off-tile memory. In the case of an indirect tile-local memory request, the process may include a process of obtaining an indirect stream descriptor and generating a list of offsets to the off-tile memory type selected by the descriptor. In the case of an indirect file register memory request, the process may include a process of generating a list of offsets from the file register read at the issuance of an indirect stream instruction.
[0215] As shown in block 1640, the process includes a process of generating a list of expanded off-tile memory requests, where each expanded request accesses a set of consecutive addresses within the off-tile memory. These requests are used to generate both tile-local memory requests and off-tile memory requests. The tile-local stride, tile-local stream type, and alignment are considered during the expansion of the requests. The process further includes a process of generating a list of partially expanded requests, where each partially expanded request accesses a set of consecutive addresses within the off-tile memory. These requests are further expanded to generate a set of requests aligned to the memory granularity selected by the off-tile memory type.
[0216] As shown in block 1650, the process includes a process of expanding a stream descriptor into a set of off-tile memory requests and tile-local memory requests.
[0217] FIG. 10 is a flowchart of an exemplary process 1700 for ordering stream transfers. The exemplary process 1700 can be executed on a system of one or more processors at one or more locations. For example, the hardware circuit 101 can execute the process 1700 as described above. Discrete stream instructions issued by cores having the same stream ID form a single stream, but ordering may not be guaranteed between separate streams. The scatter-gather engine 1722 includes a plurality of threads that can process these requests in parallel. Ordering can be guaranteed for transfers within a single stream, which can span multiple stream instructions.
[0218] As shown in block 1710, stream instructions belonging to a stream are processed in order. The corresponding requests will be issued in order by the scatter-gather engine 322.
[0219] As shown in block 1720, in the case of indirect stream instructions, the offset list is ordered. The offset elements within the offset list are processed in order. Writes are issued to the destination memory in order, but the writes may be committed out of order by the destination memory. Reads are issued to the source memory in order, but the reads may be serviced out of order by the source memory.
[0220] As shown in block 1730, the scatter-gather engine 322 updates the synchronization flag 318 to show a monotonic step-by-step progression for the stream. When the tile local memory is the source, the synchronization flag 318 tracks the reads from there. The synchronization flag value indicates the first chunk of data among the chunks of data that can be overwritten to the tile local memory. When the tile local memory is the destination, the synchronization flag 318 tracks the writes to there. Here, the synchronization flag value indicates that subsequent reads to the first chunk of data in the tile local memory will return to the requested data.
[0221] As shown in block 1740, the done bit in the synchronization flag 318 can be updated at the end of the stream. This is indicated by the set done bit in the stream descriptor. The done bit can be set after all data for requests preceding the last stream descriptor and requests including the last stream descriptor have been fully committed to memory. When the tile local memory is the source, all reads are complete, and when the tile local memory is the destination, all writes are committed.
[0222] FIG. 11 is an exemplary diagram of stream ordering. Consider the case where stream descriptor A and stream descriptor B constitute one stream. Stream descriptor B has a set done bit. The partial progress of the stream is tracked by a synchronization flag. When A0 is committed to memory either by read or write, the synchronization flag is updated to a value of 1. Even if A2 and B1 are committed before A0, the synchronization flag value is not updated to 3. When A1 is committed to memory, five consecutive chunks of data in the stream, A0, A1, A2, B0, B1, are committed, which is indicated by a synchronization flag value of 5. The done bit is not set at this point because stream descriptor A is not at the end of the stream. When B2 is committed, the synchronization flag value is set to 6. Here, the done bit can be set when all data chunks of the stream have been committed and stream descriptor B is at the end of the stream.
[0223] Cooperative Prefetch Aspects of the disclosed technology provide an instruction prefetch pipeline architecture that can be used by tiles of a sparse accelerator, which provides good performance without complicating with a full cache coherent solution deployed in a conventional CPU.
[0224] Aspects of the disclosed technology relate to methods and systems for creating a prefetch pipeline around the SPMD aspect of a programming model to reduce cold cache overhead. A prefetch response from any core is broadcast to all cores within the sparse accelerator. These prefetch responses are committed to the core's local cache. This enables other non-requesting cores to obtain a bundle of instructions or data prior to the time when the core would become available to process the instructions or data, completely avoiding process cycle drops. Additionally, there may be prefetch request filtering on the arbitration path, which is a logic and / or hardware-based path for arbitrating between requests, resulting in a task instruction memory, thereby boosting the task instruction memory bandwidth by avoiding redundant request fetches.
[0225] The task instruction memory can hold a set of programs executable by a Tile Access Core (TAC) and a Tile Execute Core (TEC). The program counter (PC) within each core is a physical offset to the task instruction memory. The task instruction memory is a software-managed memory exposed to the direct memory access system. Software can use direct memory access to populate the programs in the task instruction memory and use the appropriate program counter when issuing tasks to the tiles. Since the tiles operate in a multiple data mode of a single program, at any point in the execution of the sparse accelerator, statistically most tiles may be executing the same program. These programs may further be composed of one instruction loop or a compact instruction loop. A compact instruction loop can refer to the size of the instructions in a memory small enough to fit the tile memory. The program itself may be small in size, for example, hundreds of instruction bundles, and the program may have multiple branches that can branch within the loop or to other loops.
[0226] These and other potential features can be utilized by an instruction pipeline as described herein with reference to, for example, FIGS. 12 and 13. When an instruction bundle is received by the sparse accelerator, the sparse accelerator is configured to broadcast a prefetch response regarding the received instruction to all of the tiles within the sparse accelerator. The prefetch response is committed to the local cache of each core, enabling non-requesting tiles to obtain the bundle in advance and completely avoiding misses. Additionally, prefetch request filtering on the arbitration path connecting to the task instruction memory can be implemented in some examples to boost the task instruction memory bandwidth by avoiding redundant request fetches.
[0227] FIG. 12 shows a logic diagram of the connection between exemplary sparse accelerators tiles 1901 and 1902, according to aspects of the disclosed technology. For clarity, not all components, modules, or software blocks related to FIG. 12 are labeled. Generally, as will become apparent from the following description, by prefetching instructions, aggregating requests related to the instructions, filtering instructions or references, and holding them in the memory location closest to the request processing unit, the instructions can be provided to the processing unit or processing core more quickly, and the efficiency of the system can be increased. Additional aspects of the components related to FIG. 12 are further described below in connection with FIG. 13.
[0228] In the schematic diagram, FIG. 12 shows aspects of a task instruction memory (Timem bank or Timem bank), an instruction buffer (“iBuf”), a prefetch unit, and an instruction router. FIG. 12 shows a tile 1901 that includes a tile access core (TAC) 1910 and a tile execution core (TEC) 1920. TAC 1910 may include a prefetch 1911 and an iBuf 1912. Similarly, the TEC may include a prefetch unit 1921 and an iBuf 1922. FIG. 12 also shows a tile core 1902 that includes a TAC 1930 and a TEC 1940, which include a prefetch 1911 and an iBuf 1932, and a prefetch unit 1941 and an iBuf 342, respectively.
[0229] Figure 12 further shows Timem1951 and Timem1952 and instruction router 1960, which may be logically or physically included within floor plumbing block 1999. Timem1951 and Timem1952 can locally store instructions for faster access by each tile core as compared to tile cores that request instructions from a position further downstream from Timem. Also shown is an instruction broadcast bus that can broadcast instruction bundles downstream to floor plumbing block 1999 and the Timem banks therein. Instruction request bus 1992 can aggregate requests for instructions from various components before those instructions are requested. The deserializers and serializers can deserialize or serialize instructions for transmission along various buses, for example, to receive from instruction broadcast bus 1991 or to serialize instructions transmitted to instruction request bus 1992.
[0230] A prefetch unit such as prefetch unit 1911 or prefetch unit 1912 corresponding to a core can make read requests to Timem starting from a misprogram counter (program counter: PC) (and overlay / task ID) to the end of the prefetch window. The prefetch window is a period selectable by software having registers or other memory areas. For example, the prefetch window can be defined by a prefetch depth variable. Prefetch read requests from other tiles can be transferred by an adjacent floor plumbing block 1999. These transferred requests can be arbitrated by prefetch requests made by prefetch units within adjacent tile cores. For example, tile 1901 and tile 1902 may be adjacent to each other. In some examples, a pair of cores can be assigned to a single instruction request bus or a single instruction broadcast bus.
[0231] Several prefetch instruction request banks may exist within a tile. In some examples, there may be one bus per Timem bank, but these are arbitrable independently of each other. Independent arbitration of the buses may enable avoidance of head-of-line blocking across independent banks.
[0232] Requests sent from the prefetch window can be received at instruction router 1960. Instruction router 1960 can filter selected requests to remove duplicates before forwarding them to another instruction router or the target Timem bank. Filtering can potentially increase the instruction request bandwidth when the core is operating in SPMD mode.
[0233] Instructions read from the Timem bank can be broadcast to all tiles on the instruction broadcast bus. For example, there may be as many instruction broadcast buses as there are Timem banks. In some examples, instructions can be sent as instruction bundles. An instruction group is composed of the instructions included in the bundle. A bundle can be a sequence of instructions that starts on an aligned "boundary". Instruction bundles can be serialized on the corresponding instruction broadcast bus over a fixed number of cycles of the processor or core. In some examples, during the "steady state" operation of the system and the like, the total bandwidth of the instruction broadcast bus can be two bundles per cycle. In such a manner, instruction broadcast bus 1992 is never backpressured.
[0234] Instructions received on the broadcast bus can be deserialized by an instruction router, and one instruction is transferred to each of the iBufs. In the steady state, the system may be required to maintain up to two writes from the prefetch interface and one read from the instruction fetch interface. The prefetch processes the received instruction and determines whether it should be committed to or dropped from the ibuf.
[0235] FIG. 13 shows additional exemplary aspects of instruction router 1960. FIG. 13 shows round robin (RR) arbiter 1910, daisy chain round robin arbiter 1920, round robin arbiter 1930, filter 1940, serializers 1950 and 1951, demultiplexer (demux) 1960, and deserializers 1971 and 1972. FIG. 13 also shows other aspects and components not labeled for simplicity.
[0236] Instruction router 1960 may have an independent read request bus for each Timem bank in the system. The router 1960 can throttle the instruction bundle at a rate that matches the bandwidth of the instruction broadcast bus before transferring it to an adjacent instruction router. In the following description, it may be assumed that deserialization and serialization can be performed before the request is presented to instruction router 1960.
[0237] Instruction router 1960 can arbitrate according to the position of the core relative to the Timem bank. Instruction router 1960 can be parameterized to select the source and destination based on the instance that the instruction router 1960 is arbitrating. The demultiplexer 1960 shown in FIG. 13 can be designed according to the number of communicating time banks or serializers.
[0238] The command router 1860 can arbitrate among the following exemplary sources: namely, a prefetch readout transferred by a command router upstream or above the command router 1860, a prefetch readout transferred by a command router downstream of the command router 1860, and a prefetch readout transmitted by a core connected to the command router 1860.
[0239] demux (select (selection) is a design parameter) selects top_pre_req or bottom_pre_req to arbitrate requests originating from cores connected to the command router. This arbitration uses a daisy-chain type RR arbitration method. The daisy-chain round-robin arbiter 1920 can grant permission every "x" cycles to match the bandwidth of the command broadcast bus. If the PC matches the PC seen on the command broadcast bus, requests waiting to be arbitrated can be dropped. This can be regarded as the first level of filtering.
[0240] The winner of the daisy-chain type arbitration can be processed variously based on the position of the command router 1860 relative to the Timem bank. For example, if the Timem bank is below the command router, the winner of the daisy-chain arbitration can be transferred to the "bottommost" command router after passing through the filter 1940. If the Timem bank is above the command router 1860, the winner of the daisy-chain arbitration is transferred to the topmost command router 1860 after passing through the filter 1940.
[0241] When the Timem bank is within the instruction router 1860, the winner of the daisy chain arbitration receives arbitration at one or more levels on the requests forwarded by the bottommost instruction router. In this case, there may be two daisy chain type networks that arbitrate to reach the Timem bank. Depending on the position of the instruction router 1860, there may be an imbalance between the chains. A modified RR arbiter can be used to ensure fair access to the cores on both sides of the chain. Similar to the first level of arbitration, any requests that match the PCs on the broadcast bus will be dropped here. This can be considered as the second level of filtering.
[0242] The overall winner among the above-mentioned ones is passed to the filter 440, and the filter 440 compares the incoming request with one of the other unprocessed requests. If the request matches any of the unprocessed requests, the request is dropped. This can be considered as the third level of filtering.
[0243] Furthermore, the programmability of this system can be ensured by the fact that the filtering at each point in time can be enabled / disabled by individual programmable switches or software-controllable switches. The Timem access bus can be a bus that connects the system to all Timem banks and enables them to read and write instruction bundles to the Timem banks. The Timem access bus can have four buses such as a read request bus, a read response bus, a write request bus, and a write response bus, as further described below.
[0244] The read request bus can be an executable daisy chain type bus up to the Timem bank. Each Timem bank can forward the request to the adjacent Timem bank if the request does not address it. If the request addresses a Timem bank, the request is handled by the Timem bank.
[0245] The read response bus can be a daisy-chain type bus that can transmit an instruction bundle read from the Timem bank. In each Timem bank, there can be round-robin arbitration between an incoming instruction from an adjacent bank and an instruction bundle from the current bank. Since the instruction bundle is serialized over "n" cycles, the bus grant is held over "n" cycles.
[0246] The write request bus can be a daisy-chain type bus that is executable up to the Timem bank. The write request can be serialized, for example, over 2 cycles. Each Timem bank transfers a flit to an adjacent bank if the request does not specify an address. If the request specifies a Timem bank, the request is deserialized by the bank before being written to the Timem bank.
[0247] The write response bus can be a daisy-chain type bus that relays a write response from the Timem bank. In each Timem bank, there is arbitration between an incoming response and a response from the current bank. Simple round-robin arbitration can be used to enable one of the responses to be granted or provided.
[0248] Read requests and write requests can have a "q" -bit tag for encoding up to 2 ∧ pending read requests and write requests of q. These pending read requests and write requests are returned in the response from the bank and can be used by the overall system or component that provides an instruction to identify the request corresponding to the response.
[0249] If an endpoint cannot accept a request or a response, the bus can be "backpressured", and the bus accumulates a backlog to be transmitted over the bus if the bus cannot transfer the instructions or data it contains. Additionally, the bus can be backpressured due to arbitration losses. Since Timem access is generally low-bandwidth access, this can be acceptable throughout the system.
[0250] The tile instruction memory (Timem) can be shared by the tile core described in FIG. 12.
[0251] Aspects of the present disclosure may be implemented as one or more computer programs in a digital circuit, a computer-readable storage medium, or as one or more combinations described above. The computer-readable storage medium may be non-transitory, for example, executable by a cloud computing platform and stored as one or more instructions on a tangible storage device.
[0252] Calculation Overview As used herein, the phrase "configured to" is used in various contexts related to a computer system, hardware, or a part of a computer program, engine, or module. When a system is expressed as being configured to perform one or more operations, this means that the system has appropriate software, firmware, and / or hardware installed on the system that causes the system to perform the one or more operations during operation. When some hardware is expressed as being configured to perform one or more operations, this means that the hardware includes one or more circuits that receive an input and generate an output corresponding to the one or more operations in accordance with the input during operation. When a computer program, engine, or module is expressed as being configured to perform one or more operations, this means that the computer program includes one or more program instructions that cause one or more computers to perform the one or more operations when executed by the one or more computers.
[0253] The operations shown in the accompanying drawings and recited in the appended claims are presented in a particular order, but these operations may be performed in an order different from that shown, and some operations may be omitted, performed multiple times, and / or performed in parallel with other operations. Further, the separation of various system components configured to perform the various operations should not be understood as requiring the components to be separated. The components, modules, programs, and engines described may be integrated together as a single system or may be part of multiple systems.
[0254] As used herein, substantially any plural and / or singular terms (the term "element" is a substitute for any system, component, data, etc.), e.g., "one / the element", "one or more elements", "composite element", "plural elements", "at least one element", etc., can be converted from plural to singular and / or from singular to plural as appropriate by one of ordinary skill in the art in view of the context and / or application being described. Various singular / plural substitutions may be explicitly set forth in this specification without limitation for clarity purposes as long as they are not explicitly specified.
[0255] Aspects of the present disclosure include methods, systems, and apparatuses that use an instruction prefetch pipeline architecture that provides good performance without complicating with a full cache coherent solution deployed in a conventional CPU.
[0256] Aspects of the disclosed technology relate to components that can be used to construct an instruction prefetch pipeline including an instruction memory (TiMem), an instruction buffer (iBuf), a prefetch unit, and an instruction router.
[0257] Aspects of the present disclosure may relate to certain characteristics that may exist in conjunction with the expected or known behavior of tiles of an XPU, for example, as follows. Can a tile be expected to operate in Single Program Multiple Data (SPMD) mode? Or, at any given point in time, can it be expected that statistically most tiles are executing the same program? Can a program be composed of one or more compact loops? Can the size of the program be made small, for example, on the order of hundreds of bundles? Or, can the program have multiple branches that can diverge? The disclosed techniques can leverage these characteristics or related characteristics and have a lower complexity than full-cache-based solutions.
[0258] Aspects of the disclosed techniques include a hardware circuit. The hardware circuit may include a plurality of tiles. Each tile is configured to operate in parallel with other tiles within the plurality of tiles, and each tile of the plurality of tiles includes a processing core, a prefetch unit, and an instruction buffer. The hardware circuit further includes a plurality of data processing lanes configured to stream respective data from an upstream input to a downstream destination, and a plurality of task instruction memories. Each task instruction memory of the plurality of task instruction memories is arranged in a sequence and coupled to one or more of the plurality of tiles via an instruction router. The task instruction memory can be arranged in a downstream sequence. Each tile may include a tile access core, and the prefetch unit included in each tile may be included within the tile access core. Each tile may include a tile execution core, and the prefetch unit included in each tile may be included within the tile execution core. The hardware circuit may include an instruction broadcast bus and an instruction request bus. The instruction broadcast bus may include independent data lanes. The number of independent data lanes may correspond to the number of task instruction memories.
[0259] The command request bus may include independent data lanes. The number of independent data lanes corresponds to the number of task instruction memories. Instructions received by the task instruction memories may be broadcast to all tiles linked on the instruction broadcast bus. The prefetch may be configured to make requests for at least one task instruction memory in the prefetch window. The prefetch window may be selectable or adjustable by software. The hardware circuit may further include an instruction router. The instruction router may include a round-robin arbiter configured to arbitrate requests including prefetch read requests. The instruction buffer can store instructions for the tile access core or the tile execution core. The hardware circuit can be configured as a single instruction multiple data processor. The hardware circuit can be configured as a multiple instruction multiple data processor. The hardware circuit may include a task instruction memory access bus. The task instruction memory access bus may include a read request bus, a read response bus, a write request bus, and a write response bus.
[0260] Aspects of the disclosed technology include a TPU. The TPU may include a hardware circuit and an instruction broadcast bus coupled to the hardware circuit. The instruction broadcast bus is configured to push instructions to the hardware circuit. The hardware circuit may include a plurality of tiles, and each tile may be configured to operate in parallel with other tiles within the plurality of tiles. Each tile of the plurality of tiles may include a processing core, a prefetch unit, and an instruction buffer. The hardware circuit may further include a plurality of data processing lanes configured to stream respective data from an upstream input to a downstream destination, and a plurality of task instruction memories. Each task instruction memory of the plurality of task instruction memories is arranged in a sequence and coupled via an instruction router to one or more of the plurality of tiles. The TPU may further include an instruction request bus coupled to the hardware circuit, and the instruction request bus may be configured to receive requests for instructions.
[0261] Aspects of the disclosed technology include a method for prefetching or providing instructions by a SIMD (single instruction multiple data) processing unit. The method may include receiving requests for instructions from a plurality of tiles of the SIMD processing unit, filtering requests for instructions to eliminate duplicates for the same instruction to generate a first set of requests, generating a set of instructions in response to the first set of requests, providing the set of instructions from a computing unit to a task instruction memory of the SIMD processing unit, storing the set of instructions in the task instruction memory, and accessing instructions from the set of instructions by a prefetch unit via an instruction router. The SIMD processing unit may include a plurality of tiles. Each tile is configured to operate in parallel with other tiles within the plurality of tiles, and each tile of the plurality of tiles includes a processing core, a prefetch unit, and an instruction buffer. The receiving step may be performed in a first processing clock cycle and the providing step may be performed in a second processing clock cycle. The first processing clock cycle may be performed before the second processing clock cycle.
[0262] Generally, this specification discloses a hardware / software interface for asynchronous data movement between off-core memory and core-local memory, referred to herein as "stream transfer", and a stream ordering model. Stream transfer enables software to more efficiently represent common data movement patterns, specifically patterns seen in sparse workloads. Stream instructions belonging to a stream are processed in order. In the case of indirect stream instructions, offset elements within an offset list are processed in order. A synchronization flag is updated to indicate monotonic progress for a stream.
[0263] One aspect of the present disclosure provides a method that includes identifying, by one or more processors, the progress of data being transferred between off-core memory and core-local memory, and identifying, by one or more processors, a read from core-local memory when core-local memory is the source of the data, where the reads are issued to the source in order and are matched out of order by the source, the method further including identifying, by one or more processors, a write to core-local memory when core-local memory is the destination for the data, where the writes are issued to the destination in order and are committed out of order by the destination, the method further including accessing, by one or more processors, off-core memory based on an indirect scatter / gather memory access for a read from off-core memory when off-core memory is the source of the data and for a write to off-core memory when off-core memory is the destination for the data.
[0264] In one example, the step of identifying the progress of the transferred data further includes using a core-local synchronization flag. In another example, the method further includes selecting, by one or more processors, a memory access to a barrier based on a scalar fence instruction. In yet another example, the step of accessing off-core memory based on an indirect scatter / gather memory access further includes obtaining a source of an indirect address from a register file or core-local memory. In yet another example, the method further includes circular buffering, by one or more processors, within core-local memory.
[0265] In yet another example, the method further includes updating, by one or more processors, a core - local synchronization flag to indicate a monotonic step - by - step progress of data transfer. In yet another example, the method further includes ending, by one or more processors, data transfer when all reads from the core - local memory have been issued. In yet another example, the method further includes ending, by one or more processors, data transfer when all writes to the core - local memory have been committed.
[0266] Another aspect of the present disclosure provides a system including one or more processors and one or more storage devices coupled to the one or more processors and storing instructions. When executed by the one or more processors, the instructions cause the one or more processors to perform operations for transferring data between off - core memory and core - local memory. The operations include identifying the progress of data being transferred between off - core memory and core - local memory, and identifying reads from the core - local memory when the core - local memory is the source of the data, where the reads are issued to the source in order and are corresponded to by the source out of order. The operations further include identifying writes to the core - local memory when the core - local memory is the destination of the data, where the writes are issued to the destination in order and are committed by the destination out of order. The operations further include accessing the off - core memory based on indirect scatter / gather memory access for reads from the off - core memory when the off - core memory is the source of the data and for writes to the off - core memory when the off - core memory is the destination of the data.
[0267] In one example, the operation of identifying the progress of the transferred data further includes the operation of using a core-local synchronization flag. In another example, the operation further includes the operation of selecting a memory access to a barrier based on a scalar fence instruction. In yet another example, the operation of accessing off-core memory based on an indirect scatter / gather memory access further includes the operation of obtaining the source of an indirect address from a register file or core-local memory. In yet another example, the operation further includes the operation of performing circular buffering within core-local memory.
[0268] In yet another example, the operation further includes the operation of updating a core-local synchronization flag to indicate a monotonic stepwise progress for data transfer. In yet another example, the operation further includes the operation of terminating data transfer when all reads from core-local memory have been issued. In yet another example, the operation further includes the operation of terminating data transfer when all writes to core-local memory have been committed.
[0269] Yet another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations for transferring data between off-core memory and core-local memory. The operations include identifying the progress of data being transferred between the off-core memory and the core-local memory, and identifying a read from the core-local memory when the core-local memory is the source of the data, where the reads are issued in order to the source and are matched out of order by the source, and the operations further include identifying a write to the core-local memory when the core-local memory is the destination for the data, where the writes are issued in order to the destination and are committed out of order by the destination, and the operations further include accessing the off-core memory based on indirect scatter / gather memory accesses for a read from the off-core memory when the off-core memory is the source of the data and a write to the off-core memory when the off-core memory is the destination for the data.
[0270] In one example, the operation of accessing the off-core memory based on indirect scatter / gather memory accesses further includes obtaining the source of the indirect address from a register file or core-local memory. In another example, the operation further includes performing circular buffering within the core-local memory. In yet another example, the operation further includes updating a synchronization flag to indicate a monotonic stepwise progress of the data transfer.
[0271] Aspects of the present disclosure are directed to a cross-lane processing unit (XPU) for performing single-instruction multiple-data (SIMD) data-dependent operations across multiple data processing lanes of a processor. Instead of physically fabricating operation-specific circuitry for each data-dependent operation, the XPU can be configured to perform various operations in response to input signals that configure processing cells for performing individual operations and a crossbar arranged as a stacked network in the XPU. Each processing cell can receive and process data across multiple data processing lanes. Aspects of the present disclosure include operations for configuring the XPU to perform a vector sort while calculating duplicate counts of duplicate elements within an input vector received for sorting, eliminating the need to separately configure the XPU for sorting and duplicate counting. The XPU can be implemented as part of a hardware circuit that complements calculations on dense data structures, such as dense matrices, while accelerating processing of sparse data structures, such as sparse vectors or sparse matrices.
[0272] Aspects of the present disclosure include a hardware circuit. The hardware circuit includes a plurality of stages, each stage including a crossbar and two or more cells. The hardware circuit further includes a plurality of data processing lanes for streaming respective data from an upstream input to a downstream destination via the plurality of cells and the plurality of crossbars of the plurality of stages. The hardware circuit is configured to receive input data and a first instruction for performing a first operation from an upstream input along the plurality of data processing lanes, and in response to receiving the first instruction, to transmit, for each stage, a respective second instruction to each processing cell of the stage. Each cell is configured to perform a respective second operation in response to receiving an input from a respective data processing lane. The hardware circuit is further configured to transmit a respective third instruction to each crossbar for the stage, and the crossbar is configured to route the output from each cell of the stage to the cells of the next stage along the plurality of data processing lanes. The received input data is processed along the plurality of data processing lanes and the plurality of cells configured to perform respective second operations to perform the first operation.
[0273] The systems included in aspects of the present disclosure include hardware circuitry, the hardware circuitry including a plurality of stages each including a crossbar and two or more cells, and a plurality of data processing lanes for streaming respective data from an upstream input to a downstream destination via the plurality of cells and the plurality of crossbars of the plurality of stages, the hardware circuitry being configured to receive input data from an upstream input along the plurality of data processing lanes and a first instruction for performing a first operation, and in response to receiving the first instruction, for each stage, being configured to send a respective second instruction to each processing cell of the stage, each cell being configured to perform a respective second operation in response to receiving an input from a respective data processing lane, the hardware circuitry further being configured to send a respective third instruction to each crossbar for the stage, the crossbar being configured to route the output from each cell of the stage to cells of the next stage along the plurality of data processing lanes, and being configured to perform the first operation by processing the received input data along the plurality of data processing lanes and the plurality of cells configured to perform the respective second operations.
[0274] Aspects of the present disclosure include a method implemented by a computer. The method implemented by the computer includes a plurality of stages each including a crossbar and two or more cells, and a plurality of data processing lanes for streaming respective data from an upstream input to a downstream destination. The method includes receiving, by a hardware circuit via a plurality of cells and a plurality of crossbars of the plurality of stages, input data from an upstream input along the plurality of data processing lanes and a first instruction for performing a first operation; and in response to receiving the first instruction, for each stage, transmitting, by the hardware circuit, a respective second instruction to a respective processing cell of the stage. Each cell is configured to perform a respective second operation in response to receiving an input from a respective data processing lane. The method implemented by the computer further includes transmitting, by the hardware circuit, a respective third instruction to a respective crossbar for the stage, the crossbar being configured to route an output from each cell of the stage to a cell of a next stage along the plurality of data processing lanes. The method implemented by the computer further includes performing the first operation by processing the received input data by the hardware circuit along the plurality of data processing lanes and the plurality of cells configured to perform the respective second operations.
[0275] Aspects of the present disclosure may include one or more of the following features. In some examples, aspects of the present disclosure include all of the following features in combination.
[0276] Each cell is configured to receive a respective first input operand from a respective data processing lane passing through the cell and a respective second input operand from a respective crossbar of a stage upstream of the cell.
[0277] The destinations downstream of the data of multiple data processing lanes are vector processing units, and the vector processing units are configured to perform single-instruction multiple-data vector operations on the output data of the hardware circuits.
[0278] Each of the cells is configured to execute one or more of a plurality of predetermined primitive operations in response to one or more received instructions. The hardware circuit further includes a plurality of control cells. When transmitting respective second instructions to respective processing cells, the hardware circuit is configured to generate respective control signals by each control cell based on a first operation specified by a first instruction and transmit the respective control signals to the respective processing cells.
[0279] When generating and transmitting respective control signals by each control cell, the hardware circuit is configured to generate respective control signals for causing each processing cell to execute one of respective arithmetic operations, comparisons, and bypass operations based on at least one of a stage in which the processing cell is located and a data processing lane through which the processing cell passes.
[0280] The plurality of cells and the plurality of crossbars form a processing network of cells connected across a plurality of stages and a plurality of data processing lanes, and the processing network of the connected cells is configured to receive input data and generate respective output data in accordance with performing a first operation on the input data.
[0281] The processing network of the connected cells is configured to perform a combined vector sort and duplicate count operation, and the combined operation includes an operation of receiving an input vector of elements by the processing network and an operation of generating, as output, a sorted output vector and data specifying the count of duplicate elements in the input vector by the processing network. The input data includes sparse vector data, and the hardware circuit is configured to perform one of vector scan, vector addition, vector sort, or vector duplicate count after sending each of the second and third instructions.
[0282] Unless otherwise specified, the above alternatives are not mutually exclusive and can be implemented in various combinations to achieve specific advantages. These and other variations and combinations of the above features can be utilized without departing from the subject matter defined by the appended claims. Therefore, the above description of the examples is not intended to limit the subject matter defined by the appended claims, but should be construed as illustrative. In addition, the presentation of the examples described herein, as well as phrases such as "such as" and "including", should not be construed as limiting the subject matter of the claims to specific examples. Rather, the above examples are intended to illustrate only one of many possible implementations. Further, the same reference numbers in the various drawings may identify the same or similar elements.
Claims
1. A processor, comprising a plurality of tiles, each of the plurality of tiles including a vector core configured to generate a data-dependent address stream, and a slice of a shared software-controlled scratchpad memory, the processor further comprising a scalar core configured to dispatch tasks to the plurality of tiles, and a memory coupled to the plurality of tiles and the scalar core.
2. The processor according to claim 1, wherein each tile is configured to execute independent computations.
3. The processor according to claim 1 or 2, wherein the vector core in each of the plurality of tiles includes a plurality of single instruction, multiple data (SIMD) processing lanes.
4. The processor according to claim 1 or 2, wherein the multiple tiles of the plurality of tiles issue memory requests in parallel to a main memory.
5. The processor according to claim 1 or 2, wherein the data-dependent address stream is for any level of a memory hierarchy.
6. The processor according to claim 5, wherein each data-dependent address stream corresponds to a sequence of addresses, and the length and specific values of the addresses in the sequence are data-dependent and are only known at runtime.
7. The processor according to claim 5, wherein the vector core in each of the plurality of tiles is configured to represent the data-dependent address stream while decoupling high-performance services of the data-dependent address stream to a microarchitecture.
8. The processor according to claim 7, wherein the microarchitecture includes a scatter-gather engine for the high-performance services of the data-dependent address stream.
9. The processor according to claim 7, wherein the data-dependent address stream includes a plurality of addressing modes, a runtime-configurable transfer size, and indirect memory access with atomic arithmetic updates.
10. The processor according to claim 1 or 2, wherein the vector core in each of the plurality of tiles includes circular buffer instructions that enable transfer and access of a dynamically sized data stream over a statically sized region of memory.
11. The processor according to claim 10, further comprising a microarchitecture configured to track a runtime buffer size of the data stream of the dynamic size.
12. The vector core in each of the plurality of tiles is configured to provide a runtime configuration and an access area of the tile-local scratch pad memory as a sequential circular first-in-first-out (FIFO) access without excluding unordered access to the same area of the tile-local scratch pad memory. The processor according to claim 11.
13. The sequential circular FIFO access enables the data stream of the dynamic size on the static size area of the tile-local scratch pad memory in association with the microarchitecture. The processor according to claim 12.
14. Each tile includes a scatter-gather engine configured to manage issuance, fetch, tracking, and ordering of data streams. The processor according to claim 1 or 2.
15. Each scatter-gather engine is further configured to maintain at least 256 unprocessed read requests in-flight per tile. The processor according to claim 14.
16. Each scatter-gather engine is further configured to track and update buffer occupancy to manage flow control. The processor according to claim 14.
17. A subset of the plurality of tiles each further includes a prefetch unit configured to cooperatively prefetch data stream instructions. The processor according to claim 1 or 2.
18. The processor according to claim 1 or 2, further comprising a cross-lane processing unit configured to accelerate at least one of an irregular control flow sequence or an in-vector dependent operation.
19. Each tile is configured to support scatter from off-chip memory to its scratch pad memory and gather from its scratch pad memory to off-chip memory. The processor according to claim 1 or 2.
20. The processor according to claim 1 or 2, wherein the subset of the plurality of tiles is grouped based on a logically configurable vector width.
21. The processor according to claim 20, wherein the logically configurable vector width includes a logical SIMD width.
22. The processor according to claim 1 or 2, wherein the processor is part of a machine learning accelerator configured to execute a neural network layer exhibiting semantic sparsity.
23. The processor according to claim 22, wherein the neural network layer includes an embedded neural network or a graph neural network.
24. The processor according to claim 22, wherein the processor is connected to several other processors via a network configured to perform distributed scatter-gather and calculations required by neural network layer computations that are dynamic, irregular, and memory-bound.
Citation Information
Patent Citations
Scalable sparse matrix multiply acceleration using systolic array with feedback input
JP2021177366A
Efficient Execution of Operation Unit Graphs on User-Specific Reconfigurable Architectures
JP2022548114A
Optimizated function assignment in a multi-core processor
US20180109449A1
Accelerating dataflow signal processing applications across heterogeneous CPU / GPU systems
US20200183738A1
Efficient Execution of Operation Unit Graphs on Reconfigurable Architectures Based on User Specification
US20210081691A1