Parallel Processor Optimized for Machine Learning
By designing a multi-layer parallel processor system and using the DMA engine to perform 3-D rasterization, the problem of inefficient processing of complex arithmetic operations and tensor shape manipulation in machine learning is solved, and efficient data processing and address calculation reduction is achieved.
Patent Information
- Application Number
- CN202210796025.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-08
- Filing Date
- 2022-07-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-07-06
AI Technical Summary
The prior art is difficult to efficiently handle complex arithmetic operations and tensor shape manipulation in machine learning, resulting in inefficient large number of address calculations and processing.
A multi-layer parallel processor system is designed, including an arithmetic logic unit (ALU) array, a controller, direct memory access (DMA) block and input stream buffer, and 3-D rasterization is performed through the DMA engine to optimize data access and processing.
Through parallel processing and rasterization technology, the processing efficiency of complex arithmetic operations and tensor shape manipulation in machine learning is significantly improved, the amount of address calculation is reduced, and the overall performance of the system is improved.
Smart Images

Figure CN115599444B_ABST
Abstract
Description
Technical Field
[0001] This description generally relates to machine learning, and more particularly, to parallel processors optimized for machine learning. Background Art
[0002] Machine learning (ML) applications are typically used to compute extremely large amounts of data, the processing of which can be mapped onto large parallel programmable data path processors. ML applications operate on multi-dimensional tensors (e.g., three- and four-dimensional). ML applications operate on simple integers, quantized integers (a subset of floating point (FP) values labeled by integers), FP 32b, and half-precision FP numbers (e.g., FP16 and Brain Floating Point (BFLOAT) 16). For example, an ML network may involve a mixture of arithmetic operations, some as simple as adding two tensors, or may involve more computationally intensive operations (e.g., matrix multiplication and / or convolution) or even very complex functions (e.g., sigmoid functions, square roots, or exponential functions). ML applications also include tensor shape manipulation and can extract, compress, and reshape input tensors into another output tensor, which implies a large amount of address calculation. Summary of the Invention
[0003] In one aspect, the present application provides a parallel processor system for machine learning, the system comprising: an array of arithmetic logic units (ALUs) including a plurality of ALUs; a controller configured to provide instructions for the plurality of ALUs; and a direct memory access (DMA) block including a plurality of DMA engines configured to access an external memory to retrieve data; and an input stream buffer configured to decouple the DMA block from the ALU array and provide alignment and reordering of the retrieved data, wherein the plurality of DMA engines are configured to operate in parallel and include rasterization logic configured to perform three-dimensional (3-D) rasterization.
[0004] In another aspect, the present application provides a method, comprising: performing a first rasterization in a memory by a DMA engine to reach a memory region; and performing a second rasterization in the memory region by the DMA engine to reach a memory element address, wherein: the first rasterization is performed by defining a 3-D raster pattern via four-vector address calculation within a first cube, and the second rasterization is performed via three-vector address calculation within a second cube surrounding the memory region to reach the memory element address.
[0005] On the other hand, the present application provides a system that includes: an input stream buffer that includes a plurality of DMA engines configured to access an external memory to retrieve data; a multi-bank memory; an ALU array that includes a plurality of ALUs; and wherein: the input stream buffer is configured to decouple the DMA engines from the ALU array, and the plurality of DMA engines are configured to operate in parallel and include rasterization logic configured to perform 3-D rasterization. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Certain features of the technology are set forth in the appended claims. However, for explanatory purposes, several embodiments of the technology are set forth in the drawings.
[0007] Figure 1 is a high-level diagram illustrating an example of an architecture for optimizing a parallel processor system for machine learning in accordance with various aspects of the present technology.
[0008] Figure 2 is a schematic diagram illustrating an example of a set of heterogeneous arithmetic logic units (ALUs) fed through a folding interface in accordance with various aspects of the present technology.
[0009] Figure 3 is a schematic diagram illustrating an example of the rasterization capabilities of direct memory access (DMA) for generating complex address patterns in accordance with various aspects of the present technology.
[0010] Figure 4 is a schematic diagram illustrating an example of a technique for transposing through a stream buffer in accordance with various aspects of the present technology.
[0011] Figure 5 is a schematic diagram illustrating an example of the architecture of a complex ALU in accordance with various aspects of the present technology.
[0012] Figure 6 is a flowchart illustrating an example of a method for accessing memory in an orderly manner by a DMA engine in accordance with various aspects of the present technology. DETAILED DESCRIPTION
[0013] The detailed description set forth below is intended as a description of various configurations of the technology and is not intended to represent the only configurations in which the technology may be practiced. The drawings are incorporated herein and constitute a part of the detailed description, which includes specific details for providing a thorough understanding of the technology. However, the technology is not limited to the specific details set forth herein and may be practiced without one or more of such specific details. In some instances, structures and components are shown in block diagram form in order to avoid obscuring the concepts of the technology.
[0014] This technology relates to systems and methods for optimizing parallel processors for machine learning. The disclosed solution is a parallel processor with multiple layers of hardware (HW) specialized modules optimized to address the challenges involved in a mix of arithmetic operations and tensor shape manipulation. For example, the mix of arithmetic operations can include adding two tensors or performing more computationally intensive operations such as matrix multiplication and / or convolution or even very complex functions such as sigmoid functions, square roots, or exponential functions. Tensor shape manipulation can include processing such as extracting, compressing, and reshaping an input tensor into another output tensor, which implies a large amount of address calculation.
[0015] The parallel processor with multiple layers of HW specialized modules of this technology can be part of a computer system. The functionality of this computer system can be improved to allow execution of a mix of arithmetic operations, including adding two tensors or performing more computationally intensive operations, as explained herein. It should be noted that performing these complex operations by an ML network may involve a large amount of address calculation, which is eliminated by the disclosed solution.
[0016] Figure 1 FIG. 7 is a high-level diagram illustrating an example of the architecture of a parallel processor system 100 for optimizing for machine learning according to various aspects of the present technology. The parallel processor system 100 includes an input external memory 102, an input stream buffer 104, an input multi-bank memory 108, an input stream interface (IF) 110, an arithmetic logic unit (ALU) array 112, a controller unit 114, an output stream IF 116, an output multi-bank memory 120, an output stream buffer 122 including a set of output direct memory access (DMA) engines 124, and an output external memory 126.
[0017] The input stream buffer 104 includes a set of input DMA engines 106. The input DMA engines 106 are powerful DMA engines capable of loading segments filled with tensor elements from the external memory 102 and loading those tensor elements into the input stream buffer 104. The segment size can be as large as 32B or 64B to provide sufficient bandwidth to sustain the peak throughput of the processing engine (e.g., the ALU array 112). There are multiple DMA engines (DMA-0 to DMA-N) that can operate in parallel to achieve simultaneous prefetching of multiple tensor segments. The DMA engines 106 are equipped with complex rasterizer logic such that address calculation can be defined by a 7-tuple vector, thereby essentially allowing the definition of 3-D raster patterns, as described in more detail herein.
[0018] The input stream buffer 104 is built on the input multi-bank memory 108 and is basically for two purposes. First, by introducing data in advance and by pre-accepting multiple requests for the data to absorb (share) data bubbles on the external memory interface, the input stream buffer 104 can decouple the DMA engine 106 from the ALU array 112 which is a digital signal processing (DSP) engine. Second, the input stream buffer 104 can provide data alignment and / or reordering such that the elements of a tensor are well-aligned to feed N processing engines. In that regard, two basic orderings are implemented: direct and transposed storage.
[0019] In some embodiments, the DMA engine 106 introduces data in constant M-byte segments corresponding to E = M / 4 or M / 2 or M elements, depending on whether the input tensor consists of 32-bit floating point (fp 32), 32-bit integer (int 32), or half-floating point or quantized integer (8-bit) elements. Then, depending on how the ALU array 112 will process the E elements, the E elements are stored into the ALU array 112. In direct mode, the E elements are stored in a compact form into M / 4 consecutive banks at the same address such that later, by reading a single row across, the M / 4 banks provide up to M elements to the DSP ALU. As we can see, the segments are spread across N DSP ALUs. In transposed mode, the E elements are stored in a checkerboard pattern across K consecutive banks such that later, when reading a single row across the K banks, K elements from K different segments will be sent through the N ALUs. As we can see, in this transposed mode, a given DSP ALU is executing all samples from a single segment. Depending on the processing to be performed on the input tensor, we can choose one mode or the other to adapt the way the processing is distributed across the N ALUs in a concept that allows maximizing data processing parallelization and minimizing inter-channel data exchange.
[0020] The group of output DMA engines 124 is the same as the group of input DMA engines 106 and can read the output stream buffer data to store rows into the output external memory 126 (e.g., random access memory (RAM)). Similarly, complex raster access (explained below) can also be used herein to match the addressing capabilities that exist in the input stage. In other words, complex raster access can put the data back in the same order and / or location.
[0021] The ALU array 112 contains a group of N DSP processing units (ALU#1 to ALU#N). The group of N DSP processing units, all operating in lockstep, is driven by a controller unit 114 which has its own program memory and can execute instruction bundles consisting of M arithmetic instructions.
[0022] An output stream IF, which has the same functionality as the input stream IF but in the opposite direction, i.e., it collects data from the ALU array 112 and—after unfolding—places the data in the output stream buffer for collection by the DMA engine. Similar to the input stream IF 110 that acts as a pop interface for the ALU array 112, the output stream IF 116 acts as a push interface for the output multi-bank memory 120. The output stream buffer 122 mimics the same behavior as the input stream buffer 104. The output stream buffer 122 is based on the output multi-bank memory 120 to absorb DMA latency and is structured such that we can optionally block a set of memory banks of the received elements. This allows scenarios such as: (a) input transpose, DSP processing, and output transpose modes such that the element order and / or shape are preserved; (b) input transpose, DSP processing, and output direct mode, in which case the element order and / or shape are not preserved and can be used for geometric transformations such as tensor transpose.
[0023] Figure 2 is a schematic diagram illustrating an example architecture 200 of a set of heterogeneous ALU arrays 204 (hereinafter, ALU array 204) fed through a folding interface according to various aspects of the present technology. The example architecture 200 includes an input buffer interface block 202 and the ALU array 204. The input buffer interface block 202 is a feed-through folding interface and is similar to Figure 1 the input stream IF110. As described above, Figure 1 the input DMA engine 106 of
[0024] can introduce data through a constant M-byte segment corresponding to E = M / 4 or M / 2 or M elements, and then store the elements in the memory banks depending on how the DSP ALU will process the elements, where the DSP ALU is the same as the ALU array 204. Figure 1 The input buffer interface block 202 is located between
[0025] Figure 1 the input stream buffer 104 and the ALU array 204 of
[0025] An important function of the input buffer interface block 202 is to manage segment folding such that when a given segment may have E active elements, it can read successive segments of P elements by folding the E elements into P elements and distributing the P elements to the first PALU in the N ALU arrays 204. This folding mechanism is available in both direct and transpose modes and will generally vary according to the tensor operation to match the capabilities of the ALU arrays 204. Folding basically works by squeezing the entire tensor through a window of P elements once in a number of cycles depending on the total tensor size. A set of N ALUs in the ALU arrays 204 operate in lockstep and are driven by a controller unit (e.g., Figure 1 controller unit 114) having its own program memory and executing a bundle of instructions consisting of M arithmetic instructions.
[0026] The instruction bundle goes to all N ALUs, but only a subset of the N ALUs can be enabled at any time to execute the broadcast instructions. Note that explicit load / store instructions are not required; instead, the N ALUs can directly access each DMA channel through the input buffer interface block 202 as a direct input to the operations. The same thing also applies to the store area, where the output of the operation can be selected as an output stream. Each ALU contains additional storage via a local register bank such that intermediate computation results can be stored.
[0027] The ALU capabilities are shown in bar chart 210, which illustrates that the ALU capabilities decrease from ALU#1 to ALU#N. For example, each of ALU#1 and ALU#2 includes five functional units (FUs), ALU#3 includes three FUs, and ALU#N has only one FU. As is clearly visible from the ALU bar chart 210, not all ALUs are the same; instead, they are grouped into multiple groups of equivalent ALUs in each FU, where the capabilities of each group are a superset of the previous group. Thus, at one extreme, we find a group consisting of relatively simple ALUs that have only integer capabilities and mainly operate on 8-bit quantities with little local storage. At the other extreme, we find another group of ALUs that are extremely powerful, cover all data types, have large register files, and implement complex non-linear functions; and in between, we will find other ALU clusters with more or less capabilities (i.e., supporting only integer and half-float). This partitioning is effective for two reasons: (1) it allows balancing the total area and processing performance by allocating more or less complete ALUs depending on the type of operation; and (2) it allows balancing the memory bandwidth and the compute bandwidth (i.e., several ALUs supporting 32-bit floating point and more 8-bit ALUs for the same total input / output bandwidth). There are several key aspects behind this arrangement. Although the group of N ALUs of the ALU array 204 is actually heterogeneous, it is still controlled by a single instruction bundle that goes to all ALUs, and the ALUs only execute the instructions they support. Another key aspect of this arrangement is the folding ability of the input buffer interface block 202, which allows tensor elements to be effectively mapped individually to the ALUs that are capable of performing the required operations. The basic operations / instruction set supported by the ALUs are instances of the operations / instruction set that can be found in the DSP instruction set architecture (ISA), but are also augmented with special instructions dedicated to ML applications. It should be noted that specific instructions can be used to perform (a) quantization zero-point conversion; i.e., converting an 8-bit quantized integer to a 32-bit integer (X - Zp) x Sc can be efficiently performed in one clock cycle (and its inverse operation); (b) non-linear activation functions (e.g., rectified linear function, sigmoid function, hyperbolic tangent function) can be efficiently performed in one clock cycle; and (c) support for brain floating point 16 (BFLOAT 16) data as well as FP16 and FP32.
[0028] Figure 3 is a schematic diagram illustrating an example of the rasterization capabilities of a DMA for generating complex address patterns according to various aspects of the present technology. Figure 1The DMA engine 106 is equipped with complex rasterizer logic that enables address calculations to be defined by a 7-tuple vector. This essentially allows a 3-D raster pattern defined by, e.g., [1, sub_H, sub_W, sub_C] to step into a 4-D tensor defined by [K, H, W, C] of the cube 304. The stepping can occur every clock cycle to prefetch multiple tensor segments simultaneously. By correctly parameterizing the seven-dimensional step count, a powerful zigzag scan pattern can be created, which is required for some tensor shape manipulation operations to provide ordered access within the 4-D tensor. The combination of 3-D and 4-D coordinates (by simple coordinate addition) yields the final 4-D coordinates, which are used to address the external memory and introduce segments filled with tensor elements. To provide sufficient bandwidth to the DSP core, two DMA engines can access the external memory in parallel and introduce two segments into the input stream buffer each cycle. Although the coarse rasterization represented by the cube 302 can result in coarse sample positions 306 within the memory, an independent fine rasterization starting from the coarse sample positions 306 represented by the cube 304 can resolve the optimal positions 308 of the samples, which can be the desired memory element addresses. The coarse and fine rasterizations are independent of each other; for example, the coarse rasterization can start scanning along the C dimension, while the fine rasterization can start scanning along the H dimension first.
[0029] Figure 4 is a schematic diagram illustrating an example of a technique for transposition through a stream buffer according to various aspects of the present technology. As explained above, in the transpose mode, E elements are stored in a checkerboard pattern across K (e.g., 16) consecutive memory banks 404 such that later, when reading a row across the K consecutive memory banks 404, K elements from K different segments (memory bank 0 to memory bank 15) will be sent through N ALUs. We can see that in this transpose mode, a given DSP ALU is executing all samples from a single segment. Figure 4 The transpose is for the case where E = 16 (i.e., half-float) and the way data is stored into Figure 1 the input stream buffer 104 when configured in the transpose mode. The checkerboard pattern allows one input segment to be written into 16 memory banks (e.g., memory bank 0 to memory bank 15) and four segments to be broadcast to four ALUs during 16 cycles when reading. Due to this arrangement, a single write / read memory can be used for element transposition. This is crucial for ML applications that may require different data access orders at different times. At one time, the ALU needs to access data 402 in one order, while at some other time, it needs to access data 402 in a different order. For example, in Figure 4In the transpose, A0, B0, and C0 data are accessed in parallel in the first cycle, A1 and B1 data are accessed in parallel in the second cycle, A2 and D12 data are accessed in parallel in the third cycle, and so on.
[0030] Figure 5 FIG. 4 is a schematic diagram illustrating an example of the architecture of a complex ALU 500 according to various aspects of the present technology. The complex ALU 500 includes a vector register file (VRF) 502, a scalar register file (SRF) 504, a read network (RN) 506, a set of complex ALUs 508, a write network (WN) 510, and an instruction register 512. The VRF 502 and SRF 504 are respectively used to store vector and scalar variables, as both of them can be used in ML applications. D0 and D1 provide data from the input stage. The RN 506 transfers data to the appropriate functional units of the complex ALU 508, which includes special function units (SFUs), ALU-0, ALU-1, converter logic-0 (CONV-0), CONV-1, and CONV-2. The instruction register 512 is a long instruction register and includes instructions for the CONV-0, CONV-1, CONV-2, ALU-0, ALU-1, and SFU functional units of the complex ALU 508, and for the RN 506, WN 510, VRF 502, and SRF 504. The instruction register 512 stores the very long instruction word (VLIW) provided by Figure 1 the controller unit 114. The WN 510 is responsible for writing the processed data into Figure 1 the output stream IF 116.
[0031] Figure 6 FIG. 6 is a flowchart illustrating an example of a method 600 for the DMA engine to access memory in an orderly manner according to various aspects of the present technology. The method 600 includes the DMA engine (e.g., Figure 1 106 of Figure 1 performing a first rasterization within a memory (e.g., Figure 3 102 of Figure 3 to reach a memory area (e.g., Figure 3 306 of Figure 3 ), and the DMA engine performing a second rasterization within the memory area to reach a memory element address (e.g., Figure 3 308 of
[0032] Those skilled in the art will understand that the various illustrative blocks, modules, elements, components, memory systems, and algorithms described herein can be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, various illustrative blocks, modules, elements, components, memory systems, and algorithms have been described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application. The various components and blocks may be arranged differently (e.g., in a different order or partitioned in a different manner), but all without departing from the scope of the present technology.
[0033] It should be understood that any particular order or hierarchy of blocks in the disclosed processes is an illustration of example methods. Based on design preferences, it should be understood that the particular order or hierarchy of blocks in the processes may be rearranged, or that all of the illustrated blocks may not be performed. Any of the blocks may be executed concurrently. In one or more embodiments, multitasking and parallel processing may be advantageous. Additionally, the separation of various system components in the embodiments described above should not be understood to be required in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0034] As used in this specification and any claims of this application, the terms "base station", "receiver", "computer", "server", "processor", and "memory" all refer to electronic or other technological devices. These terms do not include a person or group of people.
[0035] The predicate "configured to" does not imply any particular tangible or intangible modification of the subject. In one or more embodiments, a processor configured to monitor and control operations or components may also mean that the processor is programmed to monitor and control operations or that the processor is operable to monitor and control operations. Similarly, a processor configured to execute code may be interpreted as a processor programmed to execute code or operable to execute code.
[0036] Phrases such as "aspect", "some embodiments", "one or more embodiments", "example", "the present technology", and other variations thereof are for convenience and do not imply that the disclosure associated with such phrases is essential to the present technology or that the disclosure applies to all configurations of the present technology. The disclosure associated with such phrases may apply to all configurations or one or more configurations. The disclosure associated with such phrases may provide one or more examples.
[0037] Any embodiment described herein as an "example" is not necessarily to be construed as preferred or advantageous over other embodiments. Further, to the extent that the terms "comprising," "having," and the like are used in the description or claims, these terms are intended to be inclusive in a manner similar to the term "including" as that term is interpreted when used as a transitional word in a claim.
[0038] All structural and functional equivalents of the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public, whether or not it is explicitly recited in the claims. A claim element should not be construed under the provisions of 35 U.S.C. § 112(f) unless the element is expressly recited using the phrase "means for" or, in the case of a claim of a memory system, the phrase "step for" to recite the element.
[0039] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein but are to be accorded the full scope consistent with the language of the claims, where the singular forms of the terms used to refer to an element do not mean "one and only one" unless specifically so stated, but rather "one or more." Unless specifically stated otherwise, the term "some" means one or more. The masculine pronouns (e.g., "his") include the feminine and neuter genders (e.g., "her" and "its") and vice versa. Headings and subheadings (if any) are used for convenience only and do not limit the disclosure.
Claims
1. A parallel processor system for machine learning, the system comprising: Arithmetic Logic Unit (ALU) array; A controller configured to provide instructions for the ALU array; An input stream buffer including a plurality of Direct Memory Access (DMA) engines configured to access an external memory to retrieve data; and wherein the input stream buffer is configured to decouple the plurality of DMA engines from the ALU array and provide alignment and reordering of the retrieved data, and wherein the plurality of DMA engines are configured to operate in parallel and include logic configured to perform rasterization via a 7-tuple vector scan.
2. The system according to claim 1, wherein the input stream buffer is further configured to absorb data bubbles on the shared interface of the external memory by accepting multiple requests for data.
3. The system according to claim 1, wherein the instructions include very long instruction words (VLIWs) containing instructions for the ALU array and for vector and scalar registers.
4. The system according to claim 1, wherein the logic includes rasterization logic configured to perform three-dimensional (3-D) rasterization.
5. The system according to claim 4, wherein the 7-tuple vector scan includes performing coarse rasterization via quadruple vector address calculation within a first cube to reach a memory region.
6. The system according to claim 5, wherein the 7-tuple vector scan further includes performing fine rasterization via triple vector address calculation within a second cube surrounding the memory region to reach a memory element address.
7. The system according to claim 1, wherein the ALU array includes multiple ALUs each including a different number of functional units with different computing capabilities.
8. The system according to claim 7, further comprising an input buffer interface configured to perform feedthrough folding to feed each of the multiple ALUs based on the corresponding functional units and computing capabilities of each ALU in the multiple ALUs.
9. The system according to claim 1, wherein the retrieved data is fed into the ALU array in one of a direct mode or a transpose mode.
10. The system according to claim 9, wherein in the transpose mode, the retrieved data is interleaved in the ALU array.
11. The system according to claim 1, further comprising a vector register file (VRF) module and a scalar register file (SRF) module configured to store vector and scalar variables respectively.
12. The system according to claim 11, wherein the ALU array includes different functional units including at least a converter logic ALU and a special function ALU, and wherein the VRF module and the SRF module are configured to feed vector or scalar variables into that ALU through a read network based on the corresponding functional units of the ALUs in the ALU array.
13. A method, comprising: Performing a first rasterization in memory by the DMA engine to reach a memory region; and Performing a second rasterization in the memory region by the DMA engine to reach a memory element address, wherein: The first rasterization is performed by defining a 3-D raster pattern via four-vector address calculation within a first cube, and The second rasterization is performed via three-vector address calculation within a second cube surrounding the memory region to reach the memory element address.
14. The method according to claim 13, wherein defining the 3-D raster pattern includes performing four-vector address calculations by stepping into a four-dimensional 4-D tensor defined by the dimensions [K, H, W, C] of the first cube.
15. The method according to claim 14, further comprising using the first rasterization to provide ordered access within the 4-D tensor.
16. The method according to claim 15, wherein the ordered access starts from scanning any one of the dimensions H, W, and C of the first cube.
17. A system, comprising: An input stream buffer including a plurality of Direct Memory Access (DMA) engines configured to access an external memory to retrieve data; Arithmetic Logic Unit (ALU) array; and An input buffer interface configured to perform feedthrough folding to feed the ALU array based on corresponding functional units and computing capabilities of the ALU array, wherein: The input stream buffer is configured to decouple the plurality of DMA engines from the ALU array, and the plurality of DMA engines are configured to operate in parallel and include rasterization logic configured to perform 3-D rasterization.
18. The system according to claim 17, further comprising a controller configured to provide instructions for the ALU array, wherein the instructions include a VLIW including instructions for the ALU array and for vector and scalar registers.
19. The system according to claim 17, wherein the input stream buffer is configured to: provide alignment and reordering of retrieved data, and absorb data bubbles on the shared interface of the external memory by accepting multiple requests for data.
20. The system according to claim 17, wherein the plurality of DMA engines are configured to operate in parallel to enable simultaneous prefetching of tensors.
Citation Information
Patent Citations
Programmable re-order buffer for decompression
US20210142438A1