A system and method for an optimized winograd convolution accelerator
By introducing the Winograd convolution accelerator into the graphics processing unit (GPU) and optimizing the hardware architecture, the memory and computational bottlenecks in deep learning model deployment on low-power devices are addressed, enabling efficient deep learning processing on low-power devices.
Patent Information
- Application Number
- CN201810737887.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-08-07
- Filing Date
- 2018-07-06
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2038-07-06
AI Technical Summary
Existing deep learning neural network models are difficult to process efficiently when deployed on low-power computing devices due to limitations in memory and computing power, which restricts their widespread use in Internet of Things (IoT) applications.
The Winograd convolution accelerator is used to improve computational efficiency and memory utilization by optimizing the hardware architecture of the graphics processing unit (GPU), thus adapting to the computing needs of low-power devices.
It improves the efficiency of low-power devices in processing deep learning tasks, reduces the requirements for memory and computing resources, and enables deep learning models to run efficiently on low-power devices.
Smart Images

Figure CN109388777B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments relate generally to data processing and more particularly to machine learning processing via general purpose graphics processing units. BACKGROUND
[0002] Machine learning has successfully solved many kinds of tasks. The computations that arise in training and using machine learning algorithms (e.g., neural networks) naturally lend themselves to efficient parallel implementations. Thus, parallel processors such as general purpose graphics processing units (GPGPUs) have played an important role in the practical implementation of deep neural networks. However, implementing machine learning systems based on deep learning can require large amounts of memory and computing power. Deep learning neural network models can be many megabytes in size and can require billions of floating point operations per second for efficient processing. Such requirements can prevent the deployment of many neural network models to low power computing devices such as devices suitable for use in the Internet of Things (IoT) application domain, which is often populated with low-end embedded devices BRIEF DESCRIPTION OF DRAWINGS
[0003] So that the manner in which the above recited features of the embodiments can be understood in detail, a brief description of embodiments can be had by reference to a brief description of embodiments, some of which are illustrated in the appended drawings. It is to be noted, however, that the appended drawings illustrate only typical embodiments and are therefore not to be considered limiting of its scope, for the embodiments can admit to other equally effective embodiments.
[0004] FIG. 1 is a block diagram of a processing system in accordance with an embodiment;
[0005] FIG. 2 is a block diagram of a processor in accordance with an embodiment;
[0006] FIG. 3 is a block diagram of a graphics processor in accordance with an embodiment;
[0007] FIG. 4 is a block diagram of a graphics processing engine of a graphics processor in accordance with some embodiments;
[0008] FIG. 5 is a block diagram of hardware logic of a graphics processor core in accordance with some embodiments described herein.
[0009] FIGS. 6A-6B thread execution logic including an array of processing elements employed in a graphics processor core is illustrated in accordance with embodiments described herein.
[0010] FIG. 7 is a block diagram illustrating a graphics processor instruction format in accordance with some embodiments;
[0011] FIG. 8 is a block diagram of a graphics processor according to another embodiment.
[0012] FIGS. 9A-9B graphics processor command formats and command sequences are shown according to some embodiments;
[0013] FIG. 10 an exemplary graphics software architecture for a data processing system is shown according to some embodiments;
[0014] FIG. 11 is a block diagram showing an IP core development system according to embodiments;
[0015] FIG. 12 is a block diagram showing an exemplary system on a chip integrated circuit according to embodiments.
[0016] FIGS. 13A-13B is a block diagram showing an exemplary graphics processor for use within a SoC according to embodiments described herein.
[0017] FIG. 14 is a block diagram showing an additional exemplary graphics processor of a system on a chip integrated circuit according to embodiments.
[0018] FIGS. 15A-15B native and Winograd-based 3D convolutions are shown
[0019] FIG. 16 architectures for F[4,3]-based 2D / 3D convolutions are shown.
[0020] FIG. 17 logic for generalizing F(4,3) Winograd convolutions to higher order kernels is shown.
[0021] FIG. 18 exemplary logic and data layout for implementing multi-span convolutions according to embodiments is shown.
[0022] FIG. 19 is a block diagram of a Winograd acceleration architecture according to embodiments described herein.
[0023] FIG. 20 input transformations according to embodiments are shown.
[0024] FIG. 21 architectures for Winograd compute blocks according to embodiments are shown.
[0025] FIG. 22 logic configurable to perform native weight transformations and optimized weight transformations is shown.
[0026] FIG. 23 An optimized Winograd weight transformation architecture according to an embodiment is presented.
[0027] FIG. 24 A process for performing hardware-based Winograd convolutions using kernels of various sizes is presented according to embodiments described herein.
[0028] FIG. 25 A process for performing hardware-based Winograd convolutions using kernels with various strides is presented according to embodiments described herein.
[0029] FIG. 26 A process for performing optimized weight transformations for hardware-based Winograd convolutions according to embodiments described herein is presented.
[0030] FIG. 27 A machine learning software stack according to an embodiment is presented.
[0031] FIGS. 28A-28B The layers of an exemplary deep neural network are shown.
[0032] FIG. 29 An exemplary recurrent neural network is shown.
[0033] FIG. 30 Demonstrates the training and deployment of deep neural networks.
[0034] FIG. 31 It is a block diagram showing distributed learning. DETAILED DESCRIPTION
[0035] In some embodiments, a graphics processing unit (GPU) is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU can be communicatively coupled to the host processor / core via a bus or another interconnect. In other embodiments, the GPU can be integrated on the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., inside the package or chip). Regardless of how the GPU is connected, the processor core can assign work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0036] In the following description, FIGS. 1-14 An overview of an exemplary data processing system and graphics processor logic is provided in conjunction with or related to various embodiments. FIG. 26Specific details of various embodiments are provided. FIGS. 27-31 An overview of machine learning hardware and software architectures is provided. Some aspects of the following embodiments are described with reference to a graphics processor, while other aspects are described with respect to a general purpose processor such as a central processing unit (CPU). Similar techniques and teachings can be applied in the context of other types of circuitry or semiconductor devices that include, but are not limited to, an integrated many-core processor, a cluster of GPUs, or one or more instances of a field programmable gate array (FPGA). In general, the teachings are applicable to any processor or machine that manipulates or transforms data represented as physical electronic quantities.
[0037] System Overview
[0038] FIG. 1 is a block diagram of a processing system 100 in accordance with an embodiment. In embodiments, the system 100 includes one or more processors 102 and one or more graphics processors 108, and can be a single processor desktop system, a multiprocessor workstation system, or a server system that includes a large number of processors 102 or processor cores 107. In one embodiment, the system 100 is a processing platform incorporated within a chip that is used in a mobile device, hand-held device, or embedded device.
[0039] In one embodiment, the system 100 can include or be incorporated within a server-based gaming platform, a game console, including a games and media console, a mobile gaming console, a handheld game console, or an online game console. In some embodiments, the system 100 is a mobile phone, a smart phone, a tablet device, or a web appliance. The processing system 100 can also include wearable devices, such as a smart watch wearable, smart glass device, augmented reality device, or virtual reality device, coupled with, or integrated into, the wearable device. In some embodiments, the processing system 100 is a television or set-top box device having one or more processors 102 and a graphical interface generated by one or more graphics processors 108.
[0040] In some embodiments, one or more processors 102 each include one or more processor cores 107 for processing instructions which, when executed, carry out the operations of system and user software. In some embodiments, each of the one or more processor cores 107 is configured to process a specific instruction set 109. In some embodiments, instruction set 109 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via a very long instruction word (VLIW). Multiple processor cores 107 can each process a different instruction set 109, which can include instructions to facilitate the emulation of other instruction sets. Processor cores 107 can also include other processing devices, such as digital signal processors (DSPs).
[0041] In some embodiments, processor 102 includes cache memory 104. Depending upon the architecture, processor 102 can have a single internal cache or multiple levels of internal caches. In some embodiments, cache memory is shared among the various components of processor 102. In some embodiments, processor 102 also uses an external cache (e.g., a level 3 (L3) cache or last level cache (LLC)) (not shown), which can be shared among processor cores 107 using known cache coherency techniques. Additionally, register file 106 is included in processor 102, which can include different types of registers to store different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). Some registers can be general-purpose registers, while other registers can be specific to the design of processor 102.
[0042] In some embodiments, one or more processors 102 are coupled with a processor bus 110 for communicating data signals, such as address signals, data signals, or control signals, between processor 102 and other components of system 100. In one embodiment, processor bus 110 is a version of a direct media interface (DMI) bus. In one embodiment, processor(s) 102 include an integrated memory controller 116 and a peripheral controller hub (PCH) 130. Memory controller 116 facilitates communication between memory devices and other components of system 100, while PCH 130 provides connectivity to I / O devices via a local I / O bus.
[0043] Memory device 120 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase- change memory device, or some other memory device having suitable performance for use as a processing memory. In one embodiment, memory device 120 can operate as system memory for system 100, for storing data 122 and instructions 121 for use when executing applications or processes by the one or more processors 102. Memory controller 116 is also coupled to an optional external graphics processor 112, which can communicate with the one or more graphics processors 108 in the processors 102 to perform graphics and media operations. In some embodiments, display device 111 can be connected to the processor(s) 102. Display device 111 can be one or more of a display device built into a mobile electronic device or laptop device, for example, or an external display device attached via a display interface (e.g., DisplayPort, etc.). In one embodiment, display device 111 can be a head-mounted display (HMD) such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.
[0044] In some embodiments, the peripheral controller 130 enables peripheral devices to connect to the memory device 120 and the processor 102 via a high-speed I / O bus. I / O peripheral devices include, but are not limited to, an audio controller 146, a network controller 134, a firmware interface 128, a wireless transceiver 126, a touch sensor 125, a data storage device 124 (e.g., hard disk drive, flash memory, etc.). The data storage device 124 can connect via a storage interface (e.g., SATA) or via a peripheral bus, such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). The touch sensor 125 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 126 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. The firmware interface 128 enables communication with system firmware, and can, for example, be a Unified Extensible Firmware Interface (UEFI). The network controller 134 can enable network connectivity to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the processor bus 110. In one embodiment, the audio controller 146 is a multi-channel high definition audio controller. In one embodiment, the system 100 includes an optional legacy I / O controller 140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system. The peripheral controller 130 can also connect to one or more Universal Serial Bus (USB) controllers 142 connecting input devices, such as a keyboard and mouse 143 combination, a camera 144, or other USB input devices.
[0045] It will be recognized that the system 100 shown is exemplary and not limiting, in that other types of data processing systems that are differently configured can also be used. For example, an instance of the memory controller 116 and the peripheral controller 130 can be integrated into a discrete external graphics processor, such as the external graphics processor 112. In one embodiment, the peripheral controller 130 and / or the memory controller 1160 can be external to the processor(s) 102. For example, the system 100 can include an external memory controller 116 and an external peripheral controller 130 that can be configured as a memory controller hub and a peripheral controller hub within a system chipset that communicates with the processor(s) 102.
[0046] FIG. 2 is a block diagram of an embodiment of a processor 200 having one or more processor cores 202A to 202N, an integrated memory controller 214, and an integrated graphics processor 208. FIG. 2Those elements of Fig. 1 having the same reference numbers (or names) as the elements of any other figure herein can operate or function in any manner similar to that described elsewhere herein, but are not limited to such. The processor 200 can include additional cores up to and including the additional core 202N represented by the dashed box. Each of the processor cores 202A to 202N includes one or more internal cache units 204A to 204N. In some embodiments, each processor core can also access one or more shared cache units 206.
[0047] The internal cache units 204A to 204N and shared cache units 206 represent a cache memory hierarchy internal to the processor 200. The cache memory hierarchy can include at least one level of instruction and data caches within each processor core and one or more levels of shared mid-level caches, such as a level 2 (L2), level 3 (L3), level 4 (L4), or other level of cache, with the highest level cache being classified as an LLC prior to external memory. In some embodiments, cache coherency logic maintains coherency between the cache units 206 and 204A to 204N.
[0048] In some embodiments, the processor 200 can also include a set of one or more bus controller units 216 and a system agent core 210. The one or more bus controller units 216 manage a set of peripheral buses, such as one or more PCI or PCI express buses. The system agent core 210 provides management functionality for the various processor components. In some embodiments, the system agent core 210 includes one or more integrated memory controllers 214 to manage access to various external memory devices (not shown).
[0049] In some embodiments, one or more of the processor cores 202A to 202N include support for simultaneous multi-threading. In such embodiments, the system agent core 210 includes components used to coordinate and operate the cores 202A to 202N during multi-threaded processing. Additionally, the system agent core 210 can also include a power control unit (PCU), including logic and components to adjust the power state of the processor cores 202A to 202N, as well as the graphics processor 208.
[0050] In some embodiments, in addition, the processor 200 includes a graphics processor 208 for performing graphics processing operations. In some embodiments, the graphics processor 208 is coupled to the shared cache unit 206 set and the system agent core 210, which includes one or more integrated memory controllers 214. In some embodiments, the system agent core 210 also includes a display controller 211 to drive graphics processor output to one or more coupled displays. In some embodiments, the display controller 211 can also be a separate module coupled with the graphics processor via at least one interconnect, or can be integrated within the graphics processor 208.
[0051] In some embodiments, a ring-based interconnect unit 212 is used to couple the internal components of the processor 200. However, alternative interconnect units such as point-to-point interconnects, switched interconnects, or other technologies can be used, including technologies well known to those of ordinary skill in the art. In some embodiments, the graphics processor 208 couples with the ring interconnect 212 via the I / O link 213.
[0052] The exemplary I / O link 213 represents at least one of a variety of I / O interconnects, including a package I / O interconnect that facilitates communication between the various processor components and a high performance embedded memory module 218, such as an eDRAM module. In some embodiments, each of the processor cores 202A-202N and the graphics processor 208 use the embedded memory module 218 as a shared last level cache.
[0053] In some embodiments, the processor cores 202A-202N are homogeneous cores executing the same instruction set architecture. In another embodiment, the processor cores 202A-202N are heterogeneous in terms of instruction set architecture (ISA), where one or more of the processor cores 202A-202N execute a first instruction set and at least one of the other cores executes a subset of the first instruction set or a different instruction set. In one embodiment, the processor cores 202A-202N are homogeneous in terms of microarchitecture, where one or more of the cores have a relatively high power consumption and one or more power cores have a lower power consumption. In addition, the processor 200 can be implemented on one or more chips or as a SoC integrated circuit with the illustrated components, among other components.
[0054] FIG. 3is a block diagram of a graphics processor 300 that can be a discrete graphics processing unit, or can be a graphics processor integrated with a multiple processing core. In some embodiments, the graphics processor communicates with the memory via a mapped I / O interface to registers on the graphics processor, and utilizes commands stored in the processor memory. In some embodiments, the graphics processor 300 includes a memory interface 314 to access a memory. The memory interface 314 can be an interface to cache memory, one or more internal caches, one or more shared external caches, and / or to system memory.
[0055] In some embodiments, the graphics processor 300 also includes a display controller 302 to drive display output data to a display device 320. The display controller 302 includes hardware for one or more overlay planes for the display and composition of multiple layers of video or user interface elements. The display device 320 can be an internal or external display device. In one embodiment, the display device 320 is a head mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. In some embodiments, the graphics processor 300 includes a video codec engine 306 to encode, decode, or transcode media
[0056] In some embodiments, the graphics processor 300 includes a block image transfer (BLIT) engine 304 to perform two-dimensional (2D) rasterizer operations including, for example, bit-boundary block transfers. However, in one embodiment, 2D graphics operations are performed using one or more components of the graphics-processing engine (GPE) 310. In some embodiments, GPE 310 is a compute engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.
[0057] In some embodiments, GPE 310 includes a 3D pipeline 312 for processing 3D operations, such as rendering three-dimensional graphics onto a display. The 3D pipeline 312 includes fixed function and programmable logic
[0058] In some embodiments, media pipeline 316 includes fixed function or programmable logic for performing one or more specialized media operations related to the processing of media data, including the decoding, encoding, pre-processing, and / or post-processing of media data. In some embodiments, media pipeline 316 includes dedicated logic that is separate from the logic
[0059] In some embodiments, 3D / media subsystem 315 includes logic to execute the threads generated by 3D pipeline 312 and media pipeline 316. In one embodiment, the pipelines send thread execution requests to 3D / media subsystem 315, which includes thread dispatch logic to arbitrate the various requests and dispatch them to available thread execution resources. The execution resources include an array of graphics execution units for processing the 3D and media threads. In some embodiments, 3D / media subsystem 315 includes one or more internal caches to cache it instructions and data used
[0060] Graphics Processing Engine
[0061] FIG. 4 FIG. 4 is a block diagram of a graphics processing engine 410 of a graphics processor, according to some embodiments. In one embodiment, the graphics processing engine (GPE) 410 is a version of the GPE 310 shown in FIG. 3. FIG. 3 FIG. 4 Those elements of FIG. 4 having the same reference numbers (or names) as the elements of any other figure herein can operate or function in any manner similar to that described elsewhere herein, but are not limited to such. For example, the GPE 310 illustrated in FIG. 3 can be used in place of or instead of the GPE 410 shown in FIG. 4. FIG. 3 of the 3D pipeline 312 and the media pipeline 316. The media pipeline 316 is optional in some embodiments of the GPE 410 and can not be explicitly included within the GPE 410. For example, and in at least one embodiment, a separate media and / or image processor is coupled to the GPE 410.
[0062] In some embodiments, the GPE 410 is coupled with or includes a command streamer 403 that provides a command stream to the 3D pipeline 312 and / or the media pipeline 316. The command streamer 403 is coupled to memory, which can be system memory, or one or more of internal cache memory and shared cache memory, in some embodiments. The command streamer 403 receives commands from the memory and sends the commands to the 3D pipeline 312 and / or media pipeline 316. The commands are instructions for the 3D pipeline 312 and media pipeline 316 on operations such as processing vertex data, rendering graphics primitives, and / or processing image data. In one embodiment, the command streamer 403 determines load requests and data supply requests to memory and translates these requests into commands to the memory controller 405. In some embodiments, the command streamer 403 is a part of the memory controller 405. In some embodiments, one or more other components of the system, such as another processor in the system, instruct the GPE 410 with instructions and / or supply data to the GPE 410. These instructions and data are transmitted from the other component to the command streamer 403.
[0063] In various embodiments, the 3D pipeline 312 includes fixed function logic and programmable logic to process one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs. The graphics core array 414 provides a unified execution platform to execute instructions specified by the shader programs. The multipurpose execution logic (e.g., execution units) within the graphics core array 414 includes support for variable length code (VLC) and multiple thresholds of parallelism to provide further control of the graphics core array 414. The graphics core array 414 includes programmable and fixed function sections.
[0064] In some embodiments, the graphics core array 414 also includes execution logic to perform media functions, such as video and / or image processing. In one embodiment, the execution units also include logic to perform parallel general-purpose computing operations FIG. 1 of the processor core(s) 107 or FIG. 2 parallel or in conjunction with the general-purpose logic within the cores 202A-202N in the GPGPU 200.
[0065] Output data generated by threads executing on the graphics core array 414 can be output to memory in a unified return buffer (URB) 418. The URB 418 can store data for a number of threads. In some embodiments, the URB 418 can be used to send data between different threads executing on the graphics core array 414. In some embodiments, the URB 418 can additionally be used for synchronization between threads on the graphics core array and fixed function logic within the shared function logic 420.
[0066] In some embodiments, the graphics core array 414 is scalable, such that the array includes varying numbers of graphics cores each having varying numbers of execution units based on target performance and power profiles of the GPGPU 400. In one embodiment, the execution resources are dynamically scalable such that execution resources can be enabled or disabled as needed.
[0067] The graphics core array 414 is coupled with shared function logic 420 that includes resources shared between the graphics cores in the graphics core array. Shared functions within the shared function logic 420 are hardware logic units that provide specialized supplemental functionality to the graphics core array 414. In various embodiments, the shared function logic 420 includes, but is not limited to, samplers 421, math 422, and inter-thread communication (ITC) 423 logic. Additionally, some embodiments implement one or more caches 425 within the shared function logic 420.
[0068] Shared functions are implemented in cases where the demand for a given special-purpose function is insufficient to be contained within the graphics core array 414. Instead, a single instance of the special-purpose function is implemented as a separate entity within the shared function logic 420 and shared among the execution resources within the graphics core array 414. The precise set of functions shared among and included within the graphics core array 414 varies between embodiments. In some embodiments, particular shared functions within the shared function logic 420 that are heavily used by the graphics core array 414 can be included within the shared function logic 416 within the graphics core array 414. In various embodiments, the shared function logic 416 within the graphics core array 414 can include some or all of the logic within the shared function logic 420. In one embodiment, all of the logical elements within the shared function logic 420 can be duplicated within the shared function logic 416 of the graphics core array 414. In one embodiment, the shared function logic 420 is executed in order to support the shared function logic 416 within the graphics core array 414.
[0069] FIG. 5 is a block diagram of hardware logic of a graphics processor core 500 in accordance with some embodiments described herein. FIG. 5 Those elements of having the same reference number (or name) in the figures herein as elements in any other figure herein can operate or function in any manner similar to the manner described elsewhere herein but are not limited to such. In some embodiments, the illustrated graphics processor core 500 includes a number of sub-cores FIG. 4 within the graphics core array 414 of In some embodiments, the graphics processor core 500 is one or more graphics cores within a modular graphics processor. An example of the graphics processor core 500 is a graphics core slice, and a graphics processor as described herein can include multiple graphics core slices based on target power and performance envelopes. Each graphics core 500 can include fixed function blocks 530 coupled with a number of sub-cores 501A-501F (also referred to as sub-slices) that include both modular general purpose logic blocks and fixed function logic blocks.
[0070] In some embodiments, the fixed function blocks 530 include a geometry / fixed function pipeline 536 that can be shared by all of the sub-cores in the graphics processor 500, for example, in low performance and / or low power graphics processor implementations. In various embodiments, the geometry / fixed function pipeline 536 includes a 3D fixed function pipeline (e.g., as in 3D pipeline 312 in FIG. 3 and FIG. 4 in the 3D pipeline 312 in FIG. 4a unified return buffer manager for the unified return buffers 418, etc.
[0071] In one embodiment, the fixed function block 530 also includes a graphics SoC interface 537, a graphics microcontroller 538, and a media pipeline 539. The graphics SoC interface 537 provides an interface between the graphics core 500 and other processor cores within a system on a chip integrated circuit. The graphics microcontroller 538 is a programmable sub-processor that is configurable to manage various functions of the graphics processor 500 including thread dispatch, scheduling, and pre-emption. The media pipeline 539 (e.g., a video pipeline) is a processing pipeline that implements media operations including image and / or video FIG. 3 and FIG. 4 The media pipeline 539 (e.g., a media pipeline 316) includes logic to facilitate decoding, encoding, pre-processing, and / or post-processing of multimedia data, including image and video data. The media pipeline 539 implements media operations via requests to compute or sample logic within the sub-cores 501-501F.
[0072] In one embodiment, the SoC interface 537 enables the graphics core 500 to communicate with general application processor cores (e.g., CPUs) and / or other components within the SoC, including memory hierarchy elements such as shared L2 cache, system RAM, and / or embedded on-chip or on-package DRAM. The SoC interface 537 can also enable communication with fixed function devices within the SoC, such as camera imaging pipelines, and enable use of and / or implementation of global memory atoms that can be shared between the graphics core 500 and a CPU within the SoC. The SoC interface 537 can also implement power management controls for the graphics core 500 and enable an interface between a clock domain of the graphics core 500 and other clock domains within the SoC. In one embodiment, the SoC interface 537 enables receiving command buffers from a command streamer and global thread dispatcher that is configured to supply commands and instructions to each of one or more graphics cores within a graphics processor. These commands and instructions can be dispatched to the media pipeline 539 when media operations are to be performed, or to the geometry and fixed function pipeline when graphics processing operations are to be performed.
[0073] The graphics microcontroller 538 can be configured to perform various scheduling and management tasks for the graphics core 500. In one embodiment, the graphics microcontroller 538 can perform graphics and / or compute workload scheduling for individual graphics processing engines within execution unit (EU) arrays 502A-502F, 504A-504F within the corelets 501A-501F. In this scheduling model, host software executing on a CPU core of a SoC including the graphics core 500 can submit workloads via one of a number of graphics processor doorbells, which invokes a scheduling operation on the appropriate graphics engine. The scheduling operation includes determining which workload to run next, submitting the workload to a command streamer, pre-empting existing workloads running on the engine, monitoring progress of the workload, and notifying host software when the workload completes. In one embodiment, the graphics microcontroller 538 can also facilitate low power or idle states for the graphics core 500, providing the graphics core 500 with the ability to save and restore registers across low power state transitions independently of operating systems and / or graphics driver software on the system.
[0074] The graphics core 500 can have more or fewer than the illustrated number of sub-cores 501A-501F, up to N modular sub-cores. For each set of N sub-cores, the graphics core 500 can also include shared function logic 510, shared memory and / or cache memory 512, geometry / fixed function pipeline 514, and additional fixed function logic 516 for accelerating various graphics and compute processing tasks. The shared function logic 510 can include logic units (e.g., sampler logic, math logic, and / or inter-thread communication logic) that can be shared by every N sub-cores within the graphics core 500. FIG. 4 The shared function logic 420 is associated with logic units (e.g., sampler logic, math logic, and / or inter-thread communication logic) that can be shared by every N sub-cores within the graphics core 500. The shared memory and / or cache memory 512 can be a last level cache for the set of N sub-cores 501A-501F within the graphics core 500, and can also act as shared memory accessible by a number of the sub-cores. The geometry / fixed function pipeline 514 can be included within the fixed function block 530 instead of the geometry / fixed function pipeline 536, and can include the same or similar logic units.
[0075] In one embodiment, graphics core 500 includes additional fixed function logic 516 which can include various fixed function acceleration logic used by graphics core 500. In one embodiment, additional fixed function logic 516 includes an additional geometry pipeline used in position only shading. In position only shading, there are two geometry pipelines: a full geometry pipeline within geometry / fixed function pipelines 516, 536; and a cull pipeline, which is an additional geometry pipeline that can be included within additional fixed function logic 516. In one embodiment, the cull pipeline is a slimmed down version of the full geometry pipeline. The full pipeline and the cull pipeline can execute different instances of the same application, each with a separate context. Position only shading can hide the long cull operation of triangles that are discarded, enabling completion of shading earlier in some instances. For example, and in one embodiment, cull pipeline logic within additional fixed function logic 516 can execute position shaders in parallel with the main application, and often generate critical results faster than the full pipeline because the full pipeline only fetches and shades position attributes of vertices, without performing rasterization and rendering of pixels to a frame buffer. The cull pipeline can use the generated critical results to compute visibility information for all triangles, without having to consider whether those triangles are culled. The full pipeline, which in this instance can be referred to as a replay pipeline, can consume the visibility information in order to skip shaded triangles that are culled, shading only the visible triangles that are ultimately passed to a rasterization stage.
[0076] In one embodiment, additional fixed function logic 516 can also include machine learning acceleration logic, such as fixed function matrix multiplication logic, for implementations including machine learning training or inferencing.
[0077] Each graphics sub-core 501A to 501F includes a set of execution resources that can be used to perform graphics operations, media operations, and compute operations in response to requests from the graphics pipeline, media pipeline, or shader programs. Graphics sub-cores 501A to 501F include: multiple EU arrays 502A to 502F, 504A to 504F; thread dispatch and inter-thread communication (TD / IC) logic 503A to 503F; 3D (e.g., texture) samplers 505A to 505F; media samplers 506A to 506F; shader processors 507A to 507F; and shared local memory (SLM) 508A to 508F. Each EU array 502A-502F and 504A-504F includes multiple execution units, which are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logical operations to service graphics, media, or compute operations, including graphics programs, media programs, or compute shader programs. TD / IC logic 503A-503F performs local thread dispatch and thread control operations for execution units within a sub-core and facilitates communication between threads executing on the execution units of the sub-core. 3D samplers 505A-505F can read textures or other 3D graphics-related data into memory. 3D samplers can read texture data in different ways based on the configured sample state and the texture format associated with a given texture. Media samplers 506A-506F can perform similar read operations based on the type and format associated with the media data. In one embodiment, each graphics sub-core 501A-501F can alternately include unified 3D and media samplers. Threads executing on execution units within each of sub-cores 501A- 501F may utilize shared local memory 508A- 508F within each sub-core to enable threads executing within a thread group to execute using a common pool of on-chip memory.
[0078] Execution Units
[0079] FIGS. 6A-6B Thread execution logic 600 is shown including an array of processing elements employed in a graphics processor core according to embodiments described herein. FIGS. 6A-6B Those elements having the same reference numbers (or names) as elements in any other figures herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein. FIG. 6A An overview of thread execution logic 600 is shown, which may include a FIG. 5 A variation of the hardware logic of each sub-core 501A to 501F. FIG. 6B Exemplary internal details of an execution unit are shown.
[0080] like FIG. 6A As shown in FIG, in some embodiments, thread execution logic 600 includes a shader processor 602, a thread dispatcher 604, an instruction cache 606, a scalable execution unit array including a plurality of execution units 608A to 608N, a sampler 610, a data cache 612, and a data port 614. In one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., execution units 608A, 608B, 608C, 608D, through any one of 608N-1 and 608N) based on the computational requirements of the workload. In one embodiment, the included components are interconnected via an interconnect structure that links to each of the components. In some embodiments, thread execution logic 600 includes one or more connections to a memory (e.g., system memory or cache memory) through the instruction cache 606, the data port 614, the sampler 610, and one or more of the execution unit arrays 608A to 608N. In some embodiments, each execution unit (e.g., 608A) is an independently programmable general-purpose computing unit capable of executing multiple simultaneous hardware threads while processing multiple data elements in parallel for each thread. In various embodiments, the array of execution units 608A to 608N is scalable to include any number of individual execution units.
[0081] In some embodiments, execution units 608A through 608N are primarily used to execute shader programs. Shader processor 602 can process various shader programs and dispatch execution threads associated with the shader programs via thread dispatcher 604. In one embodiment, the thread dispatcher includes logic for arbitrating thread initiation requests from the graphics and media pipelines and instantiating the requested threads on one or more execution units 608A through 608N. For example, the geometry pipeline can dispatch vertex processing, tessellation, or geometry processing threads to thread execution logic for processing. In some embodiments, thread dispatcher 604 can also process runtime thread generation requests from executing shader programs.
[0082] In some embodiments, execution units 608A-608N support single program multiple data (SPMD) instructions to process data elements simultaneously. The program for an SPMD instruction can be defined by flexibly compiling an original program to generate the appropriate number of threads, or by loading separate programs that are executed in parallel. In some embodiments, execution units 608A-608N support direct state addressing (DSA) and texture sampling instructions in a single instruction, multiple data (SIMD) fashion, allowing threads to process different states or texture locations in SIMD instructions.
[0083] Each of execution units 608A-608N operates on arrays of data elements. The number of data elements is the "execution size," or the number of channels that the instruction can use. An execution channel is a logical unit of execution for data element access, masking, and flow control. The number of channels may be independent of the number of physical ALUs or FPUs for a particular graphics processor. In some embodiments, execution units 608A-608N also support integer and floating-point data types.
[0084] The execution unit instruction set can include SIMD instructions. Various data elements can be stored in registers and the execution unit will process data elements based on their data size. For example, when operating on a 256-bit wide vector, the 256 bits of the vector are stored in registers and the execution unit operates on eight 32-bit data elements, sixteen 16-bit data elements, thirty-two 8-bit data elements, or sixteen 4-bit data elements (one quarter word QW size elements). However, different vector widths are possible and the execution units are not limited to a 256-bit width. For example, the data element sizes can be 16, 32, 64, 128 bits, etc. The execution units support SIMD operations and can also include floating point and / or integer arithmetic instruction support.
[0085] In one embodiment, one or more execution units can be combined in a fused execution unit 609A-609N that has thread control logic (607A-607N) common to the fused EU. Multiple EUs can be fused into an EU group. Each EU in the fused EU group can be configured to execute a separate SIMD hardware thread. The number of EUs in a fused EU group can vary from one embodiment to another. In addition, different SIMD widths can be executed per EU including, but not limited to, SIMD8, SIMD16, and SIMD32. Each fused graphics execution unit 609A-609N includes at least two execution units. For example, fused execution unit 609A includes first EU 608A, second EU 608B, and thread control logic 607A common to first EU 608A and second EU 608B. Thread control logic 607A controls threads executing on fused graphics execution unit 609A, allowing each EU within fused execution units 609A-609N to use a common instruction pointer register for execution.
[0086] One or more internal instruction caches (e.g., 606) are included in the thread execution logic 600 to cache thread instructions for the execution units. In some embodiments, one or more data caches (e.g., 612) are included to cache data for threads to be executed. In some embodiments, a sampler 610 is included to provide texture sampling for 3D operations and to provide media sampling for media operations. In some embodiments, the sampler 610 includes specialized texture or media sampling functionality to process texture or media data during sampling before being provided to the execution units.
[0087] During execution, the graphics and media pipeline sends thread initiation requests to the thread execution logic 600 via thread generation and dispatch logic. Once a set of geometry objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 602 is invoked to further calculate output information and cause results to be written to an output surface (e.g., color buffer, depth buffer, stencil buffer, etc.). In some embodiments, the pixel or fragment shader calculates values for each vertex attribute that is interpolated across the rasterized object. In some embodiments, the pixel processor logic within the shader processor 602 then executes a pixel or fragment shader program supplied by an application programming interface (API). To execute the shader program, the shader processor 602 dispatches threads to execution units (e.g., 608A) via the thread dispatcher 604. In some embodiments, the shader processor 602 uses texture sampling logic in the sampler 610 to access texture data stored in a texture map stored in memory. Arithmetic operations on the texture data and input geometry data calculate pixel color data for each geometric fragment, or discard one or more pixels without further processing.
[0088] In some embodiments, the data port 614 provides a memory access mechanism for the thread execution logic 600 to output processed data to memory for further processing on a graphics processor output pipeline. In some embodiments, the data port 614 includes or is coupled to one or more cache memories (e.g., data cache 612) to cache data for memory access via the data port.
[0089] As shown in FIG. 6B The graphics execution unit 608 can include, in one embodiment, an instruction fetch unit 637, a general register file array (GRF) 624, an architecture register file array (ARF) 626, a thread arbiter 622, a send unit 630, a branch unit 632, a set of SIMD floating point units (FPUs) 634, and in one embodiment a set of dedicated integer SIMD ALUs 635. The GRF 624 and ARF 626 include the set of general and architecture register files associated with each synchronized hardware thread that can be active in the graphics execution unit 608. In one embodiment, per-thread architecture state is maintained in the ARF 626, while data used during thread execution is stored in the GRF 624. The execution state of each thread, including the instruction pointer of each thread, can be held in thread-specific registers in the ARF 626.
[0090] In one embodiment, graphics execution unit 608 has an architecture which is a combination of a swanky multi-threaded (SMT) with fine-grained interwoven multithreading (IMT). The architecture has a modular configuration which can be tuned at design time to a target number of simultaneous threads and a target register file size per thread, in which execution unit resources are partitioned across the logic used to execute a plurality of simultaneous threads.
[0091] In one embodiment, graphics execution unit 608 can co-issue multiple instructions which can each be different instructions. Thread arbiter 622 of graphics execution unit thread 608 can dispatch instructions to one of send unit 630, branch unit 642, or SIMD FPU(s) 634 for execution. Each execution thread can have access to 128 general purpose registers within GRF 624, where each register can store 32 bytes that can be accessed as a SIMD 8-element vector with 32-bit data elements. In one embodiment, each execution unit thread has access to 4 kilobytes within GRF 624, although embodiments are not so limited as more or less register resources can be provided in other embodiments. In one embodiment, up to seven threads can execute synchronously, although the number of threads per execution unit can also vary by embodiment. In an embodiment where seven threads have access to 4 kilobytes, GRF 624 can store a total of 28 kilobytes. Flexible addressing modes can permit multiple registers to be addressed together to efficiently establish wider registers or to represent stride rectangular block data structures.
[0092] In one embodiment, memory operations, sampler operations, and other longer latency system communications are dispatched via “send” instructions executed by message passing send unit 630. In one embodiment, branch instructions are dispatched to a dedicated branch unit 632 in order to facilitate SIMD divergence and eventual convergence.
[0093] In one embodiment, graphics execution unit 608 includes one or more SIMD floating point units (FPUs) 634 to perform floating point operations. In one embodiment, FPU(s) 634 also support integer computation. In one embodiment, FPU(s) 634 can SIMD execute up to a number M of 32-bit floating point (or integer) operations, or up to 2M of 16-bit integer or 16-bit floating point operations. In one embodiment, at least one of FPU(s) 634 provides an extended math capability supporting high throughput transcendental math functions and double precision 64-bit floating point. In some embodiments, a set of 8-bit integer SIMD ALUs 635 also are present and can be specifically optimized to perform operations associated with machine learning computations.
[0094] In one embodiment, an array of multiple instances of graphics execution unit 608 can be instantiated when graphics sub-cores are grouped (e.g., sub-slices). For scalability, the product architecture can choose the exact number of execution units per sub-core grouping. In one embodiment, execution unit 608 can execute instructions across multiple execution lanes. In further embodiments, each thread executed on graphics execution unit 608 executes on a different lane.
[0095] FIG. 7 is a block diagram illustrating a graphics processor instruction format 700 according to some embodiments. In one or more embodiments, the graphics processor execution unit supports an instruction set having instructions in multiple formats. Solid-line boxes illustrate components that are typically included in execution unit instructions, while dashed lines include optional components or components that are included only in a subset of instructions. In some embodiments, the instruction format 700 described and illustrated are macroinstructions because they are instructions supplied to the execution unit, as opposed to micro-operations generated from instruction decoding (once the instruction is processed).
[0096] In some embodiments, the graphics processor execution unit natively supports instructions in the 128-bit instruction format 710. A 64-bit compact instruction format 730 may be used for some instructions based on the selected instruction, a number of instruction options, and the number of operands. The native 128-bit instruction format 710 provides access to all instruction options, while some options and operations are restricted to the 64-bit format 730. The native instructions available in the 64-bit format 730 vary depending on the embodiment. In some embodiments, instructions are partially compressed using a set of index values in the index field 713. The execution unit hardware references a set of compression tables based on the index values and uses the compression table output to reconstruct the native instruction in the 128-bit instruction format 710.
[0097] For each format, the instruction opcode 712 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a synchronous add operation across each color channel, where the color channel represents a texture element or a picture element. By default, the execution unit executes each instruction across all data channels of the operand. In some embodiments, the instruction control field 714 enables control of certain execution options, such as channel selection (e.g., prediction) and data channel sorting (e.g., blending). For instructions using the 128-bit instruction format 710, the execution size field 716 limits the number of data channels to be executed in parallel. In some embodiments, the execution size field 716 is not available for the 64-bit compact instruction format 730.
[0098] Some execution unit instructions have up to three operands, including two source operands (src0 720, src1 722) and one destination 718. In some embodiments, the execution unit supports dual destination instructions in which one of the destinations is implicit. Data operation instructions can have a third source operand (e.g., SRC2 724), where the instruction opcode 712 determines the number of source operands. The last source operand of an instruction can be an immediate (e.g., hard coded) value passed with the instruction.
[0099] In some embodiments, the 128-bit instruction format 710 includes a access / address mode field 726 that, for example, defines whether a direct register addressing mode or an indirect register addressing mode is used. When using direct register addressing mode, the register address for one or more operands is provided directly by bits in the instruction.
[0100] In some embodiments, the 128-bit instruction format 710 includes an access / address mode field 726 that specifies the address mode and / or access mode of the instruction. In one embodiment, the access mode is used to define the data access alignment for the instruction. Some embodiments support access modes including a 16-byte aligned access mode and a 1-byte aligned access mode, where the byte alignment of the access mode determines the access alignment of the instruction operands. For example, when in a first mode, the instruction can use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction can use 16-byte aligned addressing for all source and destination operands.
[0101] In one embodiment, the address mode portion of the access / address mode field 726 determines whether the instruction uses direct addressing or indirect addressing. When using direct register addressing mode, the register address for one or more operands is provided directly by bits in the instruction. When using indirect register addressing mode, the register address for one or more operands can be calculated based on an address register value and an address immediate field in the instruction.
[0102] In some embodiments, instructions are grouped based on the opcode 712 bit field to simplify opcode decoding 740. For 8-bit opcodes, bits 4, 5, and 6 allow the execution unit to determine the type of opcode. The precise opcode grouping shown is merely exemplary. In some embodiments, move and logic opcode group 742 includes data movement and logic instructions (e.g., move (mov), compare (cmp)). In some embodiments, move and logic group 742 shares five most-significant bits (MSB), with move (mov) instructions taking the form 0000xxxxb, and logic instructions taking the form 0001xxxxb. Flow control instruction group 744 (e.g., call, jmp) includes instructions that take the form 0010xxxxb (e.g., Ox20). Hybrid instruction group 746 includes a mix of instructions, including synchronization instructions (e.g., wait, send) that take the form 0011xxxxb (e.g., Ox30). Parallel mathematical instruction group 748 includes arithmetic instructions (e.g., add, mul) that take the form 0100xxxxb (e.g., Ox40) that operate on components. Parallel mathematical group 748 performs arithmetic operations across data lanes in parallel. Vector mathematical group 750 includes arithmetic instructions (e.g., dp4) that take the form 0101xxxxb (e.g., Ox50). Vector mathematical group performs arithmetic operations on vector operands, such as a dot product operation.
[0103] Graphics Pipeline
[0104] FIG. 8 is a block diagram of another embodiment of a graphics processor 800. FIG. 8 Those elements of having the same reference number (or name) in the Figures as elements identified herein are operated or function in any manner similar to that described elsewhere herein, but are not limited to such.
[0105] In some embodiments, graphics processor 800 includes geometry pipeline 820, media pipeline 830, display engine 840, thread execution logic 850, and render output pipeline 870. In some embodiments, graphics processor 800 is a graphics processor included in a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or via commands issued to the graphics processor 800 via the ring interconnect 802. In some embodiments, ring interconnect 802 couples graphics processor 800 to other processing components such as other graphics processors or general-purpose processors. Commands from the ring interconnect 802 are interpreted by a command streamer 803, which supplies instructions to individual components of graphics processor 800.
[0106] In some embodiments, command streamer 803 directs the operation of vertex fetcher 805, which reads vertex data from memory and executes vertex processing commands provided by command streamer 803. In some embodiments, vertex fetcher 805 supplies vertex data to vertex shader 807, which performs coordinate space transformation and lighting operations on each vertex. In some embodiments, vertex fetcher 805 and vertex shader 807 execute vertex processing instructions via dispatch of execution threads to execution units 852A-B by thread dispatcher 831.
[0107] In some embodiments, execution units 852A-B are vector processors that are capable of processing graphics and media instructions. In some embodiments, execution units 852A-B have an attached Ll cache 851, which is dedicated to each array or shared between arrays. The cache can be configured as a data cache, an instruction cache, or a single cache that is partitioned into separate data and instruction caches.
[0108] In some embodiments, geometry pipeline 820 includes a tessellation component for hardware-accelerated tessellation of 3D objects. In some embodiments, programmable hull shader 811 configures tessellation operations. Programmable domain shader 817 provides back-end evaluation of tessellation output. Tessellator 813 operates in the direction of hull shader 811 and contains specialized logic for generating a detailed set of geometric objects based on a coarse geometric model that is provided as input to geometry pipeline 820. In some embodiments, if tessellation is not used, the tessellation component (e.g., hull shader 811, tessellator 813, domain shader 817) can be bypassed.
[0109] In some embodiments, a complete geometric object can be processed by the geometry shader 819 via one or more threads dispatched to the execution units 852A-B, or can go directly to the clipper 829. In some embodiments, the geometry shader operates on an entire geometric object (rather than a vertex or patch of vertices as in previous stages of the graphics pipeline). If tessellation is disabled, the geometry shader 819 receives input from the vertex shader 807. In some embodiments, the geometry shader 819 can be programmed by a geometry shader program to perform geometric tessellation when the tessellation unit is disabled.
[0110] Prior to rasterization, the clipper 829 processes the vertex data. The clipper 829 can be a fixed function clipper or a programmable clipper with clip and geometry shader functionality. In some embodiments, the rasterizer and depth test components 873 in the render output pipeline 870 dispatch pixel shaders to convert a geometric object into a per-pixel representation. In some embodiments, the pixel shader logic is included in the thread execution logic 850. In some embodiments, an application can bypass the rasterizer and depth test components 873 and access un-rasterized vertex data via the outflow unit 823.
[0111] The graphics processor 800 has an interconnect bus, interconnect fabric, or some other interconnect mechanism to allow data and messages to be passed between components of the graphics processor 800. In some embodiments, the execution units 852A-B and associated circuitry (e.g., LI cache 851, sampler 854, texture cache 858, etc.) are interconnected via a data port 856 to perform memory accesses and to communicate with other processor elements and the render output pipeline components of the processor. In some embodiments, the sampler 854, cache 851, 858, and execution units 852A-B each have separate memory access ports to the data port 856. In one embodiment, the texture cache 858 can also be configured as a sampler cache.
[0112] In some embodiments, the render output pipeline 870 includes a rasterizer and depth test component 873 that converts vertex-based objects into related pixel-based representations. In some embodiments, the rasterizer logic includes a windower / masker unit for performing fixed function triangle and line rasterization. Related render cache 878 and depth cache 879 are also available in some embodiments. Pixel operation component 877 performs pixel-based operations, although in some instances pixel operations associated with 2D operations (e.g., with blended bit block image transfers) are performed by 2D engine 841 or are replaced by display controller 843 using an overlay display plane at display time. In some embodiments, a shared L3 cache 875 is available for all graphics components, allowing sharing of data without use of main system memory.
[0113] In some embodiments, graphics processor media pipeline 830 includes a media engine 837 and a video front-end 834. In some embodiments, video front-end 834 receives pipeline commands from the command streamer 803. In some embodiments, media pipeline 830 includes a separate command streamer. In some embodiments, video front-end 834 processes media
[0114] In some embodiments, graphics processor 800 includes a display engine 840. In some embodiments, display engine 840 is external to processor 800 and couples with the graphics processor through the ring interconnect 802, or some other interconnect bus or fabric. In some embodiments, display engine 840 includes a 2D engine 841 and a display controller 843. In some embodiments, display engine 840 contains special purpose logic that is, in some embodiments, capable of independent operation with a 3D pipeline 810. In some embodiments, display controller 843 couples with a display device (not illustrated), which can be a system integrated display device, as in a laptop computer, or an external display device attached via an display device connector.
[0115] In some embodiments, geometry pipeline 820 and media pipeline 830 can be configured to perform operations based on a number of graphics and media programming interfaces and not specific to any one application programming interface (API). In some embodiments, driver software for a graphics processor can translate API calls or calls that are specific to particular graphics or media libraries into commands that can be processed by the graphics processor. In some embodiments, support is provided for the Open Graphics Library (OpenGL) from the Khronos Group, the Open Computing Language (OpenCL), and / or the Vulkan graphics and compute API. In some embodiments, support can also be provided for the Direct3D library from the Microsoft Corporation. In some embodiments, combinations of these libraries can be supported. Support can also be provided for the Open Source Computer Vision Library (OpenCV). Future APIs that have a mapping from their pipeline to the pipeline of a graphics processor will also be supported if such a mapping can be made.
[0116] Graphics Pipeline Programming
[0117] FIG. 9A FIG. 9 is a block diagram illustrating a graphics processor command format 900 according to some embodiments. FIG. 9B FIG. 10 is a block diagram illustrating a graphics processor command sequence 910 according to an embodiment. FIG. 9A The solid lined boxes in FIG. 10 illustrate the components that are generally included in a graphics command while the dashed lined boxes illustrate optional implementation dependent components. FIG. 9A The exemplary graphics processor command format 900 of FIG. 9 includes data fields to identify a client 902, a command operation code (opcode) 904, and data 906 for the command. Some commands can include an optional sub-opcode 905 and a command size 908.
[0118] In some embodiments, the client 902 defines a client unit of the graphics processing unit that processes the command data. In some embodiments, the graphics processor command parser examines a client field of each command to determine the further processing to be performed on the command and routes the command data to the appropriate client unit. In some embodiments, a graphics processor client unit includes a memory interface unit, render units, a 2D unit, a 3D unit, and a media unit. Each client unit has a respective processing pipeline to process the commands. Once a command is received by a client unit, the client unit reads the operation code 904 and sub-op code 905 (if present) to determine the operation to be performed on the command. The client unit uses information in the data field 906 to perform the command. For some commands, an explicit command size 908 is expected to define the size of the command. In some embodiments, the command parser automatically determines the size of at least some of the commands based on the command opcode. In some embodiments, the commands are aligned via multiples of double-words.
[0119] FIG. 9B The flow diagram in FIG. 9 illustrates an exemplary graphics processor command sequence 910. In some embodiments, a software or firmware of a data processing system featuring embodiments of a graphics processor uses a version of the command sequence shown to set up, execute, and terminate a set of graphics operations. A sample command sequence is shown and described for illustrative purposes only. The embodiments are not limited to these specific commands or to the order of execution of the commands as shown. Moreover, the commands can be issued as batch of commands in the command sequence such that multiple commands are issued in at least partially simultaneous fashion. The sample command sequence illustrates the general process of the graphics processor to begin, execute, and end a graphics operation.
[0120] In some embodiments, the graphics processor command sequence 910 can begin with a pipeline flush command 912 to cause any active graphics pipelines to finish any current pending commands for the pipeline. In some embodiments, the 3D pipeline 922 and media pipeline 924 are not operating at the same time. The pipeline flush is performed to cause the active graphics pipeline to finish any pending commands. In response to the pipeline flush, the command parser for the graphics processor will stop processing commands until the active draw engine finishes the pending operations and invalidates the relevant read caches. Optionally, any data in the render caches that is marked as 'dirty' can be flushed to memory. In some embodiments, the pipeline flush command 912 can be used for pipeline synchronization or before placing the graphics processor in a low power state.
[0121] In some embodiments, a pipeline selection command 913 is used when the command sequence requires the graphics processor to explicitly switch between pipelines. In some embodiments, only one pipeline selection command 913 is needed in the execution context before issuing the pipeline commands, unless the context is issuing commands to two pipelines. In some embodiments, a pipeline flush clear command 912 is needed just before the pipeline switch via the pipeline selection command 913.
[0122] In some embodiments, pipeline control commands 914 configure the graphics pipeline for operation and program the 3D pipeline 922 and media pipeline 924. In some embodiments, the pipeline control commands 914 configure the pipeline state of the active pipeline. In one embodiment, the pipeline control commands 914 are used for pipeline synchronization and to clear data from one or more cache memories within the active pipeline before processing a batch of commands.
[0123] In some embodiments, a return buffer state command 916 is used to configure a set of return buffers for a respective pipeline to write data. Some pipeline operations require allocation, selection, or configuration of one or more return buffers in which to write intermediate data during processing. In some embodiments, the graphics processor also uses one or more return buffers to store output data and perform cross-thread communication. In some embodiments, the return buffer state 916 includes selecting the size and number of return buffers to use for a set of pipeline operations.
[0124] The remaining commands in the command sequence differ based on the active pipeline for operation. Based on the pipeline determination 920, the command sequence is tailored for either the 3D pipeline 922 starting in 3D pipeline state 930, or the media pipeline 924 starting at media pipeline state 940.
[0125] The commands for the 3D pipeline state 930 include 3D state setting commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables to be configured before processing 3D primitive commands. The values for these commands are determined at least in part based on the particular 3D API in use. In some embodiments, the 3D pipeline state 930 commands can also selectively disable or bypass certain pipeline elements if those elements will not be used.
[0126] In some embodiments, 3D primitive 932 commands are used to submit 3D primitives to be processed by the 3D pipeline. Commands and associated parameters passed to the graphics processor via a 3D primitive 932 command are forwarded to a vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 932 command data to generate a number of vertex data structures. The vertex data structures are stored in one or more vertex buffers. In some embodiments, 3D primitive 932 commands are used to perform vertex operations on 3D primitives via a vertex shader. To process the vertex shader, the 3D pipeline 922 dispatches shader execution threads to the graphics processor execution units.
[0127] In some embodiments, the 3D pipeline 922 is triggered via an execute 934 command or event. In some embodiments, a register write triggers command execution. In some embodiments, execution is triggered via a 'go' or 'kick' command in the command sequence. In one embodiment, a pipeline synchronization command is used to trigger command execution in order to flush the command sequence with no-ops through the graphics pipeline. The 3D pipeline will perform geometry processing for a 3D primitive. Once the processing is completed, the resulting geometry is rasterized and the pixel engine performs shading of the resulting pixels. For these operations, additional commands can also be included for controlling the pixel shading and pixel back end operations.
[0128] In some embodiments, the graphics processor command sequence 910 follows the media pipeline 924 path when performing media operations. Generally, the specific use and manner of programming for the media pipeline 924 depends on the media or compute operations to be performed. In the case of media decode, the specific media decode operations can be offloaded to the media pipeline. In some embodiments, the media pipeline can also be bypassed and media decode can be performed entirely or partially in the processor core using resources provided for general-purpose processing. In one embodiment, the media pipeline also includes elements for general-purpose graphics processor unit (GPGPU) operations, where the graphics processor is used to execute SIMD vector operations using computational shader programs that are not explicitly related to the rendering of graphics primitives.
[0129] In some embodiments, the media pipeline 924 is configured in a similar manner as the 3D pipeline 922. A set of commands to configure the media pipeline state 940 is dispatched or placed into a command queue, prior to the media object command 942. In some embodiments, the commands 940 for the media pipeline state include data to configure media pipeline elements that will be used to process the media object. This includes data to configure video decode and video encode logic within the media pipeline, such as encoding or decoding formats. In some embodiments, the commands 940 for the media pipeline state also support the use of one or more pointers to "indirect" state elements that contain a batch of state settings.
[0130] In some embodiments, the media object command 942 supplies a pointer to a media object for processing by the media pipeline. The media object includes a memory buffer that contains video data to be processed. In some embodiments, all of the media pipeline state must be valid prior to issuing the media object command 942. Once the pipeline state is configured and the media object command 942 is queued, the media pipeline 924 is triggered via an execute 944 command or equivalent execution event (e.g., register write). The output from the media pipeline 924 can then be post-processed by operations provided by the 3D pipeline 922 or the media pipeline 924. In some embodiments, GPGPU operations are configured and executed in a similar manner as media operations.
[0131] Graphics Software Architecture
[0132] FIG. 10 An exemplary graphics software architecture for a data processing system 1000 is shown in accordance with some embodiments. In some embodiments, the software architecture includes a 3D graphics application 1010, an operating system 1020, and at least one processor 1030. In some embodiments, the processor 1030 includes a graphics processor 1032 and one or more general-purpose processor cores 1034. The graphics application 1010 and operating system 1020 each execute in a system memory 1050 of the data processing system.
[0133] In some embodiments, the 3D graphics application 1010 contains one or more shader programs including shader instructions 1012. The shader language instructions can be in a high-level shader language, such as the High-Level Shader Language (HLSL) or the OpenGL Shader Language (GLSL). The application also includes executable instructions 1014 in a machine language suitable for execution by the general- purpose processor cores 1034. The application also includes graphics objects 1016 defined by vertex data.
[0134] In some embodiments, operating system 1020 is from Microsoft Corporation The operating system 1020 may be a graphics API 1022, such as a Direct3D API, an OpenGL API, or a Vulkan API. When the Direct3D API is in use, the operating system 1020 uses a front-end shader compiler 1024 to compile any shader instructions 1012 in HLSL into a lower-level shader language. The compilation may be a just-in-time (JIT) compilation, or the application may precompile the shader. In some embodiments, during the compilation of the 3D graphics application 1010, high-level shaders are compiled into low-level shaders. In some embodiments, the shader instructions 1012 are provided in an intermediate form, such as a version of the Standard Portable Intermediate Representation (SPIR) used by the Vulkan API.
[0135] In some embodiments, user-mode graphics driver 1026 includes a back-end shader compiler 1027 that converts shader instructions 1012 into a hardware-specific representation. When using the OpenGL API, shader instructions 1012 in the GLSL high-level language are passed to user-mode graphics driver 1026 for compilation. In some embodiments, user-mode graphics driver 1026 uses operating system kernel-mode functionality 1028 to communicate with kernel-mode graphics driver 1029. In some embodiments, kernel-mode graphics driver 1029 communicates with graphics processor 1032 to dispatch commands and instructions.
[0136] IP Core Implementation
[0137] One or more aspects of at least one embodiment may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, a machine-readable medium may include instructions representing the various logic within the processor. When read by a machine, the instructions may cause the machine to manufacture logic for performing the techniques described herein. This type of representation (referred to as an "IP core") is a reusable unit of logic for an integrated circuit that can be stored on a tangible, machine-readable medium as a hardware model that describes the structure of the integrated circuit. The hardware model may be supplied to each consumer or manufacturing facility that loads the hardware model on a manufacturing machine that manufactures the integrated circuit. The integrated circuit may be manufactured so that the circuit performs the operations described in association with any of the embodiments described herein.
[0138] FIG. 11 is a block diagram illustrating an IP core development system 1100 that can be used to fabricate integrated circuits to perform operations in accordance with embodiments. The IP core development system 1100 can be used to generate modular, re-usable designs that can be incorporated into larger designs or used to construct an entire integrated circuit (e.g., an SOC integrated circuit). A design facility 1130 can employ a high-level programming language (e.g., C / C++) to generate a software simulation 1110 of the IP core design. The software simulation 1110 can be used to design, test, and verify the behavior of the IP core using a simulation model 1112. The simulation model 1112 can include functional, behavioral, and / or timing simulations. The simulation model 1112 can then be used to create or synthesize a register transfer level (RTL) design 1115. The RTL design 1115 is an abstraction of the behavior of the integrated circuit (including associated logic) that models the flow of digital signals between hardware registers, including the associated logic performed thereon. In addition to an RTL design 1115, a lower- level design, such as a logic level or transistor level design, can also be
[0139] The RTL design 1115 or equivalent can be further synthesized, formatted, or prepared to produce a hardware model 1120, which can be in a hardware description language (HDL) or some other representation of the hardware design. The HDL can be further simulated or tested to verify the IP core design. The IP core design can be stored using a non-volatile storage medium 1140 (e.g., hard disk, flash memory, or any non-transitory storage medium) for delivery to a third party fabrication facility 1165. Alternatively, the IP core design can be transmitted (e.g., via the Internet) over a wired 1150 or wireless 1160 connection. The fabrication facility 1165 can then fabricate an integrated circuit based at least in part on the IP core design. The fabricated integrated circuit can be configured to perform operations in accordance with at least one embodiment described herein.
[0140] Exemplary System on a Chip Integrated Circuit
[0141] FIGS. 12-14 Exemplary integrated circuits and related graphics processors that can be fabricated using one or more IP cores in accordance with various embodiments described herein are illustrated. Other logic and circuitry can also be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0142] FIG. 121 is a block diagram illustrating an exemplary system-on-chip integrated circuit 1200 that can be manufactured using one or more IP cores according to an embodiment. The exemplary integrated circuit 1200 includes one or more application processors 1205 (e.g., CPUs), at least one graphics processor 1210, and may also include an image processor 1215 and / or a video processor 1220, any of which can be modular IP cores from the same or multiple different design facilities. The integrated circuit 1200 includes peripheral or bus logic, including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an I / O controller 1240. 2 S / I 2 The integrated circuit may also include a display device 1245 coupled to one or more of a High-Definition Multimedia Interface (HDMI) controller 1250 and a Mobile Industry Processor Interface (MIPI) display interface 1255. Storage may be provided by a flash memory subsystem 1260 (including flash memory and a flash memory controller). A memory interface may be provided via a memory controller 1265 to access SDRAM or SRAM memory devices. Some integrated circuits may also include an embedded security engine 1270.
[0143] FIGS. 13A-13B is a block diagram illustrating an exemplary graphics processor for use within a SoC according to embodiments described herein. FIG. 13A An exemplary graphics processor 1310 of a system-on-chip integrated circuit is shown that may be fabricated using one or more IP cores in accordance with an embodiment. FIG. 13B An additional exemplary graphics processor 1340 is shown that may be fabricated using one or more IP cores for a system-on-chip integrated circuit according to an embodiment. FIG. 13A Graphics processor 1310 is an example of a low-power graphics processor core. FIG. 13B The graphics processor 1340 is an example of a higher performance graphics processor core. Each of the graphics processors 1310, 1340 may be FIG. 12 A variant of the graphics processor 1210.
[0144] like FIG. 13AAs shown in FIG, the graphics processor 1310 includes a vertex processor 1305 and one or more fragment processors 1315A to 1315N (e.g., 1315A, 1315B, 1315C, 1315D, all the way to 1315N-1 and 1315N). The graphics processor 1310 can execute different shader programs via separate logic, such that the vertex processor 1305 is optimized to perform the operations of the vertex shader program, while the one or more fragment processors 1315A to 1315N perform fragment (e.g., pixel) shading operations for the fragment or pixel shader program. The vertex processor 1305 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. The fragment processor(s) 1315A to 1315N use the primitives and vertex data generated by the vertex processor 1305 to generate a frame buffer for display on the display device. In one embodiment, the fragment processor(s) 1315A through 1315N are optimized to execute fragment shader programs provided in the OpenGL API, which can be used to perform similar operations as pixel shader programs provided in the Direct 3D API.
[0145] In addition, the graphics processor 1310 also includes one or more memory management units (MMUs) 1320A to 1320B, one or more caches 1325A to 1325B, and one or more circuit interconnects 1330A to 1330B. The one or more MMUs 1320A to 1320B provide virtual to physical address mapping for the graphics processor 1310, including for the vertex processor 1305 and / or (multiple) fragment processors 1315A to 1315N. In addition to the vertex or image / texture data stored in the one or more caches 1325A to 1325B, the virtual to physical address mapping can also reference vertex or image / texture data stored in memory. In one embodiment, the one or more MMUs 1320A to 1320B can communicate with the system, including with FIG. 12 The graphics processor 1310 may be synchronized with other MMUs, including one or more MMUs associated with the one or more application processors 1205, the image processor 1215, and / or the video processor 1220, so that each processor 1205 to 1220 may participate in a shared or unified virtual memory system. In accordance with an embodiment, the one or more circuit interconnects 1330A to 1330B enable the graphics processor 1310 to interact with other IP cores within the SoC via an internal bus of the SoC or via a direct connection.
[0146] like FIG. 13B As shown in FIG, the graphics processor 1340 includes FIG. 13Athe one or more MMUs 1320A-B, caches 1325A-B, and circuit interconnect 1330A-B of the graphics processor 1310. The graphics processor 1340 includes one or more shader cores 1355A-1355N (e.g., 1355A, 1355B, 1355C, 1355D, 1355E, 1355F, up to 1355N-1, and 1355N) that provide a unified shader core architecture in which a single core or type or core can be tasked with executing all types of programmable shader code, including shader program code to implement vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present can vary in different embodiments and implementations. In addition, the graphics processor 1340 includes an inter-core task manager 1345 that acts as a thread dispatcher and task manager to accelerate tasks such as dispatching execution threads to one or more shader cores 1355A-1355N and to accelerate tessellation operations for block-based rendering in which rendering operations for a scene are subdivided in image space, for example, to exploit local spatial coherence within a scene or to optimize the use of internal caches.
[0147] FIG. 14 Additional exemplary graphics processor logic in accordance with embodiments described herein is shown. One embodiment provides graphics core 1400 that can be included within graphics processor 1210 and can be as described as FIG. 12 FIG. 13B The graphics cores 1400 include shared instruction cache 1402, texture units 1418, and cache memory / shared memory 1420 common to the execution resources within graphics cores 1400. Graphics cores 1400 can include multiple slices 1401A-1401N or partitions for each core, and a graphics processor can include multiple instances of graphics cores 1400. Slices 1401A-1401N can include support logic including a local instruction cache 1404A-1404N, a thread scheduler 1406A-1406N, a thread dispatcher 1408A-1408N, and a set of registers 1410A. To perform logical operations, slices 1401A-1401N can include a set of additional functional units (AFUs 1412A-1412N), floating point units (FPUs 1414A-1414N), integer arithmetic logic units (ALUs 1416-1416N), addressing calculation units (ACUs 1413A-1413N), double precision floating point units (DPFPUs 1415A-1415N), and matrix processing units (MPUs 1417A-1417N).
[0148] Some of these compute units operate at specific precisions. For example, FPUs 1414A-1414N can perform single precision (32-bit) and half precision (16-bit) floating point operations, while DPFPUs 1415A-1415N perform double precision (64-bit) floating point operations. ALUs 1416A-1416N can perform variable precision integer operations at 8-bit precision, 16-bit precision, and 32-bit precision, and can be configured for mixed precision operations. MPUs 1417A-1417N can also be configured for mixed precision matrix operations, including half precision floating point operations and 8-bit integer operations. MPUs 1417A-1417N can perform a variety of matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general matrix to matrix multiplication (GEMM). AFUs 1412A-1412N can perform additional logical operations not supported by floating point units or integer units, including trigonometric operations (e.g., sine, cosine, etc.).
[0149] Winograd Convolution Accelerator Architecture
[0150] Embodiments described herein provide a scalable hardware architecture for implementing a deep learning convolution accelerator with logic for performing generalized Winograd convolutions for a variety of kernel sizes and kernel spans without requiring hardware support for a variety of Winograd transforms. The deep learning convolution accelerator supports all kernel sizes by using a 1D filter Winograd F(4,3) method that supports convolutions with spans greater than 1. In one embodiment, the cost of Winograd transforms for high order kernels is minimized by reusing the transforms of low order F(4,3). In general, the architecture provided herein solves Winograd transform based CNN acceleration on multi-dimensional SIMD machines. The architecture improves convolution performance while improving compute per byte. The Winograd transform based CNN accelerator also provides about 2x improvement over other accelerators. Winograd achieves faster convolutions by reducing the number of multiplications within the compute operation by transforming the input and kernel (e.g., filter) into a form that enables higher performance for a given hardware resource. The Winograd transform is performed on the fly and broadcast to the parallel compute logic (e.g., SIMD, SIMT, etc.) without increasing the bandwidth requirements on the hardware. Additionally, the reduced number of compute operations also results in an increase in performance per watt.
[0151] Deep learning neural networks are compute intensive and large networks can be performance limited. Embodiments described herein provide custom acceleration logic for accelerating compute intensive workloads associated with deep learning neural networks such as convolutional neural networks (CNNs). In a CNN, each layer performs a 3D convolution on input feature maps, where each layer has different kernel size and span requirements. In deep learning inference applications, it is important to achieve the highest performance / watt and / or the highest performance / area. The above computations can be performed using direct convolution, but there is a new class of fast algorithms for convolutional neural networks based on the minimal filtering algorithm pioneered by Winograd. However, the Winograd method requires different compute structures for different kernel size / span requirements, which limits the efficient use of Winograd transforms in hardware accelerators for neural networks.
[0152] The embodiments described herein provide a unified hardware architecture based on F(4,3) Winograd kernel convolution and based on this foundation kernel to derive higher order convolutions. The derivation is performed in an optimized manner to achieve maximum performance per watt. The embodiments described herein utilize collinear kernel data transformation and input data transformation without incurring bandwidth increase and requiring minimal hardware complexity. The proposed hardware accelerator is configurable and scalable for efficient computation of 3D convolution using Winograd.
[0153] FIGS. 15A-15B Native convolution and Winograd based 3D convolution are shown. FIG. 15A A 3D convolution operation 1500 for a convolutional neural network is shown. FIG. 15B Details of Winograd based convolution 1510 and native convolution 1520 are shown. FIG. 15A The 3D convolution operation 1500 shown in FIG. 15 is critical for convolutional neural networks, which are applied to computer vision and speech acceleration applications. In 3D convolution, each input feature map (IFM) in a set of input feature maps 1502 (IFM1 to IFMn) is convolved 1504 with a corresponding convolution filter (K1) (e.g., K1, IFM1 to K1, IFMn). The partial results 1506 are summed to create a final output feature map (OFM) 1508. The processing operation for a convolutional neural network requires many such 3D convolutions to be performed on the same IFM to produce different OFMs.
[0154] As FIG. 15BAs illustrated in the middle, the embodiments described herein provide a hardware accelerator logic that utilizes Winograd-based convolution 1510 to implement CNN processing with reduced number of multiplications relative to native convolution 1520. In one embodiment, Winograd transform is used to compute a ID convolution for four pixels. This convolution can be repeated over the entire input feature map in order to compute the full output feature map. For example, a 1x3 kernel 1511 and a 1x6 input feature map patch 1512 can be transformed via a weight transform unit 1513 and an input transform unit 1514. The transformed 1x6 matrix can be processed via an element-wise multiplication unit 1515 that can perform the convolution. The convolution can be performed using six multiplications in order to generate a 1x6 intermediate output matrix. An output transform unit 1516 transforms the 1x6 intermediate output matrix to a 1x4 output 1518. Winograd-based convolution 1510 implements a reduced number of multiplications relative to native convolution 1520. In order to transform the 1x3 kernel 1511 and the 1x6 IMF patch 1512, native convolution 1520 utilizes a native convolution unit 1525, which requires 12 multiplications to generate the same 1x4 output matrix 1518. The acceleration architecture provided by the embodiments described herein significantly minimizes the transform cost by reusing the input transform across all output feature maps and computing the transform in parallel. In one embodiment, the transformed weights are stored in local memory in order to be reused across the entire output feature map. Different filter sizes can also be mapped to the same weight transform, which is a one-time operation per filter. Broadcasting of the transformed input amortizes the cost of the transform with negligible broadcast cost.
[0155] Winograd Support for General Purpose Kernel Sizes
[0156] FIG. 16 An architecture 1600 for 2D / 3D convolution based on F[4,3] convolution is illustrated. The illustrated architecture 1600 splits a 2D convolution filter into multiple ID kernels and performs Winograd-based convolution in order to produce a 2D convolution output. The same transform is applied to multiple IFMs in order to produce a 3D convolution output. In one embodiment, the illustrated architecture 1600 is implemented in part via hardware logic configured as a Winograd computation block 1640. The Winograd computation block 1640 uses a [1x6] tensor 1612 from an input feature map 1610 and a [1x3] tensor 1622 from a convolution kernel 1620 to perform a convolution operation 1615, resulting in a [1x4] tensor 1632 as part of an input feature map 1630.
[0157] For example, to perform a 1D convolution of size 1x5 using a 1x3 kernel, the following operations can be performed. A given 1x5 kernel K can be defined, where K = [1 2 3 4 5]. This 1x5 kernel can be decomposed into 1x3 sub-kernels K0 and K1, where K0 = [1 2 3] and K1 = [4 5 0]. Using K0 and K1, the computation logic can perform Winograd-based convolution with the respective patches of the IFM. The results can be combined in order to obtain the final result with kernel size [1X5]. Winograd convolution can be performed using 12 multiplications, yielding a 1.67x gain compared to the 20 multiplications for native convolution. The process is repeated over the 5 rows of the kernel in order to obtain the convolution result filter of size [5X5]. This way is generic and can be applied to filters of any size. The proposed hardware architecture handles the decomposition of the kernel and the computation of higher order kernels in multiple loops.
[0158] FIG. 17 Logic 1700 is shown for generalizing F(4,3) Winograd convolution to higher order kernels. Logic 1700 can be used to implement Winograd convolution for kernels of a general size using a decomposition into multiple [1x3] kernels. In one embodiment, the shown logic 1700 is implemented via a hardware processing element configured to perform Winograd convolution. Logic 1700 performs a first operation 1702 to determine values N and M. N specifies the number of Winograd operation rounds to be performed for each row, and M specifies the number of rows of operations to perform. N is computed by taking ceiling(kernel_size / 3), and M is determined by (kernel_size). For an exemplary kernel_size of 5, N = 2 and M = 5. Logic 1700 also performs a second operation 1704 to maintain values defining the next input data row (X) and the next kernel row (K). Before performing a set of operations for a row, logic 1700 performs operation 1705 to determine whether N is equal to zero. If N is not equal to zero, operation 1706 is performed to load the 1D input tensor (e.g., X(1:6)) and kernel data (X(1:3)) for processing element operations. Operation 1710 provides the input and kernel tensor data to a Winograd processing element (WPE) that performs Winograd transformation of the input and weight tensors, element-wise multiplication between the transformed input and weight, and vector accumulation operations that add the output of the multiplication (partial output feature map) to the tensor values in an accumulator register, as shown by the illustrated operations, {Wx = B n n T Xn ; Wk = Gk n ; Y n = Y n_先前 + Wx * Wk; Y n_先前 = Y n}. Wx and Wk are Winograd transforms of the input and kernel tensor data, and Wx * Wk is an element-wise multiplication of the Winograd-transformed tensors. Y n_先前 denotes the sum of the previous partial output feature maps. B T and G are F(4,3) Winograd transform matrices defined for the input and kernel data as follows:
[0159]
[0160]
[0161] The logic 1700 then performs operation 1712 to select the next set of input and weight data within the row {N = N - 1; X = X « 3; K = K « 3}. The logic 1700 returns to operation 1705 to determine whether additional rounds are needed for the row. Once the rounds for the row are complete (e.g., N == 0), the logic 1700 can move to the next row (e.g., M = M - 1 at operation 1708). The logic returns to operation 1704 until all rows are processed (e.g., M == 0 at operation 1709). Once the logic 1700 has completed processing for each row, as determined via operation 1709, the logic 1700 can perform an inverse transform operation 1714 in which the summed output feature maps Y n perform an inverse Winograd transform (e.g., Y 输出 = A T Y n ), where A T is the inverse Winograd transform of the output feature maps defined as:
[0162]
[0163] Multi-Span Convolution
[0164] Winograd transform brings efficiency to the computation by reducing the number of multiplications required to compute neighboring pixels. For convolution layers working with a stride > 1, the convolution solution will require modifications to the Winograd transform. For example, a convolution with a stride of 2 will perform convolution operations on alternating pixels. To support strides greater than one using the regular Winograd transform, the input data is modified using steering write logic, which modifies the input data as it is written to the accelerator's internal memory.
[0165] FIG. 18 An exemplary logic and data layout 1800 for implementing multi-stride convolutions according to an embodiment is shown. The exemplary logic and data layout 1800 is configured for a kernel size of five and a kernel stride of two. In one embodiment, a steering logic block (e.g., steering write logic 1806) is provided that hides alternating pixels from the Winograd transform. A 1x5 kernel 1802 having elements {k1, k2, k3, k4, k5} can be decomposed into two 1x3 kernels based on the exemplary kernel stride of two, where the first decomposed kernel 1808A includes elements {k1, K3, and K5} and the second decomposed kernel 1808B includes elements {k2, k4, 0}. An input data patch 1804 stored contiguously in memory can be written out to buffers within the Winograd computation logic via the steering write logic 1806. The steering write logic can write the input data such that stride 2 data is stored contiguously in the buffers. For example, a 1x12 input patch having elements {I1, I2, I3, I4, I5, I6, I7, I8, I9, I10, I11, I12} can be written out to the buffers such that a first input data tensor 1812A includes elements {I1, I3, I5, I7, I9, I11} and a second input data tensor 1812 includes elements {I2, I4, I6, I8, I10, I12}. A Winograd computation unit 1810 can then perform F(4,3) Winograd convolutions between the corresponding 1x3 kernels 1808A-B and 1x6 input data tensors 1812A-B, with partial outputs summed to create a final output 1814. The techniques shown can be modified as needed to implement multi-stride convolutions with different strides, as the control logic such as the steering write logic 1806 can be configured accordingly for various strides greater than one.
[0166] High Level Architecture:
[0167] FIG. 19is a block diagram of a Winograd acceleration architecture 1900 according to the embodiments described herein. The illustrated acceleration architecture 1900 has parameterizable quantities of processing tiles 1910A-1910M (e.g., Tile-0 through Tile-M), each having an array of Winograd compute blocks 1914A-1914N (WINO-Block-1 through WINO-Block-N) for performing convolution operations using Winograd transforms. In one embodiment, the Winograd acceleration logic 1900 includes a compute interface 1902, which can be one of any number of high performance compute interfaces, including network-based interfaces. The compute interface 1902 sends and receives commands and data to a local DMA unit 1901. The local DMA unit 1901 is a data fetch engine that includes a variety of memory access and transfer logic units capable of performing load and store operations to and from various memory units within the Winograd acceleration architecture 1900.
[0168] The Winograd acceleration architecture 1900 additionally includes an input write controller 1903 and a kernel write controller 1926, each of which includes steering logic for enabling support of multi-span convolutions, such as the steering write logic 1806 of FIG. 18 The kernel write controller 1926 is also configured to allocate specific kernel data to specific tiles 1910A-1910M. A set of IP registers 1924 includes accelerator configuration registers for providing topology information to the Winograd acceleration architecture 1900. The provided topology information includes kernel and input patch sizes, kernel span, input and output feature map numbers, and other information for configuring convolution operations within the Winograd acceleration architecture 1900. In one embodiment, a separate control interface is provided for configuring the IP registers 1924. The IP registers 1924 can be configured on a layer-by-layer basis, enabling each CNN layer to be configured differently. In one embodiment, the layer-by-layer topology of an entire CNN model can be preconfigured within the IP registers 1924 to enable pipelining of the CNN.
[0169] One embodiment additionally includes a Winograd controller 1922 that is a core controller that controls the compute loop and data access to and from local memory (e.g., input local memory 1904) based on topology information provided to IP registers 1924. In one embodiment, the Winograd controller 1922 is a microcontroller that controls the flow of control for operations performed by the Winograd compute blocks 1914A-1914N within each tile 1910A-1910M. The flow of control performed by the Winograd controller controls, among other things, the manner and order in which the compute portion output feature maps and the final output feature map are determined. In one embodiment, the Winograd controller 1922 manages the flow of control by issuing commands and instructions to the Winograd compute blocks 1914A-1914N within each tile 1910A-1910M. In one embodiment, the Winograd controller 1922 can perform fine-grained control of the Winograd compute blocks 1914A-1914N, including issuing specific multiply and accumulate instructions for individual stages of a Winograd convolution. The Winograd controller 1922 can also sequence and issue a series of complex logic operations in order to enable the implementation of multi-span and multi-size kernel convolutions using a single Winograd transform (e.g., a F(4,3) Winograd transform) as described herein, thereby enabling the Winograd compute architecture 1900 to support multiple kernel sizes and kernel spans without requiring hardware logic to support multiple different types of transforms. While the architecture described herein is configured for a F(4,3) Winograd transform, other embodiments can implement similar techniques using hardware designed for other Winograd transforms {F(m,r)} that use an r-tap finite impulse response (FIR) filter defined at least in part via a lx r kernel to compute m outputs.
[0170] The Winograd compute architecture 1900 includes various internal memories, such as input local memory 1904, output local memory 1916, and kernel local memory 1912. The input local memory 1904 is used to store the minimum input data required to start the computation. The data stored in the input local memory 1904 can be reused across all processing tiles 1910A-1910M. Data from the input local memory 1904 can be transformed by the input transform unit 1905 when provided to the individual processing tiles 1910A-1910M. Regarding FIG. 20 Additional details are described regarding input transformation and the input transform unit 1905.
[0171] An output local memory 1916 is included within each processing tile 1910A-1910M and stores partially computed output data. A kernel local memory 1912 stores kernel data for use in output feature map computation. The kernel local memory 1912 is included within each processing tile 1910A-1910M and stores transformed kernel data output by the weight transformation unit 1911. The transformed weight data can be reused by each Winograd computation block 1914A-1914N within a single tile, where each tile 1910A-1910M includes a separate weight transformation unit 1911. The weight transformation cost is paid per tile, as each tile computes a different output feature map, where each output feature map is associated with a different kernel. The weight transformation is done prior to writing to the kernel local memory 1912, enabling reuse across the entire output feature map computation, enabling reduced computational power for the Winograd accelerated architecture 1900. Regarding FIG. 22 Additional optimizations to the weight transformation unit 1911 are described.
[0172] Data stored in the input local memory 1904 can be reused across all processing tiles 1910A-1910M. Transformed kernel data is reused across all Winograd computation blocks 1914A-1914N within each tile. In one embodiment, each of the Winograd computation blocks 1914A-1914N includes a plurality of multiply-accumulate computation units for performing a plurality of element-wise multiplication and accumulation operations to generate partial results for all input feature maps. Regarding FIG. 21 Additional information regarding the Winograd computation blocks 1914A-1914N is provided.
[0173] In one embodiment, intermediate data is stored in output local memory 1916 before being transformed by output transform unit 1918. Output transform unit 1918 performs output transforms for all of tiles 1910A-1910M, avoiding the need to perform output transforms within each tile. In one embodiment, output transform unit 1918 can perform an output transform (e.g., an inverse Winograd transform) for a lx6 output tensor in order to generate a lx4 output tensor. The output transform can be performed in parallel for one or more of tiles 1910A-1910M. Output writes are controlled by output write control unit 1920, which reads output data and writes the output data to external memory via local DMA unit 1901. In one embodiment, output transform unit 1918 can automatically perform an inverse Winograd transform for output feature maps read out of output local memory 1916 by output write control 1920. Output write control unit 1920 reads data from output local memory 1916 in each tile 1910A-1910M and arranges the data into the proper write format from the use of local DMA unit 1901. In one embodiment, output write control 1920 serializes the output from each tile 1910A-1910M in order to reduce the amount of memory bandwidth consumed by output writes. While the output writes are serialized in this embodiment, the computational operations within each tile 1910A-1910M are performed in parallel. In the process of draining output data from local memory, any required bias addition can be performed, if needed.
[0174] FIG. 20 Input transforms are shown in accordance with an embodiment. In one embodiment, input transforms are performed by an input transform unit, such as input transform unit 1905 in FIG. 19A. FIG. 19 As needed, data of input feature map 2000 is transformed and distributed to each Winograd computation tile (e.g., Winograd computation tiles 1910A-1910M of FIG. 19B). FIG. 19
[0175] The Winograd input transform is used to compute a Winograd transform for a given input feature map 2000. For F(4,3) Winograd convolutions, the input transform B shown above is used. T The transforms are performed co-linearly, and values of input feature map 2000 can be accessed in an overlapping manner. For example, for a given input feature map 2000, a first lx6 input feature map patch 2002 partially overlaps a second lx6 input feature map 2004, which partially overlaps a third lx6 input feature map patch 2006.
[0176] FIG. 21 An architecture of a Winograd computation block 2100 according to an embodiment is shown. The Winograd computation block 2100 is a version of the Winograd computation blocks 1914A-1914N. The Winograd computation block 2100 can accept a 1x6 transform kernel and a 1x6 transform input patch as inputs 2102. The inputs 2102 are processed by an array of Winograd processing elements 2110 in order to generate a 1x6 output tensor 2104. The 1x6 output tensor 2104 is an untransformed output. An inverse Winograd transform can be applied to the output tensor 2104 in order to generate a 1x4 output tensor.
[0177] The shown array of processing elements 2110 includes six Winograd processing elements (e.g., Pe1 2112A-Pe6 2112F) for implementing element-wise multiplication and accumulation operations for F(4,3) Winograd convolutions. However, the Winograd computation block 2100 can be extended at the design stage for different Winograd convolution forms. An exemplary processing element (e.g., Pe6 2112F) includes a multiplier 2122, an adder 2126, an accumulator register 2128, and a multiplexer 2124. As specified by commands provided by a Winograd controller, such as the Winograd controller 1922, and as shown in the operations 1710 of FIG. 19 FIG. 17 Each processing element 2112F can perform a multiplication or an addition operation, as specified by commands provided by a Winograd controller, such as the Winograd controller 1922, and as shown in the operations 1710 of FIG. 17 n The partial output feature map data 2130 (e.g., Y FIG. 19 ) can be stored in the accumulator register 2128 and output to memory (e.g., the output local memory 1916 of FIG. 17 n_先前 The previous partial output feature map data 2118 (e.g., Y
[0178] In one embodiment, for each input feature map fetched from memory, the Winograd computation architecture described herein performs the computation for multiple output feature maps before fetching the next input feature map. The number of output feature maps computed in parallel depends on the number of processing tiles. The Winograd computation architecture described herein can be used to implement a computation accelerator that can be designed to work with any input memory size and can be configured for different market segments according to area budget. For example, an accelerator targeted for training or inference can be designed. The resulting accelerator can provide up to 2x improvement in performance over other neural network computation solutions. In various embodiments, the architecture can be implemented in various types of compute units, including general purpose graphics processors (e.g., GPGPUs), field programmable gate arrays (FPGAs), and hybrid accelerators using various different types of compute units.
[0179] Bandwidth Optimized Hardware for Computing Winograd Kernel Transformations
[0180] Transformations for input data and output data for Winograd convolution are relatively cheap from a computation perspective because these transformations have a computational requirement of order O(N). However, kernel / weight transformation is of order O(N 2 ) and typically requires a costly division operation. Because the kernel is known in advance, it is possible to perform the kernel transformation offline and use the transformed kernel during computation. However, this offline kernel transformation increases the memory bandwidth requirement for the system because the transformed kernel is larger in size. For example, in the case of a 3x3 kernel, the bandwidth requirement doubles. In the case of a 5x5 kernel, the transformation increases the bandwidth by 240%.
[0181] The embodiments described herein provide a hardware logic for implementing an optimized kernel / weight transformation for Winograd convolution. The hardware logic enables the optimized Winograd computation architecture to reduce the complexity of collinear kernel transformation without significantly increasing hardware resource requirements and memory bandwidth requirements. The optimized kernel / weight transformation allows for simplified hardware logic because the optimized transformation does not require a hardware divider that would otherwise be used for collinear weight transformation and does not increase the bandwidth requirement, as with offline weight transformation. The techniques described herein are computationally exact and do not introduce any precision loss. This technique can be particularly suitable for inference-optimized hardware, which is generally embedded hardware where per-watt performance or per-area performance is particularly important.
[0182] One embodiment provides hardware logic that enables a kernel transform to be split into multiple stages. A first stage includes division operations that can be performed offline without increasing memory bandwidth requirements, as the size of the kernel after the first transform stage is the same as the original kernel. A second stage can be performed in-line to generate the final transformed weights. The second stage of the transform is done in-line with the computation operations that use only addition operations, without resulting in increased memory bandwidth requirements.
[0183] In one embodiment, the Winograd convolution is performed using an F(4,3) transform that uses a kernel transform G as described above and shown below:
[0184]
[0185] The kernel transform matrix consists of divisions by 6, 12, and 24, which are not easily implemented on hardware without consuming higher cost hardware resources. Accordingly, the weight transform is modified to G' as shown below:
[0186]
[0187] The first stage of the transform can be performed offline as a divide-by-three operation on the pre-transformed kernel data. The offline-transformed weights can then be stored in system memory. The second stage of the modified weight transform can be used to perform the remaining scaling and addition transform operations of the second stage in-line during the convolution performed by the weight transform logic (e.g., weight transform logic 1911 of FIG. 19). FIG. 19
[0188] FIG. 22 Logic that can be configured to perform native weight transforms and optimized weight transforms is shown. Transformations can be performed on input weights 2202 via a weight transform unit 2204 to generate a set of output weights 2206. The following table shows a native Winograd weight transform in contrast to the optimized weight transform provided by the embodiments described herein. For a pre-transformed kernel K having elements [k0, k1, k2], a native Winograd weight transform can be performed using a weight transform configured as in Table 1 below:
[0189] Table 1 - Native F(4,3) Winograd weight transform
[0190] K'0 = k0 » 2
[0191] K'1 = -(k0 + k1 + k2) / 6
[0192] K'2 = (k1 - k0 - k2) / 6
[0193] K'3 = (K0 + 2K1 + 4K2) / 24
[0194] K'4 = (K0 - 2K1 + 4K2) / 24
[0195] K'5 = K3
[0196] One embodiment described herein uses an offline transformed kernel k based on the kernel k m with the modified kernel transformation configured as shown in Table 2 below.
[0197] Table 2 - Optimized F(4,3) Winograd weight transformation
[0198] K'0 = ((Km0 « 1) + Km0) » 2
[0199] K'1 = -(Km0 + Km1 + Km2) » 1
[0200] K'2 = (Km1 - Km0 - Km2) » 1
[0201] K'3 = (Km0 + (Km1 « 1) + (Km2 « 2)) » 3
[0202] K'4 = (Km0 - (Km1 « 1) + (Km2 « 2)) » 3
[0203] K'5 = (Km2 « 1) + Km2
[0204] Using the collinear transformation shown in Table 2, the weight transformation logic within a Winograd compute accelerator can be implemented using significantly reduced hardware logic complexity, resulting in a power and / or area reduction of the optimized weight transformation unit relative to a standard, unoptimized weight transformation unit. The weight transformations described herein can be further optimized by reusing adders to perform the calculations for different weights.
[0205] FIG. 23An optimized Winograd weight transform architecture 2300 according to an embodiment is shown. The shown architecture 2300 uses eight adders 2310A-H per weight transform. For a fixed point accelerator, this implementation is more area and power efficient compared to using divider logic, as the cost scales with the number of operations being performed (e.g., computing multiple output feature maps in parallel). The optimization provided by one embodiment allows sharing of inputs and stacking of outputs from certain adders to enable the use of a reduced number of logic units to perform the weight transforms shown in Table 2.
[0206] The optimized hardware provides five inputs to the eight adders 2310A-H in order to generate six transformed kernel elements. A first input 2301 ([Km0<<1], [Km0]) is provided to the first adder 2310A in order to generate a first intermediate output Y0. The first intermediate output Y0 is shifted right in order to generate a first transformed kernel element 2321 ([K'0=Y0»2]). A second input ([Km0], [Km2]) 2302 is provided to the second adder 2310B. The output of the second adder 2310B is provided to the third adder 2310C and the fourth adder 2310D. A third input 2303 ([Km1], [Km1<<1]) is provided to the third through sixth adders 2310C-F, where the input element [Km1] is provided to the third and fourth adders 2310C-D, and the input element [Km1<<1] is provided to the fifth and sixth adders 2310E-F. The third adder 2310C outputs a second intermediate value Y1, which is transformed to a second output element 2322 ([K'1=Y1»1]). The fourth adder 2310D outputs a third intermediate value Y2, which is transformed to a third output element 2323 ([K'2=Y2»1]). The fifth adder 2310E outputs a third intermediate value Y3, which is transformed to a fourth output element 2325 ([K'3=Y3»3]). The sixth adder 2310F outputs a fourth intermediate value Y4, which is transformed to a fifth output element 2325 ([K'4=Y4»3]). The fifth adder 2310E and the sixth adder 2310F receive output from a seventh adder 2310G, which receives a fourth input 2304 ([Km0], [Km2<<2]). A fifth input 2305 ([Km2], [Km2<<1]) is provided to the eighth adder 2310H, which generates a fifth intermediate output Y5, which is used as the sixth output 2326 ([K'5=Y5]) without transformation.
[0207] In various embodiments, the above is a hardware-based Winograd convolution acceleration architecture that enables applying a single Winograd transform method (e.g., F(4,3)) to kernels of multiple sizes and spans. The Winograd acceleration architecture includes a Winograd controller that exercises fine-grained control over Winograd processing elements in order to implement Winograd convolutions using multiple kernel sizes. The Winograd acceleration architecture also includes steering write logic that enables Winograd convolutions for multiple different kernel spans. Shared input and output transform units within the Winograd acceleration architecture allow amortization of transform costs over a large number of parallel compute operations, thereby reducing hardware complexity, power requirements, and area requirements to implement a hardware-accelerated Winograd convolution.
[0208] Additional optimizations are described that implement reduced power / area weight transforms for use during Winograd transforms. The optimized weight transforms perform a two-stage transform that includes an offline portion and an online transform that is performed in-line with Winograd compute operations. The optimized weight transform technique allows for implementing a Winograd convolution accelerator with reduced power requirements and area requirements suitable for a neural network inference accelerator.
[0209] Optimizations for Software and Firmware Implementation
[0210] The hardware design elements described herein can also be implemented using software or firmware techniques for ordering compute operations on a Winograd accelerator that supports a single kernel size and kernel span. Firmware executing on a microcontroller, such as a graphics microcontroller, can perform operations according to the process illustrated in FIG. 7. FIG. 24
[0211] As FIG. 24 As shown in FIG. 24, a process 2400 for performing hardware-based Winograd convolution using kernels having multiple sizes, in accordance with embodiments described herein, includes decomposing a high-order convolution kernel having a first kernel size into a plurality of sub-kernels having a second kernel size, as indicated at block 2402, and transforming at least one patch of an input feature map and the plurality of sub-kernels based on a Winograd transform associated with the second kernel size, as indicated at block 2404. Process 2400 additionally includes performing a plurality of successive Winograd convolution operations to generate a set of partial output feature maps, as indicated at block 2406, and accumulating the plurality of partial output feature maps into an output feature map, as indicated at block 2408. Process 2400 additionally includes performing an inverse Winograd transform on the output feature map to generate a transformed output feature map, as indicated at block 2410.
[0212] Multi-span Winograd convolution can be implemented by parsing and packing kernel data and feature map data and storing the kernel data and feature map data contiguously into a buffer based on a specified kernel span of the convolution. FIG. 25 An exemplary process is illustrated in FIG. 25.
[0213] As FIG. 25 As shown in FIG. 25, a process 2500 for performing hardware-based Winograd convolution using kernels having multiple spans, in accordance with embodiments described herein, includes loading kernel data and input feature map data into memory (e.g., system memory), where the kernel data and feature map data are to be processed via a hardware-based Winograd convolution accelerator, as indicated at block 2502. Software logic, firmware logic, or hardware logic can then write kernel data, and at least one patch of input feature map data, into a first (e.g., kernel) buffer and a second (e.g., input) buffer, the writing being performed using steering write logic to pack multiple portions of the input buffer and kernel buffer according to a specified kernel span, as indicated at block 2504. A plurality of Winograd convolution rounds can then be performed using the data in the first and second buffers, as indicated at block 2506. Intermediate outputs of these convolution rounds can be accumulated at block 2508, and then an inverse Winograd transform can be performed on the intermediate outputs to generate a transformed output feature map at block 2510.
[0214] The optimized kernel weight transformation can also be implemented in software or firmware of a data processing system in a manner independent of other techniques described herein. For example, the weight transformation operation for the Winograd kernel transformation can be divided into multiple stages, including one or more offline transformation stages, and an online transformation that is performed in-line with the Winograd computation operation. The one or more offline transformations perform operations that are computationally more expensive to implement in hardware relative to the online portion of the transformation.
[0215] In one embodiment, FIG. 26 As shown in process 2600 of FIGURE 2, a first (e.g., offline) stage of an optimized Winograd kernel weight transform may be performed, wherein the first stage includes one or more parallel division operations on weight data to be transformed, as shown at block 2602. Process 2600 further includes writing the output of the first stage of the multi-stage Winograd kernel transform to a memory accessible by a hardware-based Winograd convolution accelerator, as shown at block 2604. Process 2600 further includes performing a second stage of the multi-stage Winograd kernel transform in-line with the Winograd convolution operation using hardware adders and shift logic, as shown at block 2606. Process 2600 further includes performing at least a portion of the Winograd convolution operation using the output from the second stage of the multi-stage Winograd kernel transform, as shown at block 2608.
[0216] The second, online stage of the Winograd kernel transformation can be performed using only adders and shifters, which is easier to implement in hardware than the divider logic required for the first stage. The offline stage can be performed by general-purpose computing hardware in response to instructions provided by software or firmware. For example, a machine learning framework can perform a division operation on the weight data before it is provided to a Winograd convolution accelerator, which can perform the online transformation portion in-line with the Winograd convolution. Further optimizations can be performed in the weight transformation logic hardware portion that performs the online weight transformation. The hardware weight transformation logic can be configured so that multiple adders are shared between multiple inputs, and one or more of the shared adders are used to generate multiple transformed kernel elements.
[0217] Machine Learning Overview
[0218] Machine learning algorithms are algorithms that can learn based on a set of data. Embodiments of machine learning algorithms can be designed to model high-order abstractions within a data set. For example, image recognition algorithms can be used to determine which of several categories a given input belongs to; regression algorithms can output a numerical value given an input; and pattern recognition algorithms can be used to generate translated text or perform text-to-speech and / or speech recognition.
[0219] One example type of machine learning algorithm is a neural network. There are many types of neural networks; one simple type of neural network is a feedforward network. A feedforward network can be implemented as a directed acyclic graph, with nodes arranged in layers. Typically, a feedforward network topology includes an input layer and an output layer, separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation that is useful in generating an output in the output layer. Network nodes are fully connected to nodes in adjacent layers via edges, but there are no edges between nodes within each layer. Data received at the nodes of the input layer of a feedforward network is propagated (i.e., "fed forward") to the nodes of the output layer via an activation function that computes the state of the nodes of each successive layer in the network based on coefficients ("weights") that are respectively associated with each of the edges connecting those layers. The output from a neural network algorithm can take various forms, depending on the particular model represented by the algorithm being executed.
[0220] Before a machine learning algorithm can be used to model a particular problem, the algorithm is trained using a training data set. Training a neural network involves selecting a network topology, using a set of training data representing the problem being modeled by the network, and adjusting the weights until the network model exhibits minimal error for all instances of the training data set. For example, during a supervised learning training process for a neural network, the output produced by the network in response to an input representing an instance in the training data set is compared to the "correct" labeled output for that instance; an error signal representing the difference between the output and the labeled output is computed; and as the error signal is propagated backwards through the layers of the network, the weights associated with the connections are adjusted to minimize the error. When the error for each output generated from an instance of the training data set is minimized, the network is considered to have been "trained."
[0221] The accuracy of a machine learning algorithm can be greatly affected by the quality of the dataset used to train it. The training process can be computationally intensive and can require a significant amount of time on a conventional general-purpose processor. Therefore, many types of machine learning algorithms are trained using parallel processing hardware. This is particularly useful for optimizing the training of neural networks, as the calculations performed when adjusting the coefficients in a neural network are naturally suited to parallel implementation. In particular, many machine learning algorithms and software applications have been adapted to use parallel processing hardware within general-purpose graphics processing devices.
[0222] FIG. 27 2700 is a generalized diagram of a machine learning software stack. Machine learning applications 2702 can be configured to train a neural network using a training dataset or to use a trained deep neural network to implement machine intelligence. Machine learning applications 2702 can include training and inference functionality for the neural network and / or specialized software that can be used to train the neural network prior to deployment. Machine learning applications 2702 can implement any type of machine intelligence, including, but not limited to, image recognition, mapping and localization, autonomous navigation, speech synthesis, medical imaging, or language translation.
[0223] Hardware acceleration for machine learning applications 2702 can be achieved via a machine learning framework 2704. The machine learning framework 2704 can provide a library of machine learning primitives. Machine learning primitives are the basic operations commonly performed by machine learning algorithms. Without the machine learning framework 2704, developers of machine learning algorithms would be required to create and optimize the main computational logic associated with the machine learning algorithm, and then re-optimize the computational logic when new parallel processors are developed. Instead, machine learning applications can be configured to use the primitives provided by the machine learning framework 2704 to perform the necessary calculations. Exemplary primitives include tensor convolutions, activation functions, and pooling, which are computational operations performed when training convolutional neural networks (CNNs). The machine learning framework 2704 can also provide primitives for implementing basic linear algebra subroutines, such as matrix and vector operations, performed by many machine learning algorithms.
[0224] The machine learning framework 2704 can process input data received from the machine learning application 2702 and generate appropriate inputs to the compute framework 2706. The compute framework 2706 can abstract the underlying instructions provided to the GPGPU driver 2708 to enable the machine learning framework 2704 to leverage hardware acceleration via GPGPU hardware 2710 without the machine learning framework 2704 needing to be very familiar with the architecture of the GPGPU hardware 2710. Additionally, the compute framework 2706 can enable hardware acceleration for the machine learning framework 2704 across multiple types and generations of GPGPU hardware 2710.
[0225] Machine Learning Neural Network Implementation
[0226] The computing architecture provided by the embodiments described herein can be configured to perform these types of parallel processing that are particularly suited to training and deploying neural networks for machine learning. Neural networks can be generalized as networks of functions having graph relationships. As is well known in the art, there are multiple types of neural network implementations used in machine learning. One exemplary type of neural network is a feedforward network as previously described.
[0227] A second exemplary type of neural network is a convolutional neural network (CNN). CNNs are specialized feedforward neural networks for processing data having a known, grid-like topology, such as image data. Thus, CNNs are commonly used for computer vision and image recognition applications, but they can also be used for other types of pattern recognition, such as speech and language processing. The nodes in the input layer of a CNN are organized into groups of “filters” (feature detectors inspired by the receptive fields found in the retina), and the output of each group of filters is propagated to the nodes in successive layers of the network. The computations for a CNN include applying a convolution mathematical operation to each filter to produce the output of the filter. Convolution is a specialized mathematical operation performed by two functions to produce a third function that is a modified version of one of the original functions. In convolution network terminology, the first function with respect to the convolution can be referred to as the input, and the second function can be referred to as the convolution kernel. The output can be referred to as a feature map. For example, the input to a convolution layer can be a multidimensional data array that defines various color components of an input image. The convolution kernel can be a multidimensional array of parameters that are adapted through a training process for the neural network.
[0228] A recurrent neural network (RNN) is a class of feedforward neural network that includes feedback connections between layers. RNNs enable modeling of sequential data by sharing parameter data across different parts of the neural network. The architecture of an RNN includes loops. These loops represent the influence of a variable's current value on its own value at a future time, as at least a portion of the output data from the RNN is used as feedback for processing subsequent input in the sequence. This feature makes RNNs particularly useful for language processing due to the variable nature in which language data can be composed.
[0229] The figures described below present exemplary feedforward, CNN, and RNN networks, and describe general processes for training and deploying each of those types of networks, respectively. It will be understood that these descriptions are exemplary and non-limiting as to any particular embodiment described herein, and that the concepts illustrated can generally be applied to deep neural networks and machine learning techniques in general.
[0230] The exemplary neural networks described above can be used to perform deep learning. Deep learning is machine learning using deep neural networks. In contrast to shallow neural networks that include only a single hidden layer, deep neural networks used in deep learning are artificial neural networks composed of multiple hidden layers. More deeply neural networks are generally more computationally intensive to train. However, the additional hidden layers of the network enable multi-step pattern recognition that results in reduced output error relative to shallow machine learning techniques.
[0231] Deep neural networks used in deep learning generally include a front-end network for performing feature recognition coupled to a back-end network representing a mathematical model that can perform operations (e.g., object classification, speech recognition, etc.) based on feature representations provided to the model. Deep learning enables machine learning to be performed without hand-engineering features for the model. Instead, a deep neural network can learn features based on statistical structure or correlations within input data. The learned features can be provided to a mathematical model that can map the detected features to an output. The mathematical model used by the network is generally specific to the particular task to be performed, and different models will be used to perform different tasks.
[0232] Once a neural network is structured, a learning model can be applied to the network to train it to perform a specific task. The learning model describes how to adjust the weights within the model to reduce the network's output error. Backpropagation of error is a common method for training neural networks. An input vector is presented to the network for processing. The network's output is compared to the desired output using a loss function, and an error value is calculated for each neuron in the output layer. These error values are then propagated backward until each neuron has an associated error value that roughly represents its contribution to the original output. The network can then learn from those errors using an algorithm (such as stochastic gradient descent) to update the weights of the neural network.
[0233] FIGS. 28A-28B Shows an example convolutional neural network. FIG. 28A Show the various layers in CNN. FIG. 28A As shown in , an exemplary CNN for modeling image processing can receive input 2802, which describes the red, green, and blue (RGB) components of the input image. Input 2802 can be processed by multiple convolutional layers (e.g., convolutional layer 2804, convolutional layer 2806). Optionally, the outputs from the multiple convolutional layers can be processed by a set of fully connected layers 2808. The neurons in the fully connected layer have complete connections to all activation functions in the previous layer, as previously described for the feedforward network. The output from the fully connected layer 2808 can be used to generate output results from the network. Matrix multiplication can be used instead of convolution to calculate the activation function within the fully connected layer 2808. Not all CNN implementations use the fully connected layer DPLA08. For example, in some implementations, the convolutional layer 2806 can generate the output of the CNN.
[0234] The convolutional layers are sparsely connected, which is different from the traditional neural network configuration found in the fully connected layer 2808. Traditional neural network layers are fully connected so that every output unit interacts with every input unit. However, the convolutional layers are sparsely connected because the output of the convolution of the receptive field (rather than the corresponding state value of each node in the receptive field) is input to the nodes of the subsequent layer, as shown. The kernel associated with the convolutional layer performs a convolution operation, the output of which is sent to the next layer. The dimensionality reduction performed within the convolutional layer is one aspect that enables CNNs to scale to process large images.
[0235] FIG. 28BAn exemplary computation stage within a convolutional layer of a CNN. An input 2812 to a convolutional layer of a CNN can be processed in three stages of a convolutional layer 2814. The three stages can include a convolution stage 2816, a detector stage 2818, and a pooling stage 2820. The convolutional layer 2814 can then output data to a successive convolutional layer. The last convolutional layer of a network can generate output feature map data or provide input to a fully connected layer, for example, to generate classification values to the input to the CNN.
[0236] In the convolution stage 2816, several convolutions are performed in parallel to produce a set of linear activation functions. The convolution stage 2816 can include an affine transformation, which is any transformation that can be specified as a linear transformation plus a translation. Affine transformations include rotation, translation, scaling, and combinations of these transformations. The convolution stage computes the output (e.g., a neuron) of a function connected to a particular region in the input, which can be determined as a local region associated with the neuron. The neuron computes a dot product between the neuron's weights and the region in the local input to which the neuron is connected. The output from the convolution stage 2816 defines a set of linear activation functions that are processed by successive stages of the convolutional layer 2814.
[0237] The linear activation functions can be processed by the detector stage 2818. In the detector stage 2818, each linear activation function is processed by a non-linear activation function. The non-linear activation function increases the non-linear properties of the overall network without affecting the receptive field of the convolutional layer. Several types of non-linear activation functions can be used. One particular type is a rectified linear unit (ReLU), which uses an activation function defined as f(x) = max(0, x) such that the activation function is thresholded to zero.
[0238] The pooling stage 2820 uses a pooling function that replaces the output of the convolutional layer 2806 with summary statistics of nearby outputs. Pooling functions can be used to introduce translation invariance into a neural network such that slight translations to the input do not change the pooling output. Invariance to local translations can be useful in cases where the presence of a feature in the input data is more important than the exact position of the feature. Various types of pooling functions can be used during the pooling stage 2820, including max pooling, average pooling, and L2 norm pooling. Additionally, some CNN implementations do not include a pooling stage. Instead, such implementations substitute an additional convolutional stage with an increased stride relative to the previous convolutional stage.
[0239] The output from the convolutional layer 2814 can then be processed by a next layer 2822. The next layer 2822 can be an additional convolutional layer or one of the fully connected layers 2808. For example, FIG. 28AThe first convolutional layer 2804 can output to a second convolutional layer 2806, which can output to a first layer in a fully connected layer 2808.
[0240] FIG. 29 An example recurrent neural network 2900 is shown. In a recurrent neural network (RNN), the previous state of the network influences the output of the current state of the network. RNNs can be built in a variety of ways using a variety of functions. The use of RNNs often revolves around using a mathematical model to make predictions about the future based on a sequence of previous inputs. For example, an RNN can be used to perform statistical language modeling to predict an upcoming word given a sequence of previous words. The RNN 2900 shown can be described as having an input layer 2902 that receives an input vector, a hidden layer 2904 that implements a recurrent function, a feedback mechanism 2905 that implements a'memory' of previous states, and an output layer 2906 that outputs a result. The RNN 2900 operates based on time steps. The state of the RNN at a given time step is influenced by a previous time step via the feedback mechanism 2905. The state of the hidden layer 2904 is defined for a given time step by the previous state and the input at the current time step. An initial input (xi) at a first time step can be processed by the hidden layer 2904. A second input (x2) can be processed by the hidden layer 2904 using state information determined during processing of the initial input (xi). The given state can be computed as s t = f(Ux t + Ws t-1 ), where U and W are parameter matrices. The function f is typically non-linear, such as a hyperbolic tangent function (Tanh) or a variant of the rectified function f(x) = max(0, x). However, the particular mathematical function used in the hidden layer 2904 can vary depending on the particular implementation details of the RNN 2900.
[0241] Variations of the basic CNN and RNN networks described can also be implemented. One example RNN variation is a long short-term memory (LSTM) RNN. LSTM RNNs are capable of learning long-term dependencies that can be necessary to process longer sequences of language. A variation of a CNN is a convolutional deep belief network, which has a structure similar to a CNN and is trained in a manner similar to a deep belief network. A deep belief network (DBN) is a generative neural network composed of multiple layers of stochastic (random) variables. A DBN can be trained layer-by-layer using greedy unsupervised learning. The learned weights of a DBN can then be used to provide a pre-trained neural network by determining a set of optimal initial weights for a neural network.
[0242] FIG. 30 Training and deployment of deep neural networks is shown. Once a given network structure has been structured for a task, the neural network is trained using a training dataset 3002. Various training frameworks 3004 have been developed for implementing hardware acceleration of the training process. For example, FIG. 27 The machine learning framework 2704 of FIG. 27 can be configured as a training framework 2704. The training framework 2704 can hook into an untrained neural network 3006 and enable training of the untrained neural network using the parallel processing resources described herein to generate a trained neural network 3008.
[0243] To begin the training process, initial weights can be selected randomly or through pre-training using a deep belief network. Training loops are then performed in a supervised or unsupervised manner.
[0244] Supervised learning is a method of learning in which training is performed as an arbitration operation, such as when the training dataset 3002 includes inputs paired with expected outputs for the inputs, or in cases where the training dataset includes inputs with known outputs and the outputs of the neural network are manually graded. The network processes the inputs and the resulting outputs are compared to a set of expected or desired outputs. Errors are then backpropagated through the system. The training framework 3004 can make adjustments to adjust the weights controlling the untrained neural network 3006. The training framework 3004 can provide tools for monitoring how well the untrained neural network 3006 is converging to a model that is suitable for generating correct answers based on known input data. The training process occurs repeatedly as the weights of the network are adjusted to improve the outputs generated by the neural network. The training process can continue until the neural network reaches a statistically expected accuracy associated with the trained neural network 3008. The trained neural network 3008 can then be deployed to implement any number of machine learning operations.
[0245] Unsupervised learning is a method of learning in which the network attempts to train itself using unlabeled data. Thus, for unsupervised learning, the training dataset 3002 will include input data without any associated output data. The untrained neural network 3006 can learn groupings within the unlabeled inputs and can determine how individual inputs relate to the overall dataset. Unsupervised training can be used to generate self-organizing maps, which are a type of trained neural network 3007 that can perform operations useful in data reduction. Unsupervised training can also be used to perform anomaly detection, which allows for identification of data points in the input dataset that deviate from the normal patterns of the data.
[0246] Variations of supervised and unsupervised training can also be employed. Semi-supervised learning is a technique in which the training dataset 3002 includes a mix of labeled and unlabeled data of the same distribution. Incremental learning is a variation of supervised learning in which input data is continuously used for further training of the model. Incremental learning enables a trained neural network 3008 to adapt to new data 3012 without forgetting the knowledge rooted within the network during initial training.
[0247] Regardless of whether supervised or unsupervised, the training process for particularly deep neural networks can be too computationally intensive for a single computing node. Rather than using a single computing node, a distributed network of computing nodes can be used to accelerate the training process.
[0248] FIG. 31 is a block diagram illustrating distributed learning. Distributed learning is training a model that uses multiple distributed computing nodes to perform supervised or unsupervised training of a neural network. The distributed computing nodes can each include one or more host processors and one or more of general purpose processing nodes, such as the highly parallel general purpose graphics processing unit 2800 in FIG. 28. As illustrated, distributed learning can perform model parallelism 3102, data parallelism 3104, or a combination of model and data parallelism 3104.
[0249] In model parallelism 3102, different computing nodes in a distributed system can perform training computations for different portions of a single network. For example, each layer of a neural network can be trained by a different processing node of a distributed system. Benefits of model parallelism include the ability to scale to particularly large models. Splitting computations associated with different layers of a neural network enables training of super large neural networks in which the weights of all layers would not fit into the memory of a single computing node. In some instances, model parallelism can be particularly useful in performing unsupervised training of large neural networks.
[0250] In data parallelization 3104, different nodes of a distributed network have a complete instance of the model, and each node receives a different portion of the data. The results from the different nodes are then combined. While different approaches for data parallelization are possible, data parallel training approaches all require a technique to combine the results and synchronize the model parameters between each node. Example methods for combining the data include parameter averaging and update-based data parallelization. Parameter averaging trains each node on a subset of the training data and sets global parameters (e.g., weights, biases) to the average of the parameters from each node. Parameter averaging uses a central parameter server that holds the parameter data. Update-based data parallelization is similar to parameter averaging, except that updates to the model are passed instead of the parameters from the nodes to the parameter server. Additionally, update-based data parallelization can be performed in a decentralized manner, where the updates are compressed and passed between the nodes.
[0251] For example, the combined model and data parallelization 3106 can be implemented in a distributed system where each compute node includes multiple GPUs. Each node can have a complete instance of the model, with individual GPUs within each node used to train different portions of the model.
[0252] Distributed training has increased overhead relative to training on a single machine. However, the parallel processors and GPGPUs described herein can each implement techniques for reducing the overhead of distributed training, including techniques for implementing high-bandwidth GPU-GPU data transfer and accelerated remote data synchronization.
[0253] Exemplary Machine Learning Applications
[0254] Machine learning can be applied to solve a number of technical problems, including but not limited to computer vision, autonomous driving and navigation, speech recognition, and language processing. Computer vision has traditionally been one of the most active areas of research for machine learning applications. Applications of computer vision range from replicating human visual capabilities (e.g., recognizing human faces) to creating new classes of visual capabilities. For example, a computer vision application can be configured to recognize sound waves from vibrations induced in objects visible in a video. Parallel processor-accelerated machine learning enables training of computer vision applications using training data sets significantly larger than previously feasible, and enables deployment of inference-use systems using low-power parallel processors.
[0255] Parallel processor-accelerated machine learning has autonomous driving applications, including lane and road sign recognition, obstacle avoidance, navigation, and driving control. Accelerated machine learning techniques can be used to train driving models based on datasets that define appropriate responses to specific training inputs. The parallel processors described herein can enable rapid training of increasingly complex neural networks for autonomous driving solutions and enable the deployment of low-power inference processors in mobile platforms suitable for integration into autonomous vehicles.
[0256] Deep neural networks accelerated by parallel processors have enabled machine learning methods for automatic speech recognition (ASR). ASR involves creating a function that computes the most likely speech sequence given a sequence of input sounds. Accelerated machine learning using deep neural networks has replaced the hidden Markov models (HMMs) and Gaussian mixture models (GMMs) previously used for ASR.
[0257] Parallel processor-accelerated machine learning can also be used to accelerate natural language processing. Automatic learning programs can use statistical inference algorithms to generate models that are robust to erroneous or unfamiliar inputs. Exemplary natural language processor applications include automatic machine translation between human languages.
[0258] Parallel processing platforms for machine learning can be divided into training platforms and deployment platforms. Training platforms are typically highly parallel and include optimizations for accelerating multi-GPU single-node training and multi-node multi-GPU training. Exemplary parallel processors suitable for training include the highly parallel general-purpose graphics processing unit 2800 of FIG. 28 and FIG. 29 The multi-GPU computing system 2900. In contrast, deployed machine learning platforms typically include low-power parallel processors suitable for use in products such as cameras, autonomous robots, and autonomous vehicles.
[0259] The following clauses and / or examples refer to specific embodiments or examples thereof. Details in the examples may be used anywhere in one or more embodiments. The various features of the different embodiments or examples may be combined in various ways to suit a variety of different applications, with some features included and others excluded. Examples may include subject matter such as methods; apparatus for performing the actions of a method; at least one machine-readable medium including instructions that, when executed by a machine, cause the machine to perform the actions of a method or apparatus; or apparatus or systems according to the embodiments and examples described herein. Each component may be a means for performing the described operation or function.
[0260] One embodiment provides a computing device for performing machine learning operations, the computing device comprising a hardware accelerator comprising a compute unit for performing Winograd convolution, the compute unit configurable to perform the Winograd convolution for a first kernel size using a transform associated with a second kernel size.
[0261] In one embodiment, the hardware accelerator comprises a register for storing a value of the second kernel size, the second kernel size configurable via an interface to the register, the interface to the register provided by the hardware accelerator. The hardware accelerator can comprise a control unit for providing instructions to the compute unit for causing the compute unit to perform Winograd convolution operations based on the value of the second kernel size. In one embodiment, the hardware accelerator comprises a guide write logic for implementing Winograd convolution for a plurality of different kernel spans. In one embodiment, the compute unit is configurable to perform the Winograd convolution for a first kernel span using a transform associated with a second kernel span. In one embodiment, the hardware accelerator comprises a register for storing a value of the second kernel span, and the guide write logic is for writing input data and kernel data to memory within the hardware accelerator according to the value of the second kernel span.
[0262] In one embodiment, the hardware accelerator comprises a plurality of compute logic tiles, each tile comprising a plurality of compute units. Each of the plurality of compute logic tiles can comprise a weight transform unit for applying a Winograd transform to weights associated with a convolution kernel and storing transformed weights to a memory coupled with the weight transform unit. The transformed weights can be stored to the memory coupled with the weight transform unit shared among the plurality of compute units within the compute logic tile comprising the weight transform unit and the memory coupled with the weight transform unit. In one embodiment, the hardware accelerator comprises an input transform unit configured to apply a Winograd transform to input feature map data, the input transform unit shared among the plurality of compute logic tiles.
[0263] One embodiment provides a method for performing a machine learning operation, the method comprising: decomposing a convolutional kernel having a first kernel size into a plurality of sub-kernels having a second kernel size; transforming a portion of an input feature map and the plurality of sub-kernels based on a Winograd transform, the Winograd transform associated with the second kernel size; and performing a plurality of successive Winograd convolution operations in order to generate a set of partial output feature maps.
[0264] In one embodiment, the method additionally comprises: accumulating the plurality of partial output feature maps into an output feature map; and performing an inverse Winograd transform on the output feature map in order to generate a transformed output feature map. In one embodiment, performing the plurality of successive Winograd convolution operations in order to generate a set of partial output feature maps comprises: loading kernel data and input feature map data into memory, the kernel data and input feature map data to be processed via a hardware-based Winograd convolution accelerator; writing the kernel data into a first hardware buffer; writing at least a portion of the input feature map data into a second hardware buffer; performing a plurality of successive Winograd convolution rounds via the hardware-based Winograd convolution accelerator using data in the first hardware buffer and the second hardware buffer; accumulating intermediate outputs of the plurality of successive Winograd convolution rounds; and performing an inverse Winograd transform on the intermediate outputs in order to generate a transformed output feature map.
[0265] In one embodiment, transforming a portion of a sub-kernel based on a Winograd transform comprises: performing a first stage of a multi-stage Winograd kernel transform, the first stage comprising one or more parallel division operations; and performing a second stage of the multi-stage Winograd kernel transform, the second stage being performed in-line with a Winograd convolution operation using hardware adders and shift logic. In one embodiment, the method additionally comprises: performing at least a portion of a Winograd convolution operation using output from the second stage of the multi-stage Winograd kernel transform.
[0266] One embodiment provides a data processing system comprising: a non-transitory machine-readable medium storing instructions for execution by one or more processors of the data processing system; and a general purpose graphics processing unit comprising a hardware accelerator including a compute unit to perform a Winograd convolution, the compute unit being configurable to perform the Winograd convolution for a first kernel size using a transform associated with a second kernel size. The compute unit of the data processing system can perform any of the Winograd convolution operations described elsewhere herein.
[0267] Embodiments described herein refer to specific configurations of hardware, such as an application-specific integrated circuit (ASIC) configured to perform certain operations or have a predetermined function. Embodiments described herein can also be incorporated into hardware products, such as but not limited to FPGA, CPU, or GPU-based computer vision accelerators. These hardware and / or software technologies can be applied to various Internet of Things (IoT) solutions, including: autonomous driving, autonomous robots, and computer vision systems for augmented reality and / or virtual reality. Such electronic devices typically include a set of one or more processors coupled to one or more other components, such as one or more storage devices (non-transitory machine-readable storage media), user input / output devices (e.g., a keyboard, a touchscreen, and / or a display), and a network connection. The coupling of the set of processors and other components is typically through one or more buses and bridges (also referred to as bus controllers). The storage devices and signals carrying the network traffic respectively represent one or more machine-readable storage media and machine-readable communication media. Thus, a storage device of a given electronic device typically stores code and / or data used in connection with the execution of software on the set of one or more processors of this electronic device. Additionally, some elements can be incorporated into a software-based machine learning acceleration framework.
[0268] Of course, one or more portions of an embodiment can be implemented using different combinations of software, firmware, and / or hardware. Throughout this detailed description, for the purposes of explanation, numerous specific details were set forth in order to provide a thorough understanding of the present application. It will be apparent, however, to one skilled in the art that the present embodiments can be practiced without some of these specific details. In certain instances, well-known structures and functions have not been described in elaborate detail in order to avoid obscuring the subject matter of the present embodiments. Accordingly, the scope and spirit of the application should be judged in terms of the claims and the full breadth permitted by the patent laws, rather than the foregoing description.
Claims
1. A computing device for performing machine learning operations, the computing device comprising: A hardware accelerator, the hardware accelerator comprising a computing unit, wherein the computing unit can be configured to: Decomposing a convolution kernel having a first kernel size into a plurality of sub-kernels having a second kernel size; performing a collinear transformation on a portion of the input feature map and the plurality of sub-kernels based on a Winograd transform, the Winograd transform being associated with the second kernel size; and Perform multiple consecutive Winograd convolution operations to generate a set of partial output feature maps, wherein the hardware accelerator includes guided write logic for implementing Winograd convolution for a plurality of different kernel spans, the computation unit is configurable to perform the Winograd convolution for a first kernel span using a transformation associated with a second kernel span, the hardware accelerator includes a register for storing a value of the second kernel span, and the guided write logic is configured to write input data and kernel data to a memory within the hardware accelerator according to the value of the second kernel span, In which multiple partial output feature maps are accumulated into an output feature map, And wherein an inverse Winograd transform is performed on the output feature map to generate a transformed output feature map.
2. The computing device of claim 1 , the hardware accelerator comprising a register for storing a value of the second kernel size, the second kernel size being configurable via an interface to the register, the interface to the register being provided by the hardware accelerator. 3 . The computing device of claim 2 , wherein the hardware accelerator comprises a control unit for providing an instruction to the computing unit, the instruction for causing the computing unit to perform a Winograd convolution operation based on the value of the second kernel size. 4 . The computing device according to claim 1 , wherein the hardware accelerator comprises a plurality of computing logic blocks, each block comprising a plurality of computing units.
5. The computing device of claim 4 , wherein each of the plurality of computing logic blocks comprises a weight transformation unit configured to apply a Winograd transform to weights associated with a convolution kernel and store the transformed weights in a memory coupled to the weight transformation unit.
6. The computing device of claim 5 , wherein the transformed weights are stored in the memory coupled to the weight transformation unit, and the weight transformation unit is shared among the multiple computing units within the computing logic block including the weight transformation unit and the memory coupled to the weight transformation unit.
7. The computing device of claim 6, wherein the hardware accelerator comprises an input transformation unit, the input transformation unit being configured to apply a Winograd transformation to input feature map data, the input transformation unit being shared among the plurality of computing logic blocks.
8. A method for performing a machine learning operation, the method being performed by a computing device, the computing device comprising a hardware accelerator, the hardware accelerator comprising a computing unit, the computing unit being configurable to: Decomposing a convolution kernel having a first kernel size into a plurality of sub-kernels having a second kernel size; performing a collinear transformation on a portion of the input feature map and the plurality of sub-kernels based on a Winograd transform, the Winograd transform being associated with the second kernel size; and Perform multiple consecutive Winograd convolution operations to generate a set of partial output feature maps, wherein the hardware accelerator includes guided write logic for implementing Winograd convolution for a plurality of different kernel spans, the computation unit is configurable to perform the Winograd convolution for a first kernel span using a transformation associated with a second kernel span, the hardware accelerator includes a register for storing a value of the second kernel span, and the guided write logic is configured to write input data and kernel data to a memory within the hardware accelerator according to the value of the second kernel span, In which multiple partial output feature maps are accumulated into an output feature map, And wherein an inverse Winograd transform is performed on the output feature map to generate a transformed output feature map.
9. The method of claim 8, wherein: Performing the plurality of consecutive Winograd convolution operations to generate a set of partial output feature maps includes: Loading kernel data and input feature map data into memory, where they will be processed by a hardware-based Winograd convolution accelerator; Writing the kernel data into a first hardware buffer; writing at least a portion of the input feature map data into a second hardware buffer; performing a plurality of consecutive Winograd convolution rounds via the hardware-based Winograd convolution accelerator using data in the first hardware buffer and the second hardware buffer; Accumulating the intermediate outputs of the plurality of consecutive Winograd convolution rounds; and An inverse Winograd transform is performed on the intermediate output to generate a transformed output feature map.
10. The method of claim 8, wherein: Transforming a part of the sub-kernel based on Winograd transformation includes: Performing a first stage of a multi-stage Winograd kernel transformation, the first stage comprising one or more parallel division operations; performing a second stage of the multi-stage Winograd kernel transform, the second stage being performed in-line with the Winograd convolution operation using hardware adders and shift logic; and At least a portion of a Winograd convolution operation is performed using an output from the second stage of the multi-stage Winograd kernel transform.
11. A machine-readable medium comprising code for causing a machine to perform the method of any one of claims 8 to 10 when the code is executed.
12. A data processing system comprising: a non-transitory machine-readable medium for storing instructions for execution by one or more processors of the data processing system; as well as A general-purpose graphics processing unit, wherein the general-purpose graphics processing unit includes a hardware accelerator, wherein the hardware accelerator includes a computing unit, and wherein the computing unit can be configured to: Decomposing a convolution kernel having a first kernel size into a plurality of sub-kernels having a second kernel size; performing a collinear transformation on a portion of the input feature map and the plurality of sub-kernels based on a Winograd transform, the Winograd transform being associated with the second kernel size; and Perform multiple consecutive Winograd convolution operations to generate a set of partial output feature maps, wherein the hardware accelerator includes guided write logic for implementing Winograd convolution for a plurality of different kernel spans, the computation unit is configurable to perform the Winograd convolution for a first kernel span using a transformation associated with a second kernel span, the hardware accelerator includes a register for storing a value of the second kernel span, and the guided write logic is configured to write input data and kernel data to a memory within the hardware accelerator according to the value of the second kernel span, In which multiple partial output feature maps are accumulated into an output feature map, And wherein an inverse Winograd transform is performed on the output feature map to generate a transformed output feature map.
13. The data processing system of claim 12, the hardware accelerator comprising a register for storing a value of the second kernel size, the second kernel size being configurable via an interface to the register, the interface to the register being provided by the hardware accelerator. 14 . The data processing system of claim 13 , wherein the hardware accelerator comprises a control unit for providing instructions to the computing unit, the instructions for causing the computing unit to perform a Winograd convolution operation based on the value of the second kernel size. 15 . The data processing system according to claim 12 , wherein the hardware accelerator comprises a plurality of computing logic blocks, each block comprising a plurality of computing units.
16. The data processing system of claim 15 , wherein each of the plurality of computational logic blocks comprises a weight transformation unit configured to apply a Winograd transform to weights associated with a convolution kernel and store the transformed weights in a memory coupled to the weight transformation unit.
17. A data processing system as described in claim 16, wherein the transformed weights are stored in the memory coupled to the weight transformation unit, and the weight transformation unit is shared among the multiple computing units within the computing logic block including the weight transformation unit and the memory coupled to the weight transformation unit.
18. The data processing system of claim 17, wherein the hardware accelerator comprises an input transformation unit, the input transformation unit being configured to apply a Winograd transformation to input feature map data, the input transformation unit being shared among the plurality of computational logic blocks.
Citation Information
Patent Citations
Method for accelerative compression of deep convolutional neural networks for handwritten Chinese character recognition
CN106919942A