Creation of application-specific machine learning accelerators and global tuning
Patent Information
- Application Number
- KR1020237039235
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-03
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2041-05-03
Smart Images

Figure 112023125812356-PCT00001_ABST
Abstract
Description
Background Technology
[0001] This specification relates to integrated circuits generally used to perform machine learning calculations.
[0002] A neural network is a machine learning model that uses one or more node layers to generate an output (e.g., classification) for a received input. Some neural networks include one or more hidden layers in addition to the output layer. Some neural networks may be Convolutional Neural Networks (CNNs) configured for image processing or Recurrent Neural Networks (RNNs) configured for speech and language processing. Various types of neural network architectures can be used to perform diverse tasks related to classification or pattern recognition, prediction in data modeling, and information clustering.
[0003] A neural network layer may have a corresponding set of parameters or weights. Weights are used to process inputs (e.g., input batches) through a neural network layer to generate the corresponding output of the layer for computing neural network inference. Input batches and kernel sets can be represented as tensors, which are multidimensional arrays of inputs and weights. A hardware accelerator is an integrated circuit with a specific objective for implementing neural networks. The circuit includes memory containing locations corresponding to tensor elements that can be explored or accessed using the circuit's control logic.
[0004] Designing specialized hardware accelerators is work-intensive and time-consuming. For example, the design process often requires months of effort and may involve multiple design iterations. Furthermore, meeting application-specific performance and power objectives requires a strategy within the design process to map the target application to the underlying hardware. While the computational graph of a neural network is static, the mapping task may involve various design parameters that affect the actual performance of the circuit. Additionally, manually navigating the design space is often impossible due to the size of various settings and the interrelationships between different parameters.
[0005] This specification describes a technique for globally tuning a data processing architecture and automatically generating an application-specific machine learning (ML) accelerator based on the tuned architecture. The architecture may be a candidate architecture selected based on a set of application-level objectives. Exemplary application-level objectives may include processor utilization, power consumption, data throughput, and latency. In some cases, objectives represent performance attributes of an example ML accelerator desired by the user. Some (or all) of the objectives may be received as user input to an exemplary hardware accelerator design system. The design system may also determine one or more objectives independently of user input.
[0006] The system globally tunes and dynamically optimizes candidate architectures using application-level objectives (e.g., one or more inputs). For instance, the architecture can be tuned and optimized to run a specific type of neural network to achieve efficiency in areas such as power consumption and processor utilization. The accelerator design system tunes various aspects of the architecture using architecture-specific cost models. The output of the cost models is used to define the final configuration of the accelerator. After optimization and tuning, the system automatically generates a hardware configuration containing various architectural features, including reservation / mapping options, to create an application-specific (ML) accelerator optimized to implement the specified neural network in the hardware.
[0007] One aspect of the gist described herein may be implemented as a computer-implemented method for creating an application-specific machine learning (ML) accelerator. This method comprises selecting an architecture representing a baseline processor configuration and generating performance data for the architecture by modeling, through an ML cost model, how the architecture executes computations of a first neural network comprising at least several layers. This method includes the step of dynamically tuning the architecture based on the performance data so that the architecture meets performance objectives when implementing the first neural network and executing machine learning computations for a target application. This method also includes the step of generating a configuration of the ML accelerator in response to the dynamic tuning of the architecture. The configuration specifies a custom hardware configuration for implementing each of the several layers of the first neural network.
[0008] These and other implementations may optionally include one or more of the following features. For example, in some implementations, the method further includes a step of creating an application-specific hardware ML accelerator based on a custom hardware configuration. Additionally, the application-specific hardware ML accelerator can be optimized when executing computations for a target application using a neural network to implement each of the different layers of the neural network.
[0009] Performance objectives include multiple discrete objectives, and creating an application-specific ML accelerator may include creating an application-specific hardware ML accelerator configured to satisfy each of the multiple discrete objectives when the application-specific hardware ML accelerator executes computations for a target application. In some implementations, generating performance data may include modeling the use of the architecture for executing each layer of the multiple layers of the first neural network by an ML cost model; and generating performance parameters of the architecture for each layer through the ML cost model in response to modeling the use of the architecture for executing each layer. Performance parameters may correspond to each of the multiple discrete objectives. The multiple discrete objectives include at least one of critical processing latency, critical power consumption, critical data throughput, and critical processor utilization. In some implementations, dynamically tuning the architecture determines a mapping of computations to input tensors that causes the application-specific hardware ML accelerator to utilize a critical percentage of the hardware computing units of the hardware ML accelerator; It includes dynamically tuning the architecture based on the determined mapping.
[0010] Dynamically tuning the architecture may include dynamically tuning the architecture based on operations performed by each of the various ML cost models of the global tuner, and dynamically tuning the architecture based on operations performed by at least one of the simulated annealing tuner or random tuner of the global tuner. In some implementations, the architecture represents one or more hardware blocks of an integrated circuit, and dynamically tuning the architecture includes dynamically tuning the architecture to meet respective performance targets for each of the one or more hardware blocks when the architecture implements a first neural network to execute computations for a target application.
[0011] The configuration of the hardware ML accelerator specifies a custom software configuration for the first neural network; and creating an application-specific hardware ML accelerator involves creating an application-specific hardware ML accelerator based on the custom hardware configuration and the custom software configuration. In some implementations, the ML cost model is an architecture-ware cost model comprising one or more individual analytical models; and the architecture-ware cost model is configured to estimate the performance of the architecture based on the deterministic data flow of the data processed using the architecture.
[0012] Other implementations of this and other embodiments of the method include corresponding systems, devices, and computer programs configured to perform the operation of the method and encoded in a computer storage device. A system composed of one or more computers may be configured by software, firmware, hardware, or a combination thereof installed on the system that causes the system to perform a set of operations when operated. One or more computer programs may be configured by having instructions that cause the device to perform a set of operations when executed by a data processing device.
[0013] The gist described in this specification may be implemented in specific embodiments to realize one or more of the following advantages.
[0014] The disclosed technology provides a framework that can be used to facilitate an architecture exploration process for defining optimized hardware and software configurations, including the efficient scheduling and mapping of operations for implementing neural networks in hardware circuits. Based on this process, a hardware design system can automatically generate an output configuration that defines a system-specific optimized hardware mapping for a given set of PPA (performance, power, area) constraints. The PPA constraints may be hardware accelerator performance thresholds related to at least processor utilization, power consumption, latency, block size, and / or data throughput.
[0015] The design system can identify an exemplary network model with a fixed number of layers and determine the optimal properties of the identified hardware architecture (e.g., systolic array, compute tile, etc.), including micro-architecture properties such as block connectivity, hardware layout, or memory. In addition to these optimized hardware properties, the design system determines efficient scheduling and data allocation for layer-by-layer processing, thereby enabling the creation of an application-specific ML accelerator that consumes minimal power and circuit area while meeting (or exceeding) user or system-defined requirements for layer-by-layer processing.
[0016] Details of one or more implementations of the essence described in this specification are described in the accompanying drawings and the description below. Other potential features, aspects, and advantages of the essence will become apparent from the description, drawings, and claims. Brief explanation of the drawing
[0017] Figure 1 is a block diagram of an exemplary computing system for creating and globally tuning a machine learning accelerator. Figure 2 is a block diagram showing an exemplary system for globally tuning an application-specific machine learning accelerator. Figure 3 illustrates an exemplary framework for tuning a multilayer neural network. Figure 4 is a flowchart of an exemplary process for tuning and optimizing the graph execution schedule of a multilayer neural network. Figure 5 is a flowchart of an exemplary process used to create and globally tune a machine learning accelerator. Figure 6 is a block diagram of an exemplary application-specific hardware accelerator created using the system of Figure 1. Figure 7 illustrates an example of an input tensor, a weight tensor, and an output tensor. In various drawings, similar reference numbers and names represent similar elements. Specific details for implementing the invention
[0018] FIG. 1 is a block diagram of an exemplary hardware accelerator design system (100) ("system (100)"). Generally, the system (100) includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a special target processor, etc.), memory, and / or a data storage device that collectively forms processing resources used to execute functions for globally tuning and creating a customized hardware machine learning accelerator.
[0019] As described below, using one or more input objectives (102), the system (100) is configured to develop and output a design configuration for creating an exemplary hardware accelerator. The hardware accelerator may be implemented as a special objective or application-specific hardware circuit optimized to execute a specific type of machine learning task. For example, the application-specific circuit may be a machine learning (ML) hardware accelerator configured to implement or execute a multilayer neural network.
[0020] More specifically, application-specific circuits can be uniquely tuned and / or optimized according to various application goals, such as one or more inputs specified by the user. For example, when implementing a specific type of neural network (e.g., a multilayer CNN), a candidate data processing architecture for an application-specific ML circuit can be optimized to achieve (or exceed) threshold performance objectives related to processor utilization, power consumption, data throughput, and / or latency.
[0021] As used in this document, the data processing "architecture" may refer to a hardware circuit architecture, a software / neural architecture, or both. In this way, architecture tuning and optimization may include tuning properties of the neural architecture as well as tuning properties of the hardware architecture. The resulting architecture is optimized (e.g., fully optimized) to perform a given machine learning task according to each different application goal that may be received or determined by the system (100).
[0022] The system (100) includes control logic for configuring and managing a design space (104). The design space (104) may be configured based on a combination of hardware devices and software routines executed in the system (100). For example, the control logic may be implemented as a system controller or host device that executes programmed instructions to manage various design space operations. Operations of the design space (104) may include processing multiple design items or parameters required to tune a candidate architecture.
[0023] Generally, the system (100) uses control logic to manage activities and operations in the design space (104). In addition to optimizing the architecture for a given ML task, in some implementations, the control logic of the system (100) itself may be based on an ML model. For example, the ML model may be trained to process design inputs and control parameters necessary to tune a candidate architecture based on a set of input objectives. In some implementations, the control logic executes or applies an exemplary optimization algorithm to tune the candidate architecture according to the set of input objectives, as well as operations performed by an exemplary cost model (described below).
[0024] Candidate architectures are selected from at least the architecture repository (106) of the system (100). The system (100) may identify or select candidate architectures from the architecture repository (106) based on at least an input object (102). The architecture repository (106) contains information describing multiple different hardware architectures used to create application-specific hardware ML accelerators.
[0025] For example, a first hardware architecture accessed through the architecture repository (106) may define a systolic array architecture, while a second different hardware architecture accessed through the architecture repository (106) may define a hardware architecture based on an array of computing tiles. Similarly, a third architecture accessed through the architecture repository (106) may define a hardware architecture based on each set of tightly coupled data processing lanes forming a distinct vector processing unit (VPU), while a fourth architecture accessed through the architecture repository (106) may define a hardware architecture including at least two vector processor cores interacting with a large shared scratchpad memory and a matrix computing unit.
[0026] The candidate architecture selected for optimization and tuning may be a combination of a hardware circuit architecture and a neural architecture obtained, for example, from an architecture repository (106). The neural architecture may be obtained from a network graph module (108) containing multiple different types of neural network graphs. For example, the system (100) may select a candidate architecture based on an input target (102), an exemplary hardware layout of an integrated circuit (IC), and an exemplary neural network graph.
[0027] In some implementations, the system (100) selects candidate architectures based on one or more input goals (102) that bias the system toward the selection of a specific hardware architecture for a given neural network architecture. For example, the system (100) may select candidate architectures based on one or more hardware variables. The hardware variables may represent control parameters that restrict the architecture selection and cause the design space (104) to select a specific type of hardware architecture from the repository (106) for a given neural architecture, for example, obtained from a graph module (108).
[0028] The system (100) includes an optimization and tuning module (112) that interacts with one or more cost models to globally tune an exemplary data processing architecture. For example, the system (100) includes an architecture-aware cost model (114) that may include one or more individual data models (114). In some cases, each of these individual data models is a respective cost model (114) configured to run ML-based analysis to tune a candidate architecture based on a set of input goals. The architecture-aware cost model (114) estimates the performance of the candidate architecture based on the deterministic data flow of the data processed using the architecture.
[0029] In some implementations, the system (100) includes a cost model (114) based on one of two types of cost models: an analytical cost model or an ML-based cost model. As described in the optimization loop described below, both models can receive the same input and produce the same output. Generally, the difference between these two types of cost models is how each model internally predicts costs. There are various differences between analytical cost models and ML-based cost models.
[0030] For example, an analytical cost model may be a roofline-based model that considers various "ceilings" based on a series of hardware mapping parameters and a neural network graph. The analytical cost model does not require training data. Given an input, the analytical cost model uses "internal logic" to identify bottlenecks and output costs. A "cost module" can be shared by configuring one or more hardware blocks used internally to implement the analytical cost model.
[0031] The shared cost module can operate to generate costs by considering hardware mapping parameters and neural network computations executed on hardware blocks. In some cases, the analytical cost model generates particularly accurate cost outputs for applications with deterministic data flow.
[0032] ML-based cost models require labeled data to train machine learning models capable of predicting at least latency and throughput. For example, a machine learning model can be trained to predict cost values for various application-level objectives, including one or more PPA constraints. ML-based cost models can be implemented using supervised learning and multi-level perceptrons. In some implementations, training data for ML-based cost models is obtained through high-level synthesis and RTL simulations. To overcome the discrete nature of the inputs, the inputs to the ML-based cost model can be transformed into learned embeddings using standard techniques such as stochastic gradient descent. In some cases, ML-based cost models are trained offline. The trained ML-based cost model is used during the optimization loop (described below) to dynamically optimize candidate architectures.
[0033] The optimization and tuning module (112) and the cost model set (114) can each function as extensions of the design space (104). In some implementations, the optimization and tuning module (112) and the cost model set (114) represent a global tuner that tunes the properties of both the neural network and the hardware blocks of the candidate architecture. Control logic in the design space (104) can be used to control or manage the operation of the global tuner. For example, the global tuner can interact with various aspects (e.g., variables and constraints) of the design space (104) to tune the candidate architecture based on control signals generated using the control logic. This is described in more detail below with reference to FIG. 2.
[0034] The optimization and tuning module (112) includes an exemplary tuner (116) and an exemplary scheduler / mapper (118). In some implementations, the tuner (116) and the scheduler / mapper (118) interact to execute an exemplary tuning and optimization task (described below) of the module (112). As mentioned above, the data processing architecture may be a combination of, for example, a hardware circuit architecture obtained from the architecture repository (106) and a neural architecture obtained from the neural network graph module (108). The hardware architecture may include several individual hardware blocks, each containing hardware functions such as systolic array cells, vector processor lanes, or individual computing tiles.
[0035] The tuner (116) and the scheduler / mapper (118) cooperate to i) configure candidate mappings of neural network layers into one or more hardware blocks, and ii) tune the respective micro-architecture of each hardware block based on one or more application goals (102) for these candidate mappings. In this way, the optimization and tuning module (112) is configured to tune the respective micro-architecture of each hardware block so that a given hardware block is optimized to execute one or more layers of a neural network.
[0036] To achieve desired performance goals, the optimization and tuning module (112) may repeat the process of interacting with the architecture-wear-cost model (114) to construct candidate mappings and tune the micro-architecture of each hardware block. These tuning iterations may include, for example, signal communication through an optional data path (120) from the optimization and tuning module (112) to the design space (104). The communication may, for example, obtain new inputs, variables, constraints, or architectural features to strengthen the hardware blocks of the candidate architecture based on performance estimates generated by the cost model (114). The system (100) may include a tuning loop (122) representing the iterative process.
[0037] The system (100) generates an exemplary output configuration (130) based on the processing operations of the design space (104), the optimization and tuning module (112), and the architecture-ware cost model (114). As described below, the system (100) can automatically generate an application-specific ML hardware accelerator (e.g., an integrated circuit) based on the output configuration (130).
[0038] FIG. 2 is a block diagram illustrating an exemplary system (200) including a global tuner (202). In some cases, the system (200) is included within the system (100) as a subsystem of a software / computation module or hardware circuit having programmed instructions that can be executed by one or more processing devices.
[0039] The operation of the system (200) provides a global tuning framework for automatically generating application-specific ICs customized to perform learning tasks, such as training and inference for a target application. In some implementations, the target application (or device) is a custom hardware accelerator with a fixed hardware configuration. In some other implementations, the target application is a type of workload related to image classification, object detection, autonomous vehicle navigation, graphics processing, or scientific computing.
[0040] The global tuner (202) is configured to globally tune / optimize candidate architectures according to various application goals (102) to create application-specific ML hardware accelerators. The global tuner (202) includes a design space builder (204) that constructs a design space (104) based on one or more tuner variables and constraints (210). The design space builder (204) communicates with a design space explorer (212) and one or more cost models (214) of the global tuner (202). The cost models (214) correspond to individual models of the aforementioned architecture-wear cost models (114).
[0041] Based on the parsed neural network graph of the module (108), the design space builder (204) and the design space explorer (212) can interact to implement a neural architecture search (NAS) system for selecting a neural network architecture ("neural architecture") to be optimally performed for the target application. The NAS may employ various search techniques, such as reinforcement learning-based techniques, evolutionary search, and differential search. The design space builder (204) and the design space explorer (212) may use similar approaches to explore various hardware architectures that can be efficiently tuned and optimized for the target application.
[0042] The design space builder (204) and design space explorer (212) implement NAS and hardware architecture search techniques based on one or more tuner variables and constraints (210). The tuner variables and constraints (210) include various unroll factors, maximum mapper input / output data width, or maximum reducer input / output data width. As described above, a neural network layer may have a corresponding kernel set (e.g., weights / parameters). The kernel may be a convolution kernel with four dimensions (C - input channels; K - output channels; R - kernel height; S - kernel width). An example of a convolution operation can be represented as a nested loop using four-dimensional parameters (C, K, R, S). The kernel set is represented as a multidimensional tensor, and the various dimensions of the tensor can be explored using nested loops. In this context, the unroll factor corresponds to the unrolling of each nested loop. The global tuner (202) supports unrolling nested loops for all unroll factors and can tune candidate architectures for these factors.
[0043] The width of the mapper and reducer input / output data affects how large tensors are reduced into smaller pieces mapped to a given compute tile or cell. For example, input and output tensors can be quite large and these tensors are not generated all at once. To reduce the area and power of the hardware accelerator processing these tensors, the system (100) can utilize tensor tiling to divide the input and output tensors into several smaller pieces. For example, the system (100) can decompose (or reduce) large input tensors into smaller pieces based on mapping constraints. Mapping constraints can be tied to objectives such as power, area, latency, and / or throughput. A global tuner (202) can use these objectives to determine the configuration and size of the compute tile set for a candidate architecture. The global tuner (202) can map computations for different parts of the input tensor to a given tile of the compute tile set.
[0044] The maximum mapper input / output data width and the maximum reducer input / output data width are constraints that directly affect the data throughput of the candidate architecture. Tuner variables and constraints (210) may include other items related to exploring candidate architectures for creating a hardware ML accelerator customized to run a given neural network for a target application. In some implementations, as the tile size becomes smaller, the data transfer time becomes longer, so overall chip performance may also be affected here. All these other tuner variables and constraints (210) may result in different hardware designs that affect performance, power, and area. Therefore, the global tuner (202) maintains a balance between performance, power, and area by forming a design space from these variables / constraints and selecting optimal parameters to customize the hardware and neural architecture.
[0045] The global tuner (202) can dynamically tune the candidate architecture based on operations performed by at least each individual ML cost model (214). In some implementations, the global tuner (202) dynamically tunes the candidate architecture based on operations performed by at least one of i) a random search (exploration) tuner; ii) a simulated annealing tuner; or iii) a progressive tuner. The random search (exploration) tuner, the simulated annealing tuner, and the progressive tuner each correspond to the aforementioned tuner (116). In the case of a block partitioning model, the global tuner (202) implements a specific tuning trajectory associated with the simulated annealing tuner. Each of the random tuner, the simulated annealing tuner, and the progressive tuner may be implemented in software, hardware, or both. The functions associated with each of these tuners can be integrated into the tuner (116) implemented in the global tuner (202).
[0046] The global tuner (202) uses a random search tuner to randomly sample the search space to obtain a trial configuration such as the baseline processor configuration of the candidate architecture. It queries the performance and power cost models of the ML cost model (214) to obtain the cost of running the target application in the trial configuration / architecture.
[0047] Simulated annealing can be implemented as a tuner of a global tuner (202) and is a probabilistic technique for approximating the global optimal of a given function. At each step, this tuner considers neighboring hardware design points d' of the current hardware design point d' and probabilistically decides whether to move the current design point toward design point d' or to maintain design point d. A temperature variable is generated to control the acceptance probability. The simulated annealing tuner is configured to repeat these steps until the probabilistic result indicates that the optimal design point for the target application has been reached. For example, a probability score exceeding a threshold score may indicate that a particular design point performs optimally for the target application in relation to a given set of constraints.
[0048] Adjacent hardware design points can be generated randomly. In some implementations, neighboring hardware design points have hardware parameter choices (e.g., unrolling, tiling, mapping, or scheduling) that are similar or very similar to the current hardware design point. The similarity of parameter choices can be characterized by the degree (or percentage) of overlap in hardware parameter choices between the two design points. In some other implementations, neighboring hardware design points may have one or more of the same hardware parameter choices as the current hardware design point.
[0049] The global tuner (202) uses a progressive tuner to implement a progressive search methodology for an exemplary design space, such as the design space of a NAS. Using this progressive search methodology can reduce the time spent exploring the design space for candidate architecture tuning. In some implementations, the global tuner (202) executes the progressive search methodology to explore the design space as a step of designing and tuning ML hardware to meet (or exceed) specific throughput requirements, such as fixed data rate inputs for machine learning blocks of an integrated circuit. The progressive search methodology may include at least i) initializing a baseline design with a minimum design for all neural network layers and ii) querying a cost model (214) to identify bottleneck layers that have data throughput lower than the data rate requirements. If the cost model (214) does not identify or indicate a bottleneck and / or the global tuner (202) determines that no layer of the neural network is operating as a bottleneck, the execution of the search methodology is terminated.
[0050] The progressive search methodology may further include the steps of iii) thoroughly exploring the search space by reference to bottlenecks to determine a design configuration that minimizes bottlenecks by meeting (or exceeding) throughput requirements while minimizing the cost for overall model performance, and iv) using the design configuration determined in step iii) as a new baseline design and then returning to step ii). In some implementations, the baseline design is a baseline processor configuration containing the minimum hardware (and neural) architecture / design parameters for executing all layers of a given neural network. Thoroughly exploring the search space involves iteratively exploring various design configurations by implementing a multilayer neural network using each design configuration, evaluating the corresponding data throughput of each design configuration, and calculating the corresponding cost value for each of the various design configurations.
[0051] In the example of FIG. 2, the input goal (102) may be user-defined, system-defined, or both. For example, the input goal (102) may be received as a user configuration file or as a system-generated input file. The configuration or input file may specify various application-level goals (102) derived, for example, from a set of PPA constraints. For example, the input file may contain a set of application-level goals such as processor utilization, power consumption, data throughput, hardware block size, and / or latency. The input file also contains respective hardware accelerator performance thresholds for each application-level goal.
[0052] In some implementations, the input file includes a goal (102) indicating that the target application requires multi-vector operations. Based on this indication, control logic can trigger a hardware variable (110) in the design space (104) to be set as a vector parameter (vector_ctrl). The design space (104) can use the "vector_ctrl" parameter to restrict the selection of candidate architectures to, for example, architectures that include multi-vector processing lanes forming a tightly coupled VPU.
[0053] In the example of FIG. 2, some (or all) of the cost models (214) perform ML-based analysis to tune the candidate architecture. Depending on the input goal set (102), the global tuner (202) tunes the hardware and neural architecture of the candidate architecture based on one or more optimization algorithms. For example, the global tuner uses the cost models (214) to model the use of the candidate architecture for executing each layer of a multilayer neural network by referencing specific hardware block(s) of the neural network. In response to modeling the use of the architecture for executing each layer, the ML cost models (214) generate performance parameters that describe how the architecture performs for each layer.
[0054] In some implementations, an optimization algorithm is used to implement a cost model interaction loop, for example, an optimization loop. For example, an optimizer or global tuner (202) (e.g., simulated annealing, progressive, random, etc.) can generate a set of hardware mapping parameters such as PE number, thystolic array dimension, etc. Hardware mapping parameters are transmitted to a cost model (214) along with a neural network graph containing hierarchical dependencies and a quantization method (e.g., fixed). The cost model (214) generates costs such as latency, throughput, and power based on the input. The cost output of the cost model can be fed back to an optimizer as a step of the optimization loop. The optimizer can process the cost output and determine the next hardware mapping strategy to explore. The global tuner (202) can repeat this optimization loop until a convergence condition is satisfied or the search space is fully explored.
[0055] In some implementations, the first cost model (214) of the global tuner (202) is used to calculate performance estimates / parameters for hardware properties of the candidate architecture, while the second cost model (214) is used to calculate performance estimates / parameters for neural networks implemented in the candidate architecture. The first and second cost models (214) may be the same or different. The cost model (214) may use a single optimization algorithm to calculate performance estimates for tuning the architecture and optimizing the performance of the candidate architecture. In some other implementations, the cost model (214) uses various optimization algorithms to calculate performance estimates for optimizing various aspects of the architecture's performance.
[0056] The global tuner (202) may use at least a design space builder (204), a design space explorer (212), and a cost model (214) to implement various design space and optimization strategies for various hardware and neural network architectures being explored. For example, within each hardware block of a candidate architecture, the global tuner (202) explores various implementations specifically targeting a single layer, such as layer-by-layer tiling and tuning of systolic array dimensions. The global tuner (202) may explore layer transformations to increase parallelization. For example, the global tuner (202) may transform dense / 1 x 1 convolutions into nxn convolutions to increase the throughput and / or utilization of computing units across one or more hardware blocks.
[0057] In some implementations, based on an optimization algorithm, the cost model (214) calculates a utilization estimate from the indication that a dense convolution is assigned to a single computational unit of a hardware block containing multiple computational units. The global tuner (202) can compare the utilization estimate to a utilization threshold specified by an application objective (102) (or constraint (210)). The global tuner (202) determines whether the calculated utilization estimate is below the threshold. In response to the determination that the calculated utilization estimate is below the threshold, the global tuner (202) can convert the dense / 1 x 1 convolution to an nxn convolution to increase the utilization of the computational unit across a given hardware block. The utilization estimate is a performance parameter (or estimate) generated by the cost model (214).
[0058] In the case of a multidimensional array of processing engines (e.g., cells, tiles, or processing lanes), the global tuner (202) can determine the optimal size / area and expected power density required to achieve the desired performance goal. The global tuner (202) can change the number of processing engines (PEs) in each dimension of the array based on the determined size. The system (100, 200) is configured so that one or more deep hardware customizations for one layer of the neural network do not interfere with or adversely affect the efficient execution or operation of other layers of the neural network.
[0059] The global tuner (202) generates an output configuration (230) in response to candidate architecture tuning. The output configuration (230) is used to automatically generate an application-specific ML accelerator. The output configuration (230) may represent an ML model (or algorithm) and its corresponding architecture configuration. The system (200) uses an example code generation module (240) to convert the data representing the output configuration (230) into High-level synthesis (HLS) code. For example, the code generation module (240) can use HLS to generate a firmware implementation of an ML algorithm for a hardware accelerator.
[0060] Generally, the global tuner (202) is used to generate one or more application-specific ML accelerators that are fully customized for the target application. For example, customization may include items such as heterogeneous quantization and micro-architectures tuned to one or more neural network layers. In some implementations, the global tuner (202) and the system (200) are used to generate the customized architecture by identifying optimal hardware parameters, such as micro-architectures, spatial mappings, and temporal mappings, to optimize the entire architecture for at least a set of PPA constraints (e.g., target (102)).
[0061] Hardware functions can be separated on or inside the chip. To optimize the spatial mapping of an architecture, hardware blocks used to execute various spatially separated neural network operations are required within the chip or integrated processor block. For example, a candidate architecture can be optimized for spatial mapping by using a specific array of dedicated hardware blocks to execute dedicated operations in the neural network. This mapping allows the hardware blocks to be tuned to specific algorithms or computational patterns.
[0062] Architectures with optimized spatial mapping compared to other designs can improve performance and energy efficiency. These improvements can be realized through an array of dedicated hardware blocks customized to implement at least specific algorithms or computing patterns. In some implementations, one or more dedicated hardware blocks are configured to process fixed-dimensional tensors, support fixed-quantization schemes, and be tuned for specific neural network layers.
[0063] Optimizing the temporal mapping (307) of the architecture involves hardware blocks that are time-shared between various operations of the neural network. For example, a candidate architecture can be optimized for temporal mapping by reusing the same hardware blocks to execute various operations in the neural network. By using a given hardware block more generally, this approach can improve the programmability of the hardware. Additionally, this approach can provide application developers with more flexibility regarding the neural networks that can be executed on the hardware. In some examples, the optimized temporal mapping provides time sharing between different layers on the same hardware block and supports multiple quantization schemes.
[0064] Through customization, application-specific ML accelerators can be created that consume much less power and area compared to other processing devices that are not customized for the target application.
[0065] FIG. 3 illustrates an exemplary framework (300) for tuning a multilayer neural network. Using this framework system (100), computation nodes of a neural network graph can be iteratively mapped to various functions of a micro-architecture (or processing engine) of a given hardware block. For example, the framework (300) may be implemented in a global tuner (202) or an optimization and tuning module (112) to determine and construct dependencies between various computation nodes of the neural network graph. Dependencies may be determined, for example, when an ML cost model (214) models the execution of each layer of the neural network by a candidate architecture. The ML cost model (214) generates performance parameters that provide an evaluation of how the candidate architecture performs when executing each layer of the neural network.
[0066] In the example of FIG. 3, the neural network (302) includes five layers (L1-L5), where the first layer is L1, the second layer is L2, and so on. These five layers may have initial mappings to various hardware functions (e.g., processing engines) of the candidate architecture. For example, each of the five layers may be mapped to a different cell of a thystolic array, a different thystolic array block, a different MAC (Multiply-Accumulate Cell) of a computing tile, or a different computing tile. In some implementations, the individual cells of the thystolic array and the individual MACs of the computing tile represent aspects of the microarchitecture of the candidate architecture.
[0067] The cost model (214) can calculate a performance estimate for a candidate architecture running the neural network (302). The performance estimate includes parameters representing the time durations required for processing a specific layer, the total processing latency, and PE utilization. The cost model (214) processes the time durations to generate a neural architecture schedule (304) optimized for a set of timing constraints. Based on the performance estimate, the global tuner (202) can determine that the time required to compute layers L1 + L2 + L5 is approximately the same as the time required to compute layers L3 + L4.
[0068] Based on this decision, the global tuner (202) may remap layers L1, L2, and L5 to reuse the same hardware function B1, while layers L3 and L4 may remap to reuse the same hardware function B2 (306). In some examples, B1 and B2 are processing engines (308, 310), such as compute tiles or thystolic arrays, MACs, thystolic array cells, or even the arithmetic logic units (ALUs) of the vector processing lanes of a VPU. The global tuner (202) may perform remapping as part of a tuning operation to reduce processing latency and optimize candidate architectures to run the neural network model according to the latency requirements specified in the goal (102).
[0069] For certain neural networks, each layer may require different computation cycles. For example, after spatial remapping, some PEs may experience more idle time than others due to computational imbalance. This can be referred to as load imbalance. The system (100) can address or overcome load imbalance by utilizing tuning and optimization mechanisms that allow PE reuse across multiple layers in at least a temporary manner. For example, a tuner (116) and a scheduler / mapper (118) can detect load imbalance and tune the properties of candidate architectures to evenly balance the computation cycles of each PE.
[0070] As mentioned above, the five layers of the neural network (302) may have an initial mapping in which each layer is mapped to a different hardware function (e.g., a processing engine) of the candidate architecture. Performance estimates for this initial mapping may include utilization parameters indicating low utilization of the total computing function in each processing engine to which the layer can be mapped. Based on these estimates and parameters, the global tuner (202) may also perform remapping to increase processing utilization, for example, by remapping layers L1, L2, and L5 to reuse the same processing engine B1, and by remapping layers L3 and L4 to reuse the same processing engine B2. These remappings may be performed to increase overall utilization in B1 and B2, respectively, and to optimize the candidate architecture to run the neural network model according to the utilization (and latency) requirements specified in the goal (102).
[0071] The global tuner (202) can tune the candidate architecture to reassign different operations to any remaining PEs (e.g., B3, B4, B5). In some cases, the global tuner (202) uses the design space explorer (212) to increase the hardware layout of the candidate architecture to reduce the number of PEs (e.g., from 5 to 2). In some other cases, the global tuner (202) uses the design space explorer (212) to reconfigure the PEs to increase the amount of parallelism across at least B1 and B2. The global tuner (202) may determine that the remaining PEs (e.g., B3, B4, B5) are needed to process a smaller dataset after remapping. Based on this determination, the global tuner (202) can optimize the size and utilization of the PEs to process a smaller dataset by adjusting calculations, for example, the memory ratio of the microarchitectures of these PEs.
[0072] The framework (300) may correspond to an exemplary algorithm or computational sequence that takes a neural network graph as input, along with application-level goals (e.g., inference time, throughput, power, etc.) and applicable hardware constraints (110, 210). The global tuner (202) may use the framework (300) as a basis for performing hierarchical spatial mapping exploration for various architectural knobs. Various architectural knobs may be supported by the framework (300), and these architectural knobs may include i) design styles such as thystoletic arrays or fully unrolled designs; ii) multiple mappers (e.g., thystoletic array clusters); iii) the number of thystoletic arrays per cluster; iv) input and output tiling; and v) hardware dimension transformations for dense layers.
[0073] Each remapping or tuning to achieve optimization for a given constraint (210) may trigger corresponding adjustments to the candidate architecture in relation to other constraints. For example, remapping B1 and B2 to optimize a given timing or latency constraint (constraint) may require an increase in throughput requirements for the PE. Therefore, it is often necessary to improve various architecture knobs to meet new (or other existing) requirements. In some implementations, the system (100) iterates tuning of the candidate architecture to balance the interactions between at least latency, timing, and utilization in order to optimize the candidate architecture for each of these constraints. In some other implementations, the system (100) balances the interactions between multiple constraints, variables, and goals.
[0074] Each architecture knob can have a positive or negative impact on end-to-end application performance. Additionally, each architecture knob can also affect the effect of the architecture knob in other layer mappings. Therefore, based on at least the machine learning aspect of the control logic and the architecture-wear cost model (114), the system (100) is configured to provide a holistic view of the candidate architectures being evaluated to accurately predict these positive and negative impacts.
[0075] The candidate architecture may include multiple processing engines, and one or more layers may be mapped to separate processing engines based on predefined merging rules (e.g., conv2d + BN + activation merging, conv2d + maxpooling merging). Merging rules may be predefined, for example, as instructions or coded rules of the network graph module (108). In some implementations, two or more graph nodes (or layers) are merged if the computation of the next layer corresponds to the computation of the previous layer (e.g., conv2d (+ BN) + activation). For example, the computation for the batch normalization (BN) layer may be merged with the computation for the 2D convolution layer. Additionally, for each layer output provided as input to a subsequent layer, if the input and computation amount for the subsequent layer are of a critical size and possess specific spatial and temporal locality, this subsequent layer may be merged with the previous layer that produced the layer output. An example of this can be the layer output of a 2D convolutional layer provided as input to a pooling layer (e.g., conv2d + pooling).
[0076] In some implementations, to tune the candidate architecture, the global tuner (202) performs an initial mapping of each layer to the corresponding PE and generates a performance estimate for the initial mapping. Based on the performance estimate for the initial mapping, the global tuner (202) can tune the initial mapping by iteratively mapping various combinations of layers to the PE. The global tuner (202) generates a performance estimate for each iteration and identifies the mappings where the performance estimate matches the set of PPA constraints of the target (102).
[0077] When tuning a candidate architecture, the global tuner (202) iterates through various mappings using one or more cost models (214) and calculates performance parameters for each mapping. From the performance parameters, the system (100) identifies the mapping of computations that are optimally performed for a given set of PPA constraints (210). In some implementations, the system (100) may iteratively map computation nodes for various vector operations to a subset of vector processing lanes of the VPU using a temporal mapping that specifies a sequence of nodes operating within a processing lane.
[0078] In some implementations, the framework (300) uses an architecture-ware analysis cost model (114) to predict the cost of each trial (hardware / neural configuration) because (1) cycle accurate simulation (CAS) for each trial is time-consuming and there are often millions to billions of unique design points to evaluate; and (2) the computation of neural networks is computationally intensive and can be represented by nested loops, so the analysis model can be constructed with high fidelity. The optimization and tuning module (112) samples the search space, queries the cost of each design point in the cost model (114), and searches the design space (104) along a specific exploration trajectory. The cost of each design point and the exploration trajectory of the design space (104) are implemented to optimize candidate architectures by tuning the architecture to minimize the processing cost of at least each design point. In some cases, the search trajectory depends on different tuner algorithms used by the tuner (116).
[0079] FIG. 4 is a flowchart of an exemplary process (400) regarding a graph execution schedule of a multilayer (multiple) neural network. As previously described, the global tuner (202) generates an output configuration (230) used to automatically generate an application-specific ML accelerator. The system (200) uses an example code generation module (240) to convert data representing the output configuration (230) into an HLS code.
[0080] The neural network graph (402) is for a custom application-specific ML accelerator and represents an exemplary assignment or mapping to a set of neural network layers. In the example of FIG. 4, the first neural network layer L1 may be mapped to a given PE based on a specific hardware configuration (HW Config) (404a) and software configuration (SW Config) (404b), whereas the second other neural network layer L2 may be mapped to a given PE based on a specific hardware configuration (406a) and software configuration (406b). In some implementations, L1 and L2 may be mapped to the same PE or different PEs.
[0081] FIG. 5 is a flowchart illustrating an exemplary process (500) for creating and globally tuning an application-specific machine learning accelerator. The process (500) may be implemented or executed using the system (100) described above. The description of the process (500) may refer to the computing resources (resources) of the system (100) mentioned above. The steps or actions of the process (500) may be enabled by programmed firmware or software instructions executable by one or more processors of the devices and resources described herein.
[0082] Referring to process (500), the system (100) selects an architecture (502). For example, the controller of the system (100) may select a candidate architecture representing a baseline processor configuration. The candidate architecture may include a hardware architecture and a neural architecture corresponding to a neural network graph. In some implementations, the architecture is identified and selected based on a search operation performed by the design space builder (204) and the design space explorer (212) on the hardware layout of the architecture repository (104) and the neural architecture of the network graph module (108).
[0083] The system (200) can implement NAS and hardware architecture search techniques based on one or more tuner variables or PPA constraints (210). PPA constraints may be custom goals (102) that define performance requirements for the hardware accelerator. For example, requirements may be thresholds for processor utilization, power consumption, processing latency, and data throughput. In some implementations, architecture selection involves obtaining input criteria that specify performance objectives and identifying multiple candidate architectures for implementing a specific objective processor. For example, control logic for managing a design space (104), including a design space builder (204) and an explorer (212), may select a candidate architecture from among multiple candidate architectures based on input criteria.
[0084] The system (100) generates performance data for an architecture (504). For example, an ML cost model (214) generates performance data for a candidate architecture by modeling how the architecture executes computations of a first neural network that includes at least multiple neural network layers. In some implementations, the neural network is a known neural network, such as a multilayer ResNet-50, which is a convolutional neural network with 50 layers deep.
[0085] The system (100) dynamically tunes the architecture based on performance data (506). For example, based on performance data, the optimization and tuning module (112) dynamically tunes the candidate architecture to meet one or more performance goals. More specifically, the optimization and tuning module (112) interacts with an architecture-wear cost model (114) to model the execution of the candidate architecture for each layer of the neural network. For example, the ML cost model (214) generates performance parameters that provide an evaluation of how the candidate architecture performs when executing each layer of the neural network.
[0086] The system (100) uses a tuning loop (122) to evaluate, tune, and optimize the architecture implementation of the first neural network based on performance parameters. In some implementations, the system (100) uses global tuning (e.g., via a global tuner (202)) to discover system-specific optimized operation-specific mappings for efficient neural network execution on the target hardware platform. In some other implementations, the system (100) uses global tuning to discover optimized graph execution schedules whenever possible, such as the reuse of processing engines (PEs) across multiple layers. This was described above with reference to FIG. 3.
[0087] For example, the global tuner (202) is configured to tune a candidate architecture by remapping two or more layers (e.g., L1, L2, L5) to the same subset of MACs or compute tiles to optimize a selected neural architecture for a target application. The architecture may be optimized for exemplary applications such as a training / inference unit or an image classification workload. The control logic of the system (100) may use the timing of the clock signal to send instruction and control signals to the optimization and tuning module (112) and the architecture-wear cost model (114), respectively, at the appropriate time to generate performance data used to perform remapping. The optimization and tuning module (112) is configured to perform application-specific tuning and optimization to create a hardware layout of an integrated circuit that accelerates the ML workload. The optimization and tuning module (112) (and cost model (114)) may incorporate some (or all) of the functions of the global tuner (202) so that descriptions of operations performed by the global tuner (202) are converted into operations of the optimization and tuning module (112).
[0088] The system (100) generates a configuration of the ML accelerator in response to dynamically tuning the architecture (508). In some implementations, the tuning and optimization of step 506 is implemented as an output configuration (230) that allows for the creation of an integrated circuit of a special target with a hardware architecture that is customized by layer. Through this customization aspect, the hardware ML accelerator circuit can significantly improve energy efficiency compared to previous approaches based on a single general hardware block.
[0089] For example, after optimizing and tuning a candidate architecture, the system (100) generates a compatible hardware configuration (230) that includes various architectural features and scheduling / mapping strategies so that it can be used by at least a code generation module (240) to generate an application-specific ML accelerator. The system (200) uses the code generation module (240) to convert data representing the configuration (230) into High-level synthesis (HLS) code. The code generation module (240) can generate a firmware implementation of an ML algorithm for the hardware accelerator using High-level synthesis (HLS). Then, the system (100) can generate an application-specific hardware ML accelerator based on the firmware implementation and HLS operations (510).
[0090] FIG. 6 is a block diagram of an exemplary application-specific hardware ML accelerator (600). The hardware accelerator (600) is created using the technology disclosed herein, including exemplary operation of at least the systems (100 and 200). Using a code generator (240), the system (100) is configured to generate a hardware layout for the application-specific ML accelerator (600) that specifies each part of the hardware circuit, and each hardware circuit can be customized to run a specific layer of a neural network.
[0091] The hardware accelerator (600) may use separate hardware blocks (603a, 603b, 603c, 603d, 603e, 603f) to execute one or more layers (e.g., when they share common properties) in a streaming and pipelined manner. Each hardware block (603) is tailored specifically to these layers (e.g., quantization, layer-specific tiling, thystolic array dimensions, etc.) to enable low power and high utilization across the hardware accelerator (600), for example. In some implementations, each hardware block (103) has an association or mapping with a specific layer of the neural network, and the association between the layer of the neural network (e.g., L1, L2, L3, L4, or L5 described above) and the hardware block (103) is based in part on the functions and optimization efforts associated with the corresponding layer of the neural network.
[0092] Data flow indicators (601a, 601b, 601c, 601d, 601e, 601f) provide an exemplary sequence of neural network data communication between hardware blocks (603). In some implementations, these data flow indicators (601a, 601b, 601c, 601d, 601e, 601f) are pre-configured communication sequences based, for example, on optimization and tuning operations of a global tuner (202). The neural network data being transmitted may include computation result data such as the output of a computation unit in a specific hardware block (603), neural network input / activation, parameter weight data, and other neural network parameter-related data.
[0093] Each hardware block (603) may include a microarchitecture customized for the target application. The global tuner (202) is configured to optimize communication across different hardware blocks in a global tuning operation to balance the architectural design at the system level. Such optimizations include interface tiling for matching data transmission rates, the number of compute blocks for rating matching during computation (e.g., input channel blocking), buffer size tuning, etc. For example, hardware block (603a) may include inter-die input blocks (606a, 609b), inter-die output blocks (611a, 611b), and a host interface unit (613), and hardware block (603b) may include inter-die input blocks (621a, 621b), inter-die output blocks (623a, 623b), and a host interface unit (614).
[0094] The custom configuration of the accelerator (600) may include a first layer of the neural network mapped to a hardware block (603a) and a last layer of the neural network mapped to a hardware block (603d). The global tuner (202) may configure this architecture to incorporate a feedback layer between, for example, the hardware blocks (603a, 603d) to balance the interaction between per-op spatial mappings for efficient neural network execution and the size / area constraints of the PPA constraint (210). For example, the hardware accelerator (600) is configured to efficiently perform neural network computations using minimal hardware while matching throughput / latency according to application-specific requirements.
[0095] FIG. 7 illustrates an example of a tensor or multidimensional matrix (700) comprising variations of an input tensor (704), an output tensor (708), and a weight tensor (706). The tensor (700) is an exemplary machine learning data structure that is processed or generated using an ML hardware accelerator, such as an accelerator (600). For example, the system (100) may be used to automatically generate a custom hardware ML accelerator (600) configured to implement a neural network that receives and processes data associated with the tensors, and to tune and optimize candidate architectures for processing at least the tensors (704 and 706).
[0096] Each tensor (700) contains elements corresponding to data values for computations performed in a specific layer of the neural network. Computations may involve multiplying an input / activation tensor (704) and a parameter / weight tensor (706) in one or more clock cycles (periods) to produce an output, such as an activation / output value, that can be provided as input to another layer of the neural network. In the example of FIG. 7, each output of the output set may correspond to each element of the output tensor (708). In some examples, the input tensor (704) is an activation tensor. Multiplying the activation tensor (704) with the corresponding weight tensor (706) involves multiplying the activation from the elements of the tensor (704) by the weights from the elements of the tensor (706) to produce partial sum(s).
[0097] In some implementations, the hardware blocks (603) of the ML accelerator (600) are individual processor cores operating on vectors, which may contain multiple individual elements along the same (or different) dimensions of some multidimensional tensor. Each of the multiple elements may be represented using X,Y coordinates (2D) or X,Y,Z coordinates (3D) depending on the dimensions of the tensor. The hardware layout of the ML accelerator (600) may be optimized to compute multiple subtotals according to a given set of PPA constraints. A subtotal corresponds to a multiplication result generated by multiplying batch inputs by their corresponding weight values.
[0098] Input weight multiplication can be recorded as the sum of the products of each weight element multiplied by the discrete inputs of the input volume, such as rows or slices of the input tensor (704). These rows or slices may represent a given dimension, such as the first dimension (710) of the input tensor (704) or another second dimension (715) of the input tensor (704). Since the dimension can be mapped to various vector processing units across the hardware blocks (603), the ML accelerator (600) performs calculations periodically according to a given set of input goals (102) in a manner that prevents load imbalances and achieves critical processing utilization at each hardware block (603).
[0099] In some implementations, an output for a convolutional neural network layer may be computed using an exemplary set of computations. Computation for a CNN layer may involve performing a 2D spatial convolution between a 3D input tensor (704) and at least one 3D filter (weight tensor (706)). For example, convolving one 3D filter (706) with the 3D input tensor (704) may produce a 2D spatial plane (720 or 725). Computation may include calculating an inner product sum for a specific dimension of the input volume. For example, the spatial plane (720) may contain an output value for a sum of products computed from the input along dimension (710), whereas the spatial plane (725) may contain an output value for a sum of products computed from the input along dimension (715). Calculations for generating the sum of products for the output values of each space plane (720 and 725) can be performed using a hardware block (603) that is generated and tuned using the technique described in this document.
[0100] The gist and embodiments of functional operation described herein may be implemented in digital electronic circuits, computer software or firmware implemented in a type, computer hardware including structures disclosed herein and structural equivalents thereof, or a combination of one or more of these. The embodiments of the gist described herein may be implemented in one or more computer programs, namely, one or more modules of computer program instructions encoded in a non-transient program carrier of a type to be executed by a data processing device or to control the operation of a data processing device.
[0101] Alternatively or additionally, program instructions may be encoded in artificially generated radio signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device for execution by a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these.
[0102] The term “computing system” includes all kinds of devices, devices, and machines for processing data, including, for example, programmable processors, computers, or multiprocessors. Devices may include special-purpose logic circuits, such as, for example, field programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). Devices may also include code that, in addition to the hardware, creates an execution environment for the aforementioned computer programs, such as, for example, processor firmware, protocol stacks, database management systems, operating systems, or a combination thereof.
[0103] A computer program (also referred to as a program, software, software application, module, software module, script, or code) may be written in any form of programming language, such as a compiled language, an interpreted language, a declarative language, or a procedural language, and may be distributed in any form, including standalone programs, modules, components, subroutines, or other devices suitable for use in a computing environment.
[0104] A computer program may correspond to a file in a file system, but it does not necessarily have to be. A program is a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to that program, or multiple tuned files (e.g., a file storing one or more modules, subprograms, or parts of code). A computer program may be located on a single computer or site, or distributed across multiple sites and connected to run on multiple computers connected by a communication network.
[0105] The processes and logic flows described herein may be performed by one or more programmable computers that execute one or more computer programs that perform functions by operating on input data and generating outputs. The processes and logic flows may also be performed by a device, and the device may be implemented as a special purpose logic circuit such as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or a general purpose graphics processing unit (GPGPU).
[0106] A computer suitable for executing computer programs may be based, for example, on a general-purpose or special-purpose microprocessor, or both, or on a central processing unit of another type. Generally, the central processing unit receives instructions and data from read-only memory, random access memory, or both. Some elements of a computer are the central processing unit for executing or performing instructions and one or more memory devices for storing instructions and data. Generally, the computer will also include or be operablely combined with one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, to receive or transmit data from, or both. However, the computer does not necessarily have to have such devices. Furthermore, the computer may be embedded in other devices, such as mobile phones, PDAs (Personal Digital Assistants), mobile audio or video devices, game consoles, GPS (Global Positioning System) receivers, or portable storage devices (e.g., USB (Universal Serial Bus) flash drives).
[0107] Computer-readable media suitable for storing computer program instructions and data include semiconductor memory devices, such as, for example, EPROM, EEPROM, and flash memory devices; magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and all forms of non-volatile memory, media, and memory devices, including CD-ROM and DVD-ROM disks. Processors and memory may be complemented or integrated with special-purpose logic circuits.
[0108] To provide interaction with the user, embodiments of the gist described herein may be implemented in a computer equipped with a display device for displaying information to the user, e.g., an LCD (liquid crystal display) monitor, a keyboard for which the user can provide information, and a pointing device (e.g., a mouse or trackball for which the user can provide input to the computer). Other types of devices may also be used to provide interaction with the user (e.g., feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and user input may be received in any form including sound, voice, or tactile input). Additionally, the computer may interact with the user by exchanging documents with the device used by the user; for example, sending a web page to a web browser on the user client device in response to a request received from a web browser.
[0109] An embodiment of the essence described herein may be implemented in a computing system comprising, for example, a backend component as a data server, a middleware component (e.g., an application server), a frontend component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with the implementation of the essence described herein), and a combination of one or more of the backend, middleware, or frontend components. The components of the system may be interconnected through any form or medium of digital data communication, such as a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), such as the Internet.
[0110] A computing system may include clients and servers. Clients and servers are typically located far apart from each other and usually interact through a communication network. The relationship between a client and a server arises from computer programs running on each computer that have a client-server relationship with one another.
[0111] Although this specification contains many specific implementation details, they should not be interpreted as a limitation on the scope of any invention or claimable scope, but rather as a description of features that may be specific to a particular embodiment of a particular invention. Specific features described in this specification in relation to separate embodiments may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, while functions may be described above as operating in a specific combination and are initially claimed to be so, one or more functions of the claimed combination may be omitted from the combination in some cases, and the claimed combination may relate to a sub-combination or a variation of a sub-combination.
[0112] Similarly, although operations are depicted in a specific order in the drawings, this should not be understood as requiring that such operations be performed in the specific order depicted or in a sequential order, or that all depicted operations be performed to achieve a desired result. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the foregoing embodiments should not be understood as requiring such separation in all embodiments. It should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0113] Specific embodiments of the gist have been described. Other embodiments are within the scope of the following claims. For example, the operations cited in the claims may be performed in a different order and still obtain the desired results. As an example, the process illustrated in the accompanying drawings does not necessarily require the specific order or sequential order illustrated to achieve the desired results. In certain implementations, multitasking and parallel processing may be advantageous.
Claims
Claim 1 A computer-implemented method for generating an application-specific machine learning (ML) accelerator, comprising: selecting an architecture representing a baseline processor configuration; generating performance data for the architecture by modeling, by an ML cost model, a method for the architecture to execute computations of a first neural network comprising at least a plurality of layers; dynamically tuning the architecture based on the performance data such that the architecture implements the first neural network and satisfies a performance objective for the architecture when executing machine learning computations for a target application; determining, in response to dynamically tuning the architecture, a hardware configuration customized to implement each of the plurality of layers of the first neural network and (ii) execute machine learning computations for the target application; and generating a configuration of the ML accelerator based on the dynamically tuned architecture and the customized hardware configuration. Claim 2 The method of claim 1 further comprises the step of generating an application-specific hardware ML accelerator based on the customized hardware configuration, wherein the application-specific hardware ML accelerator is a computer-implemented method optimized to implement each of the different layers of a neural network when the neural network is used to perform computations for a target application. Claim 3 A computer-implemented method according to paragraph 2, wherein the performance objective comprises a plurality of discrete objectives, and the step of creating an application-specific ML accelerator comprises the step of creating an application-specific hardware ML accelerator configured to satisfy each of the plurality of discrete objectives when the application-specific hardware ML accelerator performs computation for a target application. Claim 4 A computer-implemented method according to claim 3, wherein the step of generating the performance data comprises: a step of modeling the use of the architecture for executing each of the plurality of layers of the first neural network by an ML cost model; and a step of generating performance parameters of the architecture for each of the plurality of layers by the ML cost model in response to modeling the use of the architecture for executing each layer. Claim 5 A computer-implemented method according to claim 4, wherein the performance parameter corresponds to each of the plurality of discrete goals; and the plurality of discrete goals include at least one of a critical processing latency, critical power consumption, critical data throughput, and critical processor utilization. Claim 6 A computer-implemented method according to claim 2, wherein the step of dynamically tuning the architecture comprises: determining a mapping of computation for an input tensor that causes the application-specific hardware ML accelerator to utilize the hardware computing unit of the hardware ML accelerator; and dynamically tuning the architecture based on the determined mapping. Claim 7 A computer-implemented method according to claim 6, wherein the step of dynamically tuning the architecture comprises: a step of dynamically tuning the architecture based on an operation performed by each of the plurality of ML cost models of a global tuner including a plurality of ML cost models—the ML cost model that generates the performance data is included in the plurality of ML cost models—; and a step of dynamically tuning the architecture based on an operation performed by at least one of a simulated annealing tuner or a random tuner of the global tuner. Claim 8 A computer-implemented method according to claim 6, wherein the architecture represents one or more hardware blocks of an integrated circuit, and the step of dynamically tuning the architecture comprises the step of dynamically tuning the architecture to satisfy a respective performance objective for each of the one or more hardware blocks when the architecture implements a first neural network to perform computation for the target application. Claim 9 In claim 6, the configuration of the hardware ML accelerator specifies a customized software configuration for the first neural network; and the step of generating the application-specific hardware ML accelerator comprises generating the application-specific hardware ML accelerator based on the customized hardware configuration and the customized software configuration, a computer-implemented method. Claim 10 In claim 6, the ML cost model is an architecture-aware cost model comprising one or more individual analysis models; and the architecture-aware cost model is a computer-implemented method configured to estimate the performance of the architecture based on the deterministic data flow of data processed using the architecture. Claim 11 A system comprising a processing unit and a non-transient machine-readable storage unit for storing instructions for creating an application-specific machine learning (ML) accelerator, wherein the instructions are executed by the processing unit to perform a set of operations, the set of operations comprising: an operation of selecting an architecture representing a baseline processor configuration; an operation of generating performance data for the architecture by modeling, by an ML cost model, a method for the architecture to execute computations of a first neural network comprising at least a plurality of layers; an operation of dynamically tuning the architecture based on the performance data such that the architecture implements the first neural network and satisfies a performance objective for the architecture when executing machine learning computations for a target application; an operation of determining, in response to the dynamic tuning of the architecture, a hardware configuration customized to implement each of the plurality of layers of the first neural network and (2) execute machine learning computations for the target application; and an operation of creating a configuration of an ML accelerator based on the dynamically tuned architecture and the customized hardware configuration. Claim 12 In claim 11, the set of operations further includes an operation to generate an application-specific hardware ML accelerator based on the customized hardware configuration, wherein the application-specific hardware ML accelerator is optimized to implement each of the different layers of the neural network when the neural network is used to perform computations for the target application, a system. Claim 13 A system according to claim 12, wherein the performance objective comprises a plurality of discrete objectives, and the operation of generating the application-specific ML accelerator comprises generating an application-specific hardware ML accelerator configured to satisfy each of the plurality of discrete objectives when the application-specific hardware ML accelerator performs computation for the target application. Claim 14 A system according to claim 13, wherein the operation of generating the performance data comprises: an operation of modeling the use of an architecture for executing each layer among a plurality of layers of the first neural network by the ML cost model; and an operation of generating performance parameters of the architecture for each of the plurality of layers by the ML cost model in response to modeling the use of an architecture for executing each layer. Claim 15 In claim 14, the performance parameter corresponds to each of the discrete goals among the plurality of discrete goals; and the plurality of discrete goals include at least one of critical processing latency, critical power consumption, critical data throughput, and critical processor utilization, a system. Claim 16 In claim 12, the operation of dynamically tuning the architecture comprises: an operation of determining a mapping of computation for an input tensor that causes the application-specific hardware ML accelerator to utilize the hardware computing unit of the hardware ML accelerator; and an operation of dynamically tuning the architecture based on the determined mapping. Claim 17 In claim 16, the operation of dynamically tuning the architecture comprises: an operation of dynamically tuning the architecture based on an operation performed by each of the plurality of ML cost models of a global tuner comprising a plurality of ML cost models—the ML cost model generating the performance data is included in the plurality of ML cost models—; and an operation of dynamically tuning the architecture based on an operation performed by at least one of a simulated annealing tuner or a random tuner of the global tuner. Claim 18 A system according to claim 16, wherein the architecture represents one or more hardware blocks of an integrated circuit, and the operation of dynamically tuning the architecture includes the operation of dynamically tuning the architecture to satisfy a respective performance objective for each of the one or more hardware blocks when the architecture implements a first neural network to execute a computation for the target application. Claim 19 In paragraph 16, the ML cost model is an architecture-aware cost model comprising one or more individual analysis models; and the architecture-aware cost model is configured to estimate the performance of the architecture based on the deterministic data flow of data processed using the architecture, a system. Claim 20 A non-transient machine-readable storage device for storing instructions for creating an application-specific machine learning (ML) accelerator, wherein the instructions are executed by a processing unit to perform a set of operations; said set of operations includes: an operation of selecting an architecture representing a baseline processor configuration; an operation of generating performance data for the architecture by modeling, by an ML cost model, a method for the architecture to execute computations of a first neural network comprising at least a plurality of layers; an operation of dynamically tuning the architecture based on said performance data such that the architecture implements the first neural network and satisfies a performance objective for said architecture when executing machine learning computations for a target application; in response to the dynamic tuning of said architecture, an operation of determining a hardware configuration customized to implement each of the plurality of layers of said first neural network and (ii) execute machine learning computations for said target application; and an operation of creating a configuration of an ML accelerator based on said dynamically tuned architecture and said customized hardware configuration.
Citation Information
Patent Citations
Electronic device and Method for controlling the electronic device thereof
KR1020210032266A