ML Accelerator Generation via Global Architecture Tuning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Designing specialized hardware accelerators for neural networks is labor-intensive and time-consuming, requiring months of effort and multiple iterations to meet application-specific performance and power targets, with manual exploration of design space being prohibitive due to the complexity of parameters and their inter relationships.
Innovation Solution
A system that globally tunes a data processing architecture and automatically generates an application-specific machine-learning accelerator by selecting a candidate architecture based on user-defined or system-determined objectives such as processor utilization, power consumption, and latency, using a cost model to optimize hardware and software configurations for efficient scheduling and mapping.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual design iteration is used to optimize hardware accelerators, then performance and power targets can be met, but the design process becomes extremely time-consuming and labor-intensive
Solution Approach 1:
The system performs automatic hardware-software co-design and optimization without requiring manual human intervention. The automated design exploration engine independently explores the design space, evaluates configurations, and selects optimal solutions, making the design process self-service and eliminating the need for months of manual iteration
Solution Approach 2:
The patent replaces manual mechanical design processes with automated computational systems. Instead of human engineers manually iterating through design configurations, an automated engine with cost models and exploration algorithms systematically evaluates design spaces, substituting human effort with machine-based optimization
2Productivity
If comprehensive design space exploration is performed to meet application-specific requirements, then optimal performance can be achieved, but the complexity of manual parameter exploration becomes prohibitive
Solution Approach 1:
The design space is segmented into distinct exploration dimensions including hardware parameters, software parameters, and their interactions. The automated engine systematically explores each segment independently and combines results, making the complex exploration manageable and structured rather than overwhelming
Solution Approach 2:
The patent introduces an automated design exploration engine as an intermediary between design requirements and final configurations. This intermediary systematically manages the complex parameter space, automatically evaluating interactions between hardware and software parameters without requiring manual navigation of the complex design landscape
3Use of energy by moving object
If specialized hardware accelerators are designed for specific neural network applications, then processing efficiency is improved, but the design process requires multiple iterations to meet power and area constraints
Solution Approach 1:
The system performs preliminary automated exploration of the design space before final implementation. The automated engine pre-evaluates multiple configurations, predicts power and area characteristics using cost models, and selects optimal designs in advance, avoiding the need for multiple physical design iterations and reducing the ease of manufacture burden
Solution Approach 2:
The patent systematically varies design parameters including hardware architecture configurations and software compilation options to optimize power efficiency. The automated engine automatically adjusts these parameters across the design space, identifying optimal combinations that achieve power targets without requiring manual trial-and-error design iterations
Data Source
AI summary
Methods, systems, and apparatus, including computer-readable media, are described for globally tuning and generating ML hardware accelerators. A design system selects an architecture representing a baseline processor configuration. An ML cost model of the system generates performance data about the architecture at least by modeling how the architecture executes computations of a neural network that includes multiple layers. Based on the performance data, the architecture is dynamically tuned to satisfy a performance objective when the architecture implements the neural network and executes machine-learning computations for a target application. In response to dynamically tuning the architecture, the system generates a configuration of an ML accelerator that specifies customized hardware configurations for implementing each of the multiple layers of the neural network.


