Co-designing Chip Architecture and Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current AI accelerators are inefficient for machine learning tasks as each task requires a specific deep neural network, and the adaptation process for their component placement and connectivity is not optimized effectively.
Innovation Solution
A method that iteratively optimizes both the chip design and the deep neural network architecture by searching through a defined architecture and hardware components space to improve hardware and machine learning task performance metrics, using optimizers to modify chip designs and neural network architectures until convergence criteria are met.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If current AI accelerators use fixed component placement and connectivity, then hardware design is simple, but adaptability to different machine learning tasks is poor
Solution Approach 1:
The patent implements dynamic component placement and connectivity adaptation through iterative optimization. The system automatically adjusts the placement of computational components and their connectivity based on the specific machine learning task requirements, transforming a static hardware architecture into a dynamically configurable one that optimizes performance for each task.
Solution Approach 2:
The system changes hardware configuration parameters including component placement positions, connectivity patterns, and architectural parameters through automated optimization algorithms. This allows the same physical hardware to be virtually reconfigured with different parameter settings to match various deep neural network requirements.
2Productivity
If each machine learning task requires a specific deep neural network with customized hardware, then task performance is optimized, but hardware resource utilization is inefficient
Solution Approach 1:
The patent creates a universal hardware platform that can execute multiple different deep neural networks through automated placement and connectivity optimization. Instead of designing dedicated hardware for each task, the system provides a single versatile accelerator that adapts its configuration to support various machine learning workloads efficiently.
Solution Approach 2:
The system uses virtualization and software-based neural network execution environments that can be instantiated and configured as needed. This allows multiple virtual instances of different neural network architectures to run on the same physical hardware, enabling efficient resource sharing and eliminating the need for physical hardware duplication for each task type.
3Extent of automation
If manual adaptation of component placement is performed, then hardware design is controllable, but development time and human intervention are high
Solution Approach 1:
The system implements self-service automation where the hardware accelerator automatically optimizes its own configuration for each machine learning task without requiring manual intervention. The optimization algorithms autonomously analyze task requirements, select appropriate components, determine their placement, and configure connectivity, eliminating time-consuming manual design processes.
Solution Approach 2:
The system performs preliminary optimization actions by pre-defining search spaces for architectural parameters and hardware configurations. The automated algorithms explore these pre-structured spaces to quickly identify optimal solutions, avoiding the need for time-consuming manual exploration of all possible configurations from scratch.
Data Source
AI summary
A method, computer program product, and system to generate a processor design via a deep neural network is provided. A processor selects an architecture search space and a hardware components space. A processor selects an initial deep neural network from the architecture search space. A processor determines an initial current chip design for executing the current deep neural network, wherein the initial chip design has a hardware performance metric for implementing the current deep neural network. A processor repeatedly executes an optimization method, the optimization method comprising modifying the chip design one or more times using components from the hardware components space and optimizing the current deep neural network by selecting a deep neural network from the architecture search space. A processor provides the optimized chip design and the specific deep neural network for performing the machine learning task.


