Processor Architecture Exploration for Deep Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current virtual prototyping methods for processor chips are limited in accurately capturing hardware nuances and performance characteristics, leading to discrepancies between virtual and physical chips, and are time-consuming and costly, especially when developing AI models that push computer chips to their limits.

Innovation Solution

A system and method that iteratively optimizes processor architecture and AI model performance without requiring a simulator, using a hardware composer, software composer, and performance calculator to ensure cycle-by-cycle accuracy, allowing for architectural exploration and unique definition of processor architecture for selected ML or AI models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If virtual prototyping is used to test processor architecture before physical chip availability, then software development can begin early, but the accuracy and realism of performance simulation deteriorates

Engineering Contradiction:
Improvesoftware development timeVSAvoidperformance simulation accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent creates a virtual prototype that copies the essential architectural features of the physical processor chip, including instruction set architecture, register files, and execution units. This virtual model enables software developers to begin development early while maintaining sufficient accuracy for performance evaluation, resolving the contradiction between early development access and simulation fidelity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The virtual prototype uses configurable parameters to adjust the fidelity of the simulation, allowing it to match the performance characteristics of the target physical chip. By dynamically adjusting simulation parameters such as clock cycles per instruction and memory access latencies, the system maintains accurate performance measurement while enabling early software development.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If virtual prototype is executed on general-purpose computers, then development cost is reduced, but performance characteristics do not match the final chip

Engineering Contradiction:
Improvedevelopment costVSAvoidperformance characteristic matching
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent replaces the need for expensive specialized emulation hardware with a software-based virtual prototype that runs on general-purpose computers. This substitution maintains cost-effectiveness while achieving reliable performance characterization through accurate architectural modeling and cycle-accurate simulation techniques.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If detailed model of chip behavior is created for virtual prototyping, then functionality can be validated, but development time and complexity increase

Engineering Contradiction:
Improvefunctionality validationVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The virtual prototype model is segmented into distinct functional components including instruction fetch, decode, execute, and write-back stages. Each component is modeled independently with appropriate detail, allowing functionality validation without requiring excessive overall model complexity. This modular approach enables targeted validation of specific processor functions while keeping the total model manageable.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240020536A1Processor architecture and model exploration system for deep learning
Publication Date: 2024.01.18 GROQ INC
  • US20240020536A1 patent drawing
  • US20240020536A1 patent drawing
  • US20240020536A1 patent drawing

AI summary

A processor architecture and model exploration system for deep learning is provided. A method of improving performance of a processor system and associated software includes selecting a set of performance parameter targets for a processor architecture having a set of functional units and an AI model. The method also includes evaluating performance of the processor architecture and the AI model and adjusting at least one of the functional units of the processor architecture to form a new processor architecture prior to iteratively evaluating the combination of the new processor architecture and the AI model. Further, the method includes repeating the evaluating step and the adjustment step until the performance evaluation of the processor architecture and AI model meets the set of performance parameter targets.