Neural Network Performance Prediction for Compiler Configuration Tuning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing neural network architectures consume significant memory, time, and computing resources due to varying configuration parameters, necessitating improved methods for optimizing compiler configurations to enhance performance.

Innovation Solution

Implementing a trained neural network to predict software performance based on compiler configurations, utilizing a software performance prediction system that simulates different hardware and compiler parameter combinations to identify optimal settings for efficient execution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional testing methods are used to determine optimal compiler parameters, then performance optimization can be achieved, but significant time and computational resources are consumed

Engineering Contradiction:
Improveperformance optimizationVSAvoidtesting time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by training a neural network model in advance using historical compiler configuration data and performance metrics. This pre-trained model can then quickly predict optimal compiler parameters for new neural network architectures without requiring extensive traditional testing, thus reducing the time loss while maintaining performance optimization capabilities

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention creates a virtual copy of the compilation and testing process through a neural network simulation model. Instead of physically testing each compiler configuration on actual hardware, the system uses the trained neural network to copy and simulate the performance outcomes, dramatically reducing the computational resources and time required while still achieving performance optimization

Inventive Principle:
Principle #26Copying

2Productivity

If extensive compiler configuration testing is performed to optimize neural network performance, then execution efficiency improves, but computational resources and costs increase

Engineering Contradiction:
Improveexecution efficiencyVSAvoidcomputational resources
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary actions by training a neural network model in advance using historical compiler configuration data and performance metrics. This pre-trained model can then quickly predict optimal compiler parameters for new neural network architectures without requiring extensive traditional testing, thus reducing the time loss while maintaining performance optimization capabilities

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention creates a virtual copy of the compilation and testing process through a neural network simulation model. Instead of physically testing each compiler configuration on actual hardware, the system uses the trained neural network to copy and simulate the performance outcomes, dramatically reducing the computational resources and time required while still achieving performance optimization

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20260056720A1Predicting neural network performance for compiler configurations
Publication Date: 2026.02.26 NVIDIA CORP
  • US20260056720A1 patent drawing
  • US20260056720A1 patent drawing
  • US20260056720A1 patent drawing

AI summary

Apparatuses, systems, and techniques to predict performance information for software to be compiled and executed on one or more integrated circuits are described. In at least one embodiment, one or more neural networks may be used to generate performance information corresonding to one ro more integrated circuits based, at least in part, on configuration parameters to configure one or more compilers to compile software to be performed by the one or more integrated circuits.