Graphics Processor Workload Partitioning for Accurate Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques do not provide for coordination between inference output and sensors, leading to inaccuracies and underutilization of graphics processors during inference operations.

Innovation Solution

A novel technique is introduced to facilitate the detection of frequently used data values using lookup tables and reduced math, along with a finite state machine that provides pointers to base addresses, enhancing the coordination and utilization of graphics processors during inference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional techniques are used for inference operations, then graphics processors remain underutilized, but accuracy of inference output deteriorates

Engineering Contradiction:
Improvegraphics processor utilizationVSAvoidinference output accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the inference workload into multiple coordinated operations distributed across graphics processor cores. Different cores handle different aspects of the inference pipeline simultaneously, enabling full utilization of GPU resources while maintaining accuracy through specialized handling of coordinate transformations and sensor data processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent makes the graphics processor universal by enabling it to handle both traditional graphics operations and machine learning inference operations simultaneously. The same GPU cores are used for both rendering and inference tasks, eliminating dedicated hardware requirements while maintaining high accuracy through coordinated multi-core operations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If graphics processors are heavily utilized for inference, then inference accuracy improves, but coordination with sensors deteriorates

Engineering Contradiction:
Improveinference output accuracyVSAvoidcoordination between inference output and sensors
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements feedback mechanisms where inference outputs are continuously coordinated with sensor inputs through the graphics processor. The system uses sensor data to adjust and refine inference results in real-time, creating a closed-loop system that improves accuracy while managing coordination complexity through the GPU's unified architecture.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent merges sensor data processing and inference operations into a single coordinated pipeline within the graphics processor. By combining these functions in the GPU, the system reduces external coordination complexity while maintaining high accuracy through integrated multi-core processing of both sensor inputs and inference computations.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12462328B2Coordination and increased utilization of graphics processors during inference
Publication Date: 2025.11.04 INTEL CORP
  • US12462328B2 patent drawing
  • US12462328B2 patent drawing
  • US12462328B2 patent drawing

AI summary

A mechanism is described for detecting, at training time, information related to one or more tasks to be performed by the one or more processors according to a training dataset for a neural network, analyzing the information to determine one or more portions of hardware of a processor of the one or more processors that is configurable to support the one or more tasks, configuring the hardware to pre-select the one or more portions to perform the one or more tasks, while other portions of the hardware remain available for other tasks, and monitoring utilization of the hardware via a hardware unit of the graphics processor and, via a scheduler of the graphics processor, adjusting allocation of the one or more tasks to the one or more portions of the hardware based on the utilization.