Machine Learning Accelerator Mechanism with Unified Execution Resources

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current graphics processors face challenges in efficiently processing diverse data operations due to limitations in programmability and parallel processing capabilities, particularly in handling machine learning tasks.

Innovation Solution

Implementing a general-purpose graphics processing unit (GPGPU) with a scalable architecture that includes a graphics processing engine (GPE) and execution units capable of performing both graphics and machine learning operations, utilizing a unified block of execution resources and shared function logic to enhance parallel processing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If fixed function computational units are used to process graphics data, then processing speed is improved, but adaptability to diverse operations deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidadaptability to diverse operations
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent implements a unified execution resource block that can be dynamically configured to perform multiple functions including graphics processing, machine learning inference, and other computational tasks. This universal processing unit replaces traditional fixed-function units, allowing the same hardware to adapt to different workloads through reconfigurable execution pipelines and instruction sets, thereby achieving both high speed and versatility

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If programmable units are implemented to support wider variety of operations, then adaptability is improved, but processing efficiency deteriorates

Engineering Contradiction:
Improvesupport for wider variety of operationsVSAvoidprocessing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent employs dynamic reconfiguration of execution resources where the processing units can switch between different operational modes and configurations based on the current workload requirements. The execution pipeline dynamically adjusts its structure and resource allocation to optimize performance for the specific task at hand, whether graphics rendering, neural network computation, or other operations, thus maintaining high efficiency across diverse workloads

Inventive Principle:
Principle #15Dynamics

3Productivity

If parallel processing is increased to maximize throughput, then productivity is improved, but device complexity deteriorates

Engineering Contradiction:
Improveparallel processing throughputVSAvoidarchitecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the processing architecture into modular execution resource blocks that can be independently configured and scaled. Each block contains standardized components that can be replicated and combined to achieve desired parallel processing throughput. This segmented modular design allows high productivity through parallel execution while managing complexity by using reusable standardized units rather than custom complex circuits for each function

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12417380B2Machine learning accelerator mechanism
Publication Date: 2025.09.16 INTEL CORP
  • US12417380B2 patent drawing
  • US12417380B2 patent drawing
  • US12417380B2 patent drawing

AI summary

An apparatus to facilitate acceleration of machine learning operations is disclosed. The apparatus comprises at least one processor to perform operations to implement a neural network and accelerator logic to perform communicatively coupled to the processor to perform compute operations for the neural network.