Tensor Core Variant Routing for Mixed-Type MXFP Calculations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing solutions for block floating point conversions in AI and machine learning workloads are inefficient, requiring software-based implementations with low performance and high startup costs, and specialized hardware like FPGA, which complicates format changes and resource usage.

Innovation Solution

Incorporating an element data type variant selector within the data stream to allow hardware to dynamically determine the appropriate execution resource for mixed type block data types during runtime, enabling efficient performance without software intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If software-based block floating point conversions are used, then compatibility and flexibility are improved, but performance is reduced and startup time increases

Engineering Contradiction:
ImprovecompatibilityVSAvoidperformance
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent replaces software-based block floating point conversions with specialized hardware circuitry. The tensor core includes dedicated logic to detect element data type variants and route operations to appropriate execution resources, substituting the mechanical/software conversion process with hardware-level detection and routing mechanisms that operate at circuit speed rather than software execution speed.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary element data type variant detector that sits between the data stream and execution resources. This detector identifies the specific variant (e.g., MXFP8, MXFP6, MXFP4) and routes operations to the appropriate execution resource, acting as a mediator that enables hardware specialization without sacrificing adaptability to different data type variants.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If specialized hardware like FPGA is used, then performance is improved, but device complexity and resource usage increase

Engineering Contradiction:
ImproveperformanceVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent designs a universal tensor core architecture that can handle multiple element data type variants (MXFP8, MXFP6, MXFP4, and future variants) through a single hardware unit. The element data type variant detector and execution resource routing logic provide multi-functionality, allowing the same hardware to efficiently process different precision requirements without requiring separate specialized circuits for each variant.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements dynamic routing within the tensor core based on the detected element data type variant. The hardware dynamically selects which execution resource to use based on the actual data being processed, allowing flexible adaptation to different precision requirements while maintaining a fixed hardware architecture. This dynamic behavior reduces the need for multiple static specialized units.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If element data type variant detection is done in software, then adaptability is improved, but execution time and resource overhead increase

Engineering Contradiction:
ImproveadaptabilityVSAvoidexecution time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent replaces software-based element data type variant detection with hardware-based detection logic integrated into the tensor core. The element data type variant detector operates at the hardware level, examining the data stream and identifying variants through circuit-level logic rather than software interpretation, thereby eliminating the time overhead associated with software detection while maintaining full adaptability to different variants.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250285207A1Tensor core hardware for mixed type microscaling floating point (MXFP) data format calculations
Publication Date: 2025.09.11 INTEL CORP
  • US20250285207A1 patent drawing
  • US20250285207A1 patent drawing
  • US20250285207A1 patent drawing

AI summary

Systems and methods for tensor core hardware for calculations involving mixed type block data types (e.g., mixed type MXFP) are described. In one example, a graphics processing unit (GPU) includes decoder circuitry and execution resources. The decoder circuitry decodes a single instruction that identifies a first source operand, a second source operand, and a destination operand, and includes an opcode indicating an operation to be performed utilizing data representing multiple numbers as scalar elements of a block data type in which a block number of an element data type has a value based on a shared scale and a corresponding scalar element. An execution resource is selected based on a variant selector included as an integral part of the block data type that is indicative of a variant of the element data type in which the execution resource is to perform the operation based on the opcode and the variant.