Tensor Core Hardware for Dynamic MXFP Type Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing solutions for block floating point conversions, such as software-based approaches and FPGA implementations, suffer from low performance and high entry/startup costs, requiring explicit consideration of element data type variants in tensor operations, which limits efficient execution.

Innovation Solution

Incorporating an element data type variant selector within the data stream allows hardware to dynamically determine the appropriate execution resource during runtime, abstracting software from specific data type considerations and enabling efficient execution of operations like convolution or matrix multiplication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If software-based approaches or FPGA implementations are used for block floating point conversions, then flexibility and adaptability are provided, but performance is low and entry/startup costs are high

Engineering Contradiction:
ImproveflexibilityVSAvoidperformance
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent replaces software-based block floating point conversion systems with a dedicated hardware tensor core system. The hardware tensor core includes specialized circuits for automatic type detection and conversion, substituting the mechanical/software approach with an electronic/hardwired solution that provides both high performance and flexibility through circuit-level parallelism and dedicated conversion pathways.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary type detection and conversion circuit within the hardware tensor core that automatically identifies element data types and performs appropriate conversions. This intermediary component bridges the gap between different data formats without requiring software intervention, enabling high-speed automatic adaptation to various block floating point formats.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If software-based approaches are used for block floating point conversions, then implementation flexibility is provided, but execution performance is low

Engineering Contradiction:
Improveimplementation flexibilityVSAvoidexecution performance
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent implements dynamic type detection and conversion capabilities directly in hardware. The tensor core automatically adapts its operation based on the detected element data type, dynamically selecting appropriate conversion pathways and computational methods without requiring software reconfiguration or compilation, thereby achieving both flexibility and high execution performance.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If element data type variants are explicitly considered in tensor operations, then precision and correctness are ensured, but software complexity increases

Engineering Contradiction:
ImproveprecisionVSAvoidsoftware complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements self-service type detection and handling within the hardware tensor core. The system automatically detects element data types, selects appropriate precision levels, and performs conversions without requiring software to explicitly manage these details. This self-service mechanism maintains precision while eliminating software complexity related to type management.

Inventive Principle:
Principle #25Self-service

4Productivity

If hardware tensor cores are used for tensor operations, then execution speed is improved, but handling mixed data type variants becomes complex

Engineering Contradiction:
Improveexecution speedVSAvoidhardware complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent designs a universal hardware tensor core that can handle multiple element data type variants (e.g., FP8, FP6, FP4) through a unified architecture. The core includes multi-functional conversion circuits and detection logic that automatically adapt to different input types, allowing a single hardware unit to perform tensor operations on various precision formats without requiring separate specialized circuits for each type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP4617864A1Tensor core hardware for mixed type microscaling floating point (MXFP) data format calculations
Publication Date: 2025.09.17 INTEL CORP
  • EP4617864A1 patent drawingFigure 1
  • EP4617864A1 patent drawingFigure 2A
  • EP4617864A1 patent drawingFigure 2B~2C

AI summary

Systems and methods for tensor core hardware for calculations involving mixed type block data types (e.g., mixed type MXFP) are described. In one example, a graphics processing unit (GPU) includes decoder circuitry and execution resources. The decoder circuitry decodes a single instruction that identifies a first source operand, a second source operand, and a destination operand, and includes an opcode indicating an operation to be performed utilizing data representing multiple numbers as scalar elements of a block data type in which a block number of an element data type has a value based on a shared scale and a corresponding scalar element. An execution resource is selected based on a variant selector included as an integral part of the block data type that is indicative of a variant of the element data type in which the execution resource is to perform the operation based on the opcode and the variant.