Tile-Based Matrix Duplicate Detection in Processors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Handling large matrices in mainstream processors is instruction intensive and inefficient, particularly when dealing with packed data registers, leading to suboptimal performance in matrix operations.

Innovation Solution

The implementation of tile-based matrix operations using 2D data structures, where matrices are divided into tiles that can be configured for specific dimensions and operations, enabling efficient execution of matrix multiplication, accumulation, and duplicate detection instructions through specialized hardware support.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If mainstream processors handle large matrices using traditional instruction-intensive methods, then they can process matrix operations, but the performance is suboptimal and instruction complexity is high

Engineering Contradiction:
Improvematrix operation performanceVSAvoidinstruction complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides large matrices into smaller tile blocks that can be processed independently and in parallel. This segmentation reduces the complexity of individual operations while maintaining overall matrix processing capability, directly addressing the instruction complexity issue in mainstream processors.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a tile-based two-dimensional processing structure where matrices are viewed and processed as collections of tile blocks arranged in a grid. This dimensional transformation enables parallel processing across multiple tiles simultaneously, improving productivity without proportionally increasing instruction complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Power

If specialized hardware for matrix multiplication is implemented, then peak compute and energy efficiency are improved, but handling of packed data registers becomes instruction intensive

Engineering Contradiction:
Improvepeak compute and energy efficiencyVSAvoidinstruction intensity for packed data
Core Design Contradiction:
PowerVSDevice complexity

Solution Approach 1:

The patent changes the data organization parameter from traditional packed data registers to tile-based structures. This parameter change allows the specialized hardware to operate on tiles directly, reducing the instruction intensity required for packed data manipulation while maintaining the high peak compute and energy efficiency benefits of specialized hardware.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If deep learning algorithms operate on low precision arithmetic, then throughput is maximized, but accuracy may be compromised without sufficient output bits

Engineering Contradiction:
Improvethroughput of deep learning algorithmsVSAvoidaccuracy of computation
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements a nested structure where low-precision arithmetic operations are performed on individual tile elements, and the results are accumulated and combined at higher levels with increased precision. This nested approach allows throughput maximization at the operational level while maintaining accuracy through hierarchical accumulation, effectively nesting low-precision operations within a high-precision framework.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS11294671B2Systems and methods for performing duplicate detection instructions on 2D data
Publication Date: 2022.04.05 INTEL CORP
  • US11294671B2 patent drawing
  • US11294671B2 patent drawing
  • US11294671B2 patent drawing

AI summary

Disclosed embodiments relate to systems and methods for performing duplicate detection instructions on two-dimensional (2D) data. In one example, a processor includes fetch circuitry to fetch an instruction, decode circuitry to decode the fetched instruction having fields to specify an opcode and locations of a source matrix comprising M×N elements and a destination, the opcode to indicate execution circuitry is to use a plurality of comparators to discover duplicates in the source matrix, and store indications of locations of discovered duplicates in the destination. The execution circuitry to execute the decoded instruction as per the opcode.