Shuffler Circuit for SIMD Lane Shuffle

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In single instruction multiple data (SIMD) architectures, such as graphics processing units (GPUs), the need for full all-lane to all-lane crossbars for data shuffling is expensive in terms of chip area and power consumption, and this cost increases quadratically with the number of processing lanes, making efficient data exchange between lanes a challenge.

Innovation Solution

A shuffler circuit that receives data from a subset of processing lanes, reorders it, and outputs the reordered data to all lanes, allowing only a subset of lanes to store the data, thereby reducing the need for physical connections between each lane and minimizing power consumption and chip area.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If full all-lane to all-lane crossbars are used for data shuffling, then data exchange capability between lanes is improved, but chip area and power consumption increase quadratically

Engineering Contradiction:
Improvedata exchange capabilityVSAvoidchip area
Core Design Contradiction:
Adaptability or versatilityVSArea of stationary object

Solution Approach 1:

The crossbar network is segmented into multiple stages, with each stage handling a subset of lane connections. Instead of one large N×N crossbar, the patent uses multiple smaller crossbars arranged in stages, where each crossbar handles only a portion of the total lane connections. This segmentation reduces the area of individual crossbars while maintaining overall data exchange capability through multi-stage routing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a temporal dimension to the data shuffling process by using multiple stages executed over multiple clock cycles. Data that would require direct lane-to-lane connections in a single cycle is instead routed through intermediate stages over multiple cycles, transforming a spatial problem into a temporal solution that reduces chip area.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If full all-lane to all-lane crossbars are used for data shuffling, then data exchange capability between lanes is improved, but power consumption increases quadratically

Engineering Contradiction:
Improvedata exchange capabilityVSAvoidpower consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by stationary object

Solution Approach 1:

The crossbar network is segmented into multiple stages, with each stage handling a subset of lane connections. Instead of one large N×N crossbar, the patent uses multiple smaller crossbars arranged in stages, where each crossbar handles only a portion of the total lane connections. This segmentation reduces the area of individual crossbars while maintaining overall data exchange capability through multi-stage routing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a temporal dimension to the data shuffling process by using multiple stages executed over multiple clock cycles. Data that would require direct lane-to-lane connections in a single cycle is instead routed through intermediate stages over multiple cycles, transforming a spatial problem into a temporal solution that reduces chip area.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Area of stationary object

If piecewise shuffling is used to reduce chip area and power, then area and power consumption are reduced, but data shuffle complexity increases

Engineering Contradiction:
Improvechip areaVSAvoidshuffle complexity
Core Design Contradiction:
Area of stationary objectVSDevice complexity

Solution Approach 1:

The patent introduces intermediate crossbars as mediator components between source and destination lanes. These intermediate crossbars act as buffering and routing points that simplify the overall shuffle logic by breaking down complex direct routing into simpler staged routing operations. The intermediaries manage the complexity of piecewise shuffling through controlled data flow between stages.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The piecewise shuffling is organized into periodic stages that execute in a regular pattern over multiple clock cycles. Each stage processes a specific subset of data with simple routing logic, and the stages repeat in a predetermined sequence. This periodic structure reduces complexity by making the shuffle operation predictable and systematic rather than requiring complex conditional routing logic.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentEP3485385B1Shuffler circuit for lane shuffle in SIMD architecture
Publication Date: 2020.04.22 QUALCOMM INC
  • EP3485385B1 patent drawingFigure 1
  • EP3485385B1 patent drawingFigure 2
  • EP3485385B1 patent drawingFigure 3

AI summary

Techniques are described to perform a shuffle operation. Rather than using an all-lane to all-lane cross bar, a shuffler circuit having a smaller cross bar is described. The shuffler circuit performs the shuffle operation piecewise by reordering data received from processing lanes and outputting the reordered data.