Vector Register Bank Rearrangement Path

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current vector processing systems face significant performance bottlenecks and increased complexity when performing rearrangement operations, such as transposes, due to the need for extensive internal storage and multiple clock cycles, especially when handling matrices, which is exacerbated by the requirement for both horizontal and vertical register access.

Innovation Solution

A data processing apparatus with a modified vector register bank featuring a write interface that includes a data rearrangement path, allowing for simultaneous rearrangement of data elements within the register bank upon receiving a rearrangement enable signal, thereby reducing the need for the vector processing unit to perform these operations and minimizing complexity and cost.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If traditional vector register bank is used for rearrangement operations, then system complexity is reduced, but processing time increases significantly

Engineering Contradiction:
Improveprocessing timeVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The write interface is pre-configured with a data rearrangement path that can immediately perform rearrangement operations when triggered. The rearrangement logic is prepared in advance within the write interface circuitry, allowing fast execution without requiring the vector processing unit to perform complex rearrangement sequences.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The write interface acts as an intermediary between the vector processing unit and the vector register bank. It receives data from the processing unit, performs rearrangement operations through its dedicated path, and writes to the register bank, thereby mediating the complexity and enabling fast rearrangement without burdening the main processing unit.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If multiple clock cycles are used for rearrangement operations, then processing accuracy is maintained, but productivity decreases

Engineering Contradiction:
Improveprocessing speedVSAvoidclock cycles
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The data rearrangement path operates continuously and immediately when triggered by a rearrangement enable signal. The rearrangement operations are performed in a single clock cycle through the dedicated write interface path, eliminating the need for multiple sequential clock cycles and maintaining continuous productive action.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The rearrangement operation is rushed through in a single clock cycle by utilizing the dedicated data rearrangement path in the write interface. This skipping of the multi-cycle process achieves the rearrangement result immediately, significantly improving productivity while maintaining accuracy through the deterministic single-cycle execution.

Inventive Principle:
Principle #21Skipping (Rushing through)

3Reliability

If extensive internal storage is provided within vector processing unit, then rearrangement operation reliability is improved, but device complexity increases

Engineering Contradiction:
Improveoperation reliabilityVSAvoidinternal storage requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The rearrangement operation functionality is extracted from the vector processing unit and relocated to the write interface circuitry. This extraction reduces the internal storage and complexity requirements within the processing unit while maintaining reliable rearrangement operation through the dedicated path in the write interface.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The write interface performs self-service by incorporating the rearrangement logic directly within its own circuitry. The write interface uses its own internal resources to perform rearrangement operations without requiring extensive external storage or complex support structures, thereby maintaining reliability while minimizing overall device complexity.

Inventive Principle:
Principle #25Self-service

4Productivity

If single write port is used in vector register bank, then device complexity is minimized, but productivity is reduced due to sequential write operations

Engineering Contradiction:
Improvewrite operation speedVSAvoidwrite interface complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The write interface is segmented into multiple functional paths: a first input for normal data writes and a second input coupled via the data rearrangement path for rearranged data. This segmentation allows simultaneous or parallel write operations through different paths, improving productivity without requiring a completely complex multi-port register bank structure.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8375196B2Vector processor with vector register file configured as matrix of data cells each selecting input from generated vector data or data from other cell via predetermined rearrangement path
Publication Date: 2013.02.12 ARM LTD
  • US8375196B2 patent drawing
  • US8375196B2 patent drawing
  • US8375196B2 patent drawing

AI summary

A data processing apparatus includes a vector register bank having a plurality of vector registers, each register including a plurality of storage cells, each cell storing a data element. A vector processing unit is provided for executing a sequence of vector instructions. The processing unit is arranged to issue a set rearrangement enable signal to the vector register bank. The write interface of the vector register bank is modified to provide not only a first input for receiving the data elements generated by the vector processing unit during normal execution, but also has a second input coupled via a data rearrangement path to the matrix of storage cells via which the data elements currently stored in the matrix of storage cells are provided to the write interface in a rearranged form representing the arrangement of data elements that would be obtained by performance of the predetermined rearrangement operation.