Unified Vector Processor Circuit for Area-Constrained Reduction Operations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vector processors face issues with circuit area bloating and power dissipation when performing vector reduction and element reduction in a fully pipelined manner, especially with larger vector register lengths, leading to congestion and timing problems.

Innovation Solution

A vector processor design that performs both vector and element reduction using the same circuit structure, allowing flexible adjustment of iterations to optimize hardware and software performance indicators, by loading operands based on state parameters and performing reduction operations within the same circuit.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If vector reduction and element reduction are implemented in a fully pipelined manner with separate circuit structures, then the reduction operation capability is improved, but circuit area bloating occurs and power dissipation increases

Engineering Contradiction:
Improvereduction operation capabilityVSAvoidcircuit area
Core Design Contradiction:
ProductivityVSArea of stationary object

Solution Approach 1:

The patent implements a unified reduction circuit that can perform both vector reduction and element reduction operations through a single structure. The circuit uses state parameters to control its behavior, allowing it to switch between different reduction modes (vector reduction, element reduction) without requiring separate dedicated circuits for each operation type, thereby reducing overall circuit area while maintaining full functionality

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The reduction circuit is designed with dynamic reconfigurability through state parameters that control the operation mode. The circuit can dynamically adjust its behavior based on the current state parameter, enabling it to perform different reduction operations sequentially rather than requiring static parallel structures for each operation type

Inventive Principle:
Principle #15Dynamics

2Productivity

If vector reduction and element reduction are implemented in a fully pipelined manner with separate circuit structures, then the reduction operation capability is improved, but power dissipation increases

Engineering Contradiction:
Improvereduction operation capabilityVSAvoidpower dissipation
Core Design Contradiction:
ProductivityVSUse of energy by stationary object

Solution Approach 1:

By using a single unified reduction circuit to handle both vector reduction and element reduction operations, the patent eliminates the need for multiple separate pipelined circuits. This consolidation reduces the total number of active components and interconnections, thereby lowering overall power dissipation while maintaining the ability to perform both reduction types

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges the vector reduction circuit and element reduction circuit into a single unified reduction circuit. This combination reduces redundant circuitry and shared resources, leading to decreased power consumption compared to having separate fully pipelined circuits for each operation type

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If the vector processor is configured for larger vector register length (VLEN) or data path length (DLEN), then the processing capacity is improved, but circuit area bloating and congestion problems are exacerbated

Engineering Contradiction:
Improveprocessing capacityVSAvoidcircuit area
Core Design Contradiction:
ProductivityVSArea of stationary object

Solution Approach 1:

The unified reduction circuit processes data in segments or stages controlled by state parameters. For larger VLEN or DLEN configurations, the circuit can process the data path in manageable segments through multiple operation cycles, avoiding the need for a single large parallel structure that would consume excessive circuit area

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The circuit uses dynamic control through state parameters to handle variable data path lengths. Rather than being statically configured for the maximum possible width, the circuit can adapt its operation to the actual data length required, reducing circuit area by only activating necessary processing elements for each specific operation

Inventive Principle:
Principle #15Dynamics

4Productivity

If the vector processor is configured for larger vector register length (VLEN) or data path length (DLEN), then the processing capacity is improved, but timing problems are exacerbated

Engineering Contradiction:
Improveprocessing capacityVSAvoidtiming problems
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The unified reduction circuit performs operations in periodic cycles controlled by state parameters. For larger data paths, rather than attempting to complete all reductions in a single cycle (which would cause timing issues), the circuit executes multiple shorter operation cycles, each handling a portion of the reduction work, thereby maintaining timing closure while achieving full processing capacity

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS20240248713A1Vector processor performing vector and element reduction method with same circuit structure
Publication Date: 2024.07.25 ANDES TECH
  • US20240248713A1 patent drawing
  • US20240248713A1 patent drawing
  • US20240248713A1 patent drawing

AI summary

A vector processor performing a vector reduction method and an element reduction method with the same circuit structure is provided. The vector processor includes a vector register file and a first lane. The first lane loads a first operand and a second operand based on a first state parameter and performs a first reduction operation on the first operand and the second operand to generate a first reduction result. The first lane performs a second reduction operation on the first and second parts of the first reduction result based on a second state parameter to generate a second reduction result.