Vector Processor Array Using Identical Circuit Slices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current vector processor architectures face challenges in efficiently executing complex data operations due to high interconnect costs and latency, as well as difficulties in splitting applications across multiple processors, leading to performance degradation in high-data-rate communications and digital signal processing tasks.

Innovation Solution

A vector processor architecture is implemented using an array of identical circuit blocks assembled by abutment, allowing for a large vector processor core to be segmented into multiple identical slices, each with shared functionality and a unique identifier for conditional logic circuit enablement, reducing interconnect complexity and enhancing layout efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a large vector processor core is implemented using traditional interconnect networks, then processing capability is improved, but interconnect cost and latency increase

Engineering Contradiction:
Improveprocessing capabilityVSAvoidinterconnect cost and latency
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The vector processor core is divided into multiple identical slices, each capable of independent operation. This segmentation eliminates the need for complex interconnect networks between processors, as each slice can process data locally. The slices are arranged in a regular grid pattern with simple nearest-neighbor connections, dramatically reducing interconnect complexity and latency while maintaining high processing capability through parallel operation of multiple slices.

Inventive Principle:
Principle #1Segmentation

2Productivity

If applications are split across multiple processors, then processing power is improved, but performance degrades due to interconnect latency

Engineering Contradiction:
Improveprocessing powerVSAvoidinterconnect latency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

Multiple processor slices are merged into a single integrated core structure where data can flow between slices through simple, fast inter-slice connections. This merging allows applications to be distributed across slices while maintaining low-latency communication, effectively combining the processing power of multiple units without the performance penalty of traditional multi-processor interconnects.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If irregular functionality is implemented in vector processor, then versatility is improved, but layout complexity increases

Engineering Contradiction:
ImprovefunctionalityVSAvoidlayout complexity
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

Each slice contains identical basic functionality, but local quality is achieved through selective enabling of specific logic circuits within each slice based on configuration values. This allows irregular and diverse functionality to be implemented across the array of slices without requiring complex custom layouts for each functional unit. The regular slice structure maintains layout simplicity while the configurable logic circuits provide versatility.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Each slice is designed as a universal building block that can perform multiple functions through configuration. The slices contain logic circuits that can be selectively enabled or disabled based on configuration values, allowing the same hardware structure to implement different functionalities. This universality provides versatility while maintaining the simplicity of a regular, repeating layout pattern.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3757812A1Apparatuses, methods, and systems for vector processor architecture having an array of identical circuit blocks
Publication Date: 2020.12.30 INTEL CORP
  • EP3757812A1 patent drawingFigure 1~2
  • EP3757812A1 patent drawingFigure 3
  • EP3757812A1 patent drawingFigure 4

AI summary

Systems, methods, and apparatuses relating to vector processor architecture having an array of identical circuit blocks are described. In one embodiment, a processor includes a single centralized circuit comprising an instruction decoder and a controller; and a plurality of circuit slices that each comprise an arithmetic logic unit, a multiplier, a register file, a local memory, and a same plurality of logic circuits and a packed data datapath in between, wherein each circuit slice includes a physical port that provides a unique identification value that identifies a circuit slice from the other circuit slices, and the controller is to broadcast a same configuration value to the plurality of circuit slices to cause a first circuit slice to enable a first logic circuit and enable a second logic circuit of the first circuit slice based on its unique identification value and the configuration value, and cause a second circuit slice to enable a same, first logic circuit and disable a same, second logic circuit of the second circuit slice based on its unique identification value and the configuration value.