Configurable Processor Pipeline for Multi-Thread Speed and Flexibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Configurable integrated circuits face a trade-off between speed and flexibility, with long interconnect paths limiting clock speed due to data transfer delays, and existing solutions require modifications to incorporate latches or pipelines, which can be cumbersome for users.

Innovation Solution

A configurable processing circuit that handles multiple threads simultaneously, featuring a thread data store, configurable execution units, a routing network, and a pipeline with sections configured per clock cycle, allowing each thread to propagate through the circuit with its associated configuration instance, and enabling dynamic reconfiguration of the routing network and execution units.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the interconnect paths are made longer to connect more function units for maximum flexibility, then the flexibility is improved, but the clock speed deteriorates due to increased data transfer delays

Engineering Contradiction:
ImproveflexibilityVSAvoidclock speed
Core Design Contradiction:
Adaptability or versatilityVSSpeed

Solution Approach 1:

The circuit is divided into multiple pipeline sections, with each section containing function units and interconnect circuits that can be independently configured. This segmentation allows data to be processed in stages across shorter interconnect paths, maintaining high clock speed while supporting flexible function unit connections through the pipelined architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The interconnect circuits are made dynamically reconfigurable through pipeline configuration, allowing the connection topology to change based on the computational task. This dynamic reconfiguration enables the system to adapt interconnect paths for different function unit combinations without being constrained by fixed physical layouts, resolving the trade-off between flexibility and speed.

Inventive Principle:
Principle #15Dynamics

2Speed

If latches are added to pipeline sections to enable pipelining, then the clock speed is improved, but the device complexity increases and user design modification is required

Engineering Contradiction:
Improveclock speedVSAvoidcircuit complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The system automatically configures pipeline sections and interconnect circuits based on thread configuration instances without requiring user intervention. The configurable logic dynamically sets up the appropriate pipeline structure and latching behavior for each computational task, eliminating the need for users to manually modify designs to incorporate latches while still achieving high clock speeds through pipelining.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10275390B2Pipelined configurable processor
Publication Date: 2019.04.30 SILICON TAILOR
  • US10275390B2 patent drawing
  • US10275390B2 patent drawing
  • US10275390B2 patent drawing

AI summary

A configurable processing circuit capable of handling multiple threads simultaneously, the circuit comprising a thread data store, a plurality of configurable execution units, a configurable routing network for connecting locations in the thread data store to the execution units, a configuration data store for storing configuration instances that each define a configuration of the routing network and a configuration of one or more of the plurality of execution units, and a pipeline formed from the execution units, the routing network and the thread data store that comprises a plurality of pipeline sections configured such that each thread propagates from one pipeline section to the next at each clock cycle, the circuit being configured to: (i) associate each thread with a configuration instance; and (ii) configure each of the plurality of pipeline sections for each clock cycle to be in accordance with the configuration instance associated with the respective thread that will propagate through that pipeline section during the clock cycle.