Digital VLSI Subscalar Arithmetic for Data-Dependent Parallelism

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing VLSI systems face challenges in maximizing processing throughput due to data-flow dependencies, which hinder the effective exploitation of parallelism, and often compromise on data width to enhance performance, leading to inefficient resource utilization.

Innovation Solution

The method involves splitting atomic data and operations into sub-atomic fragments and operations, allowing for time-multiplexed and pipelined execution to exploit latent parallelism, thereby reducing complexity and improving throughput.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If parallelism is exploited in conventional VLSI systems, then processing throughput is enhanced, but data-flow dependencies adversely impact the effectiveness of parallelism exploitation

Engineering Contradiction:
Improveprocessing throughputVSAvoiddata-flow dependencies
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments atomic data into sub-atomic data fragments and atomic operations into sub-atomic operations. This segmentation allows independent processing of data fragments through subscalar operations, enabling parallelism exploitation despite data-flow dependencies in conventional architectures. By dividing the computational work into finer granularities, the system can process multiple fragments simultaneously or in overlapped manner, improving throughput while managing complexity.

Inventive Principle:
Principle #1Segmentation

2Speed

If data width is reduced to enhance performance, then processing speed increases, but precision and resource utilization deteriorate

Engineering Contradiction:
Improveprocessing speedVSAvoiddata precision
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent employs dynamic subscalar operations where the operational granularity can be adjusted based on the specific computational requirements. By dynamically selecting appropriate sub-atomic operation types (such as partial adders, shifters, multiplexers) and configuring the number of data fragments, the system can optimize between speed and precision for different workloads. This dynamic approach allows maintaining full data width precision while achieving high-speed processing through efficient subscalar computation pipelines.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250378247A1System and method for implementation of computational logic using digital VLSI systems
Publication Date: 2025.12.11 PANDEY UMA
  • US20250378247A1 patent drawing
  • US20250378247A1 patent drawing
  • US20250378247A1 patent drawing

AI summary

Subscalar digital arithmetic computing paradigm is disclosed. The atomic data and atomic operations thereon are broken down into sub-atomic data fragments and sub-atomic partial operations. Such a break-up exposes hitherto unexploited levels of parallelism by way of allowing overlap of operations even if data-dependent. It is found that this improved exploitation of latent parallelism to enhance processing throughputs comes with a favourable impact on the area-power characteristics of corresponding computing structures. The present invention may be implemented through synthesized circuits and may result in an enhanced improvement in their area-throughput figure-of-merit (FOM).