XarMa Processor 1-to-K+1 Adjacency Network for Low Power Data Sharing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-processor architectures face challenges in reducing power consumption while maintaining high performance and scalability, particularly in accessing data from memory and sharing data between processors, due to limitations in processor architecture and memory organization.

Innovation Solution

The proposed solution involves a processing architecture with a network design based on 1 to K+1 adjacency, where execution units are interconnected in a specific matrix configuration, allowing for efficient data forwarding and reduced power usage through optimized data paths and multiplexing elements, enabling scalable and low-power data communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If traditional multi-processor architectures use large central multi-ported register files and standard interconnection networks, then data sharing between processors can be achieved, but power consumption increases and scalability is limited

Engineering Contradiction:
Improvepower consumptionVSAvoiddata sharing efficiency
Core Design Contradiction:
Use of energy by moving objectVSProductivity

Solution Approach 1:

The patent divides the traditional centralized register file into distributed local register files at each execution unit node. This segmentation eliminates the need for a large central multi-ported register file, reducing power consumption while maintaining data sharing capabilities through the interconnection network.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a two-dimensional mesh interconnection network topology with direct east-west and north-south connectivity, adding spatial dimensionality to data access. This dimensional change enables efficient data sharing between processors without relying on centralized structures, thereby reducing power consumption while maintaining productivity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If standard interconnection networks are used for processor communication, then basic data transfer is possible, but memory bandwidth is insufficient for high-performance operations

Engineering Contradiction:
Improvememory bandwidthVSAvoiddata access speed
Core Design Contradiction:
ProductivityVSSpeed

Solution Approach 1:

The interconnection network nodes are designed with multi-functionality, serving both as routing elements and as memory access points. Each node can handle data forwarding, memory access, and processor communication simultaneously, increasing memory bandwidth without sacrificing data access speed.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces buffer memory elements at each interconnection network node as intermediaries between processors and the memory system. These buffers enable high-speed data transfer by decoupling memory access from processor execution, thereby increasing memory bandwidth while maintaining fast data access.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If processors are closely coupled for efficient communication, then data sharing is improved, but the architecture becomes less scalable

Engineering Contradiction:
ImprovescalabilityVSAvoidcommunication energy
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The interconnection network employs dynamic routing algorithms that adapt communication paths based on current network conditions and data location. This dynamic behavior enables the architecture to scale efficiently while maintaining low communication energy consumption by selecting optimal paths in real-time.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the communication parameter from centralized bus arbitration to distributed token-based or dimension-ordered routing. This parameter change enables the system to scale to more processors while reducing communication energy by eliminating centralized contention and enabling direct peer-to-peer data transfer.

Inventive Principle:
Principle #35Parameter changes

4Speed

If more temporary variable storage is provided for high-performance operations, then computational speed is improved, but power consumption and device complexity increase

Engineering Contradiction:
Improvecomputational speedVSAvoidstorage structure complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent extracts the temporary variable storage function from a centralized large-capacity register file and distributes it across multiple small local register files at each execution unit. This extraction reduces device complexity and power consumption while maintaining computational speed through efficient local access and networked sharing.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11249939B2Methods and apparatus for sharing nodes in a network with connections based on 1 to k+1 adjacency used in an execution array memory array (XarMa) processor
Publication Date: 2022.02.15 PECHANEK GERALD GEORGE
  • US11249939B2 patent drawing
  • US11249939B2 patent drawing
  • US11249939B2 patent drawing

AI summary

An Execution Array Memory Array (XarMa©) processor is described for signal processing and internet of things (IoT) applications, (pronounced sharma, that means happiness in Sanskrit). The XarMa© processor uses a 1 to K+1 adjacency network in an array of execution units. The 1 to K+1 adjacency refers to connections separately made in rows and in columns of execution unit and local file nodes, where the number of Rows≥K>1 and of Columns≥K>1 and K is an odd integer. Instead of a large central multi-ported register file, a distributed set of storage files local to each execution unit is used. The instruction set architecture uses instructions that specify forwarding of execution results to execution units associated with destination instructions. This execution array is scalable to support cost effective and low power high-performance application specific processing focused on target product requirements.