XarMa Processor 1-to-K+1 Adjacency Network for Low Power Data Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multi-processor architectures face challenges in reducing power consumption while maintaining high performance and scalability, particularly in accessing data from memory and sharing data between processors, due to limitations in processor architecture and memory organization.
Innovation Solution
The proposed solution involves a processing architecture with a network design based on 1 to K+1 adjacency, where execution units are interconnected in a specific matrix configuration, allowing for efficient data forwarding and reduced power usage through optimized data paths and multiplexing elements, enabling scalable and low-power data communication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If traditional multi-processor architectures use large central multi-ported register files and standard interconnection networks, then data sharing between processors can be achieved, but power consumption increases and scalability is limited
Solution Approach 1:
The patent divides the traditional centralized register file into distributed local register files at each execution unit node. This segmentation eliminates the need for a large central multi-ported register file, reducing power consumption while maintaining data sharing capabilities through the interconnection network.
Solution Approach 2:
The patent introduces a two-dimensional mesh interconnection network topology with direct east-west and north-south connectivity, adding spatial dimensionality to data access. This dimensional change enables efficient data sharing between processors without relying on centralized structures, thereby reducing power consumption while maintaining productivity.
2Productivity
If standard interconnection networks are used for processor communication, then basic data transfer is possible, but memory bandwidth is insufficient for high-performance operations
Solution Approach 1:
The interconnection network nodes are designed with multi-functionality, serving both as routing elements and as memory access points. Each node can handle data forwarding, memory access, and processor communication simultaneously, increasing memory bandwidth without sacrificing data access speed.
Solution Approach 2:
The patent introduces buffer memory elements at each interconnection network node as intermediaries between processors and the memory system. These buffers enable high-speed data transfer by decoupling memory access from processor execution, thereby increasing memory bandwidth while maintaining fast data access.
3Adaptability or versatility
If processors are closely coupled for efficient communication, then data sharing is improved, but the architecture becomes less scalable
Solution Approach 1:
The interconnection network employs dynamic routing algorithms that adapt communication paths based on current network conditions and data location. This dynamic behavior enables the architecture to scale efficiently while maintaining low communication energy consumption by selecting optimal paths in real-time.
Solution Approach 2:
The patent changes the communication parameter from centralized bus arbitration to distributed token-based or dimension-ordered routing. This parameter change enables the system to scale to more processors while reducing communication energy by eliminating centralized contention and enabling direct peer-to-peer data transfer.
4Speed
If more temporary variable storage is provided for high-performance operations, then computational speed is improved, but power consumption and device complexity increase
Solution Approach 1:
The patent extracts the temporary variable storage function from a centralized large-capacity register file and distributes it across multiple small local register files at each execution unit. This extraction reduces device complexity and power consumption while maintaining computational speed through efficient local access and networked sharing.
Data Source
AI summary
An Execution Array Memory Array (XarMa©) processor is described for signal processing and internet of things (IoT) applications, (pronounced sharma, that means happiness in Sanskrit). The XarMa© processor uses a 1 to K+1 adjacency network in an array of execution units. The 1 to K+1 adjacency refers to connections separately made in rows and in columns of execution unit and local file nodes, where the number of Rows≥K>1 and of Columns≥K>1 and K is an odd integer. Instead of a large central multi-ported register file, a distributed set of storage files local to each execution unit is used. The instruction set architecture uses instructions that specify forwarding of execution results to execution units associated with destination instructions. This execution array is scalable to support cost effective and low power high-performance application specific processing focused on target product requirements.


