VLSI Circuit Torus Network for Parallel Computing Scalability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional parallel computing systems face challenges in scalability and high-speed data transfer between processing elements (PEs), particularly in small-scale mobile terminal devices, where data transfer between all PEs is necessary, leading to inefficiencies in processing speed and increased costs for dedicated communication networks.
Innovation Solution
A miniaturized HXNet VLSI circuit with additional buffer memories allows for data transfer between desired PEs by implementing m2 PEs and m3 buffer memories, enabling the combination of multiple VLSI circuits to form a parallel computing system with reduced bus usage and scalability, facilitating data transfer between PEs across multiple circuits.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If a conventional parallel computing system uses a transfer bus for data transfer between PEs, then data transfer is enabled, but the system cannot achieve high-speed processing when large amounts of data need to be transferred between all PEs
Solution Approach 1:
The patent transitions from a conventional two-dimensional PE arrangement to a three-dimensional torus network topology. This dimensional expansion allows data to be transferred through multiple paths (including wrap-around connections), significantly increasing bandwidth and enabling high-speed data transfer between all PEs in the parallel computing system.
Solution Approach 2:
The patent divides the parallel computing system into multiple modular PE units that can be independently configured and connected. Each PE is a self-contained module with local memory and processing units, allowing the system to be scaled by adding or removing modular segments without requiring complete system redesign.
2Speed
If a parallel computing system uses a communication network for dedicated data transfer between PEs, then high-speed data transfer is achieved, but construction cost increases significantly
Solution Approach 1:
The patent designs the PE modules and interconnection network to serve multiple functions simultaneously. The same communication infrastructure supports both data transfer and control signaling, and the modular PEs can perform different computational tasks. This multi-functionality reduces the need for dedicated specialized components, thereby lowering construction costs while maintaining high data transfer speeds.
3Adaptability or versatility
If the number of PEs is limited to m2 in HXNet, then the system can be implemented, but scalability to form larger systems is impossible
Solution Approach 1:
The patent implements a hierarchical modular structure where smaller HXNet blocks (m×m PEs) can be nested within larger torus networks. Multiple m×m blocks are interconnected to form progressively larger systems, enabling scalability from small to large configurations. Each nested level maintains the same basic topology principles, allowing systematic expansion without increasing structural complexity.
4Productivity
If each PE performs processing in a small plane for radiosity processing, then processing is simplified, but data transfer between all PEs is required, reducing processing speed
Solution Approach 1:
The patent uses the three-dimensional torus network topology to provide multiple data transfer paths between PEs performing radiosity processing. This dimensional advantage allows simultaneous data exchange along different routes, increasing overall bandwidth and maintaining high processing speeds even when all PEs need to communicate with each other for small-plane radiosity calculations.
Data Source
AI summary
Provided is a parallel computing system that has scalability and is capable of performing data transfer between desired PEs. Also provided is a computer system that utilizes the parallel computing system described above, and enables radiosity processing on small-scale mobile terminal devices. An HXNet is implemented in a VLSI, and data transfer between VLSIs is possible using additional BMs. Scalability is realized that enables selection of any number of VLSIs, and radiosity processing is enabled on small-scale mobile terminal devices.


