Multi-Processor Core Grouping for Floating Point Operations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multi-processor systems, the frequency of access to external memory increases due to the need for redundant instruction sets for large-sized vector and matrix operations, limiting processing performance and efficiency.

Innovation Solution

A method where single processor cores are grouped to share floating point units (FPUs), with one core acting as a master to load instructions and control other cores, reducing the need for redundant instruction sets and minimizing external memory access.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If vector and matrix operations are decomposed into small units and allocated to multiple single processor cores, then the operation scale can exceed the computational limit of a single core, but the frequency of access to external memory increases due to redundant instruction sets

Engineering Contradiction:
Improveoperation scaleVSAvoidfrequency of access to external memory
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

Multiple single processor cores are merged into a group that shares a common instruction set. The master core loads the instruction set once, and all cores in the group execute the same instruction set simultaneously, eliminating redundant instruction loading and reducing external memory access frequency while maintaining large-scale vector and matrix operations

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The instruction set is made universal across multiple processor cores. A single instruction set loaded by the master core can be executed by all cores in the group, allowing the same instruction to perform operations on different data partitions across multiple cores, thereby reducing the need for duplicate instruction sets

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If instruction sets are added for vector and matrix operations decomposed into small units, then the operations can be performed on multiple cores, but the overall performance of the system deteriorates

Engineering Contradiction:
Improveoperation capabilityVSAvoidoverall performance
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The instruction set loading function is merged into a single master core that serves the entire group. Instead of each core independently loading instruction sets, the master core loads once and distributes to all group members, eliminating redundant loading operations and improving overall system performance while maintaining full operational capability

Inventive Principle:
Principle #5Merging (Combining)

3Quantity of substance

If the operation scale of one vector and matrix exceeds the computational limit of a single processor core, then decomposition into small units is necessary, but additional loading of instruction sets occurs

Engineering Contradiction:
Improveoperation scaleVSAvoidinstruction set loading
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The instruction set is designed to be universal and reusable across multiple cores. A single instruction set can direct operations on different data partitions across multiple cores, allowing large-scale operations to be performed without requiring separate instruction sets for each core, thereby reducing the complexity of instruction set management and loading

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11893392B2Multi-processor system and method for processing floating point operation thereof
Publication Date: 2024.02.06 ELECTRONICS & TELECOMM RES INST
  • US11893392B2 patent drawing
  • US11893392B2 patent drawing
  • US11893392B2 patent drawing

AI summary

A method for processing floating point operations in a multi-processor system including a plurality of single processor cores is provided. In this method, upon receiving a group setting for performing an operation, the plurality of single processor cores are grouped into at least one group according to the group setting, and a single processor core set as a master in the group loads an instruction for performing the operation from an external memory, and performs parallel operations by utilizing floating point units (FUPs) of all single processor cores in the group according to the instructions.