Heterogeneous Multiprocessor Core Ratio Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing processing technologies face inefficiencies in handling tasks with variable data element utilization, as SIMD processors are less efficient when data element utilization is low, while scalar cores are more energy-efficient in such scenarios, leading to suboptimal performance and energy consumption in heterogeneous workloads.

Innovation Solution

A heterogeneous multicore processor architecture that optimizes the combination of SIMD and scalar cores based on specific workloads, dynamically partitioning tasks between SIMD and scalar cores to maximize performance and minimize energy consumption, such as in computer vision operations like cascade classifiers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If SIMD processors are used to handle tasks with variable data element utilization, then parallel processing capability is improved, but energy efficiency deteriorates when data element utilization is low

Engineering Contradiction:
Improveparallel processing capabilityVSAvoidenergy efficiency
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system dynamically configures the ratio of scalar to SIMD processor cores based on workload characteristics and data element utilization requirements. This allows the processor architecture to adapt its composition in real-time, switching between scalar and SIMD execution modes to optimize both parallel processing capability and energy efficiency for different computational tasks

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of processor core composition by configuring heterogeneous multiprocessors with specific ratios of scalar to SIMD cores (e.g., 4:1, 5:1, 6:1). This parameter optimization allows the system to achieve minimum product of execution time and die area for specific workloads like cascade classifiers, balancing parallel processing power with energy consumption

Inventive Principle:
Principle #35Parameter changes

2Productivity

If heterogeneous multiprocessors with optimized scalar to SIMD core ratios are configured, then performance per unit area is improved, but device complexity increases

Engineering Contradiction:
Improveperformance per unit areaVSAvoidprocessor architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The heterogeneous multiprocessor architecture provides multi-functionality by incorporating both scalar and SIMD processor cores that can handle different types of computational tasks. This universal design allows a single processor system to efficiently execute both data-parallel workloads (using SIMD cores) and sequential workloads (using scalar cores), improving overall performance per unit area despite increased architectural complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10891255B2Heterogeneous multiprocessor including scalar and SIMD processors in a ratio defined by execution time and consumed die area
Publication Date: 2021.01.12 INTEL CORP
  • US10891255B2 patent drawing
  • US10891255B2 patent drawing
  • US10891255B2 patent drawing

AI summary

In one embodiment, a heterogeneous multicore processor is described that is optimized to execute multi-stage computer vision algorithms such as cascade classifier workloads. In such embodiment the heterogeneous processor includes at least one SIMD core, such as a vector processor core, coupled with one or more scalar cores. In one embodiment the heterogeneous multiprocessor executes multi-stage compute operations, where the SIMD core computes a first set of stages and the one or more scalar cores compute the second set of stages. In one embodiment, a process for designing a heterogeneous multicore processor is disclosed which optimizes the ratio of scalar to SIMD cores based on execution time of the multi-stage compute operation in relation to processor die area consumed by a processor configuration having the ratio.