CXL Interconnect Workload Redistribution for FPGA Processing Delays

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing interconnect systems using Compute Express Link (CXL) switches primarily manage memory resources, leading to processing delays when applications request data processing exceeding assumed levels on devices like FPGAs or GPUs.

Innovation Solution

An information processing apparatus that includes a control unit (CU) capable of detecting processing delays in FPGAs or GPUs and dynamically redistributes workload by assigning assistant devices from a pool of idle or underutilized FPGAs or GPUs to process tasks in parallel, thereby reducing load and improving performance without involving server resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If a single device processes all tasks, then device complexity is low, but processing delay increases when workload exceeds assumed levels

Engineering Contradiction:
Improveprocessing delayVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system segments the processing workload by dividing it across multiple devices (FPGAs, GPUs, or other processors). When a processing delay is detected in one device, the control unit redistributes tasks to other available devices, effectively segmenting the processing function to reduce delays without requiring a completely complex system architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts task distribution based on real-time processing conditions. The control unit monitors processing delays and dynamically reallocates tasks from overloaded devices to underutilized devices, creating a dynamic load-balancing mechanism that reduces processing delays while maintaining manageable system complexity.

Inventive Principle:
Principle #15Dynamics

2Productivity

If multiple devices are used for processing, then productivity increases, but device complexity increases

Engineering Contradiction:
Improveprocessing throughputVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The control unit implements a universal management approach that can handle multiple types of processing devices (FPGAs, GPUs, general-purpose processors) through a single standardized interface and task distribution mechanism. This multi-functionality allows the system to leverage diverse devices for improved productivity while maintaining unified control that prevents exponential complexity growth.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The control unit acts as an intermediary between the task source and multiple processing devices. It receives tasks, determines optimal device allocation, and manages task distribution, thereby enabling multiple devices to work together for increased productivity while centralizing control logic to manage system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If workload is concentrated on one device, then device utilization is high for that device, but overall system resource utilization is low

Engineering Contradiction:
Improvesystem resource utilizationVSAvoidprocessing delay
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system implements self-service load balancing where the control unit automatically monitors processing delays and redistributes tasks without external intervention. When a device experiences processing delays, the system autonomously reallocates tasks to underutilized devices, improving overall resource utilization and reducing processing delays through self-regulating behavior.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The control unit employs feedback mechanisms by monitoring processing delays from each device and using this information to dynamically adjust task distribution. This closed-loop control ensures that devices with lower utilization receive additional tasks, thereby balancing load across the system, improving overall resource utilization, and reducing processing delays.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240212254A1Information processing apparatus, computer-readable recording medium storing program, and control method
Publication Date: 2024.06.27 FUJITSU LTD
  • US20240212254A1 patent drawing
  • US20240212254A1 patent drawing
  • US20240212254A1 patent drawing

AI summary

An information processing apparatus provided in a computer system including a first processor; an interconnect switch that conforms to an interconnect standard and a plurality of devices coupled to the processor via the interconnect switch includes a second processor, when a processing delay for a processing target is detected in a first device among the plurality of devices, configured to cause a second device among the plurality of devices that is different from the first device to process the processing target.