Remote Accelerator Interface for Multi-Pass Packet Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computing systems using hardware accelerators, such as FPGAs and GPUs, are limited in their ability to perform multi-pass packet processing, which restricts their functionality to single-pass operations in receive or transmit directions, and lack efficient integration with central processing units for enhanced computing tasks.

Innovation Solution

A remote accelerator interface system is implemented, allowing compute devices to dynamically allocate and utilize remote accelerator resources, such as FPGAs and GPUs, for pre-processing, post-processing, and multi-pass operations through an optical fabric network, enabling efficient distribution and utilization of computing resources across a data center.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If hardware accelerators are used for packet processing, then processing speed is improved, but the ability to perform multi-pass processing is limited

Engineering Contradiction:
Improvepacket processing speedVSAvoidmulti-pass processing capability
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The system segments packet processing into multiple passes, where each pass can be handled by different accelerator instances or the same accelerator processing different packet segments. This allows single-pass accelerators to achieve multi-pass processing functionality through systematic division of processing tasks across multiple stages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A control plane and data plane architecture is introduced as an intermediary layer between the CPU and hardware accelerators. The control plane manages packet routing, pass assignment, and coordinator selection, enabling complex multi-pass processing workflows while maintaining high-speed processing through dedicated data plane accelerators.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If accelerators are locally accessible via PCIe, then processing efficiency is improved, but resource sharing and pooling across data center is limited

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidresource pooling capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system creates universal accelerator resources that can be dynamically allocated to multiple compute devices and workloads. Accelerators are abstracted into poolable resources that can serve different functions and users, enabling both high local processing efficiency and broad resource sharing across the data center through a unified management framework.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The architecture transitions from traditional single-dimension local accelerator attachment to a multi-dimensional resource pool model. Accelerators are organized in pools that can be accessed by multiple compute devices simultaneously, adding the dimension of resource sharing and virtualization while maintaining low-latency local access capabilities through smart routing and caching.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Device complexity

If single-pass packet processing is implemented, then device complexity is reduced, but processing capabilities are limited

Engineering Contradiction:
Improveprocessing architecture complexityVSAvoidprocessing capability
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

Complex multi-pass processing is segmented into multiple single-pass stages, each handled by simpler accelerator instances. This segmentation allows each device to remain relatively simple while the system as a whole achieves sophisticated multi-pass processing capabilities through coordinated operation of multiple stages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Multiple single-pass processing stages are merged into a coordinated multi-pass processing workflow. The control plane merges resource management, packet routing, and performance optimization functions to enable complex processing capabilities while keeping individual accelerator devices relatively simple.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12160369B2Processor related communications
Publication Date: 2024.12.03 INTEL CORP
  • US12160369B2 patent drawing
  • US12160369B2 patent drawing
  • US12160369B2 patent drawing

AI summary

A compute device can access local or remote accelerator devices for use in processing a received packet. The received packet can be processed by any combination of local accelerator devices and remote accelerator devices. In some cases, the received packet can be encapsulated in an encapsulating packet and sent to a remote accelerator device for processing. The encapsulating packet can indicate a priority level for processing the received packet and its associated processing task. The priority level can override a priority level that would otherwise be assigned to the received packet and its associated processing task. The remote accelerator device can specify a fullness of an input queue to the compute device. Other information can be conveyed by packets transmitted between and among compute devices and remote accelerator devices to assist in determining an accelerator to use or other uses.