Programmable Storage Accelerator Pipelining for Low-Latency Data Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computational storage devices face challenges in efficiently processing large amounts of data with high latency and resource consumption, particularly when hardware-based computation modules lack flexibility and require extensive development resources.

Innovation Solution

A multi-core storage accelerator with programmable storage processing units (SPUs) executes data processing functions via software, allowing for pipelined and concurrent processing, and utilizes ping-pong buffering for efficient data transfer between SPUs, reducing latency and resource usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If hardware-based computation modules are used, then data processing efficiency is improved, but device complexity and development difficulty increase

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidhardware development complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces hardware-based computation modules with software-programmable storage processing units (SPUs). Instead of using dedicated hardware circuits for data processing, the system uses general-purpose processors that execute software instructions stored in memory. This substitution maintains data processing functionality while significantly reducing hardware development complexity and enabling flexible modification of processing functions through software updates.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent implements universal storage processing units that can execute multiple types of data processing functions through different software instruction sets. Rather than requiring separate dedicated hardware modules for each processing task, a single SPU architecture can perform various functions by loading different software programs, thereby reducing overall device complexity while maintaining processing efficiency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Speed

If hardware-based computation modules are used, then data processing speed is improved, but adaptability and flexibility deteriorate

Engineering Contradiction:
Improvedata processing speedVSAvoidflexibility of processing functions
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent introduces dynamic reconfigurability through software-based processing units. The processing functions can be dynamically changed by loading different instruction sets into the SPUs, allowing the system to adapt to different data processing requirements without physical hardware reconfiguration. This maintains high processing speed while providing full adaptability through software updates.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent enables parameter changes in processing functionality through software configuration. By changing the software instruction sets loaded into the SPUs, the system can modify processing parameters, algorithms, and functions without altering the underlying hardware architecture. This provides both speed through optimized software algorithms and flexibility through parameter reconfiguration.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If software-programmable modules are used, then ease of modification is improved, but data processing efficiency deteriorates

Engineering Contradiction:
Improveease of modificationVSAvoiddata processing efficiency
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent segments the storage system into independent processing units (SPUs) that can execute software instructions. Each SPU is a separate processing entity that can be independently programmed and configured. This segmentation allows for easy modification of individual processing functions without affecting the entire system, while maintaining efficiency through parallel execution of multiple SPUs on different data chunks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses software instruction sets that can be copied and loaded into multiple SPUs. Instead of requiring unique hardware circuits for each processing function, the same software code can be replicated across multiple processing units, enabling both ease of modification through software updates and high productivity through parallel processing of multiple data sets.

Inventive Principle:
Principle #26Copying

4Device complexity

If single-core processing is used, then device complexity is reduced, but productivity and response latency deteriorate

Engineering Contradiction:
Improveprocessing unit complexityVSAvoiddata processing throughput
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent divides the data processing task into multiple segments that can be handled by separate processing units. The data is partitioned into different data chunks that are distributed to multiple SPUs for simultaneous processing. This segmentation enables parallel execution, significantly improving throughput and reducing response latency while keeping each individual SPU relatively simple in architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from single-core sequential processing to multi-core parallel processing, adding the dimension of temporal parallelism. By executing multiple processing operations simultaneously on different data chunks across multiple SPUs, the system achieves higher productivity without proportionally increasing the complexity of individual processing units.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12619376B2Systems and methods for executing data processing functions
Publication Date: 2026.05.05 SAMSUNG ELECTRONICS CO LTD
  • US12619376B2 patent drawing
  • US12619376B2 patent drawing
  • US12619376B2 patent drawing

AI summary

Systems and methods for executing a data processing function are disclosed. A first processing device of a storage accelerator loads a first instruction set associated with a first application of a host computing device. A second processing device of the storage accelerator loads a second instruction set associated with the first application. A command is received from the host computing device. The command may be associated with data associated with the first application. The first processing device identifies at least a first criterion or a second criterion associated with the data. The first processing device processes the data according to the first instruction set in response to identifying the first criterion. The first processing device writes the data to a buffer of the second processing device in response to identifying the second criterion. The second processing device processes the data in the buffer according to the second instruction set.