In-Storage Processing for Distributed Computing Data Nodes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed computing systems, the traditional approach of using central processors to execute tasks leads to inefficiencies due to the need to move data between storage, memory, and processing units, resulting in increased power consumption and processing time, as well as limited system resource utilization.
Innovation Solution
The system employs intelligent storage mediums with controller processors that can execute tasks independently, allowing data to be processed locally within the storage medium, reducing data movement and optimizing resource usage by assigning tasks to either central processors or intelligent storage mediums based on data size and processing requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If data is processed by central processors in traditional distributed computing systems, then computational tasks can be executed, but data movement between storage, memory, and processing units increases power consumption and processing time
Solution Approach 1:
The patent segments the computing system into multiple data nodes, each containing both storage mediums and processing units. This segmentation allows computation to occur closer to where data resides, reducing the need for data movement across the system and thereby lowering power consumption while maintaining processing efficiency.
Solution Approach 2:
The patent introduces in-storage processing capabilities as an intermediary between storage and traditional central processors. This intermediary layer enables data to be processed within or near the storage medium itself, eliminating the need for extensive data movement to and from central processors, thus reducing power consumption without sacrificing productivity.
2Productivity
If data is moved between storage, memory, and processing units, then computational tasks can be completed, but processing time increases
Solution Approach 1:
By segmenting the system into data nodes with integrated storage and processing, the patent eliminates the sequential data movement steps between separate storage, memory, and processing units. Computation can occur in-place or near the data, significantly reducing data movement time and improving overall processing speed.
Solution Approach 2:
The patent implements preliminary processing actions within the storage medium or data node before data needs to be moved to central processors. By performing initial computation, filtering, or transformation operations closer to the data source, the system reduces the volume and frequency of data movement, thereby reducing processing time.
3Adaptability or versatility
If traditional central processors execute all tasks, then system architecture remains simple, but system resource utilization is limited
Solution Approach 1:
The patent makes storage mediums and data nodes multi-functional by enabling them to perform both storage and computation functions. This universality allows the system to utilize storage resources for processing tasks, significantly improving overall resource utilization without requiring a completely new architectural paradigm.
Solution Approach 2:
The patent introduces dynamic task assignment mechanisms that can flexibly allocate computational tasks between central processors and in-storage processing units based on current system conditions, data characteristics, and resource availability. This dynamic approach enables adaptive resource utilization while maintaining a relatively simple underlying architecture.
Data Source
AI summary
According to one general aspect, a scheduler computing device may include a computing task memory configured to store at least one computing task. The computing task may be executed by a data node of a distributed computing system, wherein the distributed computing system includes at least one data node, each data node having a central processor and an intelligent storage medium, wherein the intelligent storage medium comprises a controller processor and a memory. The scheduler computing device may include a processor configured to assign the computing task to be executed by either the central processor of a data node or the intelligent storage medium of the data node, based, at least in part, upon an amount of data associated with the computing task.


