Near-Memory Computing Modules for AI Platform Bottleneck
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In AI computing, centralized CPU scheduling in single-server or cloud computing scenarios leads to processing bottlenecks, limiting resource deployment flexibility and scalability, and hindering efficient computing in complex scenarios.
Innovation Solution
An AI computing platform with near-memory computing modules connected to each other and a processor, which decomposes tasks into ordered subtasks based on a network topology table, allowing each module to process specific operations and reduce unified scheduling load, enabling efficient distributed computing across single-server, multi-server, and cloud environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a CPU performs centralized scheduling of resources in NCMs and data interaction between NCMs is forwarded through the CPU, then resource management is simplified, but the CPU becomes a bottleneck in processing capability and interface communication capability
Solution Approach 1:
The patent segments the centralized CPU scheduling function into distributed scheduling across multiple NCMs. Each NCM is equipped with independent scheduling capabilities, allowing parallel task management and eliminating the CPU bottleneck. The system divides the monolithic scheduling architecture into modular, autonomous scheduling units distributed across the network.
Solution Approach 2:
The patent introduces a network-based communication intermediary layer that enables direct NCM-to-NCM data interaction without routing through the CPU. This intermediary network infrastructure allows NCMs to exchange data and coordinate tasks autonomously, reducing CPU communication overhead while maintaining system coherence.
2Ease of manufacture
If centralized CPU scheduling is used in single-server scenarios, then implementation is simple, but scalability and portability are poor when extended to multi-server or cloud computing scenarios
Solution Approach 1:
The patent designs a universal NCM architecture that can function in both single-server and multi-server/cloud scenarios. Each NCM is equipped with standardized interfaces and autonomous scheduling capabilities that allow seamless deployment across different scales. The system can adapt from a single NCM to multiple NCMs across different servers without requiring architectural changes.
Solution Approach 2:
The patent transitions from a single-dimension centralized CPU model to a multi-dimensional distributed NCM architecture. By adding the network dimension and enabling NCMs to operate autonomously across multiple servers, the system gains scalability and portability while maintaining implementation simplicity through standardized protocols and interfaces.
3Device complexity
If data interaction between NCMs is forwarded through the CPU, then data routing is simplified, but interface communication capability becomes a bottleneck
Solution Approach 1:
The patent segments the centralized data routing function into distributed routing capabilities at each NCM. Each NCM maintains its own routing table and can independently determine data paths to other NCMs, eliminating the CPU as a communication bottleneck while keeping routing management simple through standardized protocols.
Solution Approach 2:
The patent introduces a network intermediary layer that enables direct NCM-to-NCM communication. This network infrastructure acts as an intermediary, allowing fast data exchange between NCMs without CPU involvement, thereby improving interface communication capability while maintaining simplified routing through standard network protocols.
Data Source
AI summary
An AI computing platform, an AI computing method, and an AI cloud computing system, the platform including: at least one computing component, each computing component includes: a processor, configured to initiate a calculation task and decompose the calculation task into a plurality of ordered subtasks according to a network topology information table stored therein; a plurality of near-memory computing modules, the plurality of near-memory computing modules connecting in pairs with the processor, and the plurality of near-memory computing modules connecting in pairs with each other, wherein the plurality of near-memory computing modules are each configured to implement different operation types, and the plurality of near-memory computing modules complete one or more of the plurality of subtasks according to the operation types they each implement.


