Parallel Processing Network Model Operations via Segmented Modules
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional parallel computing systems for deep learning network models face high bandwidth requirements and significant energy consumption due to large memory access overhead and the need for parallel operations.
Innovation Solution
The proposed solution involves an operation device with multiple operation modules that execute computational sub-commands in parallel, where each module includes an operation unit and a storage unit for storing required data, reducing the need for external memory access.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional parallel computing systems are used for deep learning network models, then computational operations can be performed, but bandwidth requirements and energy consumption increase significantly due to large memory access overhead
Solution Approach 1:
The patent divides the computational system into multiple independent operation modules, each capable of autonomous parallel computation. By segmenting the overall computational task into sub-tasks that can be executed independently by different modules, the system reduces the need for frequent data access to external memory, thereby lowering energy consumption while maintaining high computational productivity
Solution Approach 2:
The patent transitions from a traditional hierarchical memory architecture to a distributed storage architecture where storage units are integrated at the operation module level. This dimensional change in system architecture allows data to be stored closer to where it is processed, reducing memory access overhead and energy consumption without compromising computational capability
2Productivity
If traditional parallel computing systems are used for deep learning network models, then computational operations can be performed, but memory access overhead increases leading to high bandwidth requirements
Solution Approach 1:
The patent segments the computational system into multiple operation modules with integrated storage units. Each module can independently access its local storage without requiring high-bandwidth communication with external memory, thereby maintaining parallel operation capability while significantly reducing overall data access bandwidth requirements
Solution Approach 2:
The patent introduces local storage units as intermediaries between operation units and external memory. These storage units act as buffers that hold frequently accessed data locally, reducing the need for high-bandwidth data transfer between operation units and external memory while maintaining parallel computational efficiency
3Reliability
If large-capacity storage devices are used to support parallel operations, then data availability is improved, but device cost increases
Solution Approach 1:
The patent divides the storage function into multiple small local storage units distributed across operation modules rather than using a single large-capacity storage device. This segmentation provides sufficient data availability for parallel operations while using smaller, more cost-effective storage components, thereby reducing overall device complexity and cost
Solution Approach 2:
Each operation module in the patent contains its own local storage unit, enabling self-service data access without requiring large-capacity centralized storage. This self-service architecture ensures data availability is met for parallel operations while avoiding the need for expensive large-capacity storage devices, thus reducing device complexity and cost
Data Source
AI summary
The present application relates to an operation device and an operation method. The operation device includes a plurality of operation modules. The plurality of operation modules complete an operation of a network model by executing corresponding computational sub-commands in parallel. Each operation module includes at least one operation unit configured to execute a first computational sub-command using first computational sub-data; and a storage unit configured to store the first computational sub-data. The first computational sub-data includes data needed for executing the first computational sub-command. The embodiments of the present application reduces bandwidth requirements for data access and reduces computation and equipment costs.


