CNN Training Task Segmentation for Heterogeneous Compute Scheduling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The inflexibility and speed limitations of existing hardware platforms for training convolutional neural network (CNN) models, particularly due to the inflexible migration and computing speed issues when tasks are migrated across different computing devices or processors.
Innovation Solution
A method and apparatus that split the multiply-accumulate operations in CNN model training tasks into first-place, intermediate, and last-place multiply-add operations, identify suitable computing devices based on load thresholds, and perform computations on these operations using corresponding devices to enhance flexibility and speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If CNN model training tasks are migrated on different computing devices or co-computed by different processors, then computing flexibility is improved, but the inflexible and customized computing execution granularity of existing devices causes computing speed to deteriorate
Solution Approach 1:
The patent segments the CNN training task into multiple independent multiply-accumulate operation tasks with uniform granularity. Each operation task can be independently scheduled and executed on different computing devices, enabling flexible migration while maintaining efficient parallel execution. This segmentation resolves the contradiction by creating standardized units that can be distributed across heterogeneous devices without losing execution efficiency.
2Productivity
If dedicated and customized computing execution granularity is used for CNN training tasks, then computing efficiency on single device is improved, but flexibility of task migration and cooperative computing deteriorates
Solution Approach 1:
The patent creates a universal computing execution granularity (multiply-accumulate operation task) that can be executed on multiple types of computing devices including CPUs, GPUs, FPGAs, and AI-specific processors. This universal task format enables the same task to be migrated across different device types while maintaining efficient execution, thus achieving both high productivity and adaptability.
3Measurement precision
If CNN model size increases to improve accuracy, then detection and recognition accuracy is improved, but hardware platform demands increase and reach bottleneck
Solution Approach 1:
The patent segments large-scale CNN training tasks into multiple smaller multiply-accumulate operation tasks with uniform granularity. These segmented tasks can be distributed across multiple computing devices, allowing accurate large models to be trained without requiring a single massive hardware platform. The segmentation enables parallel processing that scales horizontally across devices rather than vertically on a single device.
Data Source
Figure 1~2
Figure 3~5
Figure 6~7
AI summary
A computing method and apparatus for a convolutional neural network model. The method comprises: acquiring a computing model of a training task of a convolutional neural network model (S101); then splitting multiply-accumulate operation in a computing model of a training task of the convolutional neural network model into a plurality of multiply-add operation tasks (SI02); confirming a computing device corresponding to each multiply-add operation task according to the correlation between a preset computing model and the computing device (S103); and finally, respectively computing each multiply-add operation task by utilizing the computing device corresponding to each multiply-add operation task (S104). The purposes of improving the flexibility of migration of a CNN model training task on different computing devices or cooperative computing of different processors and improving the computing speed are achieved.