CNN Training Task Segmentation for Heterogeneous Compute Scheduling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The inflexibility and speed limitations of existing hardware platforms for training convolutional neural network (CNN) models, particularly due to the inflexible migration and computing speed issues when tasks are migrated across different computing devices or processors.

Innovation Solution

A method and apparatus that split the multiply-accumulate operations in CNN model training tasks into first-place, intermediate, and last-place multiply-add operations, identify suitable computing devices based on load thresholds, and perform computations on these operations using corresponding devices to enhance flexibility and speed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If CNN model training tasks are migrated on different computing devices or co-computed by different processors, then computing flexibility is improved, but the inflexible and customized computing execution granularity of existing devices causes computing speed to deteriorate

Engineering Contradiction:
Improveflexibility of migrationVSAvoidcomputing speed
Core Design Contradiction:
Adaptability or versatilityVSSpeed

Solution Approach 1:

The patent segments the CNN training task into multiple independent multiply-accumulate operation tasks with uniform granularity. Each operation task can be independently scheduled and executed on different computing devices, enabling flexible migration while maintaining efficient parallel execution. This segmentation resolves the contradiction by creating standardized units that can be distributed across heterogeneous devices without losing execution efficiency.

Inventive Principle:
Principle #1Segmentation

2Productivity

If dedicated and customized computing execution granularity is used for CNN training tasks, then computing efficiency on single device is improved, but flexibility of task migration and cooperative computing deteriorates

Engineering Contradiction:
Improvecomputing efficiencyVSAvoidflexibility of task migration
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal computing execution granularity (multiply-accumulate operation task) that can be executed on multiple types of computing devices including CPUs, GPUs, FPGAs, and AI-specific processors. This universal task format enables the same task to be migrated across different device types while maintaining efficient execution, thus achieving both high productivity and adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If CNN model size increases to improve accuracy, then detection and recognition accuracy is improved, but hardware platform demands increase and reach bottleneck

Engineering Contradiction:
Improveaccuracy of CNN modelVSAvoidhardware platform demand
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments large-scale CNN training tasks into multiple smaller multiply-accumulate operation tasks with uniform granularity. These segmented tasks can be distributed across multiple computing devices, allowing accurate large models to be trained without requiring a single massive hardware platform. The segmentation enables parallel processing that scales horizontally across devices rather than vertically on a single device.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4024286B1Computing method and apparatus for convolutional neural network model
Publication Date: 2026.04.08 LANGCHAO ELECTRONIC INFORMATION IND CO LTD
  • EP4024286B1 patent drawingFigure 1~2
  • EP4024286B1 patent drawingFigure 3~5
  • EP4024286B1 patent drawingFigure 6~7

AI summary

A computing method and apparatus for a convolutional neural network model. The method comprises: acquiring a computing model of a training task of a convolutional neural network model (S101); then splitting multiply-accumulate operation in a computing model of a training task of the convolutional neural network model into a plurality of multiply-add operation tasks (SI02); confirming a computing device corresponding to each multiply-add operation task according to the correlation between a preset computing model and the computing device (S103); and finally, respectively computing each multiply-add operation task by utilizing the computing device corresponding to each multiply-add operation task (S104). The purposes of improving the flexibility of migration of a CNN model training task on different computing devices or cooperative computing of different processors and improving the computing speed are achieved.