一种用于训练模型的控制方法、装置及电子设备
By encapsulating computational and communication tasks in a single kernel during deep learning model training and adjusting configuration parameters in real time, the efficiency bottleneck caused by communication latency is resolved, achieving efficient integration of computation and communication, and improving training efficiency and resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG LAB
- Filing Date
- 2026-03-26
- Publication Date
- 2026-07-17
AI Technical Summary
In large-scale deep learning model training, communication latency has become a key bottleneck restricting training efficiency and resource utilization. In particular, communication overhead accounts for a large proportion in distributed training, causing GPU computing resources to wait frequently.
By encapsulating parallel computation and communication tasks into a single fusion kernel during the model training initialization phase and collecting performance status information in real time, the running configuration parameters are dynamically adjusted to achieve efficient fusion and overlapping execution of computation and communication, reducing communication waiting overhead.
It improves the efficiency of model training and resource utilization, and realizes pipelined scheduling of computing and communication through kernel-level fusion, reducing communication waiting and kernel switching overhead, and improving the overall GPU throughput.
Smart Images

Figure CN121902075B_ABST
Abstract
Citation Information
Patent Citations
Parallel training acceleration method and system for large model parameter partitioning
CN117744838A
Training method, data processing method, electronic equipment and computer readable storage medium
CN120562507A