一种用于训练模型的控制方法、装置及电子设备

By encapsulating computational and communication tasks in a single kernel during deep learning model training and adjusting configuration parameters in real time, the efficiency bottleneck caused by communication latency is resolved, achieving efficient integration of computation and communication, and improving training efficiency and resource utilization.

CN121902075BActive Publication Date: 2026-07-17ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG LAB
Filing Date
2026-03-26
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In large-scale deep learning model training, communication latency has become a key bottleneck restricting training efficiency and resource utilization. In particular, communication overhead accounts for a large proportion in distributed training, causing GPU computing resources to wait frequently.

Method used

By encapsulating parallel computation and communication tasks into a single fusion kernel during the model training initialization phase and collecting performance status information in real time, the running configuration parameters are dynamically adjusted to achieve efficient fusion and overlapping execution of computation and communication, reducing communication waiting overhead.

Benefits of technology

It improves the efficiency of model training and resource utilization, and realizes pipelined scheduling of computing and communication through kernel-level fusion, reducing communication waiting and kernel switching overhead, and improving the overall GPU throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121902075B_ABST
    Figure CN121902075B_ABST
Patent Text Reader

Abstract

本申请一种用于训练模型的控制方法、装置及电子设备,该方法包括:在模型训练初始化阶段,根据获取的模型计算图,将可并行执行的计算和通信任务封装于单一的融合内核中;在通过训练数据集迭代训练天文大模型时:根据实时采集的性能状态信息,调整当前的运行配置参数得到目标配置参数,以生成融合内核的目标执行策略;控制融合内核,基于目标执行策略执行任务,并返回采集性能状态信息的步骤,直至达到预设迭代条件为止。由此,在模型训练前,通过内核级融合可并行执行的计算与通信任务,降低通信等待开销。在模型训练中,通过实时采集的性能状态信息,对运行配置参数进行动态调整,实现通信与计算的高效融合与重叠执行,提升模型训练效率。
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Parallel training acceleration method and system for large model parameter partitioning

    CN117744838A

  • Training method, data processing method, electronic equipment and computer readable storage medium

    CN120562507A