A method and system for parallel training of a large-scale neural network of a perception model structure

By combining semantically enhanced computation graphs and hardware topology profiles, illegal paths are eliminated. A multi-dimensional cost model and a shortest path optimization algorithm with real-time weight correction are used to solve the long-tail effect problem in heterogeneous cluster computing in existing technologies, thereby improving the throughput of model training.

CN122414248APending Publication Date: 2026-07-17NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NO 15 INST OF CHINA ELECTRONICS TECH GRP
Filing Date
2026-04-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing automation methods lack a deep understanding of the logical semantics of the model when dealing with highly integrated operators and non-Transformer structures, which leads to the inability to fully utilize the hardware potential of heterogeneous clusters and results in a long tail effect in computation.

Method used

By acquiring the original model code, it is transformed into a semantically enhanced computational graph. Combined with hardware topology profiling, constraint rule library, and fault tolerance contingency plan library, logically illegal paths and inefficient paths are eliminated. The shortest path optimization algorithm with multidimensional cost model and real-time weight correction is used to extract the best parallel strategy quadruple for training.

Benefits of technology

Ensuring the integrity of computational logic and avoiding communication spikes significantly improves model training throughput in heterogeneous architectures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122414248A_ABST
    Figure CN122414248A_ABST
Patent Text Reader

Abstract

本发明公开了一种感知模型结构的大规模神经网络并行训练方法及系统。包括:获取原始模型代码,将原始模型代码转化为语义增强计算图;基于语义增强计算图、硬件拓扑画像、约束规则库、硬件拓扑实时状态以及容错预案库剔除逻辑非法路径与低效路径,得到受约束的可选搜索空间;基于受约束的可选搜索空间、多维代价模型以及实时权重修正执行最短路径寻优,基于提取最佳并行策略四元组对模型进行训练。本发明通过语义标签映射机制识别出状态空间模型的隐藏状态依赖及融合算子的原子属性,将抽象数学语义转化为硬性切分约束;避免了在自动化切分中因“语义盲区”导致的非法拆解,确保了计算逻辑的完整性,显著提升了异构架构下的模型训练吞吐量。
Need to check novelty before this filing date? Find Prior Art