A cascaded matrix multiplication operator fusion method based on deep reinforcement learning

By employing deep reinforcement learning, the problem of cascaded matrix multiplication and fusion of computationally intensive operator chains on the domestic Ascend AI processor was solved, generating an optimal scheduling strategy that minimizes memory access overhead and maximizes hardware efficiency, adapting to the hardware boundaries of the Da Vinci architecture.

CN122413286APending Publication Date: 2026-07-17HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2026-04-14
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing compilers struggle to achieve efficient cascaded matrix multiplication operator fusion on the discrete architecture of the domestic Ascend AI processor when processing computationally intensive operator chains, especially due to insufficient consideration of hardware constraints such as 512-byte bus alignment and burst transmission length, leading to performance degradation.

Method used

A deep reinforcement learning-based approach is adopted to construct a multi-objective optimization mechanism. The scheduling strategy is searched by a reinforcement learning agent, taking into account theoretical transport volume, hardware alignment constraints, on-chip buffer fragmentation rate and pipeline parallelism, to generate the optimal cascaded matrix multiplier fusion strategy.

Benefits of technology

It achieves the minimization of memory access overhead and the maximization of hardware execution efficiency on the domestic Ascend AI processor. The generated scheduling scheme completes the inference of all scheduling parameters within 10 milliseconds, adapting to the hardware boundary requirements of the Da Vinci architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122413286A_ABST
    Figure CN122413286A_ABST
Patent Text Reader

Abstract

本发明公开了一种基于深度强化学习的级联矩阵乘算子融合方法,所述方法包括如下步骤:一、针对待优化的级联矩阵乘算子链进行拓扑分析,提取级联矩阵乘的算子数量及维度信息,同时,根据昇腾的硬件特性构建强化学习环境;二、生成调度策略并构建用于描述当前调度决策状态的向量;三、在每个决策步,系统动态生成动作掩码以过滤非法决策,确保搜索效率与方案可行性;四、当一个完整的调度策略生成后,计算综合奖励值以指导策略优化;五、利用Maskable PPO算法进行模型迭代。本发明解决了现有分析模型技术在适配昇腾分离架构AI处理器时存在的512字节总线对齐及突发传输长度等硬件约束建模缺失的技术问题。
Need to check novelty before this filing date? Find Prior Art