A cascaded matrix multiplication operator fusion method based on deep reinforcement learning
By employing deep reinforcement learning, the problem of cascaded matrix multiplication and fusion of computationally intensive operator chains on the domestic Ascend AI processor was solved, generating an optimal scheduling strategy that minimizes memory access overhead and maximizes hardware efficiency, adapting to the hardware boundaries of the Da Vinci architecture.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INST OF TECH
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-17
AI Technical Summary
Existing compilers struggle to achieve efficient cascaded matrix multiplication operator fusion on the discrete architecture of the domestic Ascend AI processor when processing computationally intensive operator chains, especially due to insufficient consideration of hardware constraints such as 512-byte bus alignment and burst transmission length, leading to performance degradation.
A deep reinforcement learning-based approach is adopted to construct a multi-objective optimization mechanism. The scheduling strategy is searched by a reinforcement learning agent, taking into account theoretical transport volume, hardware alignment constraints, on-chip buffer fragmentation rate and pipeline parallelism, to generate the optimal cascaded matrix multiplier fusion strategy.
It achieves the minimization of memory access overhead and the maximization of hardware execution efficiency on the domestic Ascend AI processor. The generated scheduling scheme completes the inference of all scheduling parameters within 10 milliseconds, adapting to the hardware boundary requirements of the Da Vinci architecture.
Smart Images

Figure CN122413286A_ABST