基于不规则卷积算子自适应分发的RISC-V模型推理加速方法、装置、设备及介质

By acquiring fine-grained shape and hardware characteristic information of the convolution operator, adjusting the operator distribution code, and conducting performance tests, the problem of low execution efficiency of irregular convolution operators on the RISC-V hardware platform was solved, and the model inference speed was improved.

CN121764697BActive Publication Date: 2026-07-17INST OF SOFTWARE - CHINESE ACAD OF SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF SOFTWARE - CHINESE ACAD OF SCI
Filing Date
2026-03-05
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing rule-based operator distribution methods cannot effectively adapt to the fine-grained shape information of irregular convolution operators and the characteristics of RISC-V hardware platforms, resulting in low computational and memory access efficiency and limiting the inference speed of deep learning models on edge devices.

Method used

By acquiring fine-grained shape information and hardware characteristic information of the convolution operator, the operator distribution code is adjusted using a large language model to generate candidate operator distribution codes. Performance tests are then conducted on the target hardware platform to select the operator distribution code with the highest execution efficiency.

Benefits of technology

It significantly improves the execution efficiency of irregular convolution operators on the RISC-V hardware platform, optimizes computation and memory access performance, and increases end-to-end model inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121764697B_ABST
    Figure CN121764697B_ABST
Patent Text Reader

Abstract

本发明涉及人工智能技术领域,提供一种基于不规则卷积算子自适应分发的RISC‑V模型推理加速方法、装置、设备及介质,包括:基于卷积算子的细粒度形状信息和目标硬件平台的硬件特性信息,对原始算子分发代码中多种卷积算法实现逻辑的选择策略进行调整,生成候选算子分发代码;对候选算子分发代码进行性能测试,并根据测试结果确定目标算子分发代码;将针对目标模型的原始算子分发代码更新为目标算子分发代码后,运行目标模型的推理应用程序。本发明通过联合利用算子细粒度形状信息与硬件特性信息对算子分发逻辑进行自适应调整,实现了针对不同应用场景和硬件环境的深度定制化推理优化,使得模型推理效率高、硬件资源利用率好且泛化能力强。
Need to check novelty before this filing date? Find Prior Art