Large model hardware deployment processing method, chip and electronic device

CN122114192AActive Publication Date: 2026-05-29JIANGSU TSINGMICRO INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU TSINGMICRO INTELLIGENT TECH CO LTD
Filing Date
2026-04-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies face the problem of low deployment efficiency when deploying large-scale model hardware, especially in terms of inference speed and hardware resource utilization. Furthermore, traditional quantization methods may lead to serious loss of quantization accuracy.

Method used

By acquiring the quantization sensitivity of each operation unit when the large model is deployed and run on the hardware device, quantization sensitive targets are identified, and a two-stage optimization strategy is adopted: determining the first quantization coefficient without relying on training, and performing layer-by-layer optimization when necessary to generate quantization coefficients suitable for the hardware device to optimize quantization accuracy.

Benefits of technology

It improves the deployment efficiency and accuracy of large models on hardware devices, and significantly enhances quantization accuracy to meet actual deployment requirements. For example, on some models, the quantization accuracy drop is reduced from 3% to about 1%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114192A_ABST
    Figure CN122114192A_ABST
Patent Text Reader

Abstract

The application discloses a large model hardware deployment processing method, a chip and an electronic device. The method comprises the following steps: acquiring the quantization sensitivity of each operation unit of a large model during deployment and running on a hardware device; identifying a quantization sensitive target which meets the set requirement for the influence on the hardware calculation error according to the quantization sensitivity; obtaining a first quantization coefficient for controlling data truncation in a manner independent of training based on the quantization sensitive target, so as to minimize the quantization output error of the hardware device; when the optimization result of improving the quantization precision based on the first quantization coefficient does not meet the requirement or the quantization precision needs to reach the set standard, performing layer-by-layer optimization based on the quantization sensitive target by using a training method to obtain a second quantization coefficient for controlling data truncation; and generating a large model hardware deployment processing file suitable for the hardware device to perform reasoning based on the first quantization coefficient or the second quantization coefficient. The application can improve the deployment efficiency of the large model.
Need to check novelty before this filing date? Find Prior Art