Large model hardware deployment processing method, chip and electronic device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU TSINGMICRO INTELLIGENT TECH CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies face the problem of low deployment efficiency when deploying large-scale model hardware, especially in terms of inference speed and hardware resource utilization. Furthermore, traditional quantization methods may lead to serious loss of quantization accuracy.
By acquiring the quantization sensitivity of each operation unit when the large model is deployed and run on the hardware device, quantization sensitive targets are identified, and a two-stage optimization strategy is adopted: determining the first quantization coefficient without relying on training, and performing layer-by-layer optimization when necessary to generate quantization coefficients suitable for the hardware device to optimize quantization accuracy.
It improves the deployment efficiency and accuracy of large models on hardware devices, and significantly enhances quantization accuracy to meet actual deployment requirements. For example, on some models, the quantization accuracy drop is reduced from 3% to about 1%.
Smart Images

Figure CN122114192A_ABST