大模型硬件部署处理方法、芯片及电子设备
By identifying quantization-sensitive targets in large-scale model hardware deployment, and employing training-independent quantization coefficient optimization and layer-by-layer training methods, the problem of low efficiency in large-scale model hardware deployment is solved, achieving efficient and accurate hardware deployment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGSU TSINGMICRO INTELLIGENT TECH CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-17
AI Technical Summary
Large models suffer from low deployment efficiency during hardware deployment, particularly in terms of inference speed and hardware resource utilization. Existing quantization methods also suffer from quantization accuracy loss and low optimization efficiency.
By acquiring the quantization sensitivity of each operation unit when the large model is deployed and run on the hardware device, identifying quantization-sensitive targets, determining the first quantization coefficient in a training-independent manner, and performing layer-by-layer optimization when necessary, a deployment processing file suitable for the hardware device is generated.
It improves the deployment efficiency and quantization accuracy of large models on hardware, reducing the quantization accuracy loss to about 1% of floating-point accuracy on the Qwen3-8B model and to about 1% on the DeepSeek R1 model, meeting practical deployment requirements.
Smart Images

Figure CN122114192B_ABST