Large model hardware deployment processing method, chip and electronic device
Patent Information
- Application Number
- CN202610580243.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-29
- Publication Date
- 2026-05-29
- Estimated Expiration
- 2046-04-29
AI Technical Summary
Existing technologies face the problem of low deployment efficiency when deploying large-scale model hardware, especially in terms of inference speed and hardware resource utilization. Furthermore, traditional quantization methods may lead to serious loss of quantization accuracy.
By acquiring the quantization sensitivity of each operation unit when the large model is deployed and run on the hardware device, quantization sensitive targets are identified, and a two-stage optimization strategy is adopted: determining the first quantization coefficient without relying on training, and performing layer-by-layer optimization when necessary to generate quantization coefficients suitable for the hardware device to optimize quantization accuracy.
It improves the deployment efficiency and accuracy of large models on hardware devices, and significantly enhances quantization accuracy to meet actual deployment requirements. For example, on some models, the quantization accuracy drop is reduced from 3% to about 1%.
Smart Images

Figure CN122114192A_ABST
Abstract
Citation Information
Patent Citations
Computer vision processing method and device, equipment and storage medium
CN119295896A
Quantization method, device and equipment of 3D target detection model, and computer program product
CN120411952A
Quantitative perception training method and device of neural network model, electronic equipment and storage medium
CN121119003A
Model optimization method, electronic equipment and storage medium
CN121745199A
Large model lightweight reasoning deployment method under limited hardware resources
CN121745311A