混合精度量化模型的比特分配方法和装置
By employing a two-stage global bit allocation process, and utilizing teacher-forced inter-layer interest information reconstruction and differentiable mask preference probability distribution, the problems of high computational overhead and high adaptation cost in existing technologies are solved, enabling the rapid generation and stable deployment of mixed-precision quantization models under extremely low bit conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 启元实验室
- Filing Date
- 2026-04-22
- Publication Date
- 2026-07-17
AI Technical Summary
Under extremely low bit conditions, existing mixed-precision quantization methods have high computational overhead and unstable results, making it difficult to quickly adapt to different budgets. This leads to high deployment costs and inconsistent performance, making it difficult to meet the deployment needs of resource-constrained environments.
A two-stage global bit allocation process is adopted. By reconstructing the inter-layer interest information under teacher-forced conditions and differentiable mask preference probability distribution, the preference scores of linear modules are learned, and discrete bit selection that strictly satisfies the average bit budget is obtained through continuous relaxation optimization.
It enables the rapid generation of reusable mixed-precision quantization models under extremely low bit conditions, reducing computational overhead and adaptation costs, and ensuring the consistency of model performance and deployment stability.
Smart Images

Figure CN122065884B_ABST