混合精度量化模型的比特分配方法和装置

By employing a two-stage global bit allocation process, and utilizing teacher-forced inter-layer interest information reconstruction and differentiable mask preference probability distribution, the problems of high computational overhead and high adaptation cost in existing technologies are solved, enabling the rapid generation and stable deployment of mixed-precision quantization models under extremely low bit conditions.

CN122065884BActive Publication Date: 2026-07-17启元实验室

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
启元实验室
Filing Date
2026-04-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Under extremely low bit conditions, existing mixed-precision quantization methods have high computational overhead and unstable results, making it difficult to quickly adapt to different budgets. This leads to high deployment costs and inconsistent performance, making it difficult to meet the deployment needs of resource-constrained environments.

Method used

A two-stage global bit allocation process is adopted. By reconstructing the inter-layer interest information under teacher-forced conditions and differentiable mask preference probability distribution, the preference scores of linear modules are learned, and discrete bit selection that strictly satisfies the average bit budget is obtained through continuous relaxation optimization.

Benefits of technology

It enables the rapid generation of reusable mixed-precision quantization models under extremely low bit conditions, reducing computational overhead and adaptation costs, and ensuring the consistency of model performance and deployment stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122065884B_ABST
    Figure CN122065884B_ABST
Patent Text Reader

Abstract

本申请涉及混合精度量化模型的比特分配方法和装置,采用两阶段全局比特分配方法,首先,采用教师强制层间感兴趣信息重建作为训练目标,使优化方法对齐后训练量化关注的表示一致性。其次,利用偏好概率分布对各候选比特输出进行连续松弛加权,在小规模校准数据上即可通过梯度优化学习偏好得分,避免在指数级组合空间中反复评估的高成本离散搜索。再次,通过求解将偏好得分转化为严格满足目标平均比特数的离散比特选择结果,从而保证生成的混合精度量化模型可直接部署且不会超过预算。最后,偏好得分可在不同目标平均比特数下复用,仅需重新执行离散投影即可快速得到新预算配置,显著降低不同预算条件下适配成本并提升工程可复现性与部署效率。
Need to check novelty before this filing date? Find Prior Art