一种适用于自回归掩码生成模型的后训练量化方法

By employing hierarchical clustering decoupling and scaling recalibration mechanisms, the problems of extreme outliers and anomalous amplification in the MAR model are solved, achieving stable deployment under low bit precision, reducing computation and storage costs, and making it suitable for edge devices.

CN121168554BActive Publication Date: 2026-07-17BEIJING JIAOTONG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING JIAOTONG UNIV
Filing Date
2025-09-15
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing post-training quantization methods cannot effectively handle extreme outliers and anomalous amplification in autoregressive mask generation models (MAR), resulting in unstable model deployment, excessively high computational and storage costs, and difficulty in application in resource-constrained environments.

Method used

Hierarchical clustering decoupling mechanism (HCD) and scaling recalibration mechanism (SR) are adopted to identify and decouple abnormal channels, dynamically adjust the quantization range, and cooperate with the full-process static quantization strategy to achieve stable deployment under low bit precision.

Benefits of technology

While maintaining generation quality, it significantly reduces storage and computing costs, enabling stable inference deployment of MAR models at low bit precision, suitable for edge devices and resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168554B_ABST
    Figure CN121168554B_ABST
Patent Text Reader

Abstract

本发明提出了一种适用于自回归掩码生成模型的后训练量化方法——Q‑MAR(Quantization for Masked Autoregressive Models),该方法创新性地设计了两个模块:分层聚类解耦模块(Hierarchical Cluster Decoupling,HCD)用于在Transformer中识别并解耦含有极大异常值(EOs)的通道,实现对正常值与异常值的分离量化;缩放重校准模块(Scaling Recalibration,SR)用于在扩散网络中对激活漂移进行动态缩放控制,稳定早期时间步的激活范围,抑制由于量化误差引起的异常值(AAs)放大导致的模型坍塌。Q‑MAR能够在无需重新训练模型的前提下,将MAR模型稳定压缩至8‑bit精度(W8A8),在保持生成质量的同时显著降低存储与计算资源消耗,为MAR模型在移动端、边缘端等资源受限设备上的部署提供了切实可行的技术支撑。
Need to check novelty before this filing date? Find Prior Art