The invention belongs to the technical field of
large model quantification, and discloses a
large model quantification method based on non-uniform grouping Adamage transform and activation distribution self-adaption, which comprises the following steps: firstly, eliminating abnormal values in an activation matrix through non-uniform grouping Adamage transform, and solving the problem of
abnormal distribution caused by uniform grouping; secondly, constructing an
equivalent transformation error
elimination algorithm, and eliminating a weight quantization error by adopting a GPTQ
algorithm; and finally, distributing an optimal quantization
data type (Int4 or Nint4) for each activation matrix through an activation distribution
adaptive algorithm so as to minimize a quantization error. According to the method, the deployment efficiency of online rotation is improved through non-uniform grouping Adamage transformation, meanwhile, the quantization error is remarkably reduced in combination with self-adaptive
data type selection, and the model quantization precision and the hardware deployment efficiency are improved. The method is suitable for compression and optimization of a large
language model, and has a wide application prospect.