The invention relates to the technical field of large
language model deployment, and discloses a large
language model weight
inverse quantization reasoning device and method.The method comprises the steps that low-precision weight data is transmitted to a high-bandwidth storage from a host and then transmitted to an on-
chip storage through the high-bandwidth storage;
data conversion from a low-precision format to a high-precision format is completed in the on-
chip memory, the data is multiplied by an
inverse quantization factor to obtain recovered high-precision weight data, and the functional unit is responsible for executing
general matrix multiplication of input data and the high-precision weight data after
inverse quantization. And pipeline parallel execution of the inverse quantization operation and the
general matrix multiplication operation is realized through a double-buffering technology. According to the invention, on the basis of a dual-path inverse quantization architecture of the vector
processing unit and a dual-buffer mechanism in the on-
chip memory, the problem of hardware
adaptation of low-precision calculation is solved, and efficient execution of a low-precision conversion
algorithm is realized under the condition of limited hardware resources.