The invention relates to a
large model mixing precision quantitative reasoning method and device based on a
rotation matrix. The method comprises the following steps: acquiring a target activation
tensor; a
rotation matrix is adopted to carry out rotation transformation on the target activation
tensor, lexical element dimension mixing precision quantification is carried out, and activation values with the same precision in the quantized activation
tensor are continuously distributed in a memory; the activation tensor of the first precision is stored in an
outlier buffer area allocated in a memory; calling a tensor core, and performing
matrix multiplication operation on the quantized activation tensor and the quantized weight tensor to obtain a quantization operation result; and carrying out
inverse quantization on a quantization operation result, and taking an
inverse quantization result as a
matrix multiplication operation result of the target activation tensor and the
target weight tensor. According to the method and the device, the activation values with the same precision are continuously distributed in the memory, and the activation tensor with the first precision is stored in the
outlier buffer area, so that the tensor core can be efficiently calculated, and the capability and relatively high accuracy of the model can be kept while model reasoning is accelerated.