The application relates to the technical field of
neural network hardware acceleration, in particular to a high-precision automatic deployment method for an
analog memory-computation integrated accelerator; the method comprises the following steps: layer-by-layer statistics of activation values and
weight distribution of a floating-
point model with a
verification set is carried out, a quantization scale factor is automatically updated to avoid information truncation; quantization parameters are constrained in a discrete
value set that can be realized by a
chip, an interlayer mathematical consistency relationship is established, and a cross-layer linkage updating mechanism is adopted to process parameter
coupling; based on an additive
Gaussian noise model, weight and activation dynamic ranges are jointly optimized from the perspective of improving end-to-end
signal-to-
noise ratio, amplitude compensation and cross-layer correction are carried out; with deployment precision or
signal-to-
noise ratio as the target, a closed-loop iteration is constructed between discrete quantization training and joint compensation until the precision meets the standard or the parameters converge; a weight
programming file, a quantization parameter configuration file and an
inference control file are generated, which are directly loaded by the
chip. The application realizes high-precision automatic deployment of the
analog memory-computation integrated accelerator.