This invention provides an in-memory computing
system, method, and
chip that supports fully parallel computation of deep
convolution channels. The
system includes: a
parallel computing module comprising k×k multi-bit data units, each multi-bit data unit comprising q storage sub-units for storing k×k q-bit activation value storage data, and performing simulated multiplication and accumulation operations on the k×k q-bit activation value storage data and input weight parameters of size (k×k,1); a weight configuration circuit connected to the
parallel computing module, used to perform bit-level
simulation recombination of the q-bit weights in the simulated multiplication and accumulation operation output by the
parallel computing module, obtaining a simulated multiplication and accumulation result of 1-bit input weight parameters and q-bit activation value storage data; an ADC quantization circuit connected to the weight configuration circuit, used to quantize the simulated multiplication and accumulation result, obtaining integer data output; and an activation value update circuit connected to the
bit line and
read bit line of the column where the parallel computing module is located, used to perform local cyclic updates of the activation value storage data within the parallel computing module. This invention can increase the parallelism of array computation and realize local cyclic updates of activation values within the array, thereby improving energy efficiency.