The present application relates to a kind of neural network global one-time structured
pruning method,
system, equipment and medium, wherein, method includes: by single
forward propagation, in the neural network to be pruned using calibration dataset and parallelly capturing the input activation
tensor of all target
layers;According to input activation
tensor, the intermediate
neuron weight of each target layer is synchronously calculated differential entropy index and
amplitude response intensity, and after normalization and fusion, static global importance atlas is formed;According to the importance threshold value determined according to preset
pruning rate, based on global importance atlas, one-time generates the index set of all global
pruning to be pruned whose mixed importance
score is below importance threshold value;Based on index set, the weight matrix of each target layer is executed one-time physical structured pruning.By doing so, the present application generates static global importance atlas by single
forward propagation and parallelly capturing activation
tensor, and completes no-
mask one-time pruning by physical structured pruning.