The invention discloses a communication efficient distributed reasoning method based on model
pruning, and relates to the technical field of distributed reasoning, and the method comprises the steps: S1, constructing a
delay prediction model, S2, determining an initial
break point in a
heuristic manner, S3, carrying out model
pruning, S4, determining the
break point through a
dynamic planning method, and S5, carrying out model
pruning. The method comprises the following steps: firstly, designing a neural
network delay prediction model, selecting according to importance by using the obtained
delay, the accuracy of the model and the communication overhead of a segmentation point, removing a part which has great influence on the
delay and has small influence on the delay in the model, and then finely adjusting the model to adjust the accuracy of the model; regularization items for communication overhead are added, the communication overhead brought in the reasoning process is further reduced, meanwhile, in order to determine optimal segmentation points of delay and communication, a
dynamic planning method can be utilized according to the prediction result of the model, finally, the pruning process and model segmentation are optimized in a cross iteration mode, and the delay and communication overhead are optimized.