The invention relates to the technical field of
deep learning, and discloses a fine-grained
pipeline scheduling method and device based on
cluster node memory difference
perception. The method comprises the following steps: estimating the ratio of the peak memory occupancy to the memory capacity of each GPU node, comparing the ratio with a preset threshold value, and selecting a re-calculation strategy and a back propagation segmentation strategy: by taking minimization of the end-to-end
training time of a model as a target, scheduling and modeling micro-batch data as a flow shop problem, analyzing the dependency constraint of each flow line stage, and calculating the flow shop problem; comprising forward-back propagation sequence dependence, stage dependence, operation dependence and memory limitation, and generating a micro-
batch operation sequence scheduling scheme meeting the memory capacity limitation; according to the scheduling scheme, an execution sequence
queue is created by executing sorting, and the
execution time sequence of calculation blocks, communication blocks and re-calculation operation is dynamically coordinated. According to the method, the GPU equipment
utilization rate can be remarkably improved, and the end-to-end training
completion time of the model is shortened.