The invention discloses a
GPU cluster scheduling method, device and equipment based on four-layer
network topology, and the method comprises the steps: building a resource portrait based on historical data, integrating a
server to a four-layer topology state of a super convergence layer, and forming
system state information; an
attack level is calculated for the task, and a double-weight
preemption relation with the unique condition that the
attack level is larger than the defense level is established; according to the
system state, candidate schemes are generated from inside to outside; verifying calculation and network resource capacity constraints layer by layer; for the feasible schemes meeting the constraint conditions, calculating a total
score obtained by summarizing the completion
score, the scale
score and the topological affinity, deducting the
preemption cost to obtain net
gain, and selecting the scheme with the maximum and positive net
gain for execution; after the task is preempted, the defense level is automatically improved by one level, and the preemptive times of the task are strictly limited. According to the method, the
utilization rate of the GPU and the HBM is remarkably improved, the cross-layer
network communication overhead is effectively reduced, and the stability and fairness of the scheduling process are guaranteed.