The invention belongs to the technical field of
machine learning training, and particularly relates to a
system and a method for training a
large model by using a heterogeneous GPU (
Graphics Processing Unit) and a dynamic instance. The
system comprises a consumption
estimation module, an automatic parallel module, a heterogeneous training module and a
dynamic monitoring module. The
system supports sensing the computing power of the heterogeneous GPU, supports the use of a dynamic instance, automatically generates a parallel strategy and updates the strategy according to the dynamic change of the GPU; according to the method, an
assembly line construction method, a micro-batch
distribution method and a
preemption distribution
processing method are designed, each GPU group is constructed into an
assembly line, and different numbers of micro-batches are distributed, so that load balancing is achieved, and
training time is minimized; when
preemption and allocation events occur, a
fast recovery mechanism is used to recover training as soon as possible; according to the method, load balancing is achieved in a heterogeneous environment, stable training is guaranteed in a dynamic scene, parallel configuration does not need to be manually set by a user, and
large model training can be efficiently and conveniently carried out in a heterogeneous GPU and dynamic instance environment.