The invention discloses a
large model distributed training method and
system based on all-in-one
machine hybrid cloud, and relates to the technical field of
resource management. A
hybrid cloud training architecture and
secure communication are established between a local all-in-one
machine cluster and a public cloud elastic cluster; dividing training data according to sensitivity, and realizing local storage of core sensitive data and cloud side storage of non-sensitive data; dividing the training task into a core task and an extension task, and scheduling the core task and the extension task to a local training node and a cloud side training node for execution; monitoring local resource load and task progress, and elastically applying and releasing cloud side training nodes according to training stages; and limiting cross-domain transmission as necessary information, updating
model parameters through gradient aggregation, and synchronously entering the next iteration. Through the technical scheme of the invention, on the premise of ensuring that sensitive data is not out of
a domain, on-demand expansion and
recovery of cloud side computing power are realized, cross-domain communication overhead is reduced, distributed training
throughput and stability are improved, a
training period is shortened, and comprehensive cost is reduced.