The application provides a
large model instance scheduling method supporting heterogeneous computing power nodes, comprising: acquiring real-time resource state data of each computing power node in a cluster, wherein the resource state data at least includes processor occupancy, memory occupancy, accelerator occupancy and
video memory occupancy; comprehensively evaluating the real-time resource state data according to a preset weight strategy, and calculating the load
score of each computing power node, wherein a
penalty factor is applied to the computing power node without an accelerator to improve the load
score; in response to a start request of a
large model instance, selecting a node with the lowest load
score from all online computing power nodes as a target deployment node; and sending a model start instruction containing parallel configuration parameters to the target deployment node to run the
large model instance on the target deployment node. The problem of
low resource utilization under a
hybrid architecture is effectively solved, and the flexibility of large model deployment and the overall performance of the cluster are improved.