The invention relates to a computing power resource dynamic scheduling and monitoring method based on a cloud platform. The method is suitable for an intelligent scheduling scene of a high-performance
GPU cluster. The method comprises seven steps of task portrait modeling, GPU node state acquisition, resource
trend prediction, SLA tracking, scheduling scoring and deployment, operation monitoring and task migration, and SLA feedback optimization. According to the
system, task
semantics are represented by constructing task vectors, node
health states, topology affinity, SLA historical performance conditions and resource prediction risks are fused, a multi-factor adjustable scheduling scoring mechanism is constructed, and second-level
perception and task thermal migration of high-temperature nodes are achieved. Compared with a traditional Kubernetes static scheduling scheme, the method has the advantages that the GPU
utilization rate, the task SLA achievement rate and the
system stability are remarkably improved, the learning ability, the self-adaptive ability and the
high availability are achieved, and the method is an intelligent scheduling closed-loop
system oriented to AI reasoning and training scenes.