The invention is suitable for the technical field of computing power
server cluster operation and maintenance and GPU
resource scheduling, and provides an intelligent operation and maintenance monitoring method and platform for a computing power
server cluster, and the method comprises the steps: determining a turning threshold value corresponding to the short-period fluctuation degree of a candidate GPU with a specified mode, and obtaining a turning threshold value corresponding to the short-period fluctuation degree of the candidate GPU; and acquiring the short-period fluctuation degree of the latest historical task sample corresponding to the candidate GPU, generating a correction factor according to the deviation degree of the latest short-period fluctuation degree relative to the turning threshold, and performing down-regulation correction on the suitability value of the candidate GPU by using the correction factor. According to the method, excessive intervention on a normal GPU is avoided, and meanwhile, the risk of
throughput degradation in the middle and later periods caused by implicit operation fluctuation accumulation when a computing task which is relatively stable in
video memory occupation requirement, relatively long in
execution time and sensitive to continuous computing power stability is processed is effectively reduced; and the stability, the accuracy and the overall operation efficiency of the scheduling decision of the computing power
server cluster are improved.