In various examples, systems, devices and methods are disclosed relating to management of
processing resources and workloads assigned thereto. A
system can obtain, from a plurality of
processing resources executing a plurality of tasks,
telemetry data and task assignment data. The
system can perform generate, using the
telemetry data and the task assignment data, a plurality of feature vectors, determine, using the plurality of feature vectors and a
machine learning model employing a graph
attention network (GAT) having a plurality of nodes, a performance state of a node of the plurality of nodes, and determine, based on the performance state of the node, an action to be taken to enhance performance of the plurality of
processing resources or mitigate node failures. Each node can represent one or more respective processing resources and each
feature vector can be associated with a respective node of the plurality of nodes.