The invention provides a
server cluster operation and maintenance method based on multi-source heterogeneous data fusion and a dynamic
knowledge graph, and the method comprises the following steps: collecting a
performance index, a log text and topological structure data of a
server cluster, splicing the performance data and the
log data based on a unified time window, and generating a multi-
modal feature sequence; and analyzing the sequence by using an unsupervised
deep learning model, constructing a dynamic health baseline, and generating a health degree portrait through the deviation with real-
time data. When an exception is detected, mapping an exception event into a dynamic
topological graph constructed based on a topological structure; analyzing a
fault propagation probability between nodes by using a graph neural network
algorithm, positioning a
root cause node, and generating a disposal strategy to execute disposal operation; and collecting the processed
recovery data as a feedback
signal, and updating the
deep learning model by using
incremental learning. The method has the beneficial effects that the fault discovery accuracy is improved, the alarm
storm is effectively inhibited, the
root cause is directly positioned, and the model self-iteration adaptability is higher.