The invention belongs to the technical field of
data processing, and discloses an active intelligent operation and maintenance monitoring method for a
data center, which comprises the following steps: step 1,
component modeling and topology configuration; 2, carrying out distributed health detection and
data acquisition; 3, performing real-time
health assessment and
anomaly detection; 4, performing intelligent alarm and
root cause analysis; 5, unified operation and maintenance and closed-
loop control are carried out; and step 6, dynamically optimizing the intelligent operation and
maintenance strategy. According to the method, a monitoring object and a dependency relationship are clarified through
component modeling and topological configuration, and specific scenes, such as multi-mode acquisition,
active detection, index pulling, log analysis, coverage
message queue theme accumulation and
database connection pool exhaustion, of distributed health detection are combined. The quantitative
health score is calculated based on the preset scoring model, and by combining with the dynamic baseline learned by the ARIMA or LSTM
algorithm, the module abnormity can be actively detected in different periods and weekly updating, and the problems of fault discovery lagging and incomplete monitoring coverage are solved.