The invention provides an intelligent operation and maintenance method and
system for an industrial computing network,
electronic equipment and a storage medium, and the method comprises the steps: deploying computing power agents at heterogeneous computing power nodes to obtain operation and maintenance events in a unified manner, and carrying out the centralized
root cause analysis of the operation and maintenance events through a cloud end, so as to match an operation and
maintenance strategy, the method comprises the following steps: acquiring an operation and
maintenance strategy library, issuing the operation and
maintenance strategy to a corresponding computing power agent for execution and feeding back a result, and finally performing closed-loop updating on the operation and maintenance strategy
library and an
artificial intelligence model based on a complete event case, thereby realizing unified
perception of heterogeneous dispersed computing power nodes through the computing power agent, and solving the problem that unified management and monitoring are difficult due to resource dispersion; through automatic
root cause analysis and strategy matching, a fault
root cause can be rapidly and accurately positioned and an operation and maintenance scheme can be generated so that serious dependence on artificial experience can be eliminated and fault
processing time can be substantially shortened. Through
continuous optimization of a closed-loop learning mechanism, the operation and maintenance
automation level is greatly improved, and
high availability and stability of an industrial computing network are guaranteed.