The invention provides an AI-based multi-node rapid
disaster recovery switching method and
system, and the method comprises the steps: obtaining multi-dimensional monitoring indexes of the CPU usage rate, the memory
occupancy rate, the network IO, the disk IO, the application
response time and the error rate of each node, and constructing a
time sequence feature matrix; extracting space correlation features among nodes by using a graph convolutional network, extracting
time sequence features through a long-short-
term memory network, constructing a deep space-time graph neural network to identify a fault precursor, and predicting a node fault 30-120 seconds in advance; generating a node
health score based on a node fault prediction result and capacity
estimation, executing a target node optimization
algorithm, and selecting an optimal
disaster recovery backup node; the method comprises the following steps: constructing a hot
backup by adopting a CRI U-based lightweight
process migration technology, and generating a
check point and an incremental snapshot of a source container; and automatic
disaster recovery switching is executed before the node fails, the
service flow is redirected, and the service is ensured not to be interrupted. According to the method, the disaster
recovery switching time can be shortened to a
millisecond level, and the
system availability is remarkably improved.