Disaster recovery system based on cloud platform
By utilizing cloud-based disaster recovery systems and modules for data acquisition, synchronization, disaster prediction, and recovery strategies, the problem of slow data recovery speed has been solved, achieving efficient data recovery and business continuity.
Patent Information
- Application Number
- CN202511116152.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-07
AI Technical Summary
Existing disaster recovery systems have bottlenecks in data recovery speed, especially when dealing with large amounts of data. Traditional backup and recovery methods take a long time, affecting the normal operation of high-availability business systems.
A cloud-based disaster recovery system is adopted. The data acquisition module acquires multi-dimensional monitoring data, the data synchronization module switches to the backup data center when the main data center fails, the disaster prediction module builds a fault prediction model, and the recovery strategy module generates data replication and recovery strategies based on the disaster prediction results and data access frequency.
It effectively reduces the amount of backup data redundancy, improves data recovery speed, reduces the risk of human error, improves disaster recovery efficiency, and ensures business continuity.
Smart Images

Figure CN120909846A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of disaster recovery technology, in particular to a cloud platform-based disaster recovery system. BACKGROUND
[0002] In the information age, with the rapid development of cloud computing, Internet of Things, big data and other technologies, the data dependency of various enterprises and institutions is increasingly enhanced, and the stability of the system and the security of the data become particularly important. Disaster recovery system refers to a system that guarantees data recovery and business continuity through a series of technical means and processes in the event of a disaster. Especially with the help of cloud computing, disaster recovery systems not only can be expanded on the basis of traditional physical machine rooms, but also can take advantage of the flexibility and efficiency of cloud platforms to achieve high availability, low cost and rapid recovery of disaster recovery systems. In the face of unpredictable situations such as natural disasters, hardware failures, network attacks, etc., cloud platforms provide key functions such as distributed backup, data redundancy, disaster switching, etc., effectively reducing the risk of data loss and business interruption. Cloud disaster recovery solution is an application solution based on cloud computing technology and disaster recovery concept, aiming to ensure data security and business continuity of enterprises in the face of catastrophic events. By deploying data backup, storage and recovery functions on the cloud, cloud disaster recovery can help enterprises quickly recover business and minimize data loss in the face of natural disasters, hardware failures, human errors and other catastrophic events.
[0003] Currently, traditional disaster recovery systems usually rely on physical servers and data centers, and need to switch data from the main data center to the backup site in the event of a disaster. Cloud platforms provide a more flexible solution, and enterprises can choose the appropriate backup method and resource scale according to actual needs. In the cloud computing environment, the disaster recovery services provided by cloud service providers include data backup, disaster recovery drills, automated recovery and other functions, thereby ensuring the high availability of services. Although cloud platforms have many advantages in disaster recovery systems, existing disaster recovery solutions still have some problems. The existing disaster recovery system has a bottleneck in data recovery speed, especially in the case of large amounts of data, traditional backup and recovery methods often take a long time to complete. This may cause a long downtime for business systems that require high availability, affecting the normal operation of the enterprise.
[0004] In order to improve these problems in existing disaster recovery systems, this paper proposes a new disaster recovery solution based on cloud platform. By optimizing the data synchronization algorithm, introducing incremental backup and real-time synchronization technology, the amount of redundant data in the backup data can be effectively reduced, and the data recovery speed can be improved. Intelligent disaster recovery management system is adopted, intelligent algorithm is used for automatic scheduling and intelligent decision of disaster recovery resources, reducing the risk of human operation and improving the efficiency of disaster recovery. SUMMARY
[0005] The application provides a cloud platform-based disaster recovery system for solving the technical problems of slow data recovery speed and insufficient system flexibility in the prior art.
[0006] In view of the above problems, the application provides a cloud platform-based disaster recovery system.
[0007] In a first aspect, the application provides a cloud platform-based disaster recovery system, comprising: a data acquisition module for acquiring multi-dimensional monitoring data and analyzing to obtain an original data set; a data synchronization module for accessing a main data center and a backup data center and connecting the backup data center when the main data center fails; a disaster prediction module for constructing a failure prediction model, importing the original data set to analyze data, generating and outputting a disaster prediction result; a data acquisition module for acquiring data access frequency; a recovery strategy module for generating a data replication strategy according to the disaster prediction result and generating a data recovery strategy according to the data access frequency; and an output module for outputting the data replication strategy and the data recovery strategy.
[0008] The one or more technical solutions provided in the application have at least the following technical effects or advantages:
[0009] The data acquisition module is used to acquire multi-dimensional monitoring data and analyze to obtain an original data set; the data synchronization module is used to access a main data center and a backup data center and connect the backup data center when the main data center fails; the disaster prediction module is used to construct a failure prediction model, import the original data set to analyze data, generate and output a disaster prediction result; the data acquisition module is used to acquire data access frequency; the recovery strategy module is used to generate a data replication strategy according to the disaster prediction result and generate a data recovery strategy according to the data access frequency; and the output module is used to output the data replication strategy and the data recovery strategy, which can effectively reduce the redundancy of backup data, improve the data recovery speed, reduce the risk of human operation, and improve the efficiency of disaster recovery.
[0010] The above description is only a summary of the technical solutions of the application. In order to more clearly understand the technical means of the application, and in order to make the above and other purposes, features and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0012] Figure 1A cloud platform-based disaster recovery system diagram provided in the present application; DETAILED DESCRIPTION
[0013] The present application provides a cloud platform-based disaster recovery system to solve the technical problems of slow data recovery speed and insufficient system flexibility in the prior art.
[0014] The technical solutions in the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application. In addition, it should be noted that, for the convenience of description, only parts related to the present application are shown in the drawings, not all.
[0015] Embodiment one
[0016] As shown in the present application, a cloud platform-based disaster recovery system is provided, comprising: Figure 1
[0017] A data collection module is configured to collect multi-dimensional monitoring data and analyze to obtain a raw data set.
[0018] In the present application, the data collection module is a core component for obtaining and analyzing multi-dimensional monitoring data. Specifically, it collects data from different sources, integrates and provides them to the subsequent analysis system, so as to perform fault diagnosis, performance evaluation and system optimization, etc. The goal of this module is to form a complete raw data set through comprehensive data collection, providing a basis for subsequent data processing, analysis and decision-making.
[0019] The data collection module includes: a historical fault data collection unit for collecting historical fault data to obtain first data parameters; a monitoring data collection unit for collecting monitoring data to obtain second data parameters; a server performance data collection unit for collecting server performance data to obtain third data parameters; and a data integration unit for integrating the first data parameters, the second data parameters and the third data parameters as a raw data set.
[0020] In the embodiments of the present application, the main function of the historical fault data collection unit is to obtain relevant data from historical fault records. These data usually include information such as the time when the system failed in the past, the type of failure, the frequency of failure, the time to repair the failure, etc. Through the analysis of these historical data, the system can identify potential failure modes, frequent failures and failure causes, etc., to provide a basis for future failure prevention and rapid response. The monitoring data collection unit is responsible for collecting monitoring data in real time during system operation. These data usually cover various indicators of the system. For example, CPU usage, memory usage, network traffic, disk space, temperature, etc. The server performance data collection unit mainly focuses on the running status of server hardware and software. This includes the health status of hardware resources, such as the working temperature of the CPU, the usage of the memory, the status of the hard disk, the health of the network interface, etc. The core task of the data processing unit is to process and integrate the raw data collected from the above three units. The data integration unit extracts valuable features and information from historical fault data, monitoring data and server performance data, and integrates them into a raw data set. This data set provides a sufficient basis for subsequent analysis and decision-making. Optionally, the data collection module can flexibly adjust the frequency and range of collection according to the scale and actual needs of the disaster recovery system, to better adapt to different application scenarios.
[0021] The data synchronization module is used to access the main data center and the backup data center, and connect the backup data center when the main data center fails.
[0022] In the embodiments of the present application, the main function of the data synchronization module is to ensure the synchronization of data between the main data center and the backup data center, and to ensure that the data of the two is always consistent. Specifically, when the main data center fails, the module can automatically connect to the backup data center to ensure business continuity. Its working principle is based on data difference identification and incremental synchronization mechanism. In terms of technical implementation, a secure transmission protocol is used to ensure the integrity and confidentiality of data during cross-center transmission, and to support efficient pushing of massive data. When the main center fails, the module activates the backup center connection channel to provide the latest data basis for subsequent business takeover, which is directly related to the recovery of the disaster recovery system.
[0023] The data synchronization module includes: a monitoring unit for monitoring the running status of the main data center and triggering a disaster recovery switching instruction when detecting a failure of the main data center; a data analysis unit for collecting and analyzing business data and generating data change records; and a process switching unit for executing a disaster recovery switching process to direct business traffic to the backup data center to take over the business.
[0024] In the embodiments of the present application, the data synchronization module has three units. Specifically, the task of the monitoring unit is to monitor the running state of the main data center in real time, including hardware performance, network condition, service health condition, etc. If it is found that the main data center fails, the unit will automatically trigger the disaster recovery switching process to switch the traffic to the backup data center. The data analysis unit is responsible for collecting and analyzing the business data in the main data center to generate data change records. These records are crucial for subsequent disaster recovery and business continuity. When the main data center fails, the change records provided by the data analysis unit can help the backup data center quickly recover to the latest state. The process switching unit is responsible for executing the disaster recovery switching process to ensure that when the main data center fails, the business traffic can be quickly and smoothly switched to the backup data center. For example, switching business systems, databases, network connections, etc. to ensure that the backup data center can take over all tasks of the main data center. Optionally, the process switching unit can support both automatic and manual switching modes. In the automatic mode, the system automatically performs switching operations according to the preset fault triggering conditions. In the manual mode, the system administrator can intervene and control the switching process according to the actual situation. In addition, the process switching unit may also need to support real-time monitoring after traffic switching to ensure that the backup data center can smoothly take over and run stably.
[0025] The disaster prediction module is used to build a fault prediction model, import and analyze the original data set, and generate and output disaster prediction results.
[0026] In the embodiments of the present application, the disaster prediction module is responsible for processing the original data set and establishing and outputting the results of disaster prediction. Specifically, through the analysis of these data, the health status of the device or system can be predicted, and potential fault risks can be identified in advance. The disaster prediction model is modeled based on deep learning technology. For example, a long short-term memory network can be used to predict anomalies or faults in time series based on past operation data to predict future states. Intelligent algorithms predict potential disasters and fault points and deploy disaster recovery resources in advance. Through monitoring system performance, network traffic, hardware health status, etc., the intelligent algorithm analyzes possible risks and triggers the recovery process in advance. Optionally, the intelligent algorithm can also optimize the backup strategy to determine the optimal backup time and frequency.
[0027] The disaster prediction module includes a data preprocessing unit for featureizing data as (x, y), where x represents an input feature vector, i.e. the observation data of the system, and y e {0, 1} is a label, 0 representing normal and 1 representing abnormal; a modeling unit for building a deep learning model to analyze data, and the model formula is:
[0028]
[0029] where θ represents the parameters of the model. is the output of the model, i.e., the predicted label of the input ; a model training unit configured to bring the plurality of groups of original data into the model to adjust model parameters θ to minimize a loss function; and a model output unit configured to set a risk threshold, wherein if the output is less than the risk threshold, the output is normal, and if the output is greater than the risk threshold, it indicates that the system may be abnormal or malfunctioning.
[0030] In the embodiments of the present application, the data preprocessing unit is responsible for cleaning, normalizing, denoising, and feature engineering preprocessing of the original data to ensure that the data is suitable for input into the model, i.e., the feature vector (x) and the label (y). The main task of the modeling unit is to build a deep learning model for classifying data and outputting a predicted label of normal or abnormal. The purpose of model training is to adjust the parameters θ in the model by inputting a large number of data samples, so that the prediction error of the model is minimized. The core of model training is to guide model optimization through a loss function. The model output unit outputs the prediction result according to the trained model, and determines the system state according to the set risk threshold. The setting of the risk threshold needs to be adjusted according to the specific application scenario, and the selection of the risk threshold can be based on the probability value output by the model. In an exemplary embodiment, the selection of the threshold usually depends on the business requirements, but the false positive and false negative risks need to be understood in advance.
[0031] A recovery strategy module configured to generate a data replication strategy according to the disaster prediction result and a data recovery strategy according to the data access frequency.
[0032] In the embodiments of the present application, the recovery strategy module generates a strategy based on the disaster prediction result and the data access frequency. Specifically, the module can predict potential disaster events and adjust the recovery strategy according to the predicted risk level. According to the data access frequency, it can identify which data is frequently accessed business critical data and ensure that these data can be recovered first in the event of a disaster. Through the prediction and analysis mechanism, the data is effectively backed up before the disaster occurs, and in the recovery process, the most critical data is recovered first, minimizing the business interruption time. Based on the disaster prediction result and the characteristics of the business data, the data replication strategy and the recovery strategy that adapt to the current risk state are dynamically generated, balancing the disaster recovery resource consumption and the business continuity guarantee capability.
[0033] The recovery strategy module comprises: a control unit for adjusting backup frequency and recovery priority based on prediction results and data access frequency; a risk level assessment unit for evaluating production environment stability through the model and outputting a risk level; a data backup unit for setting a high-risk threshold, and enabling full backup when the risk level is lower than the high-risk threshold; a data backup unit for generating a data replication mode upgrade instruction to increase backup frequency and reduce backup time interval when the risk level reaches the high-risk threshold; and a data recovery unit for analyzing business data based on data access frequency, marking data with a data access frequency higher than a frequency threshold as business critical data in combination with a preset frequency threshold, and performing a priority recovery mark.
[0034] In the embodiments of the present application, the risk level assessment unit evaluates the stability of the production environment through a predetermined model and outputs the corresponding risk level. This unit evaluates the stability of the current system by real-time monitoring of environmental indicators and combining a disaster prediction model. The risk level can be divided into three levels: low risk, medium risk, and high risk, and the backup strategy and recovery priority are automatically adjusted according to the evaluation results. High risk is set as the high-risk threshold, and when the risk level is low risk or medium risk, full backup is enabled. Full backup means that all data will be backed up, and this method is suitable for low-risk situations to ensure data integrity and security. At this time, the system can ensure that every piece of data has a backup, but it will occupy more storage space and resources. When the risk level of the system is in a high-risk state, the system will automatically increase the backup frequency, shorten the backup interval, and may use incremental backup or differential backup to reduce the impact on system performance. For example, the backup frequency is increased from once a day to once an hour. At this time, the goal is to ensure that all critical data is backed up and protected before a disaster occurs. The frequency threshold is generated by collecting database access logs and time series statistics to generate a "data heat profile", and the frequency threshold is configured differently according to business types. The data recovery module determines which data is critical to business operation according to the access frequency threshold. The system analyzes the data access pattern to determine which data is more important in actual business, so that after a disaster occurs, these high-frequency access data can be recovered first, reducing business interruption time. For example, through a machine learning model, the system can automatically detect potential failures and start an automated repair process. The cloud platform can automatically transfer traffic to a healthy node and start data recovery operations when it detects that a node may crash. Dynamically optimize disaster recovery strategies, automatically allocate resources based on business priority, intelligently switch paths in case of failure, and reduce manual intervention time.
[0035] The steps of the methods or algorithms described in this application can be directly embedded in hardware, a software unit executed by a processor, or a combination of both. The software unit can be stored in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other storage medium of any form in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and storage medium can be disposed in an ASIC, which can be disposed in a terminal. Optionally, the processor and storage medium can also be disposed in different components within the terminal. These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0036] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely illustrative examples of this application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Thus, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A cloud platform-based disaster recovery system, characterized in that, The system comprises: a data acquisition module for acquiring multi-dimensional monitoring data and analyzing to obtain an original data set; a data synchronization module for accessing a main data center and a backup data center and connecting the backup data center when the main data center fails; a disaster prediction module for constructing a failure prediction model, importing the original data set to analyze data, and generating and outputting a disaster prediction result; a data acquisition module for acquiring data access frequency; a recovery strategy module for generating a data replication strategy according to the disaster prediction result and generating a data recovery strategy according to the data access frequency; an output module for outputting the data replication strategy and the data recovery strategy.
2. The method of claim 1, wherein, The data acquisition module comprises: a historical failure data acquisition unit for acquiring historical failure data to obtain first data parameters; a monitoring data acquisition unit for acquiring monitoring data to obtain second data parameters; a server performance data acquisition unit for acquiring server performance data to obtain third data parameters; a data integration unit for taking the first data parameters, the second data parameters, and the third data parameters as an original data set.
3. The method of claim 1, wherein, The data synchronization module comprises: a monitoring unit for monitoring the running state of a main data center and triggering a disaster recovery switching instruction when detecting a failure of the main data center; a data analysis unit for acquiring and analyzing business data to generate a data change record; a process switching unit for executing a disaster recovery switching process and directing business traffic to a backup data center to take over the business.
4. The method of claim 1, wherein, The disaster prediction module comprises: a data preprocessing unit for characterizing data as (x, y), where x represents an input feature vector, i.e., observation data of the system, and y ∈ {0, 1} is a label, 0 representing normal and 1 representing abnormal; a modeling unit for constructing a deep learning model to analyze data, with a model formula being:
5. where θ denotes the parameters of the model, is the output of the model, i.e. the predicted label of the input a model training unit for bringing multiple groups of original data into the model to adjust model parameters θ to minimize a loss function; a model output unit for setting a risk threshold, outputting normal if the output is less than the risk threshold, and indicating that the system may have an abnormality or failure if the output is greater than the risk threshold.
6. The method of claim 1, wherein, The recovery strategy module comprises: a control unit for adjusting backup frequency and recovery priority based on the prediction result and data access frequency; a risk level evaluation unit for evaluating production environment stability through the model and outputting a risk level; a data backup unit for setting a high-risk threshold and enabling full backup when the risk level is below the high-risk threshold; a data backup unit for generating a data replication mode upgrade instruction, increasing the number of backups, and reducing the backup time interval when the risk level reaches the high-risk threshold; a data recovery unit for analyzing business data based on data access frequency, combining a preset frequency threshold to judge data with a data access frequency higher than the frequency threshold as business critical data, and marking for priority recovery.
Citation Information
Patent Citations
Disaster recovery system based on doris data synchronization
CN118051386B
Intelligent data protection method and system
CN118819964B
A risk prediction method and skid-mounted system for transmission pipelines based on big data
CN118821028B
Archive information management system based on cloud
CN119829817A
Backing up audio and video files across mobile devices of a user
US8805790B1