Data backup method, device, equipment, system and storage medium

By performing data detection on the first cluster and selecting the target backup mechanism in the snapshot mechanism and log mechanism, the target data is backed up, and the problem of failure to effectively reduce the impact of data loss on the business side during the data backup process in the prior art is solved, and the impact of data loss on the business is reduced during the data backup process.

CN114356650BActive Publication Date: 2025-05-02IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111397882.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-19
Publication Date
2025-05-02
Estimated Expiration
2041-11-19

AI Technical Summary

Technical Problem

The prior art fails to effectively reduce the impact of data loss on the service side during data backup, resulting in the inability to carry out the business normally when data is accidentally lost.

Method used

By responding to the backup instruction, data detection is performed on the first cluster, data detection results are obtained, and target backup mechanism is selected in the snapshot mechanism and log mechanism based on the detection results, target data is backed up, and backup it to the second cluster.

Benefits of technology

In the process of data backup, the impact of data loss on user-side services is minimized. Through independent main cluster and backup cluster design, the direct impact of data loss on business operations is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114356650B_ABST
    Figure CN114356650B_ABST
Patent Text Reader

Abstract

The present application discloses a data backup method, device, equipment, system and storage medium, wherein the data equipment method includes: in response to a backup instruction, performing data detection on a first cluster to obtain a data detection result; wherein the data detection result includes whether the target data is stored in the first cluster; based on the data detection result, selecting a target backup mechanism from a plurality of preset backup mechanisms; and the plurality of preset backup mechanisms at least include a snapshot mechanism and a log mechanism; based on the target backup mechanism, performing a backup operation on the target data to back up the target data to a second cluster. Through the above-mentioned method, the present application can minimize the impact of data loss on user-side services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data backup technology, and in particular to a data backup method, device, equipment, system and storage medium. Background Art

[0002] With the continuous development of virtualization technology, users are particularly concerned about the loss of key data in cloud hosts in virtualization scenarios. At the same time, in a hybrid multi-cloud environment, it is not limited to data backup, but also needs to consider reducing the impact on business operations while backing up and quickly switching to the backup machine when the main cluster fails, while minimizing performance impact, simplifying management, reducing storage costs, and ensuring more effective and efficient business operations.

[0003] At present, the main purpose of existing data protection solutions is data backup and recovery, but they do not focus on how to minimize the impact of data loss on the business side. Therefore, when data is accidentally lost due to force majeure and other factors, it may directly cause the business side to be unable to operate normally. In view of this, how to minimize the impact of data loss on user-side business has become an urgent problem to be solved. Summary of the invention

[0004] The main technical problem solved by the present application is to provide a data backup method, device, equipment, system and storage medium, which can minimize the impact of data loss on user-side services.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a data backup method, including: in response to a backup instruction, performing data detection on a first cluster to obtain a data detection result; and the data detection result includes whether the target data is stored in the first cluster; then based on the data detection result, selecting a target backup mechanism from a plurality of preset backup mechanisms; and the plurality of preset backup mechanisms include at least a snapshot mechanism and a log mechanism; then based on the target backup mechanism, performing a backup operation on the target data to back up the target data to a second cluster.

[0006] In order to solve the above-mentioned technical problems, the second aspect of the present application provides a data backup device, including: a detection module, which is used to respond to a backup instruction, perform data detection on a first cluster, and obtain a data detection result; the data detection result includes whether the target data is stored in the first cluster; a selection module, which is used to select a target backup mechanism from a plurality of preset backup mechanisms based on the data detection result; the plurality of preset backup mechanisms include at least a snapshot mechanism and a log mechanism; a backup module, which is used to perform a backup operation on the target data based on the target backup mechanism, so as to back up the target data to a second cluster.

[0007] In order to solve the above technical problems, the third aspect of the present application provides a data backup device, including a memory, a communication circuit and a processor, and the memory and the communication circuit are coupled to the processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the data backup method in the above first aspect.

[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a data backup system, including a data backup device and multiple clusters, the data backup device is respectively communicated with the multiple clusters, and the data backup device is used to execute the data backup method in the above first aspect to realize data backup between multiple clusters.

[0009] In order to solve the above technical problem, the fifth aspect of the present application provides a computer-readable storage medium, which stores program instructions that can be executed by a processor, and the program instructions are used to implement the data backup method in the above first aspect.

[0010] In the above scheme, in response to the backup instruction, data detection is performed on the first cluster to obtain data detection results, the data detection results include whether the target data is stored in the first cluster, and then based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and the plurality of preset backup mechanisms at least include a snapshot mechanism and a log mechanism, and then based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster. On the one hand, there is no business association between the first cluster and the second cluster, that is, the two are independent of each other, and a failure in any cluster will not affect the normal operation of the other cluster. On the other hand, based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and then the target data is backed up. The flexible selection of multiple backup mechanisms helps to reduce data loss. Therefore, the impact of data loss on user-side business can be minimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a flowchart of an embodiment of the data backup method of the present application;

[0012] Figure 2 This is an architectural diagram of an embodiment of the data backup method of the present application;

[0013] Figure 3 yes Figure 1 The architecture diagram of an embodiment of step S12;

[0014] Figure 4 yes Figure 1 The architecture diagram of another embodiment of step S12;

[0015] Figure 5 It is a schematic diagram of the framework of an embodiment of the data backup device of the present application;

[0016] Figure 6 It is a schematic diagram of the framework of an embodiment of the data backup device of the present application;

[0017] Figure 7 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0018] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.

[0019] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0020] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.

[0021] See also Figure 1 , Figure 1 It is a flowchart of an embodiment of the data backup method of the present application.

[0022] Specifically, the following steps may be included:

[0023] Step S11: In response to the backup instruction, data detection is performed on the first cluster to obtain a data detection result.

[0024] In one implementation scenario, the backup instruction may be sent by the user when the user device stores the target data in the first cluster; or, the backup instruction may be automatically sent by the user device based on a preset backup policy when the user device stores the target data in the first cluster. When the backup instruction is sent by the user device when the target data is stored in the first cluster, it means that the user is performing a write operation on the data; when the backup instruction is automatically sent by the user device when the target data is stored in the first cluster 200 based on a preset backup policy, it means that the backup policy is automatically generated and sent automatically by clicking the backup button in the data backup device 208. The backup policy can determine the backup time and method. For example, the backup policy can be set to real-time backup, which means that the data is backed up if a write operation is performed on the data. In the above manner, the sending of the two backup instructions can avoid the problem of data loss caused by the user forgetting to perform a backup operation on the data during the data storage process, further reducing the loss of data.

[0025] In a specific implementation scenario, the write operations performed by users on data can be divided into add operations, change operations, and delete operations. Among them, the add operation is the process of recording new data by the user. After the user completes adding the data, the backup operation will only be performed on the currently stored data. The change operation may occur after the data is backed up. Therefore, after the backup process of the changed data is performed, if you need to keep a copy of the data before the change, you can choose to re-save it. The re-saved data will not overwrite the data before the change. When the data before the change needs to be searched or restored, it will not affect the user's normal use and query. The deletion operation may also occur after the data is backed up, so you can choose to save directly and overwrite the previous data, or re-save and keep a copy of the data before the deletion. The specific situation is not limited here, and you can choose according to the needs of the actual application.

[0026] In a specific implementation scenario, the data backup device may include a cloud management platform, a backup agent, and a backup component. The display interface in the cloud management platform can be displayed separately on different servers through a network connection, and the cloud management platform can be installed in any server in the cluster, or in any server device connected to the cluster. The specific situation is not limited here, and can be selected according to the needs of the actual application. The instructions issued by the cloud management platform are sent out through the backup agent, and the backup agent will perform a check on the received message. If the backup agent receives a duplicate message sent by the cloud management platform, the backup agent can determine whether the received message is a resent message or a resent new message after detection. The backup component has a detection function and detects the first cluster and the second cluster. The cloud management platform, backup agent, and backup components can be installed in the same server or in different servers. The specific situation is not limited here, and can be set according to the actual situation.

[0027] In one implementation scenario, before performing data detection on the first cluster in response to a backup instruction and obtaining the data detection result, metadata of the first cluster and the second cluster may be obtained, and the metadata may include a first network interface of the first cluster and a second network interface of the second cluster, the first network interface being used to implement data transmission with the first cluster, and the second network interface being used to implement data transmission with the second cluster. The above method can, on the one hand, implement data transmission between the data backup device and the cluster, and on the other hand, can ensure that when a cluster is reconnected, no storage system in the cluster will be missed.

[0028] In a specific implementation scenario, metadata may include but is not limited to information such as name, specification, network, security group, volume mount point, etc. The metadata mainly records the definition of the first cluster and the second cluster model, the mapping relationship between each level, the data status of the monitoring data warehouse and the task running status. Obtaining metadata enables the deployment, operation and management of the first cluster and the second cluster to achieve coordination and consistency. Among them, the volume mount point in the metadata is the entry directory of the disk file system, and the entry directory can be the first network interface or the second network interface, that is, the first network interface and the second network interface can access and transmit the data in the first cluster and the second cluster, thereby realizing data transmission between the data backup device and the first cluster and the second cluster.

[0029] In an implementation scenario, the first cluster can be used as the main cluster, and the second cluster can be used as the backup cluster; of course, the first cluster can also be used as the backup cluster and the second cluster can be used as the main cluster according to business needs. The application of the first cluster and the second cluster is not specifically limited here, and can be set according to the needs of actual applications. Here, the main cluster is the cluster that performs data preservation, and the backup cluster is the cluster that performs backup preservation of the saved data. The underlying storage in the first cluster and the second cluster adopts distributed storage. Compared with the centralized storage used in the underlying storage in the prior art, the distributed storage has simpler operation and maintenance deployment, and can better and more uniformly utilize low-cost disks, ultimately providing a higher capacity storage pool. The distributed storage has better portability, which can reduce the impact of node crashes on the storage system and improve the user experience.

[0030] In one implementation scenario, data detection on the first cluster is performed by using a backup component in a data backup device to detect the data stored in the first cluster. The backup component detects the data in the first cluster to ensure that no unbacked up data is missed during the data backup process, thereby reducing data loss.

[0031] In the disclosed embodiment, the data detection result includes whether the target data is stored in the first cluster. Specifically, the target data is data that needs to be backed up, that is, if the data detection result is stored in the target data, it indicates that data that has not been backed up was previously stored, and the target data includes all data that has not been backed up; if the data detection result is not stored in the target data, it indicates that data that has not been backed up was not previously stored, and the target data is write data.

[0032] Step S12: Based on the data detection result, a target backup mechanism is selected from a plurality of preset backup mechanisms.

[0033] In an implementation scenario, the preset backup mechanism can be set according to the actual situation. For example, if the data backup method is used for a server, it can be set to a full backup mode, which can systematically back up all data; if the data backup method is used for a large trading market, it can be set to an incremental backup, which is to back up the newly changed data since the last backup operation. These newly changed data are either newly generated data or updated data. The backup event required by the incremental backup is the shortest. The incremental backup method will effectively save storage space. At the same time, when data is lost, it can be quickly restored from the backup data; differential backup can also be selected. Differential backup requires more data to be backed up. It is the result of two comparisons before and after, and different parts of the data are backed up; or it can be set to selective backup or instant backup, which can automatically perform the backup operations listed above according to the defined schedule. However, the data center operation and maintenance personnel sometimes need to back up data instantly according to the situation. The backup data should be checked frequently. When the data is found to be missing or incorrect, it is necessary to take the initiative to perform an instant backup and refresh the backup data to ensure that it is consistent with the data in the running data center.

[0034] In a specific implementation scenario, the preset backup mechanism can be a snapshot mechanism. A snapshot is a state record of data storage at a certain moment, that is, the storage state at a certain moment is recorded and restored to obtain the stored data. The data can be restored to the time point of the snapshot within a few seconds using the snapshot image, and the system administrator can selectively and quickly restore damaged or deleted files. The data snapshot function has many uses. For example, if a copy of the latest production data is needed to test a new system or provide decision support and data analysis, and the system cannot be shut down, and it takes a long time to restore a copy of data using tape backup, in this case, the backup function of the data snapshot can be used to create a snapshot copy at any time point, and the copied data can be used for testing and analysis without affecting the normal use of the system.

[0035] In a specific implementation scenario, the preset backup mechanism can be a log mechanism. The log backup mechanism will record the operations on the stored data. The stored data content can be restored based on the operation records. Moreover, since all operation records on the data are stored, data loss during the backup process is reduced as much as possible, thereby improving the efficiency of data backup.

[0036] In one implementation scenario, a target backup mechanism can be selected from a variety of preset backup mechanisms based on the data detection results. For example, when the detection results include that the target data is stored in the first cluster, a snapshot mechanism can be selected as the target backup mechanism; when the detection results include that the target data is not stored in the first cluster, a log mechanism can be selected as the target backup mechanism. In the above method, by selecting different backup mechanisms for different detection results, the data stored in the cluster is quickly backed up, further reducing the situation where data is not backed up due to data confusion, thereby causing data loss.

[0037] Step S13: Based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster.

[0038] See also Figure 2 , Figure 2 This is a diagram of the architecture of an embodiment of the data backup method of the present application. Figure 2As shown, at this time, the first cluster 200 is the main cluster, and the second cluster 212 is the backup cluster. The first cluster 200 and the second cluster 212 interact with each other through the data backup device 208 and realize data backup. The first cluster 200 includes a cloud host (1) 201, Qemu-kvm (1) 202, and a first cluster storage system 203. The cloud host (1) 201 is a virtual machine. For example, a virtual machine is a cloud host. After the user logs in, a file is created. After the file is created, data is written. After the data is written, the data needs to be stored. The data can be stored selectively, that is, the data is stored on different disks. Qemu-kvm (1) 202 is based on hardware virtualization technology and combines the device virtualization function provided by Qemu to realize the virtualization of the entire system. When the first cluster storage system 203 performs data backup, after receiving the backup instruction, it allows the underlying system to perform data backup. The first cluster storage system 203 includes a controller (1) 204, an operation log 205, a storage disk 206, a storage pool I (1) 207 and a storage pool II (1) 207. The controller (1) 204 can be used as a protocol. If new data is written, the new data can be written to a designated storage disk for storage. The operation log 205 stores operation logs. When the backup mechanism uses the log mechanism, the operation log is used to store data operation records, that is, records of new data writing, previously stored data being changed, or stored data being deleted. Data can be restored through operation records. The storage disk 206 can be used as a storage system. That is, when data is written, the data is stored in the storage disk 206 in the absence of a backup instruction. When the storage space in the storage disk 206 is insufficient, the storage space can be expanded through the storage pool I (1) 207 or the storage pool II (1) 207. The data backup device 208 includes a cloud management platform 209, a backup agent 210 and a backup component 211. The cloud management platform 209 includes a display interface, which supports user login. After successful login, the display interface presents different function selections. For example, if you click the backup button, different backup mechanisms can be selected; if the first cluster 200 and the second cluster 212 are switched, the display interface has a migration button, and the data can be backed up by selecting migration to minimize data loss. All instructions sent by the cloud management platform 209 are sent through the backup agent 210, and the backup agent 210 can receive the interaction information between the first cluster 200 and the second cluster 212; after receiving the backup instruction from the backup agent 210, the backup component 211 detects the data and executes different backup mechanisms. The second cluster 212 is similar to the first cluster 200 in structure, except that when the second cluster 212 is a backup cluster, the data storage is slightly different from that of the first cluster 200.The second cluster 212 includes a cloud host (2) 213, Qemu-kvm (2) 214, and a second cluster storage system 215. The second cluster storage system 215 includes a controller (2) 216, a storage disk I 217, a storage disk II 218, a basic storage disk 219, a storage pool I (2) 220, and a storage pool II (2) 220. Here, the controller (2) 216 can be used as a protocol. The storage disk I 217, the storage disk II 218, and the basic storage disk 219 are all storage systems. For example, the storage disk I 217 and the storage disk II 218 represent the data to be re-incrementally written later. The data written to the storage disk I 217 is "ABC", and the data written to the storage disk II 218 is "def". Different written data can be stored in different storage systems. The data backup device 208 selects different data storage methods for different backup mechanisms, sends the data to the second cluster 212, and stores it.

[0039] In a specific implementation scenario, the target backup mechanism is a snapshot mechanism. Based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster 212, including performing a snapshot mark on the target data in the first cluster 200 to obtain a data copy of the target data; then restoring the target data based on the data copy of the target data, and sending the target data to the second cluster 212. In the above manner, the target data can be quickly backed up in a short time. The snapshot backup mechanism is a snapshot backup of the data content of the storage space at a certain time point, which can minimize the problem of data omission caused by writing data while backing up, and the snapshot mechanism does not require a complete data space, thereby achieving the purpose of not only fast backup but also saving storage space.

[0040] See also Figure 2 and Figure 3 , Figure 3 yes Figure 1The architecture diagram of an embodiment of step S12 in the figure. The cloud host (1) is a virtual machine in the first cluster 200. The user can write data through the cloud host (1). When the user device sends a backup instruction when storing target data in the first cluster 200 or automatically sends a backup instruction based on a preset backup policy, the data backup device 208 sends the backup instruction to the backup component 211 through the backup agent 210. After receiving the backup instruction, the backup component 211 performs data detection on the first cluster 200. The backup component 211 detects the data stored in the storage disk 206 in the first cluster 200. If the detection result shows that the storage disk 206 contains the target data, the snapshot mechanism is selected for backup. When the snapshot mechanism is selected for the data backup mechanism, the controller (1) 204 saves the user-written data to the storage disk 206 according to the log snapshot. After the saving is successful, the storage disk 206 returns a confirmation instruction to the controller (1) 204. After receiving the confirmation instruction, the controller (1) 204 returns the confirmation instruction to the user device. The confirmation instruction presented to the user device can be "saved to **". The specific content can be set according to the actual situation. After the controller (1) 204 returns the confirmation instruction to the user device, the backup component 211 detects the storage disk 206 in the first cluster 200 based on the snapshot backup mechanism, detects that there is an unbacked up data volume in the storage disk 206, and performs a snapshot backup of the data volume content, that is, the state of the data storage at this moment is snapshot-marked, and then a data copy of the target data is obtained. The backup component 211 restores the data based on the data copy and sends it to the second cluster 212. After the backup component 211 sends the data to the second cluster 212, the data copy in the first cluster 200 is deleted; after receiving the data content sent by the backup component 211, the second cluster 212 stores the data in any storage disk such as the storage disk II. After the data backup device 208 executes the snapshot backup mechanism to back up the data, if the backup instruction is not interrupted, the user writes the data again, and the log backup mechanism can continue to perform the backup to minimize the loss of data.

[0041] In a specific implementation scenario, after the target data is recovered based on the data copy of the target data and the target data is sent to the second cluster 212, the data copy in the first cluster 200 can be deleted. The above method avoids the situation where the data is restored again and resent to the second cluster 212 by deleting the data copy, thereby improving the efficiency of data backup.

[0042] In a specific implementation scenario, the target backup mechanism is a log mechanism, and the first cluster 200 also stores an operation log, which is used to store data operation records. Based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster 212, including copying the data operation records of the target data from the first cluster 200, and restoring the target data based on the data operation records of the target data, and sending the target data to the second cluster 212. In the above manner, data is backed up through the log backup mechanism, so that data can be quickly restored, and the time required for restoring data is minimized, further reducing data loss.

[0043] See also Figure 2 and Figure 4 , Figure 4 yes Figure 1 The architecture diagram of another embodiment of step S12 in the figure. The cloud host (1) is a virtual machine in the first cluster 200. The user can write data through the cloud host (1). When the user device sends a backup instruction when storing target data in the first cluster 200 or automatically sends a backup instruction based on a preset backup policy, the data backup device 208 sends the backup instruction to the backup component 211 through the backup agent 210. After receiving the backup instruction, the backup component 211 performs data detection on the first cluster 200. The backup component 211 detects the data stored in the storage disk 206 in the first cluster 200. If the detection result shows that the target data is not stored in the storage disk 206, the log mechanism is selected for backup. When the data backup mechanism is selected as the log mechanism, the controller (1) 204 first records the user's data storage operation in the operation log according to the log mechanism. After the content of the operation log is recorded, a confirmation instruction will be returned to the controller (1) 204. After receiving the confirmation instruction, the controller (1) 204 returns the confirmation instruction to the user device. The confirmation instruction presented to the user device can be "saved to **". The specific content can be set according to the actual situation. After the controller (1) 204 returns the confirmation instruction to the user device, the original data written by the user is stored in the storage disk 206; the backup component 211 detects the operation log based on the log backup mechanism, that is, detects whether there is an operation record in the operation log 205. If there is an operation record, the backup component 211 obtains the operation record content, and restores the data content according to the operation record, and then sends the restored data to the second cluster 212. After the backup component 211 sends the restored data to the second cluster 212, the operation record of the operation log 205 in the first cluster 200 is deleted to ensure that the same operation record will not be restored multiple times; after receiving the data content sent by the backup component 211, the second cluster 212 stores the data in any storage disk such as the storage disk 1. When the user device continuously writes data, the data backup device 208 can continuously restore the data to the second cluster 212 to minimize data loss.

[0044] In a specific implementation scenario, after sending the target data to the second cluster 212, the method also includes deleting the data operation record of the target data in the first cluster 200. The above method ensures that after the target data is backed up to the second cluster 212, the data backup device 208 will not restore the target data again, thereby avoiding repeated recovery of data and reducing the space occupied by multiple storage of data, thereby maximizing the efficiency of data backup.

[0045] In the above scheme, in response to the backup instruction, data detection is performed on the first cluster to obtain data detection results, the data detection results include whether the target data is stored in the first cluster, and then based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and the plurality of preset backup mechanisms at least include a snapshot mechanism and a log mechanism, and then based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster. On the one hand, there is no business association between the first cluster and the second cluster, that is, the two are independent of each other, and a failure in any cluster will not affect the normal operation of the other cluster. On the other hand, based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and then the target data is backed up. The flexible selection of multiple backup mechanisms helps to reduce data loss. Therefore, the impact of data loss on user-side business can be minimized.

[0046] In some disclosed embodiments, the first cluster 200 is a primary cluster, and the second cluster 212 is a backup cluster. The method further includes, in response to a switching instruction, switching the first cluster 200 to a backup cluster, and switching the second cluster 212 to a primary cluster, or, in response to a migration instruction, migrating data from the backup cluster to the primary cluster based on a snapshot mechanism. In the above manner, by switching the first cluster 200 and the second cluster 212, the impact on the user side is minimized, thereby improving the user experience.

[0047] In one implementation scenario, the first cluster 200 can be the main cluster, and the second cluster 212 can be the backup cluster. When the first cluster 200 fails, the main-backup switching can be performed, that is, the first cluster 200 is switched to the backup cluster, and the second cluster 212 is switched to the main cluster. For the switching between clusters, not only can the user device switch the main-backup cluster through the cloud management platform 209, but also based on the data storage process, when the data backup device 208 fails to access the cluster or read the data, the data backup device 208 can issue a switching instruction based on the cloud management platform 209, and the backup agent 210 can switch the first cluster 200 and the second cluster 212. When the first cluster 200 and the second cluster 212 are switched, the cloud management platform 209 will call the backup agent 210 to restore the virtual machine disk. In this process, the virtual machine system disk, data disk and storage metadata (volume ID, name, size, type, etc.) will be restored, thereby reducing the normal use of the user side.

[0048] In an implementation scenario, after the virtual machine disks in the first cluster 200 and the second cluster 212 are restored, the cloud management platform 209 will create a disaster recovery virtual machine based on the stored instance cloud data information. This creation process involves a series of checks, creations, and configurations of pre-resources (such as first checking whether the instance network exists, and if so, checking whether the parameters are consistent. If not, the network creation operation will be performed). Finally, the cloud management platform 209 will call the API to create a disaster recovery virtual machine, complete the network card and disk mounting, and restore the disaster recovery virtual machine with consistent parameters. After the main cluster is restored, the user device initiates a migration operation, and the virtual machine migration operation is performed in a similar manner to switching the first cluster 200 and the second cluster 212. There is a click button in the cloud management platform 209. After clicking the button, a page for selecting whether to migrate will be displayed. The user evaluates the data content and makes a selection. If the user chooses not to migrate, the first cluster 200 and the second cluster 212 will be switched again. After the switch, the first cluster 200 will be the host cluster, and the second cluster 212 will be the backup cluster. If the user chooses to migrate, the data in the cluster will be saved, and the backup operation will be performed on the first cluster 200 and the second cluster 212 again. The data in the second cluster 212 will be backed up to the first cluster 200. The first cluster 200 and the second cluster 212 will be switched again. After the switch, the first cluster 200 will be the main cluster, and the second cluster 212 will be the backup cluster.

[0049] In some disclosed embodiments, during the execution of the backup operation, an inquiry instruction may be sent to the first cluster 200 or the second cluster 212 based on the operation instruction sent to the first cluster 200 or the second cluster 212, and the inquiry instruction is used to inquire the first cluster 200 or the second cluster 212 about the completion status of the operation instruction, and receive a feedback instruction from the first cluster 200 or the second cluster 212, and the feedback instruction includes any of the following: a completion instruction fed back based on the completed operation instruction, a confirmation instruction based on the received operation instruction but not completed. The first cluster 200 and the second cluster 212 determine whether each received instruction in the backup operation process has been completed based on the judgment module, and reply with a completion instruction if it has been completed, and reply with a confirmation instruction if it has not been completed. In the above manner, the first cluster 200 and the second cluster 212 obtain whether the query instruction content has been executed by processing the query instruction, further avoiding repeating the same operation instruction and resulting in reduced backup efficiency, thereby improving the user experience.

[0050] In one implementation scenario, after sending an operation instruction to the first cluster 200 or the second cluster 212, the first cluster 200 or the second cluster 212 performs confirmation based on the received operation instruction. For the completed operation instruction, the first cluster 200 or the second cluster 212 will not perform the operation again, and sends the completed operation instruction to the backup agent 210 in the data backup device 208, so that the operation instruction will not be sent again in the data backup device 208; if the first cluster 200 or the second cluster 212 performs confirmation based on the received operation instruction, and the first cluster 200 or the second cluster 212 has not completed the operation instruction, the first cluster 200 or the second cluster 212 sends a confirmation instruction to the data backup device 208, so that the backup agent 210 in the backup instruction confirms that the instruction is sent successfully. After the first cluster 200 or the second cluster 212 sends the confirmation instruction, the operation instruction is executed.

[0051] In one implementation scenario, the backup agent 210 confirms based on the feedback instructions received from the first cluster 200 and the second cluster 212. If the instruction sent by the first cluster 200 or the second cluster 212 is a completion instruction, the feedback instruction content received by the backup agent 210 is processed and sent to the cloud management platform 209. If the instruction sent by the first cluster 200 or the second cluster 212 is a confirmation instruction, the cloud management platform 209 confirms the instruction again and processes the feedback to the cloud management platform 209. When the cloud management platform 209 sends an instruction to the backup agent 210 again, the backup agent 210 needs to confirm the received instruction. If the received instruction is the current repeated instruction, the current feedback instruction is returned to the cloud management platform 209. If the received instruction is a new command instruction, the backup agent 210 processes the command instruction and sends it to the first cluster 200 or the second cluster 212.

[0052] In some public embodiments, the first cluster 200 includes multiple storage mirrors arranged in a distributed manner, and the data is stored in the storage mirrors. The above method uses distributed storage in the underlying storage, which makes operation and maintenance deployment simpler, and the distributed storage can better unify the use of low-cost disks, ultimately providing a higher capacity storage pool to the outside world, and reducing the impact of node crashes on the storage system. Without the need for an arbitration mechanism, it can meet the cloud host master-slave cluster data backup and minimize data loss.

[0053] Specifically, the storage image is used to store data, such as Figure 2 As described above, the storage image in the first cluster 200 may include but is not limited to the storage disk 206 , and the storage image in the second cluster 212 may include but is not limited to the storage disk I 217 , the storage disk II 218 , and the basic storage disk 219 .

[0054] See also Figure 5 , Figure 5 It is a schematic diagram of a framework of an embodiment of a data backup device 30 of the present application. The data backup device 30 includes a detection module 31, a selection module 32 and a backup module 33, wherein the detection module 31 is used to respond to the backup instruction, perform data detection on the first cluster 200, and obtain a data detection result, and the data detection result includes whether the target data is stored in the first cluster 200; the selection module 32 is used to select a target backup mechanism from a plurality of preset backup mechanisms based on the data detection result, and the plurality of preset backup mechanisms at least include a snapshot mechanism and a log mechanism; the backup module 33 is used to perform a backup operation on the target data based on the target backup mechanism, so as to back up the target data to the second cluster 212.

[0055] In the above scheme, in response to the backup instruction, data detection is performed on the first cluster to obtain data detection results, the data detection results include whether the target data is stored in the first cluster, and then based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and the plurality of preset backup mechanisms at least include a snapshot mechanism and a log mechanism, and then based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster. On the one hand, there is no business association between the first cluster and the second cluster, that is, the two are independent of each other, and a failure in any cluster will not affect the normal operation of the other cluster. On the other hand, based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and then the target data is backed up. The flexible selection of multiple backup mechanisms helps to reduce data loss. Therefore, the impact of data loss on user-side business can be minimized.

[0056] In some disclosed embodiments, based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, including selecting a snapshot mechanism as the target backup mechanism when the detection results include that the target data is stored in the first cluster 200; and selecting a log mechanism as the target backup mechanism when the detection results include that the target data is not stored in the first cluster 200.

[0057] Therefore, by selecting different backup mechanisms for different detection results, the data stored in the cluster can be quickly backed up, further reducing the situation where data is not backed up due to data confusion, thereby reducing data loss.

[0058] In some disclosed embodiments, the target backup mechanism is a snapshot mechanism. Based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster 212, including performing snapshot marking on the target data in the first cluster 200 to obtain a data copy of the target data; then restoring the target data based on the data copy of the target data, and sending the target data to the second cluster 212.

[0059] Therefore, the target data can be quickly backed up in a short time. The snapshot backup mechanism is based on a snapshot backup of the data content of the storage space at a certain point in time, which can minimize the problem of data omission caused by data writing while backing up. The snapshot mechanism does not require a complete data space, thereby achieving the purpose of not only fast backup but also saving storage space.

[0060] In a disclosed embodiment, after the target data is recovered based on the data copy of the target data and the target data is sent to the second cluster 212 , the method further includes deleting the data copy in the first cluster 200 .

[0061] Therefore, by deleting the data copy, it is avoided that the data is restored again and resent to the second cluster 212, thereby improving the efficiency of data backup.

[0062] In a disclosed embodiment, the target backup mechanism is a log mechanism. The first cluster 200 also stores an operation log, which is used to store data operation records. Based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster 212, including copying the data operation records of the target data from the first cluster 200, and restoring the target data based on the data operation records of the target data, and sending the target data to the second cluster 212.

[0063] Therefore, backing up data through a log backup mechanism can quickly restore data and minimize the time required to restore data, further reducing data loss.

[0064] In a disclosed embodiment, after sending the target data to the second cluster 212 , the method further includes deleting the data operation record of the target data in the first cluster 200 .

[0065] Therefore, after the target data is backed up to the second cluster 212, the data backup device 208 will not restore the target data again, thereby avoiding repeated data recovery, reducing the space occupied by multiple data storages, and improving the efficiency of data backup as much as possible.

[0066] In a disclosed embodiment, the backup instruction is sent by the user when the user device stores the target data in the first cluster 200 ; or the backup instruction is automatically sent by the user device based on a preset backup policy when the user device stores the target data in the first cluster 200 .

[0067] Therefore, sending the two backup instructions can avoid the problem of data loss caused by the user forgetting to perform the backup operation on the data during the data storage process, further reducing the data loss.

[0068] In a disclosed embodiment, the first cluster 200 is a primary cluster and the second cluster 212 is a backup cluster. The method further includes, in response to a switching instruction, switching the first cluster 200 to a backup cluster and switching the second cluster 212 to a primary cluster; and in response to a migration instruction, migrating data from the backup cluster to the primary cluster based on a snapshot mechanism.

[0069] Therefore, by switching the first cluster 200 and the second cluster 212 , the impact on the user side is reduced as much as possible, thereby improving the user experience.

[0070] In a disclosed embodiment, before performing data detection on the first cluster 200 in response to a backup instruction and obtaining a data detection result, the method also includes obtaining metadata of the first cluster 200 and the second cluster 212; and the metadata includes a first network interface of the first cluster 200 and a second network interface of the second cluster 212, the first network interface is used to implement data transmission with the first cluster 200, and the second network interface is used to implement data transmission with the second cluster 212.

[0071] Therefore, data transmission between the data backup device 208 and the cluster can be achieved, and when a cluster is reconnected, no storage system in the cluster will be missed.

[0072] In a disclosed embodiment, during the execution of the backup operation, the method further includes sending an inquiry instruction to the first cluster 200 and / or the second cluster 212 based on the operation instruction sent to the first cluster 200 and / or the second cluster 212; and the inquiry instruction is used to inquire the first cluster 200 and / or the second cluster 212 about the completion status of the operation instruction; receiving feedback instructions from the first cluster 200 and / or the second cluster 212; and the feedback instruction includes any of the following: a completion instruction fed back based on the completed operation instruction, a confirmation instruction based on the received operation instruction but not completed the operation instruction, the first cluster 200 and the second cluster 212 determine whether each received instruction in the backup operation process has been completed based on the judgment module, and reply with a completion instruction if completed; and reply with a confirmation instruction if not completed.

[0073] Therefore, the first cluster 200 and the second cluster 212 obtain whether the query instruction content has been executed by processing the query instruction, further avoiding repeating the same operation instruction and causing a decrease in backup efficiency, thereby improving the user experience.

[0074] In a disclosed embodiment, the first cluster 200 includes a plurality of storage mirrors arranged in a distributed manner, and data is stored in the storage mirrors.

[0075] Therefore, by using distributed storage in the underlying storage, operation and maintenance deployment is simpler, and distributed storage can better and more uniformly utilize low-cost disks, ultimately providing a higher-capacity storage pool to the outside world and reducing the impact of node crashes on the storage system. Without the need for an arbitration mechanism, it can meet the cloud host master-slave cluster data backup and minimize data loss.

[0076] See also Figure 6 , Figure 6 4 is a schematic diagram of a data backup device of the present application. The data backup device 208 includes a memory 41, a communication circuit 43 and a processor 42, and the memory 41 and the communication circuit 43 are coupled to the processor 42. The memory 41 stores program instructions, and the processor 42 is used to execute the program instructions to implement the steps in any of the above data backup method embodiments. For the specific meaning of the data backup device 208, please refer to the aforementioned disclosed embodiments and Figure 2 , I will not go into details here.

[0077] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above data backup method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.

[0078] In the above scheme, in response to the backup instruction, data detection is performed on the first cluster to obtain data detection results, the data detection results include whether the target data is stored in the first cluster, and then based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and the plurality of preset backup mechanisms at least include a snapshot mechanism and a log mechanism, and then based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster. On the one hand, there is no business association between the first cluster and the second cluster, that is, the two are independent of each other, and a failure in any cluster will not affect the normal operation of the other cluster. On the other hand, based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and then the target data is backed up. The flexible selection of multiple backup mechanisms helps to reduce data loss. Therefore, the impact of data loss on user-side business can be minimized.

[0079] In one implementation scenario, a data backup system includes a data backup device and multiple clusters. The data backup device is respectively connected to the multiple clusters for communication. The data backup device is used to execute the steps in any of the above data backup method embodiments to achieve data backup between multiple clusters.

[0080] In the above scheme, in response to the backup instruction, data detection is performed on the first cluster to obtain data detection results, the data detection results include whether the target data is stored in the first cluster, and then based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and the plurality of preset backup mechanisms at least include a snapshot mechanism and a log mechanism, and then based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster. On the one hand, there is no business association between the first cluster and the second cluster, that is, the two are independent of each other, and a failure in any cluster will not affect the normal operation of the other cluster. On the other hand, based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and then the target data is backed up. The flexible selection of multiple backup mechanisms helps to reduce data loss. Therefore, the impact of data loss on user-side business can be minimized.

[0081] See also Figure 7 , Figure 7 The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps in any of the above data backup method embodiments.

[0082] In the above scheme, in response to the backup instruction, data detection is performed on the first cluster to obtain data detection results, the data detection results include whether the target data is stored in the first cluster, and then based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and the plurality of preset backup mechanisms at least include a snapshot mechanism and a log mechanism, and then based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to the second cluster. On the one hand, there is no business association between the first cluster and the second cluster, that is, the two are independent of each other, and a failure in any cluster will not affect the normal operation of the other cluster. On the other hand, based on the data detection results, a target backup mechanism is selected from a plurality of preset backup mechanisms, and then the target data is backed up. The flexible selection of multiple backup mechanisms helps to reduce data loss. Therefore, the impact of data loss on user-side business can be minimized.

[0083] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0084] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0085] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0086] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0087] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0088] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

Claims

1. A data backup method, characterized in that: include: In response to the backup instruction, perform data detection on the first cluster to obtain a data detection result; wherein the data detection result includes whether the target data is stored in the first cluster, and the first cluster is a main cluster; In a case where the detection result includes that the target data is stored in the first cluster, selecting a snapshot mechanism as the target backup mechanism; and in a case where the detection result includes that the target data is not stored in the first cluster, selecting a log mechanism as the target backup mechanism; Based on the target backup mechanism, a backup operation is performed on the target data to back up the target data to a second cluster; wherein the second cluster is a backup cluster.

2. The method according to claim 1, characterized in that The target backup mechanism is the snapshot mechanism, and the performing a backup operation on the target data based on the target backup mechanism to back up the target data to the second cluster includes: Performing snapshot marking on the target data in the first cluster to obtain a data copy of the target data; The target data is recovered based on the data copy of the target data, and the target data is sent to the second cluster.

3. The method according to claim 2, characterized in that After the target data is recovered based on the data copy of the target data and the target data is sent to the second cluster, the method further includes: The data replica in the first cluster is deleted.

4. The method according to claim 1, characterized in that: The target backup mechanism is the log mechanism, the first cluster further stores an operation log, the operation log is used to store data operation records, and the backup operation is performed on the target data based on the target backup mechanism to back up the target data to the second cluster, including: Copying the data operation record of the target data from the first cluster, and restoring the target data based on the data operation record of the target data; The target data is sent to the second cluster.

5. The method according to claim 4, characterized in that After sending the target data to the second cluster, the method further includes: The data operation record of the target data in the first cluster is deleted.

6. The method according to claim 1, characterized in that The backup instruction is sent by the user when the user device stores the target data in the first cluster; Alternatively, the backup instruction is automatically sent by the user device based on a preset backup policy when storing the target data in the first cluster.

7. The method according to claim 1, characterized in that The method further comprises: In response to a switching instruction, switching the first cluster to the backup cluster, and switching the second cluster to the main cluster; And / or, in response to the migration instruction, data is migrated back from the backup cluster to the primary cluster based on the snapshot mechanism.

8. The method according to claim 1, characterized in that Before performing data detection on the first cluster in response to the backup instruction to obtain the data detection result, the method further includes: Obtaining metadata of the first cluster and the second cluster; The metadata includes a first network interface of the first cluster and a second network interface of the second cluster, the first network interface is used to implement data transmission with the first cluster, and the second network interface is used to implement data transmission with the second cluster.

9. The method according to claim 1, characterized in that: During the execution of the backup operation, the method further includes: Based on the operation instruction having been sent to the first cluster and / or the second cluster, sending a query instruction to the first cluster and / or the second cluster; wherein the query instruction is used to query the first cluster and / or the second cluster about the completion status of the operation instruction; receiving feedback instructions from the first cluster and / or the second cluster; wherein the feedback instructions include any one of the following: a completion instruction fed back based on the completion of the operation instruction, a confirmation instruction based on the receipt of the operation instruction but the failure to complete the operation instruction; The first cluster and the second cluster determine whether each received instruction in the backup operation process has been completed based on the judgment module, and reply with a completion instruction if it has been completed; If not completed reply confirmation command.

10. The method according to claim 1, characterized in that The first cluster includes a plurality of storage mirrors arranged in a distributed manner, and data are all stored in the storage mirrors.

11. A data backup device, characterized in that: include: a detection module, configured to respond to the backup instruction, perform data detection on the first cluster, and obtain a data detection result; wherein the data detection result includes whether the target data is stored in the first cluster, and the first cluster is a main cluster; A selection module, configured to select a snapshot mechanism as a target backup mechanism when the detection result includes that the target data is stored in the first cluster; and select a log mechanism as a target backup mechanism when the detection result includes that the target data is not stored in the first cluster; A backup module is used to perform a backup operation on the target data based on the target backup mechanism, so as to back up the target data to a second cluster; wherein the second cluster is a backup cluster.

12. A data backup device, characterized in that: It includes a memory, a communication circuit and a processor, wherein the memory and the communication circuit are coupled to the processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the data backup method according to any one of claims 1 to 10.

13. A data backup system, characterized in that: It comprises a data backup device and a plurality of clusters, wherein the data backup device is respectively communicatively connected to the plurality of clusters, and the data backup device is used to execute the data backup method according to any one of claims 1 to 10 to realize data backup between the plurality of clusters.

14. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the data backup method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Data migration method, device and apparatus and computer storage medium

    CN111538719A