A method, device, equipment and storage medium for backing up data of a storage cluster

By establishing communication connections and replication links between the master cluster and the slave cluster, and using directory-level snapshot technology for data replication, the problem of inability to cope with the entire cluster failure in the existing technology is solved, and asynchronous disaster recovery of distributed file storage clusters is realized to ensure the persistence and recovery of data.

CN114996054BActive Publication Date: 2025-07-04JINAN INSPUR DATA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210611162.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-07-04
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

In the prior art, disaster recovery solutions for distributed storage clusters usually can only deal with the failure of nodes in the cluster, and cannot deal with the failure of the entire cluster, resulting in a high risk of data loss.

Method used

By establishing a communication connection between the master cluster and the slave cluster, creating a replication link, and using directory-level snapshot technology to replicate data, the master cluster data is realized asynchronously backed up in the slave cluster, ensuring that data can be recovered in the event of a cluster failure.

Benefits of technology

It realizes asynchronous disaster recovery of distributed file storage clusters, ensures data continuity and recovery, avoids data loss caused by cluster failure, and does not affect business operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114996054B_ABST
    Figure CN114996054B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and storage medium for backing up data of a storage cluster, relating to the field of data backup. The method includes: creating a communication connection between a local cluster and the target remote cluster according to the parameter information of the target remote cluster, so as to construct a communication connection between the master cluster and the slave cluster; creating a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information; creating a master directory snapshot corresponding to the master directory, copying the master directory data corresponding to the master directory snapshot to the slave directory through the replication link, and creating a corresponding slave directory snapshot in the slave directory, so as to realize the backup of the data of the master cluster in the slave cluster. Asynchronous disaster tolerance of a distributed file storage cluster can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data backup, and particularly to a method, device, equipment and storage medium for backing up data of a storage cluster. Background Art

[0002] At present, distributed storage clusters have been increasingly applied. With the advancement of digitalization, data has gradually become the core of the operation of enterprises and institutions, and users have higher and higher requirements for the stability of the storage system that bears the data. In order to ensure the persistence, recoverability and high availability of data storage, remote disaster recovery solutions have emerged, and remote replication technology is one of the key technologies of remote disaster recovery solutions. Its core idea is to synchronously replicate our data multiple times to various places, so as to avoid data loss caused by natural disasters or human damage as much as possible. However, currently, it is usually for disaster recovery of nodes within the cluster, and this method cannot cope with the situation where the entire cluster fails. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for backing up data of a storage cluster, which can perform asynchronous disaster recovery for a distributed file storage cluster. The specific scheme is as follows:

[0004] In a first aspect, the present application discloses a method for backing up data of a storage cluster, including:

[0005] Create a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster, so as to construct a communication connection between the master cluster and the slave cluster;

[0006] Create a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information;

[0007] Create a master directory snapshot corresponding to the master directory, copy the master directory data corresponding to the master directory snapshot to the slave directory through the replication link, and create a corresponding slave directory snapshot in the slave directory to realize the backup of the data of the master cluster in the slave cluster.

[0008] Optionally, the creating a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster includes:

[0009] Add a link to the target remote cluster in the local cluster according to the parameter information of the target remote cluster;

[0010] Add a link to the local cluster in the target remote cluster according to the parameter information of the local cluster through the Secure Shell protocol, so as to create a communication connection between the local cluster and the target remote cluster.

[0011] Optionally, before creating a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster, the method further includes:

[0012] Sending a key acquisition request to the target remote cluster to obtain the key and connection name of the target remote cluster;

[0013] Verifying the key, and if the verification is successful, sending an IP acquisition request to the target remote cluster to obtain the IP bound to the management node of the target remote cluster;

[0014] Sending a key ring acquisition request to the target remote cluster to obtain the key ring corresponding to the target remote cluster;

[0015] Obtaining the parameter information of the target remote cluster based on the connection name, the IP, and the key ring.

[0016] Optionally, creating a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information includes:

[0017] Obtaining backup information; the backup information includes the cluster information of the target remote cluster, the directory information of the target remote cluster, and the directory information of the local cluster;

[0018] Creating a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information.

[0019] Optionally, copying the master directory data corresponding to the master directory snapshot to the slave directory through the replication link includes:

[0020] Traversing the slave directory in sequence according to each sub-directory included in the master directory snapshot to determine the differential files between the master directory and the slave directory; the differential files include the first type of files that exist in the slave directory but do not exist in the master directory, and the second type of files with the same file name but different file types in the master directory and the slave directory;

[0021] Deleting the differential files, and then copying the data corresponding to each sub-directory included in the master directory snapshot to the slave directory through the replication link in sequence.

[0022] Optionally, copying the data corresponding to each sub-directory included in the master directory snapshot to the slave directory through the replication link in sequence includes:

[0023] If there is no directory corresponding to the sub-directory of the master directory in the slave directory, creating a corresponding sub-directory in the slave directory according to the sub-directory of the master directory;

[0024] If there is a directory in the slave directory corresponding to the sub-directory of the master directory, modify the directory information of the corresponding sub-directory of the slave directory according to the sub-directory of the master directory; the directory information includes format, access time, and modification time.

[0025] Optionally, the storage cluster data backup method further includes:

[0026] If the master directory snapshot contains target data that is not a directory, determine the synchronization method of the target data, and synchronize it to the slave directory based on the target data according to the synchronization method; the target data includes data to be backed up and corresponding metadata;

[0027] Among them, determining the synchronization method of the target data includes:

[0028] If there is no file in the slave directory, determine that the synchronization method is to synchronize the target data;

[0029] If the file existing in the slave directory is of a different file type from the target data, determine that the synchronization method is to synchronize the target data;

[0030] If the file existing in the slave directory is different from the target data in file size or modification time, determine that the synchronization method is to synchronize only the data to be backed up;

[0031] If the file existing in the slave directory is different from the target data in the status change time, determine that the synchronization method is to synchronize only the metadata.

[0032] In a second aspect, the present application discloses a storage cluster data backup device, including:

[0033] A connection creation module, configured to create a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster, so as to construct a communication connection between the master cluster and the slave cluster;

[0034] A link creation module, configured to create a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information;

[0035] A backup module, configured to create a corresponding master directory snapshot of the master directory, copy the master directory data corresponding to the master directory snapshot to the slave directory through the replication link, and create a corresponding slave directory snapshot in the slave directory, so as to implement the backup of the data of the master cluster in the slave cluster.

[0036] In a third aspect, the present application discloses an electronic device, including:

[0037] A memory, configured to store a computer program;

[0038] A processor for executing the computer program to implement the foregoing storage cluster data backup method.

[0039] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein when the computer program is executed by a processor, the foregoing storage cluster data backup method is implemented.

[0040] In the present application, a communication connection between a local cluster and the target remote cluster is created according to the parameter information of the target remote cluster to construct a communication connection between the master cluster and the slave cluster; according to the backup information, a replication link for remote replication is created between the master directory of the master cluster and the slave directory of the slave cluster; a master directory snapshot corresponding to the master directory is created, and the master directory data corresponding to the master directory snapshot is copied to the slave directory through the replication link, and a corresponding slave directory snapshot is created in the slave directory to implement the backup of the data of the master cluster in the slave cluster. It can be seen that on the basis of establishing a communication connection between clusters, a replication link between the master directory of the master cluster and the slave directory of the slave cluster is established, and data synchronization is performed using this replication link based on snapshots to achieve the backup of the entire cluster, and remote replication is based on directory-level snapshot technology to achieve asynchronous data replication between the master cluster and the slave cluster without affecting business operations, that is, asynchronous disaster recovery of the distributed file storage cluster is realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings according to the provided drawings without creative efforts.

[0042] Figure 1 It is a flowchart of a storage cluster data backup method provided by the present application;

[0043] Figure 2 It is a timing diagram of a specific cluster connection creation provided by the present application;

[0044] Figure 3 It is a schematic diagram of a storage cluster data remote replication structure provided by the present application;

[0045] Figure 4 It is a timing diagram of a specific storage cluster data backup provided by the present application;

[0046] Figure 5 It is a flowchart of another storage cluster data backup method provided by the present application;

[0047] Figure 6 Schematic structural diagram of a storage cluster data backup device provided by this application;

[0048] Figure 7 Structural diagram of an electronic device provided by this application. Specific implementation manners

[0049] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] In the prior art, disaster tolerance is usually performed for nodes within a cluster, and this method cannot cope with the situation where the entire cluster fails. To overcome the above technical problems, this application proposes a storage cluster data backup method, which can achieve asynchronous disaster tolerance for a distributed file storage cluster.

[0051] An embodiment of this application discloses a storage cluster data backup method. Refer to Figure 1 as shown, this method may include the following steps:

[0052] Step S11: Create a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster, so as to construct a communication connection between the master cluster and the slave cluster.

[0053] In this embodiment, first, a connection is established between clusters. When creating a connection, the IP of the peer cluster, the secret key of the peer cluster, and the connection name need to be input. The peer cluster is the cluster that wants to create a connection with the local cluster. In this embodiment, the local cluster can be either the master cluster or the slave cluster. The purpose of this step is to create a connection between the master cluster and the slave cluster, so as to synchronize the data of the master cluster to the slave cluster. The master cluster is the master storage system, which is the source file system for synchronous snapshots, and the slave cluster is the slave storage system, which is the destination file system for synchronous snapshots.

[0054] In this embodiment, creating a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster may include: adding a link to the target remote cluster in the local cluster according to the parameter information of the target remote cluster; adding a link to the local cluster in the target remote cluster according to the parameter information of the local cluster through the Secure Shell protocol, so as to create a communication connection between the local cluster and the target remote cluster. That is, first perform key verification to ensure the security of the cluster, then add a link to the peer cluster in the local cluster, and then through the Secure Shell protocol (SSH), add the local cluster as a link to the remote cluster in the remote cluster.

[0055] In this embodiment, before creating a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster, it may further include: sending a key acquisition request to the target remote cluster to obtain the key and connection name of the target remote cluster; verifying the key, and if the verification is successful, sending an IP acquisition request to the target remote cluster to obtain the IP bound to the management node of the target remote cluster; sending a keyring acquisition request to the target remote cluster to obtain the keyring corresponding to the target remote cluster; obtaining the parameter information of the target remote cluster based on the connection name, the IP, and the keyring. Specifically, for example Figure 2 As shown, after the local cluster sends a key acquisition request to the remote cluster, it receives the key and connection name fed back by the remote cluster, and then verifies the key. If the verification fails, it fails and returns an error prompt. If the verification is successful, it sends an IP acquisition request to the remote cluster to obtain the IP bound to the management node (monitor) of the remote cluster, and sends a keyring acquisition request to the remote cluster to receive the corresponding keyring fed back by the remote cluster. Then, the local cluster calls the remote connection addition interface, and then calls the interface in the remote cluster through SSH to create a remote connection in the remote cluster. Furthermore, the remote cluster calls the remote connection addition interface, and finally returns the result to the local cluster to complete the connection establishment.

[0056] Step S12: Create a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information.

[0057] In this embodiment, after the connection between the master and slave clusters is established, a replication relationship needs to be established, that is, a replication link (pair) for remote replication. The pair is a link from the local cluster to the remote cluster, and the pair specifies the remote replication directory.

[0058] In this embodiment, creating a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information may include: obtaining the backup information; the backup information includes the cluster information of the target remote cluster, the directory information of the target remote cluster, and the directory information of the local cluster; creating a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information. Creating a pair can be operated on the master cluster. It is necessary to select the remote cluster, local directory, and remote directory. After creation, the local directory is the master directory, and the remote directory is the slave directory, and the data is synchronized from the master directory to the slave directory.

[0059] Step S13: Create a master directory snapshot corresponding to the master directory, copy the master directory data corresponding to the master directory snapshot to the slave directory through the replication link, and create a corresponding slave directory snapshot in the slave directory to implement the backup of the data of the master cluster in the slave cluster.

[0060] In this embodiment, after the master directory and the slave directory establish a remote replication relationship, synchronization will be started first, as Figure 3 shown. Before synchronization, a snapshot of the master directory is created, and the master directory data at the snapshot time point is copied to the slave directory to ensure that the data in the slave directory is consistent with the data in the master directory (at a certain snapshot time point). The snapshot of a directory is a backup image of the directory at a certain time point.

[0061] In this embodiment, copying the master directory data corresponding to the master directory snapshot to the slave directory through the replication link may include: traversing the slave directory in sequence according to each sub-directory included in the master directory snapshot to determine the different files between the master directory and the slave directory; the different files include the first type of files that exist in the slave directory but do not exist in the master directory, and the second type of files with the same file name but different file types in the master directory and the slave directory; deleting the different files, and then copying the data corresponding to each sub-directory included in the master directory snapshot to the slave directory through the replication link. That is, for example Figure 4 shown, the master directory snapshot needs to be traversed twice. The purpose of the first traversal is to ensure that there are no redundant file directories in the slave directory, and the second traversal is to synchronize the data in the slave directory. Before synchronization, a master directory snapshot, called s1, is created in the master directory, and then the snapshot directory is traversed and processed. For each sub-directory in the master directory, if there is a sub-directory with the same path in the slave directory, the sub-directory of the slave directory is traversed first, and the directory files that exist in the slave directory but do not exist in the master directory, and the files with the same file name but different file types in the master directory and the slave directory are deleted, with the purpose of ensuring that there are no redundant file directories in the slave directory.

[0062] In this embodiment, the step of copying the data corresponding to each sub-directory included in the master directory snapshot to the slave directory in sequence through the replication link may include: if there is no directory corresponding to the sub-directory of the master directory in the slave directory, creating a corresponding sub-directory in the slave directory according to the sub-directory of the master directory; if there is a directory corresponding to the sub-directory of the master directory in the slave directory, modifying the directory information of the corresponding sub-directory in the slave directory according to the sub-directory of the master directory; the directory information includes format, access time, and modification time. That is, if there is already a corresponding sub-directory in the slave directory, the directory metadata, that is, the directory parameters, can be directly modified. The directory parameters include, but are not limited to, format (mode), access time (mtime), and modification time (atime); if there is no directory corresponding to the sub-directory of the master directory currently traversed in the slave directory, a corresponding sub-directory is created in the slave directory according to the sub-directory of the master directory.

[0063] In this embodiment, creating a corresponding slave directory snapshot in the slave directory may include: based on the data copied to the slave directory, creating a slave directory snapshot with the same name as the master directory snapshot in the slave directory, and saving the snapshot ID of the master directory snapshot to the metadata of the slave directory snapshot. That is, after all traversal processing is completed, the data in the slave directory is consistent with the master directory snapshot directory. A snapshot s1 with the same name is created in the slave directory, so that the data under s1 in the master and slave directories is consistent, and data synchronization is completed. That is, after replication, a snapshot with the same name as the synchronized snapshot is created to ensure that when a failure occurs in the master cluster, the data in the slave cluster can be used normally. Moreover, through the directory timing snapshot function, timed remote replication can be achieved.

[0064] The above asynchronous disaster recovery solution based on a massive distributed file system can be applied to disaster recovery backup scenarios and data distribution scenarios. The disaster recovery backup scenarios include: (1) point-to-point disaster recovery: deploying a production site and a disaster recovery site, and the disaster recovery site is used as the data backup of the production site; (2) centralized disaster recovery: deploying multiple production sites and a disaster recovery site, and the master directories in different production sites copy and back up data to different slave directories in the disaster recovery site. That is, the snapshot data is asynchronously replicated to a single or multiple slave clusters, and the master and slave ends of the remote replication can be placed in the same place or in different locations, so as to achieve the disaster recovery function. Data distribution scenario: Data distribution means that the data of the master site is periodically replicated to different slave sites, which is mainly applied to scenarios such as the headquarters regularly distributing data to branches.

[0065] As can be seen from the above, in this embodiment, a communication connection between the local cluster and the target remote cluster is created according to the parameter information of the target remote cluster, so as to construct a communication connection between the master cluster and the slave cluster; according to the backup information, a replication link for remote replication is created between the master directory of the master cluster and the slave directory of the slave cluster; a master directory snapshot corresponding to the master directory is created, and the master directory data corresponding to the master directory snapshot is copied to the slave directory through the replication link, and a corresponding slave directory snapshot is created in the slave directory, so as to realize the backup of the data of the master cluster in the slave cluster. It can be seen that on the basis of establishing a communication connection between clusters, a replication link between the master directory of the master cluster and the slave directory of the slave cluster is established to realize the backup of the entire cluster, and the remote replication is based on the snapshot technology at the directory level, realizing asynchronous replication of data between the master cluster and the slave cluster, without affecting the business operation, that is, realizing asynchronous disaster tolerance of the distributed file storage cluster.

[0066] Based on the above embodiments, a data synchronization method is also disclosed. Refer to Figure 5 as shown, including:

[0067] Step S21: If the master directory snapshot contains target data that is not a directory, determine the synchronization method of the target data; the target data includes the data to be backed up and the corresponding metadata;

[0068] Step S22: If there is no file in the slave directory, determine that the synchronization method is to synchronize the target data;

[0069] Step S23: If the file existing in the slave directory is of a different file type from the target data, determine that the synchronization method is to synchronize the target data;

[0070] Step S24: If the file existing in the slave directory is different from the target data in terms of file size or modification time, determine that the synchronization method is to synchronize only the data to be backed up;

[0071] Step S25: If the file existing in the slave directory is different from the target data in terms of the status change time, determine that the synchronization method is to synchronize only the metadata;

[0072] Step S26: Synchronize the target data to the slave directory based on the synchronization method.

[0073] It can be understood that the master directory snapshot obtained by snapshot may contain data in a non-directory form. This kind of data includes the data itself and the metadata, that is, the above-mentioned data to be backed up and the metadata. At this time, in order to save resources and improve the backup speed, first determine the synchronization method of the target data, and then synchronize the target data to the slave directory based on the synchronization method. For example Figure 4As shown in the figure, the synchronization method is to determine whether the data or metadata to be backed up needs to be synchronized. If the file does not exist in the slave directory, both the data and metadata to be backed up need to be synchronized; if the file types are different, both the data and metadata to be backed up need to be synchronized; if the size or mtime of the file is different, the data to be backed up needs to be synchronized; if the change time (ctime) of the file status is different, the metadata needs to be synchronized.

[0074] Furthermore, when synchronizing data, if hard links are used for synchronization, data is read from the source file in the master directory according to the maximum size read each time, and then written to the destination file in the slave directory. If soft links are used, the original link is released and a new link is established before synchronization. When synchronizing metadata, it includes basic metadata such as mode, mtime, atime, user ID (User ID, abbreviated as UID), and group ID (Group ID, abbreviated as GID). At the same time, extended attributes can also be included, such as directory configuration, configuration of ACL (Access Control Lists) to avoid network worms such as WORM.

[0075] As can be seen from the above, if the master directory snapshot contains target data that is not a directory, the synchronization method of the target data is determined; the target data includes the data to be backed up and the corresponding metadata; if the file does not exist in the slave directory, the synchronization method is determined to be synchronizing the target data; if the file type of the file existing in the slave directory is different from that of the target data, the synchronization method is determined to be synchronizing the target data; if the file size or modification time of the file existing in the slave directory is different from that of the target data, the synchronization method is determined to be only synchronizing the data to be backed up; if the change time of the file status of the file existing in the slave directory is different from that of the target data, the synchronization method is determined to be only synchronizing the metadata; based on the target data, it is synchronized to the slave directory according to the synchronization method. It can be seen that for the data in the form of non-directory contained in the master directory snapshot, the corresponding data synchronization method is determined through multiple steps, and then the obtained synchronization method is used for synchronization, which not only increases the flexibility of data backup but also improves the efficiency of data content backup.

[0076] Correspondingly, the embodiment of the present application also discloses a storage cluster data backup device. Refer to Figure 6 As shown in the figure, the device includes:

[0077] A connection creation module 11, configured to create a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster, so as to construct a communication connection between the master cluster and the slave cluster;

[0078] A link creation module 12, configured to create a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information;

[0079] A backup module 13 is used to create a corresponding master directory snapshot of the master directory, copy the master directory data corresponding to the master directory snapshot to the slave directory through the replication link, and create a corresponding slave directory snapshot in the slave directory to implement the backup of the data of the master cluster in the slave cluster.

[0080] As can be seen from the above, in this embodiment, a communication connection between the local cluster and the target remote cluster is created according to the parameter information of the target remote cluster to construct a communication connection between the master cluster and the slave cluster; according to the backup information, a replication link for remote replication is created between the master directory of the master cluster and the slave directory of the slave cluster; a master directory snapshot corresponding to the master directory is created, and the master directory data corresponding to the master directory snapshot is copied to the slave directory through the replication link, and a corresponding slave directory snapshot is created in the slave directory to implement the backup of the data of the master cluster in the slave cluster. It can be seen that on the basis of establishing a communication connection between clusters, a replication link between the master directory of the master cluster and the slave directory of the slave cluster is established to implement the backup of the entire cluster, and the remote replication is based on the directory-level snapshot technology to achieve asynchronous replication of data between the master cluster and the slave cluster, without affecting the business operation, that is, to achieve asynchronous disaster tolerance of the distributed file storage cluster.

[0081] In some specific embodiments, the connection creation module 11 may specifically include:

[0082] A local link adding unit is used to add a link of the target remote cluster to the local cluster according to the parameter information of the target remote cluster;

[0083] A remote link adding unit is used to add a link of the local cluster to the target remote cluster according to the parameter information of the local cluster through the Secure Shell protocol to create a communication connection between the local cluster and the target remote cluster.

[0084] In some specific embodiments, the storage cluster data backup device may specifically include:

[0085] A connection name obtaining unit is used to send a key obtaining request to the target remote cluster to obtain the key and connection name of the target remote cluster;

[0086] An IP obtaining unit is used to verify the key, and if the verification is successful, send an IP obtaining request to the target remote cluster to obtain the IP bound to the management node of the target remote cluster;

[0087] A key ring obtaining unit is used to send a key ring obtaining request to the target remote cluster to obtain the key ring corresponding to the target remote cluster;

[0088] A parameter information determination unit, configured to obtain parameter information of the target remote cluster based on the connection name, the IP, and the key ring.

[0089] In some specific embodiments, the link creation module 12 may specifically include:

[0090] A backup information acquisition unit, configured to acquire backup information; the backup information includes cluster information of the target remote cluster, directory information of the target remote cluster, and directory information of the local cluster;

[0091] A replication link creation unit, configured to create a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information.

[0092] In some specific embodiments, the backup module 13 may specifically include:

[0093] A differential file determination unit, configured to traverse the slave directory according to each sub-directory included in the master directory snapshot in sequence to determine differential files between the master directory and the slave directory; the differential files include a first type of files that exist in the slave directory but do not exist in the master directory, and a second type of files with the same file name but different file types in the master directory and the slave directory;

[0094] A deletion unit, configured to delete the differential files, and then copy the data corresponding to each sub-directory included in the master directory snapshot to the slave directory through the replication link in sequence.

[0095] In some specific embodiments, the deletion unit may specifically include:

[0096] A sub-directory creation unit, configured to create a corresponding sub-directory in the slave directory according to the sub-directory of the master directory if there is no directory corresponding to the sub-directory of the master directory in the slave directory;

[0097] A directory information modification unit, configured to modify the directory information of the sub-directory corresponding to the sub-directory of the master directory in the slave directory according to the sub-directory of the master directory if there is a directory corresponding to the sub-directory of the master directory in the slave directory; the directory information includes format, access time, and modification time.

[0098] In some specific embodiments, the storage cluster data backup device may specifically include:

[0099] A synchronization method determination unit, configured to, if the master directory snapshot includes target data that is not a directory, determine the synchronization method of the target data, and synchronize it to the slave directory based on the target data according to the synchronization method; the target data includes data to be backed up and corresponding metadata;

[0100] Among them, the synchronization mode determination unit is configured to determine that the synchronization mode is to synchronize target data if the file does not exist in the slave directory; determine that the synchronization mode is to synchronize target data if the file type of the file existing in the slave directory is different from that of the target data; determine that the synchronization mode is to only synchronize the data to be backed up if the file size or modification time of the file existing in the slave directory is different from that of the target data; and determine that the synchronization mode is to only synchronize metadata if the status change time of the file existing in the slave directory is different from that of the target data.

[0101] Furthermore, an embodiment of the present application also discloses an electronic device. Refer to Figure 7 As shown, the content in the figure cannot be regarded as any limitation on the scope of use of the present application.

[0102] Figure 7 It is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the storage cluster data backup method disclosed in any of the foregoing embodiments.

[0103] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and specific limitations are not imposed here.

[0104] In addition, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc. The resources stored thereon include an operating system 221, a computer program 222, and data 223 including backup information, etc., and the storage method can be short-term storage or permanent storage.

[0105] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, so as to implement the operation and processing of the massive data 223 in the memory 22 by the processor 21. It can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the storage cluster data backup method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0106] Furthermore, an embodiment of the present application also discloses a computer storage medium. When the computer executable instructions stored in the computer storage medium are loaded and executed by a processor, the steps of the storage cluster data backup method disclosed in any of the foregoing embodiments are implemented.

[0107] In this specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0108] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of both. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0109] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0110] The above has introduced in detail a method, apparatus, device and medium for backing up data of a storage cluster. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for backing up data of a storage cluster, characterized in that, Including: Create a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster, so as to build a communication connection between the master cluster and the slave cluster; Create a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information; Create a master directory snapshot corresponding to the master directory, copy the master directory data corresponding to the master directory snapshot to the slave directory through the replication link, and create a corresponding slave directory snapshot in the slave directory to realize the backup of the data of the master cluster in the slave cluster; Among them, the copying the master directory data corresponding to the master directory snapshot to the slave directory through the replication link includes: Traverse the slave directory in sequence according to each sub-directory included in the master directory snapshot to determine the differential files between the master directory and the slave directory; the differential files include the first type of files that exist in the slave directory and do not exist in the master directory, and the second type of files with the same file name and different file types in the master directory and the slave directory; Delete the differential files, and then copy the data corresponding to each sub-directory included in the master directory snapshot to the slave directory through the replication link in sequence; Among them, the copying the data corresponding to each sub-directory included in the master directory snapshot to the slave directory through the replication link in sequence includes: If there is no directory corresponding to the sub-directory of the master directory in the slave directory, create a corresponding sub-directory in the slave directory according to the sub-directory of the master directory; If there is a directory corresponding to the sub-directory of the master directory in the slave directory, modify the directory information of the corresponding sub-directory of the slave directory according to the sub-directory of the master directory; the directory information includes format, access time and modification time; Among them, the creating a corresponding slave directory snapshot in the slave directory includes: based on the data copied to the slave directory, create a slave directory snapshot with the same name as the master directory snapshot in the slave directory, and save the snapshot ID of the master directory snapshot to the metadata of the slave directory snapshot; Among them, the creating a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information includes: Obtain the backup information; the backup information includes the cluster information of the target remote cluster, the directory information of the target remote cluster and the directory information of the local cluster; Create a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information; Among them, the method for storing cluster data backup further includes: If the master directory snapshot contains target data that is not a directory, judge the synchronization method of the target data, and synchronize it to the slave directory based on the target data according to the synchronization method; the target data includes the data to be backed up and the corresponding metadata; Among them, the judging the synchronization method of the target data includes: If there is no file in the slave directory, determine that the synchronization method is to synchronize the target data; If the file types of the files existing in the slave directory are different from those of the target data, it is determined that the synchronization method is to synchronize the target data; If the file sizes or modification times of the files existing in the slave directory are different from those of the target data, it is determined that the synchronization method is to synchronize only the data to be backed up; If the status change times of the files existing in the slave directory are different from those of the target data, it is determined that the synchronization method is to synchronize only the metadata; when the synchronization method is to synchronize only the metadata, the synchronized metadata includes basic metadata and metadata of extended attributes, and the metadata of the extended attributes includes directory configuration and configured access control list; Among them, synchronizing the target data to the slave directory according to the synchronization method includes: If hard links are used to synchronize data, read data from the source file in the master directory and write it to the destination file in the slave directory according to the maximum read size each time; If soft links are used to synchronize data, release the original link and establish a new link for synchronization; Among them, before creating a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster, it further includes: Sending a key acquisition request to the target remote cluster to obtain the key and connection name of the target remote cluster; Verifying the key, and if the verification is successful, sending an IP acquisition request to the target remote cluster to obtain the IP bound to the management node of the target remote cluster; Sending a keyring acquisition request to the target remote cluster to obtain the keyring corresponding to the target remote cluster; Obtaining the parameter information of the target remote cluster based on the connection name, the IP, and the keyring.

2. The method for backing up storage cluster data according to claim 1, wherein Creating a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster includes: Adding a link to the target remote cluster in the local cluster according to the parameter information of the target remote cluster; Through the Secure Shell protocol, adding a link to the local cluster in the target remote cluster according to the parameter information of the local cluster to create a communication connection between the local cluster and the target remote cluster.

3. A storage cluster data backup device, characterized in that, It includes: A connection creation module, configured to create a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster to construct a communication connection between the master cluster and the slave cluster; A link creation module, configured to create a replication link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information; A backup module, configured to create a master directory snapshot corresponding to the master directory, copy the master directory data corresponding to the master directory snapshot to the slave directory through the replication link, and create a corresponding slave directory snapshot in the slave directory to implement the backup of the data of the master cluster in the slave cluster; Among them, the backup module is used to traverse the slave directory in sequence according to each sub-directory included in the master directory snapshot to determine the differential files between the master directory and the slave directory; the differential files include the first type of files that exist in the slave directory but do not exist in the master directory, and the second type of files with the same file name but different file types in the master directory and the slave directory; delete the differential files, and then copy the data corresponding to each sub-directory included in the master directory snapshot to the slave directory in sequence through the copy link; The backup module is used to create a corresponding sub-directory in the slave directory according to the sub-directory of the master directory if there is no directory corresponding to the sub-directory of the master directory in the slave directory; if there is a directory corresponding to the sub-directory of the master directory in the slave directory, modify the directory information of the corresponding sub-directory of the slave directory according to the sub-directory of the master directory; the directory information includes format, access time, and modification time; The backup module is used to create a slave directory snapshot with the same name as the master directory snapshot in the slave directory based on the data copied to the slave directory, and save the snapshot ID of the master directory snapshot to the metadata of the slave directory snapshot; Among them, the link creation module is used to obtain backup information; the backup information includes the cluster information of the target remote cluster, the directory information of the target remote cluster, and the directory information of the local cluster; create a copy link for remote replication between the master directory of the master cluster and the slave directory of the slave cluster according to the backup information; Among them, the storage cluster data backup device is used to judge the synchronization method of the target data if the master directory snapshot contains target data that is not a directory, and synchronize it to the slave directory based on the target data according to the synchronization method; the target data includes the data to be backed up and the corresponding metadata; Among them, the storage cluster data backup device is used to determine that the synchronization method is to synchronize the target data if there is no file in the slave directory; determine that the synchronization method is to synchronize the target data if the file type of the file existing in the slave directory is different from that of the target data; determine that the synchronization method is to only synchronize the data to be backed up if the file size or modification time of the file existing in the slave directory is different from that of the target data; determine that the synchronization method is to only synchronize the metadata if the status change time of the file existing in the slave directory is different from that of the target data; Among them, the storage cluster data backup device is configured to send a key acquisition request to the target remote cluster before creating a communication connection between the local cluster and the target remote cluster according to the parameter information of the target remote cluster, so as to obtain the key and connection name of the target remote cluster; verify the key, and if the verification is successful, send an IP acquisition request to the target remote cluster to obtain the IP bound to the management node of the target remote cluster; send a key ring acquisition request to the target remote cluster to obtain the key ring corresponding to the target remote cluster; and obtain the parameter information of the target remote cluster based on the connection name, the IP, and the key ring.

4. An electronic device, characterized in that, It includes: a memory for storing a computer program; a processor for executing the computer program to implement the storage cluster data backup method according to claim 1 or 2.

5. A computer-readable storage medium, characterized in that, for storing a computer program; wherein the computer program, when executed by the processor, implements the storage cluster data backup method according to claim 1 or 2.

Citation Information

Patent Citations

  • A disaster recovery platform and a disaster recovery method

    CN109597718A

  • Disaster recovery backup method and device, equipment and storage medium

    CN113672436A

  • Data synchronization method and device

    CN113821490A