Data backup method, device and equipment based on a distributed storage system
By slicing and parallel uploading and verification of data in a distributed storage system, the problem of low data upload and verification efficiency in the prior art is solved, and a more efficient data backup verification process is realized.
Patent Information
- Application Number
- CN202210528420.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-16
AI Technical Summary
In the prior art, the methods of data upload and verification are relatively low, resulting in low data verification efficiency.
The data backup method based on the distributed storage system is adopted, and data is uploaded in parallel by calculating the password hash value of the sharded data for verification, ensuring that the data is successfully backed up in the off-site storage device.
Through sharded parallel upload and verification, the verification time is significantly shortened and the verification efficiency of data backup is improved.
Smart Images

Figure CN114816858B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data backup method, device and equipment in a distributed storage system. Background Art
[0002] Data backup and recovery is a very important part of the game business operation process. It can be used for data file checking, data scanning, data rollback, and business function gameplay testing in the game operation process. Data backup and recovery is generally divided into three stages: data backup, data upload and verification, and data recovery.
[0003] In the prior art, when executing the data upload and verification process, the upload verification is generally implemented based on the backup tool provided by the Simple Storage Service (S3).
[0004] However, since the data uploading and verification methods in the existing simple storage service are relatively simple, there is a problem of low verification efficiency. Summary of the invention
[0005] The purpose of this application is to provide a data backup method, device and equipment based on a distributed storage system to improve data verification efficiency in view of the deficiencies in the above-mentioned prior art.
[0006] To achieve the above purpose, the technical solution adopted in the embodiment of the present application is as follows:
[0007] In a first aspect, the present invention provides a data backup method based on a distributed storage system, which is applied to a first cluster, wherein the first cluster includes multiple servers to be backed up, the first cluster is communicatively connected to a second cluster, the second cluster includes multiple block devices, a first server to be backed up in the first cluster is mounted with at least one first block device among the multiple block devices, and the first block device includes first backup data that has been backed up for first data to be backed up in the first server to be backed up, the method comprising:
[0008] In response to a first backup request for the first server to be backed up, sharding the first backup data through a first client to obtain a plurality of sharded data;
[0009] Uploading a plurality of the shard data in parallel to an off-site storage device and calculating a first cryptographic hash value of each of the shard data, wherein the off-site storage device calculates a second cryptographic hash value of each of the shard data in parallel after receiving each of the shard data in parallel, and the second cryptographic hash value of each of the shard data is synchronously stored in a first server corresponding to the first client;
[0010] Receive each of the second password hashes sent by the first server through the first client;
[0011] If it is determined that each piece of data is successfully verified based on each of the first password hashes and each of the second password hashes, it is determined that the first backup data is successfully backed up to the off-site storage device.
[0012] In an alternative embodiment, the first server to be backed up determines at least one first block device mounted according to a preset binding relationship, where the preset binding relationship includes: the mapping relationship between the server identifier corresponding to each server to be backed up and the device identifier corresponding to each block device.
[0013] In an alternative embodiment, the method further includes:
[0014] In response to a second backup request for the first server to be backed up, generate a backup instruction according to the second backup request, where the second backup request includes first data to be backed up;
[0015] According to the backup instruction, back up the first data to be backed up to at least one first block device mounted by the first server to be backed up.
[0016] In an alternative embodiment, the method further includes:
[0017] Perform a snapshot operation on the first backup data to generate backup snapshot data, and obtain the backup snapshot address corresponding to the backup snapshot data, where the backup snapshot data includes: the data identifier of the first backup data, the snapshot backup time of the first backup data;
[0018] Generate a snapshot backup mapping relationship according to the backup snapshot data and the backup snapshot address, where the snapshot backup mapping relationship includes: the mapping relationship between the data identifier of the first backup data, the snapshot backup time of the first backup data, and the backup snapshot address.
[0019] In an alternative embodiment, after generating the snapshot backup mapping relationship according to the backup snapshot data and the backup snapshot address, the method further includes:
[0020] In response to a first data recovery request for the first server to be backed up, obtain the snapshot backup mapping relationship, where the data recovery request includes: the data identifier of the first data to be recovered, the first request recovery time;
[0021] Determine the backup snapshot address corresponding to the first data to be recovered according to the data identifier of the first data to be recovered, the first request recovery time, and the snapshot backup mapping relationship;
[0022] Mount the backup snapshot address corresponding to the first data to be restored to the first server to be backed up, and perform logical restoration on the first request data corresponding to the first data restoration request through the target block device corresponding to the backup snapshot address.
[0023] In an alternative embodiment, the method further includes:
[0024] In response to a second data restoration request for the first server to be backed up, obtain a preset binding relationship, where the data restoration request includes: the data identifier of the second data to be restored, and the second requested restoration time;
[0025] Determine the target block device identifier corresponding to the second data to be restored according to the data identifier of the second data to be restored, the second requested restoration time, and the preset binding relationship;
[0026] Perform copy restoration on the second request data corresponding to the second data restoration request through the target block device corresponding to the target block device identifier.
[0027] In an alternative embodiment, after determining that each piece of sharded data is successfully backed up to the off-site storage device, the method further includes:
[0028] In response to a third data restoration request for the first server to be backed up, forward the third data acquisition request to the off-site storage device, where the third data restoration request includes: the data identifier of the third data to be restored, and the third requested restoration time;
[0029] Receive the third request data returned by the off-site storage device according to the third data acquisition request.
[0030] In an alternative embodiment, the preset binding relationship is created according to the distance between the first geographical location information of each block device and the second geographical location information of each server to be backed up.
[0031] In an alternative embodiment, the step of sharding the first backup data through the first client to obtain multiple sharded data includes:
[0032] In response to first upload parameters set through the first client, the first upload parameters include: the size of each piece of sharded data;
[0033] According to the first upload parameters, shard the first backup data through the first client to obtain multiple sharded data.
[0034] In an alternative embodiment, the step of uploading multiple pieces of sharded data to the off-site storage device in parallel includes:
[0035] In response to second upload parameters set by the first client, the second upload parameters include: the number of shard data for parallel upload;
[0036] According to the second upload parameters, multiple pieces of the shard data are uploaded in parallel to the off-site storage device.
[0037] In an optional implementation manner, the method further includes:
[0038] In response to an initial backup request, a logical volume management snapshot is created in the logical volume management partition of the first server to be backed up based on the copy-on-write technology;
[0039] The backing up of the first data to be backed up to at least one first block device mounted by the first server to be backed up includes:
[0040] Backing up the first data to be backed up in the logical volume management partition to at least one first block device mounted by the first server to be backed up.
[0041] In a second aspect, the present invention provides a data backup method based on a distributed storage system, which is applied to an off-site storage device. The method includes:
[0042] Receiving multiple pieces of shard data uploaded in parallel by a first cluster, the multiple pieces of shard data are obtained by the first cluster sharding first backup data through a first client in response to a first backup request. The first cluster includes multiple servers to be backed up, the first cluster is communicatively connected to a second cluster, the second cluster includes multiple block devices, at least one first block device among the multiple block devices is mounted by a first server to be backed up in the first cluster, and the first block device includes first backup data that has been backed up for the first data to be backed up in the first server to be backed up;
[0043] Parallel computing the second cryptographic hash values of the respective pieces of shard data, and storing the respective second cryptographic hash values in a first server corresponding to the first client;
[0044] Sending the second cryptographic hash values to the first client through the first server, so that when the first cluster determines that each piece of shard data passes the verification according to each first cryptographic hash value and each second cryptographic hash value, it is determined that the first backup data is successfully backed up;
[0045] Wherein, each of the first cryptographic hash values is obtained by the first cluster calculating the cryptographic hash value of each piece of shard data.
[0046] Thirdly, the present invention provides a data backup device based on a distributed storage system, which is applied to a first cluster. The first cluster includes multiple servers to be backed up. The first cluster is communicatively connected to a second cluster. The second cluster includes multiple block devices. At least one first block device among the multiple block devices is mounted on a first server to be backed up in the first cluster. The first block device includes first backup data that has been backed up for first backup data in the first server to be backed up. The data backup device includes:
[0047] A sharding module, configured to shard the first backup data through a first client in response to a first backup request, and obtain multiple sharded data;
[0048] A processing module, configured to upload the multiple sharded data to a remote storage device in parallel and calculate a first cryptographic hash value for each of the sharded data. Wherein, after receiving each of the sharded data in parallel, the remote storage device calculates a second cryptographic hash value for each of the sharded data, and the second cryptographic hash value for each of the sharded data is synchronously stored in a first server corresponding to the first client;
[0049] A sending module, configured to receive each of the second cryptographic hash values sent by the first server through the first client;
[0050] A determining module, configured to determine that the first backup data has been successfully backed up to the remote storage device if it is determined that each of the sharded data passes the verification according to each of the first cryptographic hash values and each of the second cryptographic hash values.
[0051] In an optional embodiment, the first server to be backed up determines at least one first block device to be mounted according to a preset binding relationship. The preset binding relationship includes: a mapping relationship between a server identifier corresponding to each server to be backed up and a device identifier corresponding to each block device.
[0052] In an optional embodiment, the data backup device further includes: a backup module, configured to generate a backup instruction according to a second backup request in response to the second backup request for the first server to be backed up. The second backup request includes first backup data;
[0053] According to the backup instruction, back up the first backup data to at least one first block device mounted on the first server to be backed up.
[0054] In an alternative embodiment, the data backup device further includes: a generation module, configured to perform a snapshot operation on the first backup data to generate backup snapshot data, and obtain a backup snapshot address corresponding to the backup snapshot data, where the backup snapshot data includes: a data identifier of the first backup data, and a snapshot backup time of the first backup data;
[0055] Generate a snapshot backup mapping relationship according to the backup snapshot data and the backup snapshot address, where the snapshot backup mapping relationship includes: a mapping relationship among the data identifier of the first backup data, the snapshot backup time of the first backup data, and the backup snapshot address.
[0056] In an alternative embodiment, the data backup device further includes: a first recovery module, configured to, in response to a first data recovery request for the first server to be backed up, obtain the snapshot backup mapping relationship, where the data recovery request includes: a data identifier of the first data to be recovered, and a first requested recovery time;
[0057] Determine a backup snapshot address corresponding to the first data to be recovered according to the data identifier of the first data to be recovered, the first requested recovery time, and the snapshot backup mapping relationship;
[0058] Mount the backup snapshot address corresponding to the first data to be recovered to the first server to be backed up, and perform logical recovery on the first requested data corresponding to the first data recovery request through a target block device corresponding to the backup snapshot address.
[0059] In an alternative embodiment, the data backup device further includes: a second recovery module, configured to, in response to a second data recovery request for the first server to be backed up, obtain a preset binding relationship, where the data recovery request includes: a data identifier of the second data to be recovered, and a second requested recovery time;
[0060] Determine a target block device identifier corresponding to the second data to be recovered according to the data identifier of the second data to be recovered, the second requested recovery time, and the preset binding relationship;
[0061] Perform copy recovery on the second requested data corresponding to the second data recovery request through the target block device corresponding to the target block device identifier.
[0062] In an alternative embodiment, the data backup device further includes: a third recovery module, configured to, in response to a third data recovery request for the first server to be backed up, forward the third data acquisition request to the remote storage device, where the third data recovery request includes: a data identifier of the third data to be recovered, and a third requested recovery time;
[0063] Receive the third request data returned by the remote storage device according to the third data acquisition request.
[0064] In an alternative embodiment, the preset binding relationship is created according to the distance between the first geographical location information of each block device and the second geographical location information of each backup server to be backed up.
[0065] In an alternative embodiment, the sharding module is specifically configured to respond to a first upload parameter set by the first client, where the first upload parameter includes: the size of each shard data;
[0066] According to the first upload parameter, the first client shards the first backup data to obtain a plurality of shard data after sharding.
[0067] In an alternative embodiment, the processing module is specifically configured to respond to a first upload parameter set by the first client device, where the first upload parameter includes: the size of each shard data;
[0068] According to the first upload parameter, the first client shards the first backup data to obtain a plurality of shard data after sharding.
[0069] In an alternative embodiment, the data backup device further includes: a creation module, configured to respond to an initial backup request, and create a logical volume management snapshot based on the copy-on-write technology in the logical volume management partition of the first backup server to be backed up; the backup module is specifically configured to back up the first backup data in the logical volume management partition to at least one first block device mounted by the first backup server.
[0070] In a fourth aspect, the present invention provides a data backup device based on a distributed storage system, which is applied to a remote storage device. The data backup device includes:
[0071] A receiving module, configured to receive a plurality of shard data uploaded in parallel by a first cluster. The plurality of shard data are obtained by the first cluster sharding first backup data through a first client in response to a first backup request. The first cluster includes a plurality of backup servers to be backed up. The first cluster is communicatively connected to a second cluster. The second cluster includes a plurality of block devices. At least one first block device among the plurality of block devices is mounted by a first backup server in the first cluster. The first block device includes first backup data that has been backed up for the first backup data in the first backup server;
[0072] A calculation module, configured to calculate the second cryptographic hash values of the respective shard data in parallel and store the respective second cryptographic hash values in the first server corresponding to the first client;
[0073] A sending module, configured to send the second cryptographic hash values to the first client through the first server, so that when the first cluster determines that each shard data is successfully verified according to each first cryptographic hash value and each second cryptographic hash value, it is determined that the first backup data is successfully backed up, where each of the first cryptographic hash values is obtained by the first cluster calculating the cryptographic hash value of each shard data.
[0074] In a fifth aspect, the present invention provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the data backup method based on a distributed storage system according to any one of the foregoing embodiments.
[0075] In a sixth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it performs the steps of the data backup method based on a distributed storage system according to any one of the foregoing embodiments.
[0076] The beneficial effects of this application are:
[0077] A data backup method, device, and equipment based on a distributed storage system provided by an embodiment of the present application. This method can be applied to a first cluster. The first cluster includes multiple servers to be backed up. The first cluster is communicatively connected to a second cluster. The second cluster includes multiple block devices. At least one first block device among the multiple block devices is mounted on a first server to be backed up in the first cluster. The first block device includes first backup data that has been backed up for the first server to be backed up. This method may include: in response to a first backup request for the first server to be backed up, sharding the first backup data through a first client to obtain multiple sharded data; uploading the multiple sharded data to a remote storage device in parallel and calculating a first cryptographic hash value for each sharded data. Among them, after the remote storage device receives each sharded data in parallel, it calculates a second cryptographic hash value for each sharded data in parallel, and the second cryptographic hash value of each sharded data is synchronously stored in a first server corresponding to the first client; receiving, through the first client, each second cryptographic hash value sent by the first server; if it is determined that each sharded data passes the verification according to each first cryptographic hash value and each second cryptographic hash value, it is determined that the first backup data has been successfully backed up to the remote storage device, realizing that the first backup data can be uploaded and verified in parallel in a sharding manner. Compared with the prior art method of first uploading and then verifying the entire first backup data, the verification duration can be effectively shortened, and the verification efficiency of the backup data can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0079] Figure 1 It is a schematic diagram of the architecture of a data backup system provided by an embodiment of the present application;
[0080] Figure 2 It is a data backup method based on a distributed storage system provided by an embodiment of the present application;
[0081] Figure 3 It is another data backup method based on a distributed storage system provided by an embodiment of the present application;
[0082] Figure 4 It is another data backup method based on a distributed storage system provided by an embodiment of the present application;
[0083] Figure 5 It is yet another data backup method based on a distributed storage system provided by an embodiment of the present application;
[0084] Figure 6 Another data backup method based on a distributed storage system provided by an embodiment of the present application;
[0085] Figure 7 Yet another data backup method based on a distributed storage system provided by an embodiment of the present application;
[0086] Figure 8 Another data backup method based on a distributed storage system provided by an embodiment of the present application;
[0087] Figure 9 Yet another data backup method based on a distributed storage system provided by an embodiment of the present application;
[0088] Figure 10 Another data backup method based on a distributed storage system provided by an embodiment of the present application;
[0089] Figure 11 Schematic diagram of the functional modules of a data backup device based on a distributed storage system provided by an embodiment of the present application;
[0090] Figure 12 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0091] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0092] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0093] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0094] Data backup and recovery generally consists of three stages: data backup, data upload and verification, and data recovery. In the prior art, when performing data backup, it is mainly implemented based on a redundant array of independent disks (RAID). For example, multiple hard disks on a single machine can be configured in a RAID5 mode as the database backup medium to provide the data backup function. When performing the data upload and verification process, it is generally implemented based on the backup tool (S3 object storage upload backup tool bsctl) provided by the Simple Storage Service (S3). During the data recovery process, the corresponding data needs to be downloaded from the local disk backup data to the target recovery machine.
[0095] It can be seen that in the prior art, for the process of database backup, since the data backup medium is mainly the local disk, where various types of resource disks share the RAID card resources, the overall performance is affected by the throughput of the RAID card, and it is difficult to ensure the running performance of the data disk during the backup peak period. For the data upload and verification process, the existing method is to calculate the MD5 message-digest algorithm on the local backup file, and then upload it to S3 through the bsctl tool. After the upload, MD5 is required to verify the integrity of the data. This method needs to go through two stages of data upload and MD5 calculation. The overall time is related to the size of the data volume and the network speed. The larger the data volume, the longer the recovery time. Therefore, for the calculation of MD5 for large files, there is a problem of low backup efficiency. For the data recovery process, the duration of data recovery is related to the size of the business data volume. When the business data urgently needs backup and recovery for data verification within a limited maintenance period, there will be a problem of long recovery time, and in some scenarios, the downloaded data may be incomplete, resulting in the situation that the recovery is unavailable.
[0096] In view of this, the embodiments of the present application provide a data backup method. By applying this method during the database backup process, the diversity of backup methods can be realized, and the running performance of the first cluster can be ensured; during the data recovery process, the diversity of data recovery methods can be realized, the reliability of data recovery can be ensured, and the data recovery efficiency can be improved; during the data upload and verification process, the data upload and verification efficiency can be effectively improved.
[0097] Figure 1 The architecture schematic diagram of a data backup system provided by the embodiments of the present application is as Figure 1As shown, the data backup system may include a first cluster 120, a second cluster 130, and a storage resource scheduling platform 110. Among them, the first cluster 120 and the second cluster 130 may be communicatively connected through a B network 150. The first cluster 120 includes multiple servers to be backed up, and the second cluster 130 includes multiple block devices. At least one first block device 131 among the multiple block devices is mounted on a first server to be backed up 121 in the first cluster 120. The first block device 131 includes first backup data that has been backed up for the first data to be backed up in the first server to be backed up.
[0098] Optionally, the storage resource scheduling platform 110 may be deployed in a preset server for centrally managing the first cluster 120 and / or the second cluster 130, and it may be communicatively connected to the first cluster 120 and the second cluster 130 respectively through an A network 140. In some embodiments, configuration may be specifically implemented through a preset server. Among them, the preset server may be configured to include: a preset interface, for example, a restful interface, through which the management of storage resources (such as storage content, storage space size) in the first cluster is realized.
[0099] Among them, the first cluster 120 may be a DB cluster, and the servers to be backed up may be database servers. The database servers may include data to be backed up, such as game service data, video surveillance data, payment service data, etc., which is not limited herein. For better understanding of this application, the following method embodiments will be described by taking a game service scenario as an example. In some embodiments, according to the type of data to be backed up, different types of data may be stored in different servers to be backed up. Of course, the functions of each server to be backed up are not limited herein.
[0100] The second cluster 130 may be a distributed storage cluster, such as a Ceph cluster. Ceph is a distributed file system that can support object storage, block device storage, and file system storage, and can achieve horizontal dynamic expansion and apply for the size of resources as needed. Of course, the type of the second cluster is not limited thereto. For example, the second cluster may also be a Hadoop cluster. The block devices in the second cluster 130 may be distributed consistent storage devices (RADOS Block Device, RBD). Of course, they may also be other block devices, which is not limited herein. Among them, a block is an ordered byte, and the common block size is 512 bytes, such as a hard disk.
[0101] Optionally, the first server to be backed up 121 can be any one of the multiple servers to be backed up, and the first block device 131 can be any one of the multiple block devices, which is not limited herein. Based on the above description, it can be seen that the first server to be backed up 121 can mount one or more first block devices 131, and the first block device 131 contains the first backup data in the first server to be backed up, that is, the first block device 131 backs up some of the backup data in the first server to be backed up.
[0102] Figure 2 A data backup method based on a distributed storage system provided by an embodiment of the present application can be applied to the first cluster of the above data backup system, such as Figure 2 shown, the method may include:
[0103] S101. In response to a first backup request for the first server to be backed up, the first client slices the first backup data to obtain multiple sliced data.
[0104] Among them, the first client can be installed in the first cluster. In response to the first backup request for the first server to be backed up, the first cluster can slice the first backup data through the first client to slice the first backup data into multiple sliced data. Optionally, the size of each sliced data can be 16MB / block, 100MB / block, etc., which is not limited herein. Optionally, the above first client can be a bsctl client.
[0105] S102. Parallelly upload the multiple sliced data to the off-site storage device and calculate the first cryptographic hash value of each sliced data.
[0106] Among them, after the off-site storage device parallelly receives each sliced data, it parallelly calculates the second cryptographic hash value of each sliced data, and the second cryptographic hash value of each sliced data is synchronously stored in the first server corresponding to the first client.
[0107] The off-site storage device may refer to a storage device deployed off-site from the first server to be backed up. Optionally, the off-site storage device may be based on object storage. In some embodiments, object storage software may be installed in the off-site storage device to support object storage functions. The off-site storage device may deploy the first server corresponding to the first client. Based on the above description, to ensure the reliability of data backup, the first cluster may parallelly upload the multiple sliced data to the off-site storage device, and during the upload process, the first cryptographic hash value of each sliced data may be calculated and stored in parallel; for the off-site storage device, after the off-site storage device parallelly receives each sliced data, it may parallelly calculate the second cryptographic hash value corresponding to each sliced data and synchronously store it in the first server.
[0108] Among them, the first password hash value and the second password hash value can be calculated based on the same password hash function. For example, they can be calculated based on the Message Digest Algorithm MD5. Of course, the specific acquisition method is not limited to this. Optionally, the first server corresponding to the first client can be the bsctl server. In some embodiments, the bsctl server realizes storage through an S3 bucket.
[0109] S103. Receive each second password hash value sent by the first server through the first client.
[0110] After obtaining the second password hash values corresponding to each shard data, the first server can send each second password hash value to the first client through the communication connection with the first client. It can be understood that at this time, the first cluster can receive each second password hash value sent by the first server through the first client.
[0111] S104. If it is determined that each shard data passes the verification according to each first password hash value and each second password hash value, it is determined that the first backup data is successfully backed up to the off-site storage device.
[0112] Based on the above description, it can be seen that at this time, the first client not only includes the first password hash values of each shard data, but also includes the second password hash values of each shard data. Then, at this time, each first password hash value and each second password hash value can be compared. It can be understood that if the password hash functions corresponding to the first password hash value and the second password hash value are the same, then for the same shard data, the values of the two should be the same. If the values of the two are the same, it can be determined that the corresponding shard data passes the verification, and it can also be determined that the first backup data is successfully backed up to the off-site storage device, realizing the consistency verification of the first backup data in a sharded manner. Through experimental verification, compared with the prior art, the embodiment of the present application can save 0.12 minutes / GB of verification time; otherwise, if the values of the two are different, it can be determined that the corresponding shard data fails the verification, that is, it can be determined that the first backup data backup fails. Compared with the prior art, without losing the data upload consistency, the embodiment of the present application can improve the upload efficiency and eliminate the time for separately calculating MD5. In particular, for the first backup data with a large amount of data, the embodiment of the present application can effectively improve the upload verification efficiency, and thus can improve the data backup efficiency.
[0113] In summary, the embodiment of the present application provides a data backup method based on a distributed storage system. This method can be applied to a first cluster, which includes multiple servers to be backed up. The first cluster is communicatively connected to a second cluster, which includes multiple block devices. At least one first block device among the multiple block devices is mounted on a first server to be backed up in the first cluster. The first block device contains first backup data that has been backed up for the first server to be backed up. The method may include: in response to a first backup request for the first server to be backed up, fragmenting the first backup data through a first client to obtain multiple fragmented data; uploading the multiple fragmented data to a remote storage device in parallel and calculating the first cryptographic hash value of each fragmented data. Among them, after receiving each fragmented data in parallel, the remote storage device calculates the second cryptographic hash value of each fragmented data in parallel, and the second cryptographic hash value of each fragmented data is synchronously stored in the first server corresponding to the first client; receiving, through the first client, each second cryptographic hash value sent by the first server; if it is determined that each fragmented data passes the verification according to each first cryptographic hash value and each second cryptographic hash value, it is determined that the first backup data has been successfully backed up to the remote storage device, realizing parallel uploading and verification of the first backup data in a fragmented manner. Compared with the prior art method of uploading the entire first backup data first and then verifying, the verification duration can be effectively shortened, and the verification efficiency of the backup data can be improved.
[0114] Based on the above description, in some embodiments, the above-mentioned first cluster can be stored based on a distributed cloud disk and is configured to include a Quality of Service (QoS) policy to avoid resource preemption between cloud disks, which can improve the overall utilization rate of each server to be backed up in the first cluster and reduce the resource cost of service backup.
[0115] Optionally, the first server to be backed up can determine at least one first block device to be mounted according to a preset binding relationship. The preset binding relationship may include: the mapping relationship between the server identifier corresponding to each server to be backed up and the device identifier corresponding to each block device.
[0116] Among them, the number of first block devices mounted on the first server to be backed up can be flexibly set according to the actual application scenario, or can be configured according to human experience, which is not limited herein.
[0117] In some embodiments, the preset binding relationship can be set according to the usage conditions of each block device in the second cluster, the requirements of each server to be backed up for backup space, etc. Of course, when setting, it can also be set for the purpose of balancing the second cluster, which is not limited herein.
[0118] In some embodiments, the preset binding relationship may be specifically created according to the distance between the first geographical location information of each block device and the second geographical location information of each server to be backed up. Optionally, the server identifier corresponding to each server to be backed up may include the geographical location identifier corresponding to the server to be backed up, and the device identifier corresponding to each block device may include the geographical location identifier corresponding to each block device. Optionally, when binding, the geographical location relationship between each server to be backed up and each block device may be used for proximity binding, so as to achieve proximity allocation and access of storage resources, optimize network traffic and reduce latency, and ensure the minimum consumption of network bandwidth during data backup and recovery. For example, the first server to be backed up and at least one first block device determined according to the preset binding relationship may be distributed in adjacent cabinets in the same physical computer room.
[0119] Based on the above description, optionally, the above preset binding relationship may be obtained through a storage resource scheduling platform, and the first cluster may be configured to include a preset client. In some embodiments, the first cluster may receive the preset binding relationship sent by a preset server in the storage resource scheduling platform through the preset client, and establish a mounting association relationship between each server to be backed up and each block device through the preset client.
[0120] Of course, it should be noted that in some embodiments, the block devices corresponding to each server to be backed up may also be updated according to the actual application scenario requirements to achieve the purpose of expansion or contraction.
[0121] Figure 3 Another data backup method based on a distributed storage system provided by an embodiment of the present application. Optionally, as Figure 3 shown, the above method further includes:
[0122] S201. In response to a second backup request for a first server to be backed up, generate a backup instruction according to the second backup request.
[0123] The second backup request includes an identifier of the first data to be backed up. Optionally, the identifier of the first data to be backed up may include: the identifier of the first server to be backed up, the storage address of the first data to be backed up, the project identifier corresponding to the first data to be backed up, etc., which are not limited herein.
[0124] S202. According to the backup instruction, back up the first data to be backed up to at least one first block device mounted by the first server to be backed up.
[0125] In some embodiments, to meet the backup requirements in different scenarios, according to the backup instruction, the first cluster can obtain the first data to be backed up, and then can back up the first data to be backed up to at least one first block device mounted on the first server to be backed up, so that at least one first block device includes the first data to be backed up, and the backup process of the first data to be backed up is completed. Optionally, when performing the backup, the specified backup can be performed according to the identifiers of the first block devices.
[0126] Figure 4 Another data backup method based on a distributed storage system provided by an embodiment of the present application. Optionally, taking the application of the data backup method provided by the embodiment of the present application to the game field as an example for illustration, in order to speed up the rapid recovery of different data backups during game maintenance and meet the timeliness requirements of game rollback and file scanning, the embodiment of the present application can further create a snapshot backup of the first backup data backed up to at least one first block device. As Figure 4 shown, the above method further includes:
[0127] S301. Perform a snapshot operation on the first backup data to generate backup snapshot data, and obtain the backup snapshot address corresponding to the backup snapshot data.
[0128] Among them, by performing a snapshot operation on the first backup data, the backup snapshot data at the current snapshot moment of the first backup data can be generated. It can be understood that in order to distinguish the backup snapshot data, the backup snapshot data can include: the data identifier of the first backup data, the snapshot backup time of the first backup data. Optionally, the obtained backup snapshot data can be stored at a preset location or a custom location. It can be understood that at this time, the address corresponding to the preset location or the custom location is the backup snapshot address corresponding to the backup snapshot data.
[0129] S302. Generate a snapshot backup mapping relationship according to the backup snapshot data and the backup snapshot address.
[0130] By binding and mapping the backup snapshot data and the backup snapshot address, the snapshot backup mapping relationship can be obtained. The obtained snapshot backup mapping relationship can be stored at a specified location. The snapshot backup mapping relationship may include: the mapping relationship between the data identifier of the first backup data, the snapshot backup time of the first backup data, and the backup snapshot address. Based on the above description, it can be understood that the first backup data can be quickly restored based on the snapshot backup mapping relationship. Applying the embodiments of the present application retains the data of the first backup data at the current snapshot moment, avoids modifying the first backup data when writing data to the block device later, and also facilitates subsequent data restoration based on the backup snapshot data corresponding to the first backup data. Of course, in some embodiments, based on the data retention requirements of the game service, the snapshot can also be periodically cleaned to avoid the invalid storage of expired backup snapshot data.
[0131] Figure 5 Another data backup method based on a distributed storage system provided by the embodiments of the present application. Optionally, as Figure 5 shown, after generating the snapshot backup mapping relationship according to the backup snapshot data and the backup snapshot address, the method further includes:
[0132] S401. In response to a first data recovery request for the first server to be backed up, obtain the snapshot backup mapping relationship.
[0133] Wherein, the data recovery request includes: the data identifier of the first data to be recovered, and the first requested recovery time.
[0134] S402. Determine the backup snapshot address corresponding to the first data to be recovered according to the data identifier of the first data to be recovered, the first requested recovery time, and the snapshot backup mapping relationship.
[0135] S403. Mount the backup snapshot address corresponding to the first data to be recovered to the first server to be backed up, and perform logical recovery on the first requested data corresponding to the first data recovery request through the target block device corresponding to the backup snapshot address.
[0136] The first data recovery request is used to request a data recovery operation. In response to the first data recovery request for the first server to be backed up, the first cluster can obtain a snapshot backup mapping relationship. It can be understood that further, according to the data identifier of the first data to be recovered and the first requested recovery time, the backup snapshot address corresponding to the first data to be recovered can be determined from the snapshot backup mapping relationship. Then, the backup snapshot address can be mounted to the first server to be backed up, so that the corresponding target block device can be obtained based on the backup snapshot address. Through the target block device, logical recovery can be performed on the first requested data corresponding to the first data recovery request. Applying the embodiments of the present application, since there is no need to copy the entire data, only cloning and mounting based on the snapshot technology are required. Therefore, compared with the physical copy of data, the embodiments of the present application can effectively improve the data recovery efficiency. Taking 300GB of data as an example, the physical copy takes 1 hour, while applying the embodiments of the present application only takes 10 minutes to complete, which can significantly improve the copy efficiency. In addition, compared with the method of data recovery through off-site storage devices, the embodiments of the present application can support temporary backup recovery through the snapshot function. Since there is no need to download backup data from off-site storage devices for recovery, the data recovery efficiency can be greatly improved.
[0137] For example, taking the application scenario as a game scenario and the first cluster including a mongodb database, if it is necessary to recover the data of the mongodb database at the first requested recovery time, referring to the above embodiments, optionally, after determining the backup snapshot address corresponding to the mongodb database, the snapshot of the target block device of each shard data corresponding to the mongodb database can be cloned according to the snapshot of the target block device disk corresponding to the backup snapshot address, and then the cloned disk can be mounted to the first server to be backed up and started as the data directory of the mongodb process. Further, after the mongodb process is started, the metadata of the data in the mongodb database can be modified, and the data integrity can be verified. Applying the embodiments of the present application, the proprietary temporary backup snapshot function for the game scenario is supported based on the snapshot clone technology, and the temporary backup data can be quickly recovered, meeting the high timeliness requirements for data recovery during game project maintenance.
[0138] Figure 6 Another data backup method based on a distributed storage system provided by the embodiments of the present application. Optionally, in some embodiments, in order to further achieve the diversity of data recovery methods, of course, physical copy recovery of data can also be achieved through block devices. As Figure 6 shown, the above method further includes:
[0139] S501. In response to a second data recovery request for a first server to be backed up, obtain a preset binding relationship.
[0140] The data recovery request includes: the data identifier of the second data to be recovered, and the second requested recovery time.
[0141] S502. Determine the target block device identifier corresponding to the second data to be recovered according to the data identifier of the second data to be recovered, the second requested recovery time, and the preset binding relationship.
[0142] S503. Copy and recover the second requested data corresponding to the second data recovery request through the target block device corresponding to the target block device identifier.
[0143] For the content of steps S501 and S502, reference can be made to the descriptions of S401 and S402 above, and details will not be elaborated here.
[0144] Optionally, after determining the target block device identifier corresponding to the second data to be recovered, the second requested data corresponding to the second data recovery request in the target block device can be copied to the first server to be backed up through a file synchronization service (such as Rsync service), so as to achieve data copy and recovery.
[0145] Figure 7 Another data backup method based on a distributed storage system provided by an embodiment of the present application. Optionally, in some embodiments, in order to further achieve the diversity of data recovery methods, data recovery can of course also be achieved through off-site storage devices. As Figure 7 shown, after determining that each shard data is successfully backed up to the off-site storage device, the above method further includes:
[0146] S601. In response to a third data recovery request for a first server to be backed up, forward a third data acquisition request to the off-site storage device.
[0147] The third data recovery request includes: the data identifier of the third data to be recovered, and the third requested recovery time.
[0148] S602. Receive the third requested data returned by the off-site storage device according to the third data acquisition request.
[0149] The first cluster forwards the third data recovery request to the off-site storage device. After receiving the third data recovery request, the off-site storage device can query and obtain the third requested data corresponding to the third data recovery request from the off-site storage device according to the data identifier of the third data to be recovered and the third requested recovery time, and send the third requested data to the first server to be backed up. It can be understood that at this time, the first server to be backed up can perform data recovery according to the third requested data.
[0150] In summary, it can be seen that the embodiments of the present application provide three data backup methods. Correspondingly, three data recovery methods are also provided. In this way, the reliability and adaptability of the first cluster can be improved, and the backup requirements and data recovery requirements in different scenarios can be met.
[0151] Figure 8 Another data backup method based on a distributed storage system provided by the embodiments of the present application. Optionally, as Figure 8 shown, the step of fragmenting the first backup data by the first client to obtain multiple fragmented data after fragmentation may include:
[0152] S701. Respond to the first upload parameter set by the first client.
[0153] Among them, the first upload parameter includes: the size of each fragmented data.
[0154] S702. According to the first upload parameter, fragment the first backup data by the first client to obtain multiple fragmented data after fragmentation.
[0155] Optionally, the first client can be configured to set the first upload parameter of the fragmented data. The size of each fragmented data can be 16MB, 100MB, etc., and can be flexibly set according to the actual application scenario, which is not limited herein.
[0156] In response to the setting operation of the first upload parameter, the first cluster can fragment the first backup data by the first client according to the first upload parameter. Among them, the size of each fragmented data after fragmentation should be less than or equal to the setting of the first upload parameter. For example, in some embodiments, the size of the first backup data is not an integer multiple of the first upload parameter, then there will be fragmented data with a data volume less than the first upload parameter among the multiple fragmented data after fragmentation.
[0157] Figure 9 Another data backup method based on a distributed storage system provided by the embodiments of the present application. Optionally, as Figure 9 shown, the parallel uploading of multiple fragmented data to the off-site storage device includes:
[0158] S801. Respond to the second upload parameter set by the first client. The second upload parameter includes: the number of fragmented data to be uploaded in parallel.
[0159] S802. According to the second upload parameter, parallel upload multiple fragmented data to the off-site storage device.
[0160] Optionally, the first client can be configured to set the number of sharded data for parallel upload. For example, the number of sharded data for parallel upload can be 10, 20, etc., which can be flexibly set according to the operating conditions of the first cluster and / or the off-site storage device, and is not limited herein.
[0161] Among them, in response to the second upload parameter set by the first client, the first cluster can set the number of multiple sharded data to be uploaded in parallel to the off-site storage device according to the second upload parameter. For example, the number of sharded data to be uploaded in parallel to the off-site storage device can be set to at most 5.
[0162] Optionally, the above method further includes: in response to an initial backup request, creating a logical volume management snapshot based on the copy-on-write technology in the logical volume management partition of the first server to be backed up.
[0163] In some embodiments, the initial backup request can be triggered according to a preset backup frequency (for example, at 8:00 am every day, the first cluster can generate an initial backup request according to a preset configuration file) or triggered by a user, and is not limited herein. The first server to be backed up can create an LVM partition based on the Logical Volume Manager (LVM) mechanism.
[0164] In response to the initial backup request, the first cluster can create an LVM snapshot on the LVM partition of the first server to be backed up based on the copy-on-write technology, so as to achieve a consistent backup of the status of the LVM partition. In this way, the data of the first backup data at the current snapshot moment can be retained, avoiding the modification of the first backup data when writing data to the block device later, and also facilitating the subsequent data recovery based on the backup snapshot data corresponding to the first backup data.
[0165] Correspondingly, the step of backing up the first data to be backed up to at least one first block device mounted by the first server to be backed up may include:
[0166] Backing up the first data to be backed up in the logical volume management partition to at least one first block device mounted by the first server to be backed up.
[0167] Based on the above description, further, at this time, the first data to be backed up in the LVM partition of the first server to be backed up can be backed up to at least one first block device mounted by the first server to be backed up. For example, the first data to be backed up in the LVM snapshot can be copied to at least one associated mounted first block device through the rsync service.
[0168] Figure 10 Another data backup method based on a distributed storage system provided by the embodiments of the present application. Optionally, asFigure 10 As shown in the figure, an embodiment of the present application provides a data backup method based on a distributed storage system. This method can be applied to the off-site storage device in the above data backup system. The method includes:
[0169] S901. Receive multiple shard data uploaded in parallel by the first cluster. The multiple shard data are obtained by the first cluster through the first client sharding the first backup data in response to the first backup request.
[0170] Among them, the first cluster includes multiple servers to be backed up. The first cluster is communicatively connected to the second cluster. The second cluster includes multiple block devices. At least one first block device among the multiple block devices is mounted on the first server to be backed up in the first cluster. The first block device includes the first backup data that has been backed up for the first backup data in the first server to be backed up.
[0171] S902. Calculate the second cryptographic hash value of each shard data in parallel and store each second cryptographic hash value in the first server corresponding to the first client.
[0172] S903. Send the second cryptographic hash value to the first client through the first server, so that when the first cluster determines that each shard data passes the verification based on each first cryptographic hash value and each second cryptographic hash value, it is determined that the first backup data is successfully backed up.
[0173] Among them, each first cryptographic hash value is obtained by the first cluster calculating the cryptographic hash value of each shard data.
[0174] For this part of the content, reference can be made to the relevant part mentioned above, and details will not be repeated here. Applying the embodiment of the present application, it is realized that the first backup data can be uploaded and verified in parallel by sharding. Compared with the method of uploading the entire first backup data first and then verifying in the prior art, the verification duration can be effectively shortened and the verification efficiency of the backup data can be improved.
[0175] Figure 11 This is a schematic diagram of the functional modules of a data backup device provided by an embodiment of the present application. This device can be applied to the first cluster. The first cluster includes multiple servers to be backed up. The first cluster is communicatively connected to the second cluster. The second cluster includes multiple block devices. At least one first block device among the multiple block devices is mounted on the first server to be backed up in the first cluster. The first block device includes the first backup data that has been backed up for the first backup data in the first server to be backed up. The basic principle and the technical effects generated by this device are the same as those of the corresponding method embodiment. For a brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the method embodiment. As Figure 11 shown, the data backup device 200 includes:
[0176] The sharding module 210 is configured to, in response to the first backup request, shard the first backup data through the first client to obtain multiple sharded data pieces;
[0177] The processing module 220 is configured to upload multiple sharded data pieces to the off-site storage device in parallel and calculate the first cryptographic hash value of each sharded data piece. Wherein, after receiving each sharded data piece in parallel, the off-site storage device calculates the second cryptographic hash value of each sharded data piece in parallel, and the second cryptographic hash value of each sharded data piece is synchronously stored in the first server corresponding to the first client;
[0178] The sending module 230 is configured to receive, through the first client, each second cryptographic hash value sent by the first server;
[0179] The determining module 240 is configured to determine that the first backup data is successfully backed up to the off-site storage device if it is determined that each sharded data piece passes the verification according to each first cryptographic hash value and each second cryptographic hash value.
[0180] In an alternative embodiment, the first server to be backed up determines at least one first block device mounted according to a preset binding relationship, where the preset binding relationship includes: the mapping relationship between the server identifier corresponding to each server to be backed up and the device identifier corresponding to each block device.
[0181] In an alternative embodiment, the data backup device further includes: a backup module, configured to, in response to a second backup request for the first server to be backed up, generate a backup instruction according to the second backup request, where the second backup request includes the first data to be backed up;
[0182] According to the backup instruction, back up the first data to be backed up to at least one first block device mounted by the first server to be backed up.
[0183] In an alternative embodiment, the data backup device further includes: a generating module, configured to perform a snapshot operation on the first backup data to generate backup snapshot data, and obtain the backup snapshot address corresponding to the backup snapshot data, where the backup snapshot data includes: the data identifier of the first backup data, the snapshot backup time of the first backup data;
[0184] According to the backup snapshot data and the backup snapshot address, generate a snapshot backup mapping relationship, where the snapshot backup mapping relationship includes: the mapping relationship between the data identifier of the first backup data, the snapshot backup time of the first backup data, and the backup snapshot address.
[0185] In an alternative embodiment, the data backup device further includes: a first recovery module, configured to obtain a snapshot backup mapping relationship in response to a first data recovery request for the first server to be backed up, where the data recovery request includes: a data identifier of the first data to be recovered, and a first requested recovery time;
[0186] Determine a backup snapshot address corresponding to the first data to be recovered according to the data identifier of the first data to be recovered, the first requested recovery time, and the snapshot backup mapping relationship;
[0187] Mount the backup snapshot address corresponding to the first data to be recovered to the first server to be backed up, and perform logical recovery on the first requested data corresponding to the first data recovery request through a target block device corresponding to the backup snapshot address.
[0188] In an alternative embodiment, the data backup device further includes: a second recovery module, configured to obtain a preset binding relationship in response to a second data recovery request for the first server to be backed up, where the data recovery request includes: a data identifier of the second data to be recovered, and a second requested recovery time;
[0189] Determine a target block device identifier corresponding to the second data to be recovered according to the data identifier of the second data to be recovered, the second requested recovery time, and the preset binding relationship;
[0190] Perform copy recovery on the second requested data corresponding to the second data recovery request through a target block device corresponding to the target block device identifier.
[0191] In an alternative embodiment, the data backup device further includes: a third recovery module, configured to forward the third data acquisition request to the off-site storage device in response to a third data recovery request for the first server to be backed up, where the third data recovery request includes: a data identifier of the third data to be recovered, and a third requested recovery time;
[0192] Receive the third requested data returned by the off-site storage device according to the third data acquisition request.
[0193] In an alternative embodiment, the preset binding relationship is created according to the distance between the first geographical location information of each block device and the second geographical location information of each server to be backed up.
[0194] In an alternative embodiment, the sharding module is specifically configured to respond to first upload parameters set by the first client, where the first upload parameters include: the size of each sharded data;
[0195] According to the first upload parameter, the first client shards the first backup data to obtain multiple sharded data.
[0196] In an alternative embodiment, the processing module is specifically configured to respond to the first upload parameter set by the first client device, where the first upload parameter includes: the size of each sharded data.
[0197] According to the first upload parameter, the first client shards the first backup data to obtain multiple sharded data.
[0198] In an alternative embodiment, the data backup device further includes: a creation module, configured to create a logical volume management snapshot based on the copy-on-write technology in the logical volume management partition of the first server to be backed up in response to an initial backup request; the backup module is specifically configured to back up the first data to be backed up in the logical volume management partition to at least one first block device mounted on the first server to be backed up.
[0199] Optionally, the present invention provides a data backup device based on a distributed storage system, which is applied to a remote storage device. The basic principle and the technical effects generated by this device are the same as those of the corresponding method embodiment described above. For the sake of brief description, for the parts not mentioned in this embodiment, reference may be made to the corresponding content in the method embodiment. Optionally, the data backup device may include:
[0200] A receiving module, configured to receive multiple sharded data uploaded in parallel by the first cluster. The multiple sharded data are obtained by the first cluster sharding the first backup data through the first client in response to a first backup request. The first cluster includes multiple servers to be backed up, the first cluster is communicatively connected to a second cluster, the second cluster includes multiple block devices, at least one first block device among the multiple block devices is mounted on the first server to be backed up in the first cluster, and the first block device includes the first backup data that has been backed up for the first data to be backed up in the first server to be backed up.
[0201] A calculation module, configured to calculate the second cryptographic hash value of each sharded data in parallel and store each second cryptographic hash value in the first server corresponding to the first client.
[0202] A sending module, configured to send the second cryptographic hash value to the first client through the first server, so that when the first cluster determines that each sharded data passes the verification based on each first cryptographic hash value and each second cryptographic hash value, it is determined that the first backup data backup is successful, where each first cryptographic hash value is obtained by the first cluster calculating the cryptographic hash value of each sharded data.
[0203] The above device is used to execute the method provided in the foregoing embodiment, and its implementation principle and technical effects are similar, so details are not described herein again.
[0204] The above modules may be one or more integrated circuits configured to implement the above method. For example: one or more application specific integrated circuits (ASICs), or, one or more microprocessors, or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element dispatching program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0205] Figure 12 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 12 shown, the electronic device may include: a processor 310, a storage medium 320, and a bus 330. The storage medium 320 stores machine-readable instructions executable by the processor 310. When the electronic device runs, the processor 310 communicates with the storage medium 320 through the bus 330, and the processor 310 executes the machine-readable instructions to execute the steps of the above method embodiment. The specific implementation manner and technical effects are similar, and details are not described herein again.
[0206] Optionally, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the above method embodiment. The specific implementation manner and technical effects are similar, and details are not described herein again.
[0207] In several embodiments provided by the present application, it should be understood that the disclosed device and method may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some interfaces. The indirect coupling or communication connection of devices or units may be in an electrical, mechanical or other form.
[0208] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0209] In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, may exist physically alone for each unit, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0210] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit stored in a storage medium includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods of each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (English: Read-Only Memory, abbreviated as: ROM), random access memories (English: Random Access Memory, abbreviated as: RAM), magnetic disks or optical discs that can store program codes.
[0211] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0212] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A data backup method based on a distributed storage system, characterized in that, it is applied to a first cluster, the first cluster includes multiple servers to be backed up, the first cluster is communicatively connected to a second cluster, the second cluster includes multiple block devices, at least one first block device among the multiple block devices is mounted on a first server to be backed up in the first cluster, and the first block device includes first backup data that has been backed up for the first backup data in the first server to be backed up. The method includes: In response to a first backup request for the first server to be backed up, the first client slices the first backup data to obtain multiple sliced data; Parallel upload the multiple sliced data to a remote storage device and calculate the first cryptographic hash value of each sliced data. Among them, after the remote storage device receives each sliced data in parallel, it calculates the second cryptographic hash value of each sliced data in parallel, and the second cryptographic hash value of each sliced data is synchronously stored in the first server corresponding to the first client; Receive, by the first client, each of the second cryptographic hash values sent by the first server; If it is determined that each sliced data passes the verification according to each first cryptographic hash value and each second cryptographic hash value, it is determined that the first backup data is successfully backed up to the remote storage device. Among them, the first server to be backed up determines at least one first block device it mounts according to a preset binding relationship, and the preset binding relationship is created according to the distance between the first geographical location information of each block device and the second geographical location information of each server to be backed up.
2. The method according to claim 1, characterized in that, the preset binding relationship includes: the mapping relationship between the server identifier corresponding to each server to be backed up and the device identifier corresponding to each block device.
3. The method according to claim 1, characterized in that, the method further includes: In response to a second backup request for the first server to be backed up, generate a backup instruction according to the second backup request, and the second backup request includes first backup data; According to the backup instruction, back up the first backup data to at least one first block device mounted on the first server to be backed up.
4. The method according to claim 1, characterized in that, the method further includes: Perform a snapshot operation on the first backup data to generate backup snapshot data, and obtain the backup snapshot address corresponding to the backup snapshot data. Among them, the backup snapshot data includes: the data identifier of the first backup data, the snapshot backup time of the first backup data; Generate a snapshot backup mapping relationship according to the backup snapshot data and the backup snapshot address. The snapshot backup mapping relationship includes: the mapping relationship between the data identifier of the first backup data, the snapshot backup time of the first backup data, and the backup snapshot address.
5. The method according to claim 4, characterized in that, after generating the snapshot backup mapping relationship according to the backup snapshot data and the backup snapshot address, the method further includes: In response to a first data recovery request for the first server to be backed up, obtain a snapshot backup mapping relationship, where the data recovery request includes: a data identifier of first data to be recovered, and a first requested recovery time; Determine a backup snapshot address corresponding to the first data to be recovered according to the data identifier of the first data to be recovered, the first requested recovery time, and the snapshot backup mapping relationship; Mount the backup snapshot address corresponding to the first data to be recovered to the first server to be backed up, and perform logical recovery on first requested data corresponding to the first data recovery request through a target block device corresponding to the backup snapshot address.
6. The method according to claim 2, wherein, the method further includes: In response to a second data recovery request for the first server to be backed up, obtain a preset binding relationship, where the data recovery request includes: a data identifier of second data to be recovered, and a second requested recovery time; Determine a target block device identifier corresponding to the second data to be recovered according to the data identifier of the second data to be recovered, the second requested recovery time, and the preset binding relationship; Perform copy recovery on second requested data corresponding to the second data recovery request through a target block device corresponding to the target block device identifier.
7. The method according to claim 1, wherein, after determining that each of the shard data is successfully backed up to the off-site storage device, the method further includes: In response to a third data recovery request for the first server to be backed up, forward the third data acquisition request to the off-site storage device, where the third data recovery request includes: a data identifier of third data to be recovered, and a third requested recovery time; Receive third requested data returned by the off-site storage device according to the third data acquisition request.
8. The method according to claim 1, wherein, the obtaining, by the first client, of multiple shard data after sharding the first backup data includes: In response to first upload parameters set by the first client, the first upload parameters including: the size of each shard data; Shard the first backup data through the first client according to the first upload parameters to obtain multiple shard data after sharding.
9. The method according to claim 1, wherein, the parallel uploading of multiple shard data to the off-site storage device includes: In response to second upload parameters set by the first client, the second upload parameters including: the number of shard data for parallel uploading; Parallel upload multiple shard data to the off-site storage device according to the second upload parameters.
10. The method according to claim 3, wherein, the method further includes: In response to an initial backup request, create a logical volume management snapshot in a logical volume management partition of the first server to be backed up based on the copy-on-write technology; Backing up the first data to be backed up to at least one first block device mounted by the first server to be backed up includes: backing up the first data to be backed up in the logical disk volume management partition to at least one first block device mounted by the first server to be backed up.
11. A data backup method based on a distributed storage system, characterized in that it is applied to a remote storage device, and the method includes: Receiving a plurality of sharded data uploaded in parallel by a first cluster, the plurality of sharded data being obtained by the first cluster sharding first backup data through a first client in response to a first backup request. The first cluster includes a plurality of servers to be backed up, the first cluster is communicatively connected to a second cluster, the second cluster includes a plurality of block devices, and at least one first block device among the plurality of block devices is mounted on a first server to be backed up in the first cluster. The first block device includes first backup data that has been backed up for the first data to be backed up in the first server to be backed up. Wherein, the first server to be backed up determines at least one first block device to be mounted according to a preset binding relationship, and the preset binding relationship is created according to the distance between the first geographical location information of each block device and the second geographical location information of each server to be backed up; Parallel computing the second cryptographic hash value of each of the sharded data and storing each of the second cryptographic hash values in a first server corresponding to the first client; Sending the second cryptographic hash value to the first client through the first server, so that when the first cluster determines that each of the sharded data passes the verification according to each first cryptographic hash value and each of the second cryptographic hash values, it is determined that the first backup data backup is successful, where each of the first cryptographic hash values is obtained by the first cluster calculating the cryptographic hash value of each of the sharded data.
12. A data backup device based on a distributed storage system, characterized in that it is applied to a first cluster, the first cluster includes a plurality of servers to be backed up, the first cluster is communicatively connected to a second cluster, the second cluster includes a plurality of block devices, and at least one first block device among the plurality of block devices is mounted on a first server to be backed up in the first cluster. The first block device includes first backup data that has been backed up for the first data to be backed up in the first server to be backed up. The data backup device includes: A sharding module, configured to shard the first backup data through a first client in response to a first backup request to obtain a plurality of sharded data after sharding; A processing module, configured to upload a plurality of the sharded data to a remote storage device in parallel and calculate the first cryptographic hash value of each of the sharded data. Wherein, the remote storage device calculates the second cryptographic hash value of each of the sharded data in parallel after receiving each of the sharded data in parallel, and each of the second cryptographic hash values of the sharded data is synchronously stored in a first server corresponding to the first client; A sending module, configured to receive each of the second cryptographic hash values sent by the first server through the first client; A determination module, configured to determine that the first backup data is successfully backed up to the off-site storage device if it is determined that each piece of shard data passes the verification according to each of the first password hashes and each of the second password hashes; Wherein, the first server to be backed up determines at least one first block device it mounts according to a preset binding relationship, and the preset binding relationship is created according to the distance between the first geographical location information of each block device and the second geographical location information of each server to be backed up.
13. An electronic device, Characterized in that, It includes: A processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the data backup method based on a distributed storage system according to any one of claims 1-10.
14. A computer-readable storage medium, Characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it performs the steps of the data backup method based on a distributed storage system according to any one of claims 1-10.
Citation Information
Patent Citations
Storage method under cloud computing platform
CN106294585A
Data backup method, device and system
CN109408280A
Backup recovery method and device of OpenStack virtualization platform based on Ceph storage
CN113806145A