Incremental object backup and recovery method, system and storage medium for object storage
By establishing a backup agent in the object store, monitoring incremental logs and parsing operation data, the object-level backup problem in the object store is solved, and the effect of rapid recovery and cost reduction is achieved.
Patent Information
- Application Number
- CN202211237339.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-10-11
AI Technical Summary
The prior art is difficult to realize object-level backup of object storage, and the backup and recovery process consumes time and resources, which cannot effectively reduce costs.
By establishing a connection between the backup agent to the object store, monitoring and obtaining incremental logs, parsing incremental operation data, obtaining incremental object data, and transmitting it to the backup store, using a hash index table to achieve rapid recovery.
It realizes object-level backup and rapid recovery of object storage, reduces backup and recovery costs, and improves backup efficiency.
Smart Images

Figure CN115658382B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of object storage, and in particular relates to an incremental object backup and recovery method, system and storage medium for object storage. Background Art
[0002] Object storage is often used to store unstructured data such as images, audio, and video. It does not have the namespace, file directory, and other structures of traditional file storage. Instead, it stores all data as objects at the same level of the flat address space and has good scalability.
[0003] Object storage data has the characteristics of large volume, diverse types, low value density, and fast access speed, so traditional backup methods such as tape replication cannot be used. The existing data multiple copies and erasure coding technologies of cloud storage can only guarantee the reliability of data in a single data center, and cannot provide data backup capabilities. Some cloud service providers provide cross-region replication or data migration to another cloud service provider to achieve object storage synchronization, but compared with the storage cost of backup, the cost of using and maintaining redundant data clusters is much more expensive. The prior art discloses a snapshot method based on object storage buckets (publication number CN110515543A), which uses the method of taking snapshots of object storage buckets to back up object storage and provides bucket-level backup, but the method cannot directly restore a single object. The object backup method of third-party backup software often uses methods such as comparing the modification time of the object and comparing the hash value of the object to determine the incremental object, which requires traversing the entire object storage and comparing the objects one by one. This backup solution takes a long time and consumes a lot of resources when backing up massive objects.
[0004] Therefore, how to provide users with object-level backup of object storage, quickly obtain incremental objects in object storage, and reduce the backup and recovery costs of object storage has become a technical problem that urgently needs to be solved. Summary of the invention
[0005] In order to solve the technical problems in the above background technology, the present invention provides an incremental object backup and recovery method, system and storage medium for object storage. The technical solution is as follows:
[0006] In a first aspect, a method for incremental object backup and recovery of object storage is provided, the method comprising the steps of:
[0007] Establish a connection step to create a connection between the backup proxy and the object storage;
[0008] In the step of obtaining an incremental log, the backup agent monitors the operation log of the object storage and obtains the incremental log, wherein the incremental log is the additional write data of the operation log;
[0009] The backup agent parses the incremental log to obtain incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name and storage bucket name;
[0010] The step of obtaining incremental object data, wherein the backup agent obtains incremental object data from the object storage according to the object name and bucket name in the incremental operation data;
[0011] Incremental object data transmission step, the backup agent creates a hash index table, then transmits the incremental object data to the backup storage, and updates the incremental operation data into the hash index table;
[0012] In the recovery step, the backup agent queries the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object to recover the data of the specified object.
[0013] In one embodiment, the step of obtaining the incremental log includes:
[0014] Initialize the operation log monitoring table step, create an operation log monitoring table containing the creation time record and the read pointer offset;
[0015] Determine the operation log rotation step, and compare whether the creation time of the operation log A in the working state is the same as the creation time record in the operation log monitoring table;
[0016] In the step of writing incremental logs, if yes, read the data in operation log A starting from the read pointer offset, and write the data read from operation log A into the local file of the backup proxy; if no, traverse the rotation log file, read the operation log B whose creation time is equal to the creation time record starting from the read pointer offset, and read the operation log C whose creation time is later than the creation time record from the beginning, and then write the data read from operation log B and operation log C into the local file of the backup proxy;
[0017] In the step of updating the operation log monitoring table, the read pointer offset in the operation log monitoring table is updated to the end position of the operation log A, and the creation time record is updated to the creation time of the operation log A;
[0018] The step of obtaining incremental logs at regular intervals is to repeatedly execute the step of determining the rotation of the operation logs to the step of updating the operation log monitoring table within a time interval that is less than the operation log rotation period set by the object storage to obtain incremental logs.
[0019] The step of constructing the object operation regular expression is to analyze the operation log structure of the object storage and construct the object operation regular expression;
[0020] To obtain a single incremental log, read the incremental log line by line and obtain a single incremental log.
[0021] The first judgment step is to judge whether a single incremental log contains an object operation keyword, wherein the object operation keywords include: PUT and DELETE;
[0022] Parse a single incremental log step. If yes, parse the single incremental log according to the object operation regular expression to obtain incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name, and bucket name; if no, repeat the steps of obtaining a single incremental log to parsing a single incremental log;
[0023] A second judgment step is to judge whether the operation status in the incremental operation data is a success status;
[0024] an incremental operation data processing step, if yes, retaining the incremental operation data; if no, not retaining the incremental operation data;
[0025] Repeat the steps of obtaining a single incremental log to processing incremental operation data until the incremental log is parsed and all incremental operation data is obtained.
[0026] In one embodiment, the step of transmitting the incremental object data includes:
[0027] Create a hash index table step, create a hash index table containing keys and values;
[0028] The backup data index obtaining step is to transfer the incremental object data to the backup storage and obtain the backup data index;
[0029] The step of updating the hash index table uses the object name and bucket name in the incremental object information as keys, and the operation timestamp, operation type and backup data index in the incremental object information as values to insert into the hash index table.
[0030] In one embodiment, the recovery step includes:
[0031] The backup data index obtaining step queries the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object to obtain the backup data index of the specified object;
[0032] The step of restoring the designated object data is to restore the designated object data from the backup storage according to the backup data index of the designated object.
[0033] In a second aspect, an incremental object backup and recovery system for object storage is also provided, the system comprising:
[0034] Establish a connection module for establishing a connection between the backup proxy and the object storage;
[0035] An incremental log acquisition module is used for the backup agent to monitor the operation log stored in the object and acquire the incremental log, wherein the incremental log is the additional write data of the operation log;
[0036] An incremental operation data acquisition module is used for the backup agent to parse the incremental log to acquire incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name and storage bucket name;
[0037] An incremental object data acquisition module is used for the backup agent to acquire incremental object data from the object storage according to the object name and bucket name in the incremental operation data;
[0038] An incremental object data transmission module is used for the backup agent to create a hash index table, then transmit the incremental object data to the backup storage, and update the incremental operation data into the hash index table;
[0039] The recovery module is used for the backup agent to query the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object, and recover the data of the specified object.
[0040] In one embodiment, the incremental log acquisition module includes:
[0041] Initialize the operation log monitoring table unit, which is used to create an operation log monitoring table containing a creation time record and a read pointer offset;
[0042] The operation log rotation unit is used to compare whether the creation time of the operation log A in the working state is the same as the creation time record in the operation log monitoring table;
[0043] The write incremental log unit is used to read the data in the operation log A starting from the read pointer offset if yes, and write the data read from the operation log A into the local file of the backup agent; if no, traverse the rotation log file, read the operation log B whose creation time is equal to the creation time record starting from the read pointer offset, and read the operation log C whose creation time is later than the creation time record from the beginning, and then write the data read from the operation log B and the operation log C into the local file of the backup agent;
[0044] An operation log monitoring table unit is updated to update the read pointer offset in the operation log monitoring table to the end position of operation log A, and the creation time record is updated to the creation time of operation log A;
[0045] The incremental log acquisition unit is used to repeatedly execute the operation log rotation determination step to the operation log monitoring table update step within a time interval less than the operation log rotation period set by the object storage to obtain the incremental log.
[0046] In one embodiment, the module for obtaining incremental operation data includes:
[0047] An object operation regular expression building unit is used to analyze the operation log structure of the object storage and build an object operation regular expression;
[0048] Get a single incremental log unit, which is used to read incremental logs by line and get a single incremental log;
[0049] The first judgment unit is used to judge whether a single incremental log contains an object operation keyword, wherein the object operation keyword includes: PUT and DELETE;
[0050] Parsing a single incremental log unit, for if yes, parsing a single incremental log according to the object operation regular expression to obtain incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name, bucket name; if no, repeating obtaining a single incremental log unit to parsing a single incremental log unit;
[0051] A second judgment unit, used to judge whether the operation status in the incremental operation data is a success status;
[0052] The incremental operation data processing unit is used to retain the incremental operation data if yes; if no, not retain the incremental operation data.
[0053] In one embodiment, the recovery module includes:
[0054] A backup data index obtaining unit is used to query the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object to obtain the backup data index of the specified object;
[0055] The designated object data recovery unit is used to recover the designated object data from the backup storage according to the backup data index of the designated object.
[0056] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method for incremental object backup and recovery of the above-mentioned object storage is implemented.
[0057] Beneficial effects of the present invention:
[0058] (1) Provide object-level backup of object storage and back up objects to backup storage without additional cluster maintenance overhead, thus reducing the backup cost of object storage;
[0059] (2) Quickly obtain incremental objects in object storage to improve backup efficiency;
[0060] (3) Object-level recovery of object storage is realized, which can quickly restore specified objects and reduce the recovery cost of object storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0062] Figure 1 This is a flow chart of an incremental object backup and recovery method for object storage in Embodiment 1 of the present invention;
[0063] Figure 2 This is a structural diagram of an incremental object backup and recovery system for object storage in Embodiment 2 of the present invention;
[0064] Figure 3 This is a structural diagram of an incremental log acquisition module in Embodiment 2 of the present invention;
[0065] Figure 4 This is a structural diagram of a module for obtaining incremental operation data in Embodiment 2 of the present invention;
[0066] Figure 5 This is a structural diagram of the recovery module in the second embodiment of the present invention.
[0067] In the accompanying drawings, the components represented by the reference numerals are listed as follows:
[0068] 1001. Establish connection module, 1002. Obtain incremental log module, 1003. Obtain incremental operation data module, 1004. Obtain incremental object data module, 1005. Transmit incremental object data module, 1006. Recovery module, 10021. Initialize operation log monitoring table unit, 10022. Determine operation log rotation unit, 10023. Write incremental log unit, 10024. Update operation log monitoring table unit, 10025. Timed acquisition of incremental log unit, 10031. Build object operation regular expression unit, 10032. Obtain single incremental log unit, 10033. First judgment unit, 10034. Parse single incremental log unit, 10035. Second judgment unit, 10036. Incremental operation data processing unit, 10061. Obtain backup data index unit, 10062. Restore specified object data unit. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0070] The method provided by the present invention has an application scope including but not limited to the following environment: the object storage is Ceph object storage, and the operating system of the backup agent is Ubuntu-20.04.
[0071] Embodiment 1
[0072] like Figure 1 A method for incremental object backup and recovery of object storage is provided, the method comprising the steps of:
[0073] S1. Create a connection from the backup proxy to the object storage.
[0074] It is understandable that the backup proxy is responsible for performing backup and recovery tasks and transferring data between the source end and the backup storage. In order to facilitate the acquisition of operation logs, the gateway node in the object storage cluster is used as the backup proxy.
[0075] S2. The backup agent monitors the operation log of the object storage and obtains the incremental log, wherein the incremental log is the additional write data of the operation log.
[0076] It should be understood that before step S2, the log function of the object storage will be configured so that it can record the operation information of the object storage.
[0077] Optionally, the step S2 includes:
[0078] S21. Create an operation log monitoring table containing a creation time record and a read pointer offset.
[0079] It is understandable that the operation log monitoring table is used to record processed operation log parameters to facilitate rapid location of newly added operation log data.
[0080] S22. Compare the creation time of the operation log A in the working state with the creation time record in the operation log monitoring table to see whether they are the same.
[0081] S23. If so, read the data in operation log A starting from the read pointer offset, and write the data read from operation log A into the local file of the backup agent; if not, traverse the rotation log file, read the operation log B whose creation time is equal to the creation time record starting from the read pointer offset, and read the operation log C whose creation time is later than the creation time record from the beginning, and then write the data read from operation logs B and C into the local file of the backup agent.
[0082] It is understandable that in order to prevent the operation log from occupying too much disk capacity, the Linux system manages the operation log through its own log management program logrotate. When the size of the operation log reaches the set maximum value, logrotate will rename the operation log and use it as a rotation log file, and then create a new empty log file. Therefore, the backup agent determines whether the data in the operation log has been processed based on the creation time of the operation log.
[0083] It is also understandable that the creation time of operation log A may be equal to or later than the creation time record in the operation log monitoring table. If it is equal, logrotate has not rotated the operation log, so the data read from the read pointer offset position of operation log A is the newly added operation log data. If it is later, logrotate has rotated the operation log, so it is necessary to traverse the rotated log file to find the operation log B with the same creation time record. The data read from the read pointer offset position of operation log B, plus the data in all operation logs C with creation times later than the creation time record, is the newly added operation log data.
[0084] S24. The read pointer offset in the operation log monitoring table is updated to the end position of operation log A, and the creation time record is updated to the creation time of operation log A.
[0085] For ease of understanding, specifically, an operation example is provided for steps S22-S24: the creation time of the log file rgw.log in the working state is 2022 / 5 / 12, which is later than the creation time record 2022 / 5 / 10 in the operation log monitoring table. Therefore, traverse the rotation log files, the creation time of the log file rgw.log.1 is 2022 / 5 / 10, which is equal to the creation time record, and the creation time of the log file rgw.log.2 is 2022 / 5 / 11, which is later than the creation time record. Therefore, the data of rgw.log.1 is read from the read pointer offset position in the operation log monitoring table, and the data of rgw.log.2 and rgw.log are read from the beginning of the file, and the read data is written to the local file of the backup agent. Finally, the read pointer offset in the operation log monitoring table is updated to the end position of rgw.log, and the creation time record is updated to the creation time of rgw.log 2022 / 5 / 12.
[0086] S25. Repeat steps S22 to S24 within a time interval that is less than the operation log rotation period set by the object storage to obtain incremental logs.
[0087] It is understandable that in order to prevent the log from being too large, logrotate can set the log rotation period and the number of logs to be retained. Therefore, it is necessary to repeatedly process the operation log within the time interval of the rotation period to effectively avoid missing data.
[0088] S3. The backup agent parses the incremental log to obtain incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name and storage bucket name.
[0089] Optionally, the S3 step includes:
[0090] S31. Analyze the operation log structure of the object storage and construct the object operation regular expression.
[0091] S32. Read the incremental log line by line to obtain a single incremental log.
[0092] S33. Determine whether a single incremental log contains an object operation keyword, where the object operation keywords include: PUT and DELETE.
[0093] It is understandable that object storage provides a RESTful interface, and operations related to incremental objects include: PUT / api / rgw / bucket / {bucket} and DELETE / api / rgw / bucket / {bucket}. Therefore, the object operation keywords include: PUT and DELETE.
[0094] S34. If yes, parse the single incremental log according to the object operation regular expression to obtain incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name, bucket name; if no, repeat steps S32 to S34.
[0095] S35. Determine whether the operation status in the incremental operation data is a success status.
[0096] S36. If yes, retain the incremental operation data; if no, do not retain the incremental operation data.
[0097] It is understandable that the operation status is returned in the form of an HTTP status code: a PUT operation status of 200 and a DELETE operation status of 204 indicate a successful operation, while other operation statuses indicate a failed operation.
[0098] For ease of understanding, specifically, an operation example is provided for steps S33-S36: an operation log is as follows: 2022-04-08T14:37:04.926+0800 7f5f9afcd700 1 beast: Ox7f604c2d1810:192.168.100.142- - [2022-04-08T14:37:04.92622+0800]"PUT / my-bucket-name / 3.txtHTTP / 1.1"200 9 - - -. After regular matching, it can be parsed that the operation timestamp is 2022-04-08T14:37:04.92622+0800, the operation type is PUT, the operation status is 200, the object name is 3.txt, and the bucket name is my-bucket-name. The operation type of this operation log is PUT, and the operation status is 200, which means that the operation is executed successfully and the operation data is retained.
[0099] S37. Repeat steps S32 to S36 until the incremental log is parsed and all incremental operation data is obtained.
[0100] S4. The backup agent obtains the incremental object data from the object storage according to the object name and bucket name in the incremental operation data.
[0101] It is understandable that object storage uses object name and bucket name to uniquely identify an object. Therefore, the object data can be obtained from object storage through the object name and bucket name.
[0102] S5. The backup agent creates a hash index table, transfers the incremental object data to the backup storage, and updates the incremental operation data to the hash index table.
[0103] Optionally, the step S5 includes:
[0104] S51. Create a hash index table containing keys and values.
[0105] It is understandable that the hash index table can quickly obtain values through keys, which can improve search efficiency. When the amount of data is small, a hash table can be used to implement the hash index table. When the amount of data is large, the key-value database redis can be used to implement the hash index table.
[0106] S52. Transfer the incremental object data to the backup storage and obtain the backup data index.
[0107] It can be understood that the backup data index here, that is, the location where the incremental object data is stored in the backup storage end, can be used to obtain specific data of the corresponding incremental object.
[0108] S53. Use the object name and bucket name in the incremental object information as keys, and use the operation timestamp, operation type and backup data index in the incremental object information as values, and insert them into the hash index table.
[0109] It is understandable that the object name and bucket name can uniquely identify an object, so the object name and bucket name are used as keys, and then the value records the operation timestamp and operation type that determine the object status and the backup data index that can obtain the specific data of the object.
[0110] S6. The backup agent queries the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object, and restores the data of the specified object.
[0111] Optionally, the step S6 includes:
[0112] S61. According to the object name of the specified object and the storage bucket name of the specified object, query the updated hash index table to obtain the backup data index of the specified object.
[0113] S62. Restore the designated object data from the backup storage according to the backup data index of the designated object.
[0114] The technical solution of this embodiment can quickly obtain incremental objects in the object storage by analyzing incremental logs, thereby improving backup efficiency. The technical solution of this embodiment also realizes object-level recovery of the object storage, can quickly recover the specified object, and reduces the recovery cost of the object storage.
[0115] Embodiment 2
[0116] like Figure 2 In one embodiment, a system for incremental object backup and recovery of object storage is provided, the system comprising:
[0117] Establishing connection module 1001, used to establish a connection between the backup proxy and the object storage;
[0118] An incremental log acquisition module 1002 is used for the backup agent to monitor the operation log stored in the object and acquire the incremental log, wherein the incremental log is the additional write data of the operation log;
[0119] An incremental operation data acquisition module 1003 is used for the backup agent to parse the incremental log to acquire incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name and storage bucket name;
[0120] An incremental object data acquisition module 1004 is used to acquire incremental object data from the object storage according to the object name and bucket name in the incremental operation data;
[0121] The incremental object data transmission module 1005 is used to create a hash index table, then transmit the incremental object data to the backup storage, and update the incremental operation data into the hash index table;
[0122] The recovery module 1006 is used to query the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object, and recover the data of the specified object.
[0123] Optional, such as Figure 3 As shown, based on this embodiment, the incremental log acquisition module 1002 includes:
[0124] Initialize the operation log monitoring table unit 10021, used to create an operation log monitoring table containing a first creation time and a read pointer offset;
[0125] An operation log rotation determination unit 10022 is used to compare whether a second creation time of an operation log A in a working state is equal to a first creation time in the operation log monitoring table;
[0126] The write incremental log unit 10023 is used to read the data in the operation log A starting from the read pointer offset, and write the data read from the operation log A to the local file of the backup agent; if not, traverse the rotation log file, read the operation log B whose creation time is equal to the first creation time starting from the read pointer offset, and read the operation log C whose creation time is later than the first creation time from the beginning, and then write the data read from the operation log B and the operation log C to the local file of the backup agent;
[0127] An update operation log monitoring table unit 10024 is used to update the read pointer offset in the operation log monitoring table to the end position of the operation log A, and update the first creation time to the second creation time of the operation log A;
[0128] The timed incremental log acquisition unit 10025 is used to repeatedly execute the operation log rotation determination step to the operation log monitoring table update step within a time interval less than the operation log rotation period set by the object storage to obtain the incremental log.
[0129] Optional, such as Figure 4 As shown, based on this embodiment, the module 1003 for obtaining incremental operation data includes:
[0130] An object operation regular expression building unit 10031 is used to analyze the operation log structure of the object storage and build an object operation regular expression;
[0131] The single incremental log acquisition unit 10032 is used to read the incremental log line by line and acquire a single incremental log;
[0132] The first judgment unit 10033 is used to judge whether a single incremental log contains an object operation keyword, where the object operation keyword includes: PUT and DELETE;
[0133] The single incremental log parsing unit 10034 is used to parse the single incremental log according to the object operation regular expression to obtain incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name, and bucket name; if not, repeat the step of obtaining the single incremental log;
[0134] The second judging unit 10035 is used to judge whether the operation status in the incremental operation data is a success status;
[0135] The incremental operation data processing unit 10036 is used to retain the incremental operation data if yes; if no, not retain the incremental operation data.
[0136] Optional, such as Figure 5 As shown, based on this embodiment, the recovery module 1006 includes:
[0137] The backup data index obtaining unit 10061 is used to query the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object to obtain the backup data index of the specified object;
[0138] The designated object data recovery unit 10062 is used to recover the designated object data from the backup storage according to the backup data index of the designated object.
[0139] The technical solution of this embodiment is to establish a connection module 1001, which is used to create a connection between the backup agent and the object storage; obtain an incremental log module 1002, which is used for the backup agent to monitor the operation log of the object storage and obtain the incremental log, wherein the incremental log is the appended write data of the operation log; obtain an incremental operation data module 1003, which is used for the backup agent to parse the incremental log and obtain the incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name and storage bucket name; obtain an incremental object data module 1004, which is used to obtain the incremental object data from the object storage according to the object name and bucket name in the incremental operation data; transmit the incremental object data module 1005, which is used to create a hash index table, and then transmit the incremental object data to the backup storage, and update the incremental operation data to the hash index table; restore the module 1006, which is used to query the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object, and restore the specified object data. This embodiment provides object-level backup of object storage, and backs up objects to backup end storage, without additional cluster maintenance overhead, reducing the backup cost of object storage.
[0140] Embodiment 3
[0141] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for incremental object backup and recovery of object storage described in Embodiment 1 is implemented.
[0142] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0143] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. An incremental object backup and recovery method for object storage, It is characterized in that The method comprises the steps of: Establish a connection step to create a connection between the backup proxy and the object storage cluster; In the step of obtaining an incremental log, the backup agent monitors the operation log of the object storage cluster and obtains the incremental log, wherein the incremental log is the additional write data of the operation log; The backup agent parses the incremental log to obtain incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name and bucket name, and an object is uniquely identified by the object name and bucket name; The step of obtaining incremental object data, wherein the backup agent obtains incremental object data from the object storage cluster according to the object name and bucket name in the incremental operation data; Incremental object data transmission step, the backup agent creates a hash index table, then transmits the incremental object data to the backup storage, and updates the incremental operation data into the hash index table; In a recovery step, the backup agent queries the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object to recover the data of the specified object; The step of obtaining incremental logs includes: Initialize the operation log monitoring table step, create an operation log monitoring table containing the creation time record and the read pointer offset; Determine the operation log rotation step, and compare whether the creation time of the operation log A in the working state is the same as the creation time record in the operation log monitoring table; In the step of writing incremental logs, if yes, read the data in operation log A starting from the read pointer offset, and write the data read from operation log A into the local file of the backup agent; if no, traverse the rotation log file, read the operation log B whose creation time is equal to the creation time record starting from the read pointer offset, and read the operation log C whose creation time is later than the creation time record from the beginning, and then write the data read from operation log B and operation log C into the local file of the backup agent; In the step of updating the operation log monitoring table, the read pointer offset in the operation log monitoring table is updated to the end position of the operation log A, and the creation time record is updated to the creation time of the operation log A; The step of obtaining incremental logs at regular intervals is to repeatedly execute the step of determining the rotation of the operation logs to the step of updating the operation log monitoring table within a time interval that is less than the operation log rotation period set by the object storage cluster to obtain incremental logs.
2. The incremental object backup and recovery method of object storage according to claim 1, It is characterized in that The step of obtaining incremental operation data includes: The step of constructing the object operation regular expression is to analyze the operation log structure of the object storage cluster and construct the object operation regular expression; To obtain a single incremental log, read the incremental log line by line and obtain a single incremental log. The first judgment step is to judge whether a single incremental log contains an object operation keyword, wherein the object operation keywords include: PUT and DELETE; Parse a single incremental log step. If yes, parse the single incremental log according to the object operation regular expression to obtain incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name, and bucket name; if no, repeat the steps of obtaining a single incremental log to parsing a single incremental log; A second judgment step is to judge whether the operation status in the incremental operation data is a success status; an incremental operation data processing step, if yes, retaining the incremental operation data; if no, not retaining the incremental operation data; Repeat the steps of obtaining a single incremental log to processing incremental operation data until the incremental log is parsed and all incremental operation data is obtained.
3. The incremental object backup and recovery method of object storage according to claim 1, It is characterized in that The step of transmitting the incremental object data comprises: Create a hash index table step, create a hash index table containing keys and values; The backup data index obtaining step is to transfer the incremental object data to the backup storage and obtain the backup data index; The step of updating the hash index table uses the object name and bucket name in the incremental object information as keys, and the operation timestamp, operation type and backup data index in the incremental object information as values to insert into the hash index table.
4. The incremental object backup and recovery method of object storage according to claim 1, It is characterized in that The recovery step comprises: The backup data index obtaining step queries the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object to obtain the backup data index of the specified object; The step of restoring the designated object data is to restore the designated object data from the backup storage according to the backup data index of the designated object.
5. An incremental object backup and recovery system for object storage, It is characterized in that The system includes: a connection establishment module, configured to establish a connection between a backup proxy and an object storage cluster; An incremental log acquisition module is used for the backup agent to monitor the operation log of the object storage cluster and acquire the incremental log, wherein the incremental log is the additional write data of the operation log; An incremental operation data acquisition module is used for the backup agent to parse the incremental log to acquire incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name and storage bucket name, and an object is uniquely identified by the object name and storage bucket name; An incremental object data acquisition module is used for the backup agent to acquire incremental object data from the object storage cluster according to the object name and bucket name in the incremental operation data; An incremental object data transmission module is used for the backup agent to create a hash index table, then transmit the incremental object data to the backup storage, and update the incremental operation data into the hash index table; A recovery module, used for the backup agent to query the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object, and restore the specified object data; The incremental log acquisition module includes: Initialize the operation log monitoring table unit, which is used to create an operation log monitoring table containing a creation time record and a read pointer offset; The operation log rotation unit is used to compare whether the creation time of the operation log A in the working state is the same as the creation time record in the operation log monitoring table; The write incremental log unit is used to read the data in the operation log A starting from the read pointer offset if yes, and write the data read from the operation log A into the local file of the backup agent; if no, traverse the rotation log file, read the operation log B whose creation time is equal to the creation time record starting from the read pointer offset, and read the operation log C whose creation time is later than the creation time record from the beginning, and then write the data read from the operation log B and the operation log C into the local file of the backup agent; An operation log monitoring table unit is updated to update the read pointer offset in the operation log monitoring table to the end position of operation log A, and the creation time record is updated to the creation time of operation log A; The incremental log acquisition unit is used to repeatedly execute the operation log rotation determination step to the operation log monitoring table update step within a time interval less than the operation log rotation period set by the object storage cluster to obtain the incremental log.
6. The incremental object backup and recovery system for object storage according to claim 5, It is characterized in that The module for obtaining incremental operation data includes: The object operation regular expression building unit is used to analyze the operation log structure of the object storage cluster and build the object operation regular expression; Get a single incremental log unit, which is used to read incremental logs by line and get a single incremental log; The first judgment unit is used to judge whether a single incremental log contains an object operation keyword, wherein the object operation keyword includes: PUT and DELETE; Parsing a single incremental log unit, for if yes, parsing a single incremental log according to the object operation regular expression to obtain incremental operation data, wherein the incremental operation data includes: operation timestamp, operation type, operation status, object name, bucket name; if no, repeating obtaining a single incremental log unit to parsing a single incremental log unit; A second judgment unit, used to judge whether the operation status in the incremental operation data is a success status; The incremental operation data processing unit is used to retain the incremental operation data if yes; if no, not retain the incremental operation data.
7. The incremental object backup and recovery system for object storage according to claim 5, It is characterized in that The recovery module comprises: A backup data index obtaining unit is used to query the updated hash index table according to the object name of the specified object and the storage bucket name of the specified object to obtain the backup data index of the specified object; The designated object data recovery unit is used to recover the designated object data from the backup storage according to the backup data index of the designated object.
8. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method for incremental object backup and recovery of object storage as described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Snapshot method, device and system based on object storage bucket
CN110515543A
Distributed file level backup method and system based on object storage
CN113946471A
System and method of handling journal space in a storage cluster with multiple delta log instances
US11042296B1