Data reconstruction method and device, data storage system, and storage medium

By creating access configuration information and reconstructing metadata using call enumeration interfaces, the problem of batch reconstruction in the existing technology cannot be achieved, the rapid and efficient reconstruction of data is achieved, and the efficiency of data migration and disaster recovery is improved.

CN115017153BActive Publication Date: 2025-05-02BEIJING XSKY TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210493625.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-07
Publication Date
2025-05-02
Estimated Expiration
2042-05-07

AI Technical Summary

Technical Problem

In the prior art, the return method cannot be used to realize batch reconstruction of data, especially in multi-version scenarios, seamless migration and efficient reconstruction of data cannot be achieved.

Method used

By responding to the data reconstruction instruction, we create access configuration information to the preset storage platform, index the object list in the source station bucket to which the target data belongs, determine the source station information of multiple storage objects, and use the pre-configured call enumeration interface to reconstruct the metadata of each storage object, write the metadata to the local cluster, and generate data reconstruction tasks to achieve reconstruction of the target data.

Benefits of technology

It realizes rapid batch reconstruction of data, solves the problem that batch reconstruction cannot be achieved through source return, and improves the efficiency of data migration and disaster recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115017153B_ABST
    Figure CN115017153B_ABST
Patent Text Reader

Abstract

The present invention discloses a data reconstruction method and device, a data storage system, and a storage medium. The method includes: responding to a data reconstruction instruction, creating access configuration information for accessing a preset storage platform, wherein the access configuration information is used to access the preset storage platform to index an object list of each storage object in a source station storage bucket to which the target data belongs; based on the object list, determining the source station information of multiple storage objects; based on the source station information, using a pre-configured call enumeration interface to rebuild the metadata of each storage object, and writing the metadata to a local cluster; based on the source station information and the metadata, generating a data reconstruction task, wherein the data reconstruction task is used to reconstruct the target data. The present invention solves the technical problem in the related art that batch reconstruction of data cannot be achieved by using a return-to-source method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage, and in particular to a data reconstruction method and device, a data storage system, and a storage medium. Background Art

[0002] In related technologies, with the rapid development of Internet applications, the increasing amount of unstructured data needs to be stored. The current commonly used data storage solution is object storage, which can provide a massive storage solution and support product specifications of tens of billions or hundreds of billions of objects. Usually in data disaster recovery scenarios, for the same product, data backup is achieved by configuring data site synchronization or bucket replication functions, and for different products, data migration is usually used to actively trigger writing to another cluster, which is similar to re-uploading. However, in multi-version scenarios, due to the limitations of the mechanisms of different products, it is impossible to achieve seamless data migration.

[0003] In order to meet the needs of data disaster recovery or management of other products, the current data migration method is to implement data migration manually. That is, at the data site to be migrated, a third-party tool is used to list the objects under the storage bucket, and then the bucket objects are traversed to write the data to another site. This migration method has low efficiency.

[0004] Another solution is to use the back-to-source method, that is, to configure third-party access rules for the local storage bucket. After accessing the local data, if the 404 error message is triggered, the data is retrieved from the source station in real time and written to the local server. However, this back-to-source method can only trigger a single point for an object, and cannot perform batch data reconstruction, which is a defect.

[0005] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0006] The embodiments of the present invention provide a data reconstruction method and device, a data storage system, and a storage medium, so as to at least solve the technical problem that batch reconstruction of data cannot be achieved by returning to the source in the related art.

[0007] According to one aspect of an embodiment of the present invention, there is provided a data reconstruction method, which is applied to a server platform, and includes: in response to a data reconstruction instruction, creating access configuration information for accessing a preset storage platform, wherein the access configuration information is used to access the preset storage platform to index an object list of each storage object in a source station storage bucket to which target data belongs; based on the object list, determining source station information of multiple storage objects; based on the source station information, using a pre-configured call enumeration interface to rebuild metadata of each storage object, and writing the metadata to a local cluster; based on the source station information and the metadata, generating a data reconstruction task, wherein the data reconstruction task is used to reconstruct the target data.

[0008] Optionally, the data reconstruction method is applied to a preset data reconstruction mode, wherein the preset data reconstruction mode is one of the following: a normal reconstruction mode, an upgrade reconstruction mode, and a disaster reconstruction mode.

[0009] Optionally, when the preset data reconstruction mode is the normal reconstruction mode, the step of determining the source station information of multiple storage objects based on the object list includes: traversing the object list to obtain multiple object information; reconstructing the storage object based on the object information; judging whether the reconstructed storage object exists in the local cluster; if the reconstructed storage object does not exist in the local cluster, accessing the source station to obtain metadata information and object tags of the storage object; and representing the metadata information and the object tags as the source station information.

[0010] Optionally, when the preset data reconstruction mode is the normal reconstruction mode, the step of reconstructing the metadata of each storage object using a pre-configured call enumeration interface based on the source site information also includes: obtaining the storage category supported by the server platform; calling a pre-configured call enumeration interface according to the storage category; and using the call enumeration interface to reconstruct the metadata of each storage object.

[0011] Optionally, in the case where the preset data reconstruction mode is a disaster reconstruction mode, the step of determining the source site information of multiple storage objects based on the object list includes: in the event of a data center failure, obtaining the tiered data that is pre-tiered in the public cloud of the failed data center, wherein the tiered data includes multiple tiered storage buckets; switching the failed data center to a business data center; creating a mirror storage bucket corresponding to the tiered storage bucket in the business data center; and determining the source site information of multiple mirror storage buckets based on the object list.

[0012] Optionally, when the preset data reconstruction mode is an upgrade reconstruction mode, the step of determining metadata of source site information of multiple storage objects based on the object list includes: in a cross-version upgrade scenario, traversing the object list to obtain multiple object information; based on the object information, using an incremental synchronization strategy to rebuild the storage object; and determining metadata of the source site information of each storage object.

[0013] Optionally, an incremental synchronization strategy is adopted to rebuild the storage object, including: obtaining the platform type of the second-level platform under the server platform; based on the platform type of the second-level platform, fully rebuilding the storage object under the server platform; after completing the full reconstruction of the storage object, incrementally rebuilding each storage object under the second-level storage bucket.

[0014] Optionally, the step of determining the source site information of multiple storage objects based on the object list also includes: for multi-version storage buckets, obtaining the bucket identifier of each version in the source site storage bucket to which the target data belongs; based on the bucket identifier of each version, indexing the object list of each storage object in the source site storage bucket; based on the object list, determining the source site information of multiple storage objects.

[0015] According to another aspect of an embodiment of the present invention, there is also provided a data reconstruction device, which is applied to a server platform and includes: a creation unit, which is used to respond to a data reconstruction instruction and create access configuration information for accessing a preset storage platform, wherein the access configuration information is used to access the preset storage platform to index an object list of each storage object in a source station storage bucket to which target data belongs; a determination unit, which is used to determine source station information of multiple storage objects based on the object list; a reconstruction unit, which is used to reconstruct metadata of each storage object based on the source station information using a pre-configured call enumeration interface and write the metadata to a local cluster; and a generation unit, which is used to generate a data reconstruction task based on the source station information and the metadata, wherein the data reconstruction task is used to reconstruct the target data.

[0016] Optionally, the data reconstruction method is applied to a preset data reconstruction mode, wherein the preset data reconstruction mode is one of the following: a normal reconstruction mode, an upgrade reconstruction mode, and a disaster reconstruction mode.

[0017] Optionally, in the case where the preset data reconstruction mode is the normal reconstruction mode, the determination unit includes: a first traversal module, used to traverse the object list to obtain multiple object information; a first reconstruction module, used to reconstruct the storage object based on the object information; a first judgment module, used to judge whether the reconstructed storage object exists in the local cluster; a first access module, used to access the source station to obtain the metadata information and object tag of the storage object if the reconstructed storage object does not exist in the local cluster; a first determination module, used to represent the metadata information and the object tag as the source station information.

[0018] Optionally, when the preset data reconstruction mode is the normal reconstruction mode, the reconstruction unit also includes: a first acquisition module, used to obtain the storage category supported by the server platform; a first calling module, used to call a pre-configured calling enumeration interface according to the storage category; and a first reconstruction module, used to use the calling enumeration interface to reconstruct the metadata of each of the storage objects.

[0019] Optionally, in the case where the preset data reconstruction mode is a disaster reconstruction mode, the determination unit includes: a second acquisition module, used to obtain the layered data of the failed data center in the public cloud in advance in the event of a data center failure, wherein the layered data includes multiple layered storage buckets; a first switching module, used to switch the failed data center to a business data center; a first creation module, used to create a mirror storage bucket corresponding to the layered storage bucket in the business data center; and a second determination module, used to determine the source site information of multiple mirror storage buckets based on the object list.

[0020] Optionally, when the preset data reconstruction mode is an upgrade reconstruction mode, the determination unit includes: a second traversal module, used to traverse the object list in a cross-version upgrade scenario to obtain multiple object information; a second reconstruction module, used to reconstruct the storage object based on the object information and adopt an incremental synchronization strategy; a third determination module, used to determine the metadata of the source site information of each storage object.

[0021] Optionally, the second reconstruction module includes: an acquisition submodule, used to obtain the platform type of the second-level platform under the server platform; a first reconstruction submodule, used to fully rebuild the storage objects under the server platform based on the platform type of the second-level platform; and a second reconstruction submodule, used to incrementally rebuild each storage object under the second-level storage bucket after completing the full reconstruction of the storage objects.

[0022] Optionally, the determination unit also includes: a third acquisition module, used to obtain the bucket identifier of each version in the source station storage bucket to which the target data belongs for multi-version storage buckets; an indexing module, used to index the object list of each storage object in the source station storage bucket based on the bucket identifier of each version; and a fourth determination module, used to determine the source station information of multiple storage objects based on the object list.

[0023] According to another aspect of an embodiment of the present invention, a data storage system is also provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above-mentioned data reconstruction methods by executing the executable instructions.

[0024] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the data reconstruction methods described above.

[0025] In an embodiment of the present invention, a response data reconstruction instruction is used to create access configuration information for accessing a preset storage platform, wherein the access configuration information is used to access the preset storage platform to index the object list of each storage object in the source station storage bucket to which the target data belongs, and based on the object list, the source station information of multiple storage objects is determined, and based on the source station information, the metadata of each storage object is rebuilt using a pre-configured call enumeration interface, and the metadata is written to the local cluster, and a data reconstruction task is generated based on the source station information and metadata, wherein the data reconstruction task is used to reconstruct the target data. In this embodiment, when creating an access configuration signal for the storage platform, the platform configuration information can be used to index the object list of the source station bucket, and the metadata information can be written to the local cluster to decouple the metadata and the target data, and the source station metadata and data information can be quickly rebuilt locally by creating and issuing tasks, thereby realizing rapid reconstruction of batch data, thereby solving the technical problem that batch reconstruction of data cannot be realized by using a return-to-source method in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0027] Figure 1 is a flow chart of an optional data reconstruction method according to an embodiment of the present invention;

[0028] Figure 2 is a schematic diagram of an optional relationship between a local cluster and a source station cluster according to an embodiment of the present invention;

[0029] Figure 3 is a schematic diagram of an optional method of performing data reconstruction in a common reconstruction mode according to an embodiment of the present invention;

[0030] Figure 4 is a schematic diagram of an optional method of performing data reconstruction in a disaster reconstruction mode according to an embodiment of the present invention;

[0031] Figure 5 is a schematic diagram of an optional method of performing data reconstruction in an upgrade reconstruction mode according to an embodiment of the present invention;

[0032] Figure 6 is a schematic diagram of an optional data reconstruction device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0033] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0035] To facilitate those skilled in the art to understand the present invention, some terms or nouns involved in the embodiments of the present invention are explained below:

[0036] Object storage gateway, Rados Gateway, referred to as RGW.

[0037] Storage Class, a collection of storage with different storage media or different redundancy.

[0038] Bucket, Bucket.

[0039] Network File System, NFS for short, is an application based on UDP / IP protocol. Its implementation mainly adopts remote procedure call (RPC) mechanism. RPC provides a set of operations for accessing remote files that are independent of the machine, operating system and low-level transmission protocol.

[0040] SMB, Server Message Block, is a network file system used for Web connections and information communication between clients and servers.

[0041] S3, Simple Storage System, is a global storage area network (SAN) that appears as an oversized hard drive where digital assets are stored and retrieved.

[0042] The present invention can be applied to various data storage systems\platforms to realize distributed data storage services. The data storage system may include but is not limited to: a cache layer (interconnected to the client\front end, including high-speed cache disks such as SSD) and a data layer (including low-speed data disks such as HDD).

[0043] Different from the traditional source return method, this invention provides three usage methods and application scenarios: ordinary reconstruction, upgrade reconstruction and disaster recovery reconstruction. By creating a storage platform access configuration, using the platform configuration information to index the object list of the source station bucket, writing metadata information in this cluster, and establishing an association with the source station through the storage category.

[0044] The present invention provides a data disaster recovery and reconstruction strategy, which is applied to distributed storage services\distributed storage systems, decouples metadata and data, and quickly rebuilds source station metadata and data information locally by creating and issuing tasks.

[0045] The data reconstruction strategy of the present invention provides three modes including normal reconstruction, upgrade reconstruction and disaster reconstruction, which are applied to different usage scenarios. For normal reconstruction, it is applied to source stations such as public cloud, private cloud, NFS, etc., and the metadata information, tagging information, etag (a mark associated with Web resources), mtime (time status) and other information of the source station are rebuilt locally, so as to take over the source station and establish the association between local metadata and third parties; and provide a unified access portal for this site, and quickly manage the third-party platform through metadata reconstruction. For disaster reconstruction, the original key and acl and other information of the object can be rebuilt; for upgrade reconstruction, it is applied to fully rebuild the object to avoid incremental objects generated during the reconstruction process. After the reconstruction is completed, the data object can be rebuilt.

[0046] The present invention provides a method for rebuilding multi-version objects for various implementation scenarios of storage backends (e.g., S3 scenarios). The method provides a method for rebuilding metadata indexes, provides a flexible reconstruction strategy, solves the reconstruction problem of different demand fields, and helps customers achieve fast switching of services.

[0047] The data reconstruction strategy of the present invention creates storage categories, and the included data backend platform supports different storage categories, such as S3, NFS or SMB types.

[0048] In the present invention, user service switching is achieved by quickly rebuilding metadata. At the same time, in order to provide a data acceleration function, data is subsequently pulled back through flexible configuration.

[0049] The present invention is described in detail below in conjunction with various embodiments.

[0050] Embodiment 1

[0051] According to an embodiment of the present invention, a data reconstruction method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0052] According to one aspect of an embodiment of the present invention, a data reconstruction method is provided, which is applied to a server platform.

[0053] Figure 1 is a flow chart of an optional data reconstruction method according to an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:

[0054] Step S102, in response to the data reconstruction instruction, creates access configuration information for accessing the preset storage platform, wherein the access configuration information is used to access the preset storage platform to index an object list of each storage object in the source station storage bucket to which the target data belongs.

[0055] By creating a storage platform access configuration, use the platform configuration information to index the object list of the source server bucket.

[0056] Step S104: determining source site information of multiple storage objects based on the object list.

[0057] Step S106: Based on the source site information, the metadata of each storage object is rebuilt using a pre-configured call enumeration interface, and the metadata is written to the local cluster.

[0058] In this embodiment, the metadata index can be rebuilt first, and then the target data can be rebuilt based on the rebuilt metadata index, and the metadata and data of the source station can be rebuilt locally. The data reconstruction method in this embodiment includes creating a secondary storage class, and the data backend storage is mounted to the corresponding secondary storage class.

[0059] Figure 2 is a schematic diagram of an optional relationship between a local cluster and a source station cluster according to an embodiment of the present invention, such as Figure 2 As shown, the corresponding local storage category in the local cluster is the local backend, and the secondary storage category provides S3 / NFS / SMB. Under the local cluster, multiple storage objects are stored through buckets, and the data is stored in each storage object; at the same time, the source cluster can exist as a source station. There are multiple source station targets in the source cluster, and the source station target can also store data in the form of buckets.

[0060] By creating a reconstruction configuration and issuing a reconstruction task, the reconstruction process is divided into the reconstruction of metadata and data. By calling the enumeration interface according to different backend types, such as the corresponding S3 scenario, calling the ListBucket object interface to rebuild the metadata. After the metadata reconstruction of a single object is completed, a scheduled task for rebuilding the data is issued.

[0061] Among them, the local cluster creates different storage categories pointing to the local data storage pool or secondary storage platform, where the secondary storage target is the connection method of the source station target, such as S3 by providing the access path, user name, password and target bucket and other source station information. After the object source station information is rebuilt locally, the data storage points to the source station. During the process of accessing and obtaining the object, if the data is not rebuilt locally, the agent is triggered to read the data.

[0062] Step S108: Generate a data reconstruction task based on the source station information and metadata, wherein the data reconstruction task is used to reconstruct the target data.

[0063] Through the above steps, it is possible to respond to the data reconstruction instruction and create access configuration information for accessing the preset storage platform, wherein the access configuration information is used to access the preset storage platform to index the object list of each storage object in the source station storage bucket to which the target data belongs, determine the source station information of multiple storage objects based on the object list, and based on the source station information, use a pre-configured call enumeration interface to rebuild the metadata of each storage object, and write the metadata to the local cluster, and generate a data reconstruction task based on the source station information and metadata, wherein the data reconstruction task is used to rebuild the target data. In this embodiment, when creating an access configuration signal for the storage platform, the platform configuration information can be used to index the object list of the source station bucket, and the metadata information can be written to the local cluster to decouple the metadata and the target data, and the source station metadata and data information can be quickly rebuilt locally by creating and issuing tasks, thereby realizing rapid reconstruction of batch data, thereby solving the technical problem that batch reconstruction of data cannot be realized by using the return-to-source method in the related technology.

[0064] An embodiment of the present invention provides a method for metadata index reconstruction, which can decouple metadata and data, and quickly rebuild source station metadata and data information locally by creating and issuing tasks.

[0065] The present invention is described in detail below in conjunction with various implementation steps.

[0066] Optionally, the data reconstruction method of this embodiment is applied to a preset data reconstruction mode, wherein the preset data reconstruction mode is one of the following: normal reconstruction mode, upgrade reconstruction mode, and disaster reconstruction mode. Three modes including normal reconstruction, upgrade reconstruction, and disaster reconstruction are provided for different usage scenarios; by creating storage categories, the included data backend platform supports different categories, such as S3, NFS, or SMB types.

[0067] The three data reconstruction modes are described below.

[0068] 1. Normal reconstruction mode

[0069] In this embodiment, optionally, when the preset data reconstruction mode is the normal reconstruction mode, the step of determining the source station information of multiple storage objects based on the object list includes: traversing the object list to obtain multiple object information; reconstructing the storage object based on the object information; determining whether the reconstructed storage object exists in the local cluster; if the reconstructed storage object does not exist in the local cluster, accessing the source station to obtain metadata information and object tags of the storage object; and representing the metadata information and object tags as source station information.

[0070] Another optional step, when the preset data reconstruction mode is the normal reconstruction mode, based on the source site information, using a pre-configured call enumeration interface to rebuild the metadata of each storage object, also includes: obtaining the storage category supported by the server platform; according to the storage category, calling the pre-configured call enumeration interface; using the call enumeration interface to rebuild the metadata of each storage object.

[0071] In this embodiment, the common reconstruction application scenario refers to the common management of a secondary platform.

[0072] Figure 3 is a schematic diagram of an optional data reconstruction using a common reconstruction mode according to an embodiment of the present invention, such as Figure 3 As shown in the figure, the implementation process includes: first, creating a reconstruction configuration, which includes the platform information of the source station; then issuing a reconstruction task, in which the reconstruction task records the current execution progress; judging whether the reconstruction is completed, if not, it is necessary to access the source station to obtain the object list for reconstruction (i.e. Figure 3 ); traverse and rebuild all objects by traversing the object list. If the object exists locally, skip it. Otherwise, go to the source object metadata information; determine whether there is tagging information. If so, obtain the tagging information. After completion, update the object metadata information, where the metadata information includes the secondary information that needs to be recorded; if not, rebuild the tagging and store the metadata; then issue the data reconstruction task to asynchronously rebuild the data part.

[0073] The ordinary reconstruction in this embodiment can be applied to source stations such as public clouds, private clouds, and NFS, and the metadata information, tagging information, etag, mtime and other information of the source station can be rebuilt locally to take over the source station and establish an association relationship between local metadata and a third party.

[0074] Reconstruction supports different types of backends. By creating storage categories and flexibly configuring the backend, the source station can be rebuilt. Whether to rebuild data is a configuration item of the reconstruction type. There is a need to manage third-party platforms but not rebuild data. In order to ensure data reliability, all metadata information of the object needs to be rebuilt locally, including meta information, tagging information, etag, and mtime.

[0075] For the source station, if there is no etag information, special processing is required to generate etag information generated by an empty object.

[0076] 2. Disaster reconstruction mode

[0077] Optionally, when the preset data reconstruction mode is a disaster reconstruction mode, the step of determining the source site information of multiple storage objects based on the object list includes: in the event of a data center failure, obtaining the tiered data that is pre-tiered in the public cloud of the failed data center, wherein the tiered data includes multiple tiered storage buckets; switching the failed data center to a business data center; creating a mirror storage bucket corresponding to the tiered storage bucket in the business data center; and determining the source site information of the multiple mirror storage buckets based on the object list.

[0078] For disaster reconstruction, there are abnormal scenarios where a data center fails and another data center is required to take over the business.

[0079] Figure 4 FIG. 1 is a schematic diagram of an optional method of using a disaster reconstruction mode to perform data reconstruction according to an embodiment of the present invention. Figure 4 As shown, in the event of a cluster failure, the local storage category corresponding to Data Center 1 is the local backend, and the secondary storage category provides S3 / NFS / SMB; its implementation process includes: the data of Data Center 1 is tiered to the public cloud (tiering here means that the data of Data Center 1 is written to the public cloud, the local data is deleted, and then the local metadata information is updated to point to the public cloud, where the metadata of the public cloud object records the original object's ACL information and the object's original key information). When Data Center 1 fails, you can switch the business, create a bucket with the same name in Data Center 2, and then create a disaster reconstruction for reconstruction.

[0080] Table 1 defines an optional disaster reconstruction framework.

[0081] Table 1

[0082] Source site key definition x-amz-meta-source-etag Rebuild to local as etag x-amz-meta-source-key Rebuild locally as key x-amz-meta-source-mtime Rebuild to local as mtime x-amz-meta-source-acl Rebuild locally as acl

[0083] In this embodiment, disaster reconstruction will rebuild the metadata and restore it to the original information. For example, object 1.txt in the bucket of data center 1 is layered to the bucket bucket_test of the public cloud, and the object name uses a random tag to restore the original information of the object.

[0084] 3. Upgrade and rebuild mode

[0085] Optionally, when the preset data reconstruction mode is the upgrade reconstruction mode, the step of determining the metadata of the source site information of multiple storage objects based on the object list includes: in a cross-version upgrade scenario, traversing the object list to obtain multiple object information; based on the object information, using an incremental synchronization strategy to rebuild the storage object; and determining the metadata of the source site information of each storage object.

[0086] As an optional implementation method of this embodiment, an incremental synchronization strategy is adopted to rebuild the storage objects, including: obtaining the platform type of the second-level platform under the server platform; based on the platform type of the second-level platform, fully rebuilding the storage objects under the server platform; after completing the full reconstruction of the storage objects, incrementally rebuilding each storage object under the second-level storage bucket.

[0087] For cross-version upgrade scenarios, an upgrade reconstruction mode is provided. Since the reconstruction process is to traverse and list the source site objects for reconstruction, there is a possibility that objects will be written ahead of the progress during the traversal process, resulting in omissions. Therefore, an incremental synchronization process is required. An incremental reconstruction mechanism is provided for internal products.

[0088] Figure 5 FIG. 1 is a schematic diagram of an optional data reconstruction using an upgrade reconstruction mode according to an embodiment of the present invention. Figure 5 As shown, the implementation process includes: creating an upgrade and rebuilding; then checking the secondary platform type, and then determining whether it is a preset object storage system ( Figure 5 XEOS is used as an example in the figure); if yes, the secondary opens the bucket log and performs a full rebuild, and then incrementally rebuilds the secondary bucket log; if no, a full rebuild can be performed directly.

[0089] As an optional implementation of this embodiment, after the reconstruction is completed, the data object can be reconstructed.

[0090] The process of rebuilding data objects takes a long time. In order to improve the ability of rapid management, metadata is usually rebuilt first. After the metadata is rebuilt, a task to rebuild the data will be generated. The task execution time is a configurable item. For example, after the metadata reconstruction is completed, it is set to delay for 1 hour and then rebuild the target data.

[0091] Optionally, the step of determining the source site information of multiple storage objects based on the object list also includes: for multi-version storage buckets, obtaining the bucket identifier of each version in the source site storage bucket to which the target data belongs; based on the bucket identifier of each version, indexing the object list of each storage object in the source site storage bucket; based on the object list, determining the source site information of multiple storage objects.

[0092] In an embodiment of the present invention, a multi-version reconstruction method is provided to make the multi-versions of the secondary correspond one-to-one with the cluster, so as to realize seamless switching and takeover of the business. For different secondary storage categories, for example, there is a secondary storage target of S3. If the storage bucket has multiple versions, it is necessary to rebuild the multi-version id of the source site to the current site. Since the length of the versionid of different manufacturers is different, the reconstruction process must not only ensure the order, but also ensure the one-to-one correspondence of the versionid (version number). Therefore, in the reconstruction process, it is necessary to generate a link key (keyword) locally to match the local versionid with the secondary to meet the access of getobject and listbucket.

[0093] The above two reconstruction modes (disaster reconstruction and upgrade reconstruction) are two full-scale management modes. Disaster reconstruction can rebuild the original key, acl and other information of the object, and upgrade reconstruction is applied to the full reconstruction object to avoid the incremental objects generated during the reconstruction process.

[0094] Through the above embodiments, three reconstruction strategies and application scenarios are provided: ordinary reconstruction, upgraded reconstruction and disaster recovery reconstruction. Ordinary reconstruction is applied to source stations such as public clouds, private clouds, and NFS. The metadata information, tagging information, etag, mtime and other information of the source station are rebuilt locally to take over the source station and establish an association between local metadata and third parties. A unified access portal for this site is provided to quickly manage third-party platforms through metadata reconstruction. Two full-scale management modes are provided: disaster reconstruction and upgraded reconstruction. Disaster reconstruction can rebuild the original key and acl and other information of the object; upgraded reconstruction is applied to fully reconstructed objects to avoid incremental objects generated during the reconstruction process.

[0095] Through the above embodiments, data index reconstruction can be achieved, metadata and data can be decoupled, and the object list of the source station bucket can be indexed by creating a storage platform access configuration, using the platform configuration information to write metadata information to the cluster, and at the same time, an association relationship with the source station can be established through the storage category. Then, metadata can be quickly rebuilt to achieve user business switching, and source station metadata and data information can be quickly rebuilt locally by creating and issuing tasks.

[0096] The present invention is described in detail below in conjunction with another optional embodiment.

[0097] Embodiment 2

[0098] This embodiment provides a data reconstruction device, which is applied to a server platform. The data reconstruction device includes multiple implementation units, each of which corresponds to each implementation step in the above-mentioned embodiment 1.

[0099] Figure 6is a schematic diagram of an optional data reconstruction device according to an embodiment of the present invention. Figure 6 As shown, the device may include: a creation unit 61, a determination unit 63, a reconstruction unit 65, and a generation unit 67, wherein:

[0100] A creation unit 61 is used to create access configuration information for accessing a preset storage platform in response to a data reconstruction instruction, wherein the access configuration information is used to access the preset storage platform to index an object list of each storage object in a source station storage bucket to which the target data belongs;

[0101] A determination unit 63, configured to determine source station information of multiple storage objects based on the object list;

[0102] A reconstruction unit 65 is used to reconstruct the metadata of each storage object based on the source station information by using a pre-configured call enumeration interface, and write the metadata to the local cluster;

[0103] The generating unit 67 is used to generate a data reconstruction task based on the source station information and metadata, wherein the data reconstruction task is used to reconstruct the target data.

[0104] The above-mentioned data reconstruction device can create access configuration information for accessing the preset storage platform in response to the data reconstruction instruction through the creation unit 61, wherein the access configuration information is used to access the preset storage platform to index the object list of each storage object in the source station storage bucket to which the target data belongs, and the source station information of multiple storage objects is determined based on the object list by the determination unit 63, and the metadata of each storage object is reconstructed based on the source station information by the reconstruction unit 65 using the pre-configured call enumeration interface, and the metadata is written to the local cluster, and the data reconstruction task is generated based on the source station information and metadata by the generation unit 67, wherein the data reconstruction task is used to reconstruct the target data. In this embodiment, when creating an access configuration signal for the storage platform, the object list of the source station bucket can be indexed using the platform configuration information, and the metadata information is written to the local cluster, and the metadata and the target data are decoupled, and the source station metadata and data information can be quickly reconstructed to the local by creating and sending tasks, so as to realize the rapid reconstruction of batch data, thereby solving the technical problem that the batch reconstruction of data cannot be realized by using the back-to-source method in the related technology.

[0105] Optionally, the data reconstruction method is applied to a preset data reconstruction mode, wherein the preset data reconstruction mode is one of the following: a normal reconstruction mode, an upgrade reconstruction mode, and a disaster reconstruction mode.

[0106] Optionally, when the preset data reconstruction mode is the normal reconstruction mode, the determination unit includes: a first traversal module, used to traverse the object list to obtain multiple object information; a first reconstruction module, used to reconstruct the storage object based on the object information; a first judgment module, used to judge whether there is a reconstructed storage object in the local cluster; a first access module, used to access the source station to obtain metadata information and object tags of the storage object if there is no reconstructed storage object in the local cluster; a first determination module, used to represent the metadata information and object tags as source station information.

[0107] Optionally, when the preset data reconstruction mode is the normal reconstruction mode, the reconstruction unit also includes: a first acquisition module, used to obtain the storage category supported by the server platform; a first calling module, used to call a pre-configured calling enumeration interface according to the storage category; and a first reconstruction module, used to use the calling enumeration interface to reconstruct the metadata of each storage object.

[0108] Optionally, when the preset data reconstruction mode is a disaster reconstruction mode, the determination unit includes: a second acquisition module, used to obtain the layered data of the failed data center in the public cloud in advance in the event of a data center failure, wherein the layered data includes multiple layered storage buckets; a first switching module, used to switch the failed data center to a business data center; a first creation module, used to create a mirror storage bucket corresponding to the layered storage bucket in the business data center; and a second determination module, used to determine the source site information of multiple mirror storage buckets based on the object list.

[0109] Optionally, when the preset data reconstruction mode is the upgrade reconstruction mode, the determination unit includes: a second traversal module, used to traverse the object list in a cross-version upgrade scenario to obtain multiple object information; a second reconstruction module, used to reconstruct the storage object based on the object information and adopt an incremental synchronization strategy; a third determination module, used to determine the metadata of the source site information of each storage object.

[0110] Optionally, the second reconstruction module includes: an acquisition sub-module, used to obtain the platform type of the second-level platform under the server platform; a first reconstruction sub-module, used to fully rebuild the storage objects under the server platform based on the platform type of the second-level platform; and a second reconstruction sub-module, used to incrementally rebuild each storage object under the second-level storage bucket after completing the full reconstruction of the storage objects.

[0111] Optionally, the determination unit also includes: a third acquisition module, used to obtain the bucket identifier of each version in the source station storage bucket to which the target data belongs for multi-version storage buckets; an indexing module, used to index the object list of each storage object in the source station storage bucket based on the bucket identifier of each version; and a fourth determination module, used to determine the source station information of multiple storage objects based on the object list.

[0112] The above-mentioned data reconstruction device may also include a processor and a memory. The above-mentioned creation unit 61, determination unit 63, reconstruction unit 65, generation unit 67, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.

[0113] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the kernel parameters are adjusted to reconstruct the metadata of each storage object based on the source station information using a pre-configured call enumeration interface, and the metadata is written to the local cluster. Based on the source station information and the metadata, a data reconstruction task is generated, wherein the data reconstruction task is used to reconstruct the target data.

[0114] The above-mentioned memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one storage chip.

[0115] According to another aspect of an embodiment of the present invention, there is also provided a data storage system, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any of the above-mentioned data reconstruction methods by executing the executable instructions.

[0116] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium including a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned data reconstruction methods.

[0117] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program that is initialized with the following method steps: in response to a data reconstruction instruction, creating access configuration information for accessing a preset storage platform, wherein the access configuration information is used to access the preset storage platform to index an object list of each storage object in a source station storage bucket to which the target data belongs; based on the object list, determining the source station information of multiple storage objects; based on the source station information, using a pre-configured call enumeration interface to rebuild the metadata of each storage object, and writing the metadata to a local cluster; based on the source station information and the metadata, generating a data reconstruction task, wherein the data reconstruction task is used to reconstruct the target data.

[0118] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0119] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0120] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0121] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0122] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0123] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0124] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A data reconstruction method, characterized in that: Applied to the server platform, the data reconstruction method is applied to a preset data reconstruction mode, wherein the preset data reconstruction mode is one of the following: a normal reconstruction mode, an upgrade reconstruction mode, and a disaster reconstruction mode, including: In response to the data reconstruction instruction, create access configuration information for accessing a preset storage platform, wherein the access configuration information is used to access the preset storage platform to index an object list of each storage object in a source station storage bucket to which the target data belongs; Based on the object list, determining source site information of multiple storage objects; In the case where the preset data reconstruction mode is the common reconstruction mode, the step of determining the source station information of the plurality of storage objects based on the object list includes: traversing the object list to obtain the plurality of object information; reconstructing the storage object based on the object information; determining whether the reconstructed storage object exists in the local cluster; if the reconstructed storage object does not exist in the local cluster, accessing the source station to obtain the metadata information and the object tag of the storage object; and characterizing the metadata information and the object tag as the source station information; Based on the source station information, a pre-configured call enumeration interface is used to rebuild the metadata of each storage object, and the metadata is written to the local cluster; A data reconstruction task is generated based on the source station information and the metadata, wherein the data reconstruction task is used to reconstruct the target data.

2. The method according to claim 1, characterized in that In the case where the preset data reconstruction mode is the common reconstruction mode, the step of reconstructing the metadata of each storage object by using a preconfigured calling enumeration interface based on the source station information further includes: Obtain storage categories supported by the server platform; According to the storage category, calling a preconfigured call enumeration interface; The metadata of each storage object is rebuilt by calling the enumeration interface.

3. The method according to claim 1, characterized in that: In the case where the preset data reconstruction mode is a disaster reconstruction mode, the step of determining source station information of a plurality of storage objects based on the object list includes: In the event of a data center failure, obtaining pre-tiered tiered data of the failed data center in a public cloud, wherein the tiered data includes a plurality of tiered storage buckets; Switching the faulty data center to a service data center; Creating a mirror storage bucket corresponding to the hierarchical storage bucket in the business data center; Based on the object list, source site information of the plurality of image storage buckets is determined.

4. The method according to claim 1, characterized in that In the case where the preset data reconstruction mode is the upgrade reconstruction mode, the step of determining metadata of source station information of a plurality of storage objects based on the object list includes: In a cross-version upgrade scenario, traverse the object list to obtain information of multiple objects; Based on the object information, an incremental synchronization strategy is adopted to rebuild the storage object; Metadata that identifies the origin information for each storage object.

5. The method according to claim 4, characterized in that The step of using an incremental synchronization strategy to rebuild the storage object includes: Obtaining the platform type of the second-level platform under the server platform; Based on the platform type of the second-level platform, fully rebuild the storage object under the server platform; After the full reconstruction of the storage object is completed, each storage object under the second-level storage bucket is incrementally rebuilt.

6. The method according to claim 1, characterized in that The step of determining source station information of a plurality of storage objects based on the object list further includes: For buckets with multiple versions, obtain the bucket ID of each version in the source site bucket to which the target data belongs; Based on the bucket identifier of each version, index an object list of each storage object in the source station storage bucket; Based on the object list, source site information of a plurality of storage objects is determined.

7. A data reconstruction device, characterized in that: Applied to the server platform, applied to a preset data reconstruction mode, wherein the preset data reconstruction mode is one of the following: a normal reconstruction mode, an upgrade reconstruction mode, and a disaster reconstruction mode, including: A creation unit, configured to respond to the data reconstruction instruction and create access configuration information for accessing a preset storage platform, wherein the access configuration information is used to access the preset storage platform to index an object list of each storage object in a source station storage bucket to which the target data belongs; a determining unit, configured to determine source station information of a plurality of storage objects based on the object list; In the case where the preset data reconstruction mode is the common reconstruction mode, the determination unit includes: a first traversal module, used to traverse the object list to obtain multiple object information; a first reconstruction module, used to reconstruct the storage object based on the object information; a first judgment module, used to judge whether the reconstructed storage object exists in the local cluster; a first access module, used to access the source station to obtain metadata information and object tags of the storage object if the reconstructed storage object does not exist in the local cluster; a first determination module, used to characterize the metadata information and the object tag as the source station information; A reconstruction unit, configured to reconstruct the metadata of each storage object based on the source station information by using a preconfigured call enumeration interface, and write the metadata into a local cluster; A generating unit is used to generate a data reconstruction task based on the source station information and the metadata, wherein the data reconstruction task is used to reconstruct the target data.

8. A data storage system, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to perform the data reconstruction method described in any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the data reconstruction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Replication of data objects from a source server to a target server

    CN103765817A