Method and system for copying production database storage to non-native backup storage
The method and system optimize non-native backup storage by replicating filesystem layout and using key-value pairs to ensure data consistency and integrity, addressing storage inefficiencies and improving data recovery efficiency.
Patent Information
- Application Number
- PCT/EP2024/050700
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-17
AI Technical Summary
Conventional database management systems face challenges in optimizing storage space utilization for non-native backups, leading to suboptimal performance and lack of efficient data reliability during restoration operations.
A method and system for copying production database storage to non-native backup storage by replicating the filesystem layout, data, and metadata using key-value pairs, ensuring data consistency and integrity, and enabling live-mounting for efficient data recovery.
The solution reduces storage costs by eliminating redundant data, maintains data consistency and integrity, and facilitates easy data retrieval with reduced computational overhead, enhancing data availability and recovery efficiency.
Smart Images

Figure EP2024050700_17072025_PF_FP_ABST
Abstract
Description
[0001] METHOD AND SYSTEM FOR COPYING PRODUCTION DATABASE STORAGE TO NON-NATIVE BACKUP STORAGE
[0002] TECHNICAL FIELD
[0003] The present disclosure relates generally to the field of data management and more specifically, to a method and a system for copying a production database storage to a non-native backup storage.
[0004] BACKGROUND
[0005] Generally, a database management system includes a database that is used by different organizations to store critical and large amounts of data (or assets), for example, the data related to multiple subjects, potential customers, and the like. The data is typically saved on a filesystem, for example, in the form of one or more files that include data, metadata (e.g., lock files, or data mappings), or the combination of the data with the metadata (e.g., logfiles, or data files including indexing). Moreover, the database management system is required to be resilient to handle data node failures in order to allow data retrieval even if any of the data nodes of the database fails. Therefore, the database management system conducts data replication within the database to ensure data availability across multiple nodes, even in the event of a failure in any of the database nodes. A non- native backup can be used to store replicated data that can be further used for data recovery. Moreover, such non-native backup is used to perform data storage at different storages, or a storage namespace creating a border between production data used by the database and the non-native backup. The data of the non-native backup can be accessed via a database application programming interface (API) or by copying the files on a production filesystem of the database.
[0006] Conventionally, the database management systems pose significant challenges, such as data striping, non-availability of enough resources that are required to store the data, and the like due to which a conventional database management system fails to retrieve the relevant data. However, certain attempts have been made to provide a space-efficient backup for databases with replication, such as mapping duplicate files / blocks to a single copy, exposing a virtual filesystem showing only the data that is relevant to the data subject, and the like. Moreover, such attempts lack optimization for database replication data, leading to suboptimal performance, which is not desirable. In addition, the conventional database management system also lacks an efficient and effective storage space allocation, that lacks data reliability for the entire backup process, which is again not desirable. Thus, there exists a technical problem of how to reduce the storage space utilization for non-native backups, while retaining an ease of restoration operation with an ability to live-mount the backup.
[0007] Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with the conventional methods and systems to provide a space-efficient backup for databases with replication.
[0008] SUMMARY
[0009] The present disclosure provides a method and a system for copying a production database storage to a non-native backup storage. The present disclosure provides a solution to the existing problem of how to reduce the storage space utilization for non-native backups, while retaining an ease of restoration operation with an ability to live-mount the backup. An objective of the present disclosure is to provide a solution that overcomes at least partially the problems encountered in the prior art and provides an improved method and an improved system for copying the production database storage to the non-native backup storage, such as by providing a space-efficient backup for databases with replication.
[0010] One or more objectives of the present disclosure are achieved by the solutions provided in the enclosed independent claims. Advantageous implementations of the present disclosure are further defined in the dependent claims. In one aspect, the present disclosure provides a method of copying a production database storage to a non-native backup storage comprising steps of copying a filesystem layout from a production database filesystem of the production database storage to a backup database filesystem of the non-native backup storage. Furthermore, the method includes copying the data and metadata from the production database storage to the non-native backup storage and populating a backup deduplication database of the non-native backup storage with a plurality of key -value pairs. Moreover, each key is a combination of a file and an offset for respective data and metadata stored in the production database storage. Each value is a location in the non-native backup storage where the respective data has been saved. Furthermore, the method includes copying replicated metadata from the production database storage to the non-native backup storage and further populating the backup deduplication database of the non-native backup storage with a plurality of key -value pairs. Moreover, each key is a combination of a file and an offset for respective replicated metadata and each value is a location in the backup storage where the respective metadata has been saved. The method further includes populating the backup deduplication database of the non-native backup storage with a plurality of keyvalue pairs, where each key is a combination of a file and an offset for respective replicated data, and each value is a location in the backup storage where the respective data has already been saved when previously copying the data. Moreover, the order of copying the data and metadata from the production database storage to the non-native backup storage and copying the replicated metadata from the production database storage to the non-native backup storage can be reversed.
[0011] Advantageously, the method enables copying the production database storage to the non-native backup storage for providing an efficient, accurate, and reliable data replication while maintaining data consistency, data integrity, and data availability. The copying of the filesystem layout from the production database filesystem to the backup database filesystem is performed to ensure that the structure and organization of files and directories within the production database storage are exactly replicated in the non-native backup storage. Furthermore, the copying of the data and the metadata to the non-native backup storage while populating the backup deduplication database with the plurality of key -value pairs is performed to identify and eliminate the redundant data that reduces the storage cost of the non-native backup storage. Furthermore, the copying of the replicated metadata to the non-native backup storage and populating the backup deduplication database with the plurality of key -value pairs provides an organized and efficient management of the duplicated metadata within the non-native backup storage, contributing to data consistency and reliability. The population of the backup deduplication database with the plurality of keyvalue pairs for the replicated data is performed to ensure that a comprehensive data record is maintained, contributing to an efficient tracking and retrieval of the data. In addition, the combination of copying the filesystem layout, the data, the metadata, and the replicated metadata optimizes the non-native backup storage, such as by exposing the exact filesystem of the production database storage (i.e., the production filesystem) on the non-native backup storage. Furthermore, the method is used to enhance the capability of the non-native backup storage for live -mounting, while also ensuring the ease of the restoration operation. As a result, a reliable, efficient, and secure data recovery and data replication can be performed while maintaining the data consistency and data integrity with reduced cost utilization.
[0012] In an implementation form, after the first copying step, the data and metadata from the production database storage are located using an open database format.
[0013] Advantageously, the open database format facilitates the data retrieval and data processing to allow an accurate, efficient, and reliable recognition of the data and the metadata.
[0014] In another implementation form, the respective data is located using predefined offsets.
[0015] In another implementation form, the copying steps are performed by a copy agent. In such an implementation form, the copy agent is used to ensure a seamless and an accurate execution of the copy operation for an efficient copying of the production storage database to the non-native backup storage.
[0016] In another implementation form, a virtual file system is used to expose the data stored on the backup storage system.
[0017] By virtue of using the virtual file system, the data on the non-native backup storage can be presented and managed in a consistent, organized, and easily navigable manner, providing a user-friendly interface for accessing and manipulating the stored information.
[0018] In another implementation form, the file system layout includes file / directory relations and file sizes.
[0019] Advantageously, the file / directory relations and file sizes are used to provide an identical representation of the structure of the corresponding files and directories stored in the production database storage that is further used to ensure data consistency and reliability while copying the production database storage to the non-native backup storage.
[0020] In another aspect, the present disclosure provides a system comprising means adapted for carrying out all the steps of the method.
[0021] The system achieves all the advantages and technical effects of the method of the present disclosure.
[0022] It is to be appreciated that all the aforementioned implementation forms can be combined.
[0023] It has to be noted that all devices, elements, circuitry, units, and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof. It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims.
[0024] Additional aspects, advantages, features, and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative implementations constmed in conjunction with the appended claims that follow.
[0025] BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The summary above, as well as the following detailed description of illustrative embodiments, is better understood when read in conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions of the disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods and instrumentalities disclosed herein. Moreover, those in the art will understand that the drawings are not to scale. Wherever possible, like elements have been indicated by identical numbers.
[0027] Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein: FIG. 1 is a block diagram that depicts a system configured to copy a production database storage to a non-native backup storage, in accordance with an embodiment of the present disclosure;
[0028] FIG. 2 is a flowchart depicting a method of copying the production database storage to the non-native backup storage, in accordance with an embodiment of the present disclosure; and
[0029] FIG. 3 is an exemplary diagram that depicts the copying of the production database storage to the non-native backup storage, in accordance with an embodiment of the present disclosure.
[0030] In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing.
[0031] DETAILED DESCRIPTION OF EMBODIMENTS
[0032] The following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art would recognize that other embodiments for carrying out or practicing the present disclosure are also possible.
[0033] FIG. 1 is a block diagram that depicts a system configured to copy a production database storage to a non-native backup storage, in accordance with an embodiment of the present disclosure. With reference to FIG. l, there is shown a block diagram that includes a system 100. The system 100 includes a production database storage 102, a copy agent 104, and a non-native backup storage 106. The production database storage 102 includes a production database filesystem 108, a filesystem layout 110, a first processor 118, a first memory 120, a first network interface 122, and a communication network 136. Moreover, the non- native backup storage 106 includes a backup database filesystem 124, a virtual filesystem 126, a backup deduplication database 128, a second processor 130, a second memory 132, and a second network interface 134.
[0034] The production database storage 102 refers to a database storage that is configured to store data, such as tabular data, documents, images, audio files, video files, HTML files, and the like. Moreover, the non-native backup storage 106 refers to a storage, which is external to a database engine and is used to create a backup database. Examples of the production database storage 102 and the non-native backup storage 106 may include but are not limited to, dedicated servers equipped with storage systems, network-attached storage (NAS), storage area network (SAN) devices, and the like.
[0035] In accordance with an embodiment, the copying steps are performed by the copy agent 104. The copy agent 104 is configured to perform the copy operation for copying the production database storage 102 to the non-native backup storage 106. In an example, the copy agent 104 is configured to copy the filesystem layout 110 from the production database filesystem 108 of the production database storage 102 to the backup database filesystem 124 of the non-native backup storage 106. In another example, the copy agent 104 is configured to copy the data and metadata from the production database storage 102 to the non- native backup storage 106. In yet another example, the copy agent 104 is configured to copy the replicated metadata from the production database storage 102 to the non-native backup storage 106. In an implementation, the copy agent 104 is communicatively coupled with the production database storage 102 and the non-native backup storage 106 to perform the copy operations through the communication network 136. In another implementation, the copy agent 104 is located in the production database storage 102 to perform the copy operations without affecting the scope of the present disclosure. As a result, the copy agent 104 is used to ensure a seamless and accurate execution of the copy operation for an efficient copying of the production database storage 102 to the non-native backup storage 106. The first processor 118 is configured to execute all the necessary operations of the production database storage 102 and the second processor 130 is configured to execute all the necessary operations of the non-native backup storage 106. Examples of the first processor 118 and the second processor 130 may include, but are not limited to, a microcontroller, a microprocessor, a central processing unit (CPU), a complex instruction set computing (CISC) processor, an application-specific integrated circuit (ASIC) processor, a reduced instruction set (RISC) processor, a very long instruction word (VLIW) processor, a data processing unit, and other processors or control circuitry.
[0036] The production database filesystem 108 refers to a filesystem, which is associated with the production database storage 102 and the backup database filesystem 124 refers to a filesystem that is associated with the non-native backup storage 106. Furthermore, the virtual filesystem 126 refers to an abstraction layer that is configured to provide a unified and consistent interface for interacting with the data, such as by exposing the data stored in the non-native backup storage 106. In other words, the virtual filesystem 126 acts as the abstraction layer of the non-native backup storage 106, which uses the backup deduplication database 128 in orderto access any file. Examples of the production database filesystem 108, the backup database filesystem 124, and the virtual filesystem 126 may include but are not limited to, a global file system (GFS), hierarchical file system (HFS), universal disk format (UDF), or any other file system.
[0037] The backup deduplication database 128 corresponds to a collection of data stored in an organized manner, which can be accessed electronically. Moreover, the backup deduplication database 128 acts as an abstraction layer that is used to map the file and offset to relevant data even if the data is replicated. The backup deduplication database 128 is used to describe the data placement in the non-native backup storage 104. Examples of the backup deduplication database 128 may include, but are not limited to, hierarchical databases, network databases, object-oriented databases, relational databases, NoSQL databases, and the like.
[0038] The first memory 120 and the second memory 132 are configured to store the instructions executable by the first processor 118 and the second processor 130 respectively. Examples of the first memory 120 and the second memory 132 may include, but are not limited to, an Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read-Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, Solid-State Drive (SSD), persistent memory, remote direct memory access (RDMA), or CPU cache memory.
[0039] The first network interface 122 may include hardware or software that is configured to establish communication among the first processor 118, the first memory 120, and the production database filesystem 108. Furthermore, the second network interface 134 may include hardware or software that is configured to establish communication among the second processor 130, the second memory 132, the backup database filesystem 124, the virtual filesystem 126, and the backup deduplication database 128. Examples of the first network interface 122 and the second network interface 134 may include but are not limited to, a computer port, a network socket, a network interface controller (NIC), and any other network interface device.
[0040] The communication network 136 includes a medium (e.g., a communication channel) through which the production database storage 102, the copy agent 104, and the non-native backup storage 106 communicate with each other. The communication network 136 may be a wired or wireless communication network. Examples of the communication network 136 may include, but are not limited to, a Local Area Network (LAN), a wireless personal area network (WPAN), a Wireless Local Area Network (WLAN), a wireless wide area network (WWAN), a cloud network, a Long-Term Evolution (LTE) network, a plain old telephone service (POTS), a Metropolitan Area Network (MAN), and / or the Internet.
[0041] In operation, the system 100 is configured to copy the production database storage 102 to the non-native backup storage 106. The production database storage 102 is copied to create a data backup in the form of the non-native backup storage 106 while retaining an ease of data recovery operation along with an ability to live-mount the backup in order to ensure data redundancy, data security, and data availability in case of unexpected events (e.g., data node failure) affecting the data retrieval from the production database storage 102. Moreover, the non-native backup storage 106 uses a database engine application programming interface (API) to read and write the data, especially when data stripping is very complex. The data stripping corresponds to a technique of segmenting logically sequential data, such as a file, so that consecutive segments are stored on different physical storage devices. In non-structured query language databases (non-SQL DBs), key values or table rows are striped, and each key value is distributed across the different storage locations or multiple storage locations. The consecutive identical parts of such keys are too short due to which conventional deduplication algorithms fail to identify duplicate values. However, in conventional database management systems, the replicated data resides within a same file with either alignment or stripe size, which is not sufficient to locate for an unaware deduplication algorithm. Such issue of locating the data is resolved through the population of the backup deduplication database 128. Furthermore, the ability of the system 100 to live-mount the backup allows the system 100 to check the backup before the execution of the data recovery operation. For example, when a user wants to perform the data recovery operation in order to recover from ransomware. In such a situation, the last backup copy, which is the backup copy before the ransomware hit, is not easy to find, therefore, by using the live-mount ability of the system 100, the user can perform the data recovery operation efficiently with reduced cost and with reduced resource utilization.
[0042] Furthermore, the system 100 is configured to copy the filesystem layout 110 from the production database filesystem 108 of the production database storage 102 to the backup database filesystem 124 of the non-native backup storage 106. The filesystem layout 110 refers to a hierarchical structure and organization of files and directories that are stored in the production database filesystem 108. In other words, the filesystem layout 110 acts as a blueprint for the data, which is stored in the production database filesystem 108. Beneficially, the copying of the filesystem layout 110 from the production database filesystem 108 of the production database storage 102 to the backup database filesystem 124 of the non-native backup storage 106 ensures an exact replication of the structure and organization of files and directories in the non-native backup storage 106, which can be further used to facilitate an efficient and reliable exposing of the data in the non-native backup storage 106. In addition, by copying the filesystem layout 110 to the non-native backup storage 106, the data stored in the non-native backup storage 106 can be easily accessible with reduced computational cost and overall processing time that is required to locate the data.
[0043] In accordance with an embodiment, the filesystem layout 110 includes file / directory relations and file sizes. The filesystem layout 110 provides a hierarchical relationship between the files and directories, such as through the files / directory relations and the file sizes. Moreover, the transfer of the file / directory relations to the non-native backup storage 106 provides an identical representation of the structure of the corresponding files and directories stored in the production database storage 102. Additionally, files or directories that are linked to a NULL data bulk are indicated as empty files. The file sizes indicate the amount of storage space that each file occupies within the production database filesystem 108. As a result, the file / directory relations and the file sizes are used to ensure data consistency and reliability while copying the production database storage 102 to the non-native backup storage 106.
[0044] In accordance with an embodiment, after the first copying step, the data, and the metadata from the production database storage 102 are located using an open database format. In other words, the open database format is employed to access the data and metadata copied from the production database storage 102. Therefore, the open database format facilitates the data retrieval and data processing to allow an accurate, efficient, and reliable recognition of the data and the metadata.
[0045] In accordance with an embodiment, the open database format is used to read each file, parse the metadata, and locate the respective data. In an implementation, the open database format is used to read files and further parse the metadata for locating the data. In another implementation, the open database format is used to read the files and further go to pre-defined offsets in order to locate the data. An example of the filesystem layout 110 for the production database storage 102 (i.e., rocksdb) is given in Table 1 provided below:
[0046] Table 1
[0047] The above-mentioned filesystem layout 110 (i.e., a filesystem layout for a database, for example, a rocksdb database) represent file names, file size, metadata, and the like. For example, a fifth column of Table 1 represents the file size (e.g., 37, 4. OK, 5.2K, 42M, and the like), and a ninth column of Table 1 represents the file names (e.g., 000006. sst). In addition, the file named 000006. sst includes the data and rest of the files (e.g., 000003.log) includes the metadata that may be used by the system 100 to manage data accessibility and optimize data retrieval operations. As a result, the overall time required to locate the data is reduced while maintaining the overall data consistency and data integrity. In accordance with an embodiment, the respective data is located by using the predefined offsets. Firstly, the system 100 is configured to read the files and then locate the respective data by directly using the pre-defined offsets. By locating the respective data through the predefined offset the system 100 is configured to ensure a systematic and accurate data retrieval while reducing the overall processing time that is required to locate the respective data during the execution of the data retrieval or data recovery operation.
[0048] In accordance with an embodiment, a saved data format differs between data and metadata. For example, the saved data format for the data is different from the saved data format for the metadata. By having different saved data formats for the data and the metadata, the system 100 can manage and organize the data and the metadata separately thereby, ensuring an efficient data storage, data retrieval, and data processing.
[0049] Furthermore, the system 100 is configured to copy the data and the metadata from the production database storage 102 to the non-native backup storage 106 and populate the backup deduplication database 128 of the non-native backup storage 106 with a plurality of key -value pairs. Each key is a combination of a file and an offset for respective data and metadata stored in the production database storage 102 and each value is a location in the non-native backup storage 106 where the respective data has been saved. In other words, the copy agent 104 of the system 100 is configured to copy the data and the metadata from the production database storage 102 to the non-native backup storage 106. Thereafter, the system 100 is configured to populate the backup deduplication database 128 with the plurality of key -value pairs. In an implementation, the files include the data, the metadata, or a combination of the data and the metadata (e.g., indexed data, or logfiles). Moreover, the data refers to the content of the file while the metadata refers to information about the data, such as a structure, properties, or other descriptive details of the data that are required to locate, access, and identify the data. Furthermore, the backup deduplication database 128 is populated with the plurality of key -value pairs, such as through the second processor 130 of the non-naive backup storage 106. Each key -value pair includes a key, which is the combination of the file and the offset. In an implementation, the file represents the data, which is stored in the non-native backup storage 106 while the offset represents the exact position of the corresponding data within the corresponding file. The offset is used to pinpoint the exact location of the data and the metadata within the non- native backup storage 106, such as by referring to the backup deduplication database 128. Moreover, the value depicts the location of the corresponding data in the non-native backup storage 106. As a result, the data is easily accessible during the execution of the data retrieval operation with reduced cost utilization and reduced overall processing time, which is required to locate the data and further retrieve the data.
[0050] The system 100 is further configured to copy replicated metadata from the production database storage 102 to the non-native backup storage 106 and further populate the backup deduplication database 128 of the non-native backup storage 106 with the plurality of key -value pairs, where each key is the combination of the file and the offset for respective replicated metadata, and each value is the location of in the non-native backup storage 106 where the respective metadata has been saved. In other words, the system 100 is configured to copy the replicated metadata (e.g., file descriptions, timestamps, or other relevant details) that is stored in the production database storage 102 to the non-native backup storage 106, such as through the copy agent 104. After that, the system 100 is configured to populate the backup deduplication database 128 of the non-native backup storage 106 with the plurality of key-value pairs, such as through the second processor 130 of the non-native backup storage 106. Moreover, each key of the plurality of key -value pairs depicts a combination of the file and the offset that uniquely identifies each of the replicated metadata. In addition, each value of the plurality of key -value pairs represents the location within the non-native backup storage 106 where the corresponding replicated metadata is stored. In an implementation, the second processor 130 of the non-native backup storage 106 is configured to populate the backup deduplication database 128 of the non-native backup storage 106 with the plurality of key -value pairs for providing a unique identifier to the replicated metadata. The backup deduplication database 128 acts as an abstraction layer that allows the mapping of the file and the offset with the relevant data even if the data is replicated. Advantageously, the copying of the replicated metadata and further populating the backup deduplication database 128 of the non-native backup storage 106 with the plurality of key-value pairs provides an organized and efficient management of the duplicated metadata within the non-native backup storage 106, contributing to data consistency and reliability.
[0051] Furthermore, the system 100 is configured to populate the backup deduplication database 128 of the non-native backup storage 106 with the plurality of key -value pairs, where each key is the combination of the file and the offset for respective replicated data. Moreover, each value is the location in the non-native backup storage 106 where the respective data has already been saved when previously copying the data. In an implementation, the second processor 130 of the non-native backup storage 106 is configured to populate the backup deduplication database 128 of the non-native backup storage 106 with the plurality of keyvalue pairs for providing a unique identifier to the respective replicated data. However, the order of copying the data and the metadata does not affect the scope of the present disclosure. As a result of populating the backup deduplication database 128, the system 100 is configured to maintain an organized record of replicated data, utilizing key -value pairs to efficiently track the specific files and the location of the respective replicated data in the non-native backup storage 106 thereby, maintaining the overall data consistency, data integrity, and the data accessibility.
[0052] In an implementation, the system 100 is configured to copy the data and the metadata from the production database storage 102 to the non-native backup storage 106 and populate the backup deduplication database 128 of the non-native backup storage with a plurality of key-value pairs. Subsequently, the system 100 is configured to copy the replicated metadata from the production database storage 102 to the non-native backup storage 106 and then populate the backup deduplication database 128 of the non-native backup storage 106 with the plurality of key -value pairs. In another implementation, the system 100 is configured to copy the replicated metadata from the production database storage 102 to the non-native backup storage 106 and then populate the backup deduplication database 128 of the non-native backup storage 106 with the plurality of key -value pairs. Furthermore, the system 100 is configured to copy the data and the metadata from the production database storage 102 to the non-native backup storage 106 and further populate the backup deduplication database 128 of the non-native backup storage 106 with the plurality of key -value pairs. In other words, the order of executing the copy operation for copying the data and the metadata and the copy operation for coping the replicated metadata can be reversed without affecting the scope of the present disclosure.
[0053] In accordance with an embodiment, the virtual filesystem 126 is used to expose the data stored on the non-native backup storage 106. In other words, the virtual filesystem 126 is configured to map each of the files to the relevant data and the metadata that had been originally contained in the file. By using the virtual filesystem 126, the data on the non-native backup storage 106 can be presented and managed in a consistent, organized, and easily navigable manner, providing a user-friendly interface for accessing and manipulating the stored information.
[0054] The system 100 allows an efficient backup and storage management for distributed databases with multiple nodes as compared to conventional distributed database management systems that keep multiple copies of the data while mirroring the multiple nodes. For example, a distributed database that maintains multiple copies of the data is listed in Table 2 that is provided below.
[0055] Table 2
[0056] The system 100 is configured to optimize data backup by selectively storing the five rows along with the metadata describing the file structure of each node. For example, the first node (nl) includes files with KI, K4, and K5 rows. Similarly, the second node (n2) includes the files with KI, K2, and K5 rows. The key value lies in the backup process, which stores each row once and constructs metadata around it. In contrast, traditional deduplication methods would back up the entire five nodes, resulting in three times more data. Even if deduplication engines were modified to comprehend row boundaries, the present system (i.e., the system 100) provides more efficient and effective data backup with reduced overall processing time. With the metadata constructs, the system 100 is configured to expose each of the nodes in its native format. Moreover, the system 100 never reads all the data from the node as it is written thereby, ensuring that each data piece is backed up only once. As a result, the system 100 enhances the backup efficiency and provides an organized and efficient data management as compared to the conventional database management systems.
[0057] Advantageously, the system 100 is configured to copy the production database storage 102 to the non-native backup storage 106 for providing efficient, accurate, and reliable data replication while maintaining data consistency, data integrity, and data availability. The copying of the filesystem layout 110 from the production database filesystem 108 to the backup database filesystem 124 is performed to ensure that the structure and organization of files and directories within the production database storage 102 are exactly replicated in the non-native backup storage 106. Furthermore, the copying of the data and the metadata to the non-native backup storage 106 while populating the backup deduplication database 128 with the plurality of key -value pairs is performed to identify and eliminate the redundant data that reduces the storage cost of the non-native backup storage 106. Furthermore, the copying of the replicated metadata to the non-native backup storage 106 and populating the backup deduplication database 128 with key -value pairs provides an organized and efficient management of the duplicated metadata within the non-native backup storage 106, contributing to data consistency and reliability.
[0058] The population of the backup deduplication database 128 with the plurality of key-value pairs for the replicated data is performed to ensure that a comprehensive data record is maintained, contributing to an efficient tracking and retrieval of the data. In addition, the combination of copying the filesystem layout 110, the data, the metadata, and the replicated metadata optimizes the non-native backup storage 106, such as by exposing the exact filesystem of the production database storage 102 (i.e., the production database filesystem 108) on the non-native backup storage 106. Furthermore, the system 100 is configured to enhance the capability of the non-native backup storage 106 for live -mounting, while also ensuring the ease of the restoration operation. As a result, a reliable, efficient, and secure data recovery data replication can be performed while maintaining the data consistency and integrity with reduced cost utilization.
[0059] FIG. 2 is a flowchart depicting a method of copying the production database storage to the non-native backup storage, in accordance with an embodiment of the present disclosure. With reference to FIG. 2, there is shown a flowchart of a method 200 for copying the production database storage 102 to the non-native backup storage 106. The method 200 includes steps 202 to 210.
[0060] There is provided the method 200 of copying the production database storage 102 to the non-native backup storage 106. The production database storage 102 is copied for creating the data backup in the form of the non-native backup storage 106 while retaining an ease of data recovery operation along with an ability to live-mount the backup in order to ensure data redundancy, data security, and data availability in case of unexpected events (e.g., data node failure) affecting the data retrieval from the production database storage 102. The ability to live-mount the backup allows to check the backup before the execution of the data recovery operation. For example, when a user wants to perform the data recovery operation in order to recover from the ransomware. In such a situation, the last backup copy, which is the backup copy before the ransomware hit, is not easy to find, therefore, by using the live-mount ability of the system 100, the user can perform the data recovery operation efficiently with reduced cost and with reduced resource utilization.
[0061] In operation, at step 202, the method 200 includes copying the filesystem layout 110 from the production database filesystem 108 of the production database storage 102 to the backup database filesystem 124 of the non-native backup storage 106. The filesystem layout 110 refers to a hierarchical structure and organization of files and directories that are stored in the production database filesystem 108. In other words, the filesystem layout 110 acts as a blueprint for the data, which is stored in the production database filesystem 108. Beneficially, the copying of the filesystem layout 110 from the production database filesystem 108 of the production database storage 102 to the backup database filesystem 124 of the non-native backup storage 106 ensures an exact replication of the structure and organization of files and directories in the non-native backup storage 106, which can be further used to facilitate an efficient and reliable exposing of the data in the non-native backup storage 106. In addition, by copying the filesystem layout 110 to the non-native backup storage 106, the data stored in the non-native backup storage 106 can be easily accessible with reduced computational cost and overall processing time that is required to locate the data.
[0062] At step 204, the method 200 further includes copying data and metadata from the production database storage 102 to the non- native backup storage 106 and populating the backup deduplication database 128 of the non-native backup storage 106 with a plurality of key -value pairs. Moreover, each key is a combination of a file and an offset for respective data and metadata stored in the production database storage 102, and each value is a location in the non-native backup storage 106 where the respective data has been saved. In other words, the copy agent 104 is configured for copying the data and the metadata from the production database storage 102 to the non-native backup storage 106. Thereafter, the method 200 is used to populate the backup deduplication database 128 with the plurality of key -value pairs. In an implementation, the files include the data, metadata, or a combination of the data and the metadata (e.g., indexed data, or logfiles). Moreover, the data refers to the content of the file while the metadata refers to information about the data, such as a structure, properties, or other descriptive details of the data that is required to locate, access, and identify the data. Furthermore, the backup deduplication database 128 is populated with the plurality of key-value pairs, such as through the second processor 130 of the non-naive backup storage 106. Each key -value pair includes a key, which is the combination of the file and the offset. In an implementation, the file represents the data, which is stored in the non-native backup storage 106 while the offset represents an exact position of the corresponding data within the corresponding file. The offset is used to pinpoint the exact location of the data and the metadata within the non-native backup storage 106, such as by referring to the backup deduplication database 128. Moreover, the value depicts the location of the corresponding data in the non-native backup storage 106. As a result, the data is easily accessible during the execution of the data retrieval operation with reduced cost utilization and reduced overall processing time, which is required to locate the data and further retrieve the data.
[0063] At step 206, the method 200 further includes copying replicated metadata from the production database storage 102 to the non- native backup storage 106 and further populating the backup deduplication database 128 of the non-native backup storage 106 with a plurality of key -value pairs. Moreover, each key is a combination of a file and an offset for respective replicated metadata, and each value is a location in the non-native backup storage 106 where the respective metadata has been saved. Advantageously, the copying of the replicated metadata and further populating the backup deduplication database 128 of the non-native backup storage 106 with the plurality of key-value pairs provides an organized and efficient management of the duplicated metadata within the non-native backup storage 106, contributing to data consistency and reliability.
[0064] At step 208, the method 200 further includes populating the backup deduplication database 128 of the non-native backup storage 106 with a plurality of key-value pairs. Moreover, each key is a combination of a file and an offset for respective replicated data, and each value is a location in the non-native backup storage 106 where the respective data has already been saved when previously copying the data. As a result of populating the backup deduplication database 128, the method 200 ensures that the backup deduplication database 128 maintains an organized record of replicated data, utilizing key-value pairs to efficiently track the specific files and the location of the respective replicated data in the non-native backup storage 106 thereby, maintaining the overall data consistency, data integrity, and the data accessibility.
[0065] Furthermore, at step 210, the method 200 includes locating the data and the metadata in the non-native backup storage 106. Moreover, the backup database filesystem 124 is configured to locate the data (replicated or non-replicated) and the metadata in the non-native backup storage 106 with reduced processing time and efficient resource allocation.
[0066] Advantageously, the method 200 enables copying of the production database storage 102 to the non-native backup storage 106 for providing an efficient, accurate, and reliable data replication while maintaining data consistency, data integrity, and data availability. The copying of the filesystem layout 110 from the production database filesystem 108 to the backup database filesystem 124 is performed to ensure that the structure and organization of files and directories within the production database storage 102 are exactly replicated in the non-native backup storage 106. Furthermore, the copying of the data and the metadata to the non-native backup storage 106 while populating the backup deduplication database 128 with the plurality of key -value pairs is performed to identify and eliminate the redundant data that reduces the storage cost of the non-native backup storage 106. Furthermore, the copying of the replicated metadata to the non-native backup storage 106 and populating the backup deduplication database 128 with key -value pairs provides an organized and efficient management of the duplicated metadata within the non-native backup storage, contributing to data consistency and reliability. The population of the backup deduplication database 128 with the plurality of key-value pairs for the replicated data is performed to ensure that a comprehensive data record is maintained, contributing to an efficient tracking and retrieval of the data. In addition, the combination of copying the filesystem layout 110, the data, the metadata, and the replicated metadata optimizes the non-native backup storage 106, such as by exposing the exact filesystem of the production database storage 102 (i.e., the production database filesystem 108) on the non-native backup storage 106. Furthermore, the method 200 is used to enhance the capability of the non-native backup storage 106 for live-mounting, while also ensuring the ease of the restoration operation. As a result, a reliable, efficient, and secure data recovery data replication can be performed while maintaining the data consistency and data integrity with reduced cost utilization. The steps 202 to 210 are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein.
[0067] There is provided a computer program comprising instructions that, when executed by a computer system, cause the computer system to implement the method 200. In an example, the instructions are implemented on the computer-readable media, which include, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read-Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory. In an example, the instructions are generated by a computer program, which is implemented in view of the method 200 for copying the production database storage 102 to the non-native backup storage 106.
[0068] FIG. 3 is an exemplary diagram that depicts the copying of the production database storage to the non-native backup storage, in accordance with an embodiment of the present disclosure. FIG. 3 is described in conjunction with elements from FIG. 1. With reference to FIG. 3 there is shown an exemplary diagram 300 that depicts the production database storage 102, the production database filesystem 108, the copy agent 104, the non-native backup storage 106, and the backup database filesystem 124
[0069] In an exemplary scenario, there is shown the copying of the production database storage 102 (of FIG. 1) to the non-native backup storage 106 (of FIG. 1). The copy agent 104 is configured to perform the copying operations that are performed to copy the production database storage 102 to the non-native backup storage 106. At operation 104A, the copy agent 104 is configured to copy the filesystem layout 110 from the production database filesystem 108 of the production database storage 102 to the backup database filesystem 124 of the non-native backup storage 106. Furthermore, at operation 104B, the copy agent 104 is configured to copy the data and the metadata from the production database storage 102 to the non-native backup storage 106. After that, the backup deduplication database 128 is populated with the plurality of key-value pairs, where each key is a combination of a file and an offset for respective data and metadata stored in the production database storage 102 and each value is a location in the non-native backup storage 106 where the respective data has been saved. At operation 104C, the copy agent 104 is configured to copy the replicated metadata from the production database storage 102 to the non-native backup storage 106. Thereafter, the backup deduplication database 128 of the non-native backup storage 106 is populated with the plurality of key -value pairs, where each key is a combination of the file and the offset for the respective replicated metadata, and each value is a location of in the non-native backup storage 106 where the respective metadata has been saved. After the execution of the copying operation, such as by the copy agent 104, the non-native backup storage 106 can locate any data or the metadata (replicated or non-replicated), which is copied by the production database storage 102 to the non-native backup storage 106. As a result, the copying of the production database storage 102 to the non-native backup storage 106 provides an efficient, accurate, and reliable data replication while maintaining data consistency, data integrity, and data availability for secure data recovery data replication with reduced cost utilization and reduced resource utilization.
[0070] Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe, and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be constmed to relate to the plural. The word "exemplary" is used herein to mean "serving as an example, instance or illustration". Any embodiment described as “exemplary” is not necessarily to be constmed as preferred or advantageous over other embodiments or to exclude the incorporation of features from other embodiments. The word "optionally" is used herein to mean "is provided in some embodiments and not provided in other embodiments". It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the present disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable combination or as suitable in any other described embodiment of the disclosure.
Claims
CLAIMS1. A method (200) of copying a production database storage (102) to a non-native backup storage (106) comprising steps of:(a) copying a filesystem layout (110) from a production database filesystem (108) of the production database storage (102) to a backup database filesystem (124) of the non-native backup storage (106);(b) copying data and metadata from the production database storage (102) to the non-native backup storage (106), and populating a backup deduplication database (128) of the non-native backup storage (106) with a plurality of key -value pairs, where each key is a combination of a file and an offset for respective data and metadata stored in the production database storage (102), and each value is a location in the non-native backup storage (106) where the respective data has been saved; and(c) copying replicated metadata from the production database storage (102) to the non-native backup storage (106) and further populating the backup deduplication database (128) of the non-native backup storage (106) with a plurality of key -value pairs, where each key is a combination of a file and an offset for respective replicated metadata, and each value is a location of in the non-native backup storage (106) where the respective metadata has been saved; and(d) further populating the backup deduplication database (128) of the non-native backup storage (106) with a plurality of key -value pairs, where each key is a combination of a file and an offset for respective replicated data, and each value is a location in the non-native backup storage (106) where the respective data has already been saved when previously copying the data; wherein the order of steps (b) and (c) in the method can be reversed.
2. The method (200) of claim 1 wherein, after the first copying step, the data, and metadata from the production database storage (102) are located using an open database format.
3. The method (200) of claim 2 wherein the open database format is used to read each file, parse the metadata, and locate the respective data.
4. The method (200) of claim 3 wherein the respective data is located using predefined offsets.
5. The method (200) of claim 1 wherein the copying steps are performed by a copy agent (104).
6. The method (200) of claim 1 wherein a saved data format differs between data and metadata.
7. The method (200) of claim 1 wherein a virtual file system (126) is used to expose the data stored on the non-native backup storage (106).
8. The method (200) of claim 1 wherein the filesystem layout (110) includes file / directory relations and file sizes.
9. A system (100) comprising means adapted for carrying out all the steps of the method (200) according to any preceding method (200) claim.
10. A computer program comprising instructions for carrying out all the steps of the method according to any preceding method (200) claim, when said computer program is executed on a computer system.
Citation Information
Patent Citations
Method and system for auto live-mounting database golden copies
US20200349017A1
Systems and methods for optimizing restoration of deduplicated data stored in cloud-based storage resources
US20210173744A1
Optimize backup from universal share
US20210303408A1
Creating file recipes for copy overwrite workloads in deduplication file systems
US20230409438A1