A method and device for processing object storage metadata

By using technical means of multiple database clusters cooperating with each other in object storage metadata management, the problems of unstable performance of stand-alone databases and complexity of database and table division are solved, and stable processing and easy expansion of 100 billion objects are achieved.

CN112115206BActive Publication Date: 2025-05-16BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910531925.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-19
Publication Date
2025-05-16
Estimated Expiration
2039-06-19

AI Technical Summary

Technical Problem

In the prior art, when storing metadata through relational database management objects, the performance of stand-alone machines is unstable, especially when the number of objects reaches 10 billion, and the database and table-dividing solution will lead to complexity in scope query and high cost.

Method used

The technical means of multiple database clusters cooperating with each other are used to store object storage metadata in the write library. When the write library is full or fails, it is downgraded to a read-only library, and the metadata in the read-only library is merged into the stable library at the set time interval, and the read-only library is cleared as a backup library to achieve expansion and fault tolerance.

Benefits of technology

It solves the problem of unstable performance of stand-alone databases, realizes stable processing of metadata when the number of objects is extremely large, and easily expands to support 100 billion objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112115206B_ABST
    Figure CN112115206B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for processing object storage metadata, and relates to the field of computer technology. A specific implementation of the method includes: storing the current object storage metadata in a write library, and when the write library is full or fails, downgrading the write library to a read-only library; merging the object storage metadata in the read-only library into a stable library at a set time interval; clearing the read-only library, and using the cleared read-only library as a backup library, and when the write library is full or fails, using the backup library as a new write library. Because this implementation adopts technical means of multiple database clusters cooperating with each other, it overcomes the technical problem of unstable performance of a single-machine database when the number of objects is large, thereby achieving the technical effect of stable processing of metadata even when the number of objects is extremely large, and at the same time can achieve the purpose of easily expanding to support hundreds of billions of objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for processing object storage metadata. Background Art

[0002] In recent years, object storage has become the most important supporting service for public clouds, supporting a variety of services such as live broadcast, on-demand, and pictures. In particular, with the development of various video services in recent years, the amount of video and picture data has increased. Object storage requires a large-scale metadata management system to manage metadata. Generally speaking, an object generally includes the object name, object data (unformatted data), and some attributes attached to the object (Meta, such as the modification time of the object). Object storage is a distributed storage service used to store objects. The name of the object, the object's attributes, and the location where the object data is stored are the object's metadata (i.e., object storage metadata). There are two main ways to access object storage metadata: one is to obtain the object's data and Meta information through the object's name, and the other is to query all objects behind a certain object. Currently, the common way is to use a relational database to manage object storage metadata. The database technology itself is very mature and major companies have rich maintenance experience. It is a very suitable metadata storage system.

[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:

[0004] 1. Object storage metadata is managed through a relational database. The storage capacity is heavily dependent on the service capabilities of a single machine. When the number of objects is large (reaching the tens of billions), the performance of a single-machine relational database is unstable.

[0005] 2. Solving the problem of unstable single-machine performance by sharding databases and tables will make range queries complicated or even infeasible. Business growth will require further splitting of the database, which is extremely costly. Summary of the invention

[0006] In view of this, an embodiment of the present invention provides a method and device for processing object storage metadata, which can solve the problem of unstable performance of a single-machine database when the number of objects is large, is easy to expand, and can be easily expanded to support hundreds of billions of objects.

[0007] To achieve the above-mentioned purpose, according to one aspect of an embodiment of the present invention, a method for processing object storage metadata is provided, comprising: storing current object storage metadata in a write library, and when the write library is full or fails, downgrading the write library to a read-only library; merging the object storage metadata in the read-only library into a stable library at a set time interval; clearing the read-only library, and using the cleared read-only library as a spare library, and when the write library is full or fails, using the spare library as a new write library.

[0008] Optionally, at a set time interval, the object storage metadata in the read-only repository is merged into the stable repository, including: at a set time interval, randomly selecting a set number of object storage metadata from the current stable repository, and calculating a split point based on the object storage metadata; generating a metadata table and a metadata range record table based on the split point; wherein the metadata range record table records the storage range of the metadata table; based on the metadata table and the metadata range record table, merging the object storage metadata in the read-only repository into the stable repository.

[0009] Optionally, based on the metadata table and the metadata range record table, the object storage metadata in the read-only library is merged into the stable library, including: traversing and reading the object storage metadata of the current stable library and the object storage metadata of the read-only library in sequence, and writing them into the metadata table; the metadata table and the metadata range record table constitute a new stable library; performing an offline comparison and verification on the object storage metadata in the current stable library and the read-only library with the object storage metadata in the new stable library; after the verification passes, modifying the configuration in the cluster configuration pointing to the current stable library to point to the new stable library, and deleting the current stable library.

[0010] Optionally, the method further includes: based on the object name, querying object storage metadata corresponding to N object names; wherein N is a positive integer; and returning a query result set according to the object storage metadata corresponding to the object name.

[0011] Optionally, querying object storage metadata corresponding to N object names includes: randomly reading N object storage metadata from a write library, a read-only library, and a stable library to obtain three read result sets; merging and traversing the three read result sets in turn according to the object names; after the merging and traversing is completed, if the amount of data in the query result set does not reach a set threshold, re-executing the process of merging and traversing the three read result sets starting from the last object name until the amount of data in the query result set reaches the set threshold.

[0012] Optionally, three read result sets are merged and traversed in turn according to the object name, including: setting priorities, and the priorities of the read result sets of the write library, the read-only library, and the stable library are decreased in turn; three read result sets are merged and traversed in turn according to the object name according to the priority: if an object name appears once in total in the three read result sets, determine whether the object storage metadata corresponding to the object name carries a deletion mark, if it does, jump out and start merging and traversing with the next object name, if not, return the object storage metadata corresponding to the object name to the query result set; if an object name appears multiple times in total in the three read result sets, return the object storage metadata corresponding to the object name read from the read result set with the highest priority to the query result set.

[0013] According to one aspect of an embodiment of the present invention, there is provided a device for processing object storage metadata, including: a downgrade processing module, used to store current object storage metadata in a write library, and when the write library is full or fails, downgrade the write library to a read-only library; a merge processing module, used to merge the object storage metadata in the read-only library into a stable library at a set time interval; a standby processing module, used to clear the read-only library and use the cleared read-only library as a standby library, and when the write library is full or fails, use the standby library as a new write library.

[0014] Optionally, the merge processing module is also used to: randomly select a set number of object storage metadata from the current stable library at a set time interval, and calculate a split point based on the object storage metadata; generate a metadata table and a metadata range record table based on the split point; wherein the metadata range record table records the storage range of the metadata table; based on the metadata table and the metadata range record table, merge the object storage metadata in the read-only library into the stable library.

[0015] Optionally, the merge processing module is also used to: traverse and read the object storage metadata of the current stable repository and the object storage metadata of the read-only repository in sequence, and write them into the metadata table; the metadata table and the metadata range record table constitute a new stable repository; perform offline comparison and verification on the object storage metadata in the current stable repository and the read-only repository with the object storage metadata in the new stable repository; after the verification is passed, modify the configuration in the cluster configuration that points to the current stable repository to point to the new stable repository, and delete the current stable repository.

[0016] Optionally, the device further comprises a list query module, which is used to: based on the object name, query the object storage metadata corresponding to N object names; wherein N is a positive integer; and return a query result set according to the object storage metadata corresponding to the object name.

[0017] Optionally, the list query module is also used to: randomly read N object storage metadata from the write library, the read-only library, and the stable library to obtain three read result sets; merge and traverse the three read result sets in turn according to the object names; after the merge and traversal is completed, if the amount of data in the query result set does not reach a set threshold, re-execute the process of merging and traversing the three read result sets starting from the last object name until the amount of data in the query result set reaches the set threshold.

[0018] Optionally, the list query module is also used to: randomly read N object storage metadata from the write library, the read-only library, and the stable library to obtain three read result sets; set priorities, and the priorities of the read result sets of the write library, the read-only library, and the stable library decrease in sequence; merge and traverse the three read result sets according to the object names in sequence according to the priorities: if an object name appears once in the three read result sets, determine whether the object storage metadata corresponding to the object name carries a deletion mark, if it does, jump out and start merging and traversing with the next object name, if not, return the object storage metadata corresponding to the object name to the query result set; if an object name appears multiple times in the three read result sets, return the object storage metadata corresponding to the object name read from the read result set with the highest priority to the query result set.

[0019] According to one aspect of an embodiment of the present invention, there is provided an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for storing metadata of a processing object as provided in the aforementioned embodiment.

[0020] According to one aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for processing object storage metadata provided in the above-mentioned embodiment is implemented.

[0021] An embodiment of the above invention has the following advantages or beneficial effects: because of the technical means of using multiple database clusters to cooperate with each other, the technical problem of unstable performance of a single-machine database when the number of objects is large is overcome, and it is not limited by the service capabilities of a single database, thereby achieving the technical effect of stably processing metadata even when the number of objects is extremely large, and at the same time, it can also achieve the purpose of easily expanding to support hundreds of billions of objects.

[0022] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with the specific implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.

[0024] Figure 1 is a schematic diagram of a basic process of a method for processing object storage metadata according to an embodiment of the present invention;

[0025] Figure 2 is a schematic diagram of basic modules of an apparatus for processing object storage metadata according to an embodiment of the present invention;

[0026] Figure 3 is a schematic diagram of a system architecture for processing object storage metadata according to an embodiment of the present invention;

[0027] Figure 4 is a schematic diagram of implementing a merging function in a system for processing object storage metadata according to an embodiment of the present invention;

[0028] Figure 5 is a schematic diagram of architecture evolution failure rollback in a system for processing object storage metadata according to an embodiment of the present invention;

[0029] Figure 6 is an exemplary system architecture diagram to which embodiments of the present invention may be applied;

[0030] Figure 7 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0032] Figure 1 FIG. 1 is a schematic diagram of the basic process of a method for processing object storage metadata according to an embodiment of the present invention. Figure 1 As shown, an embodiment of the present invention provides a method for processing object storage metadata, including:

[0033] Step S101. Store the current object storage metadata in the write library. When the write library is full or fails, the write library is downgraded to a read-only library. When the current object storage metadata is stored in the write library, the object storage metadata at this time is incremental data (readable and writable). When the write library is full or fails, the incremental data is converted to stock data, and the stock data is set to read-only. At this time, the stock data is data that will not be modified.

[0034] Step S102: Merge the object storage metadata in the read-only repository into the stable repository at a set time interval;

[0035] Step S103: Clear the read-only library, and use the cleared read-only library as a spare library. When the write library is full or fails, use the spare library as a new write library.

[0036] Because the embodiment of the present invention adopts the technical means of multiple database clusters cooperating with each other, it overcomes the technical problem of unstable performance of a single-machine database when the number of objects is large, and thus achieves the technical effect of stably processing metadata even when the number of objects is extremely large; the write point of the embodiment of the present invention can be migrated at will to solve the problem of single cluster failure, and at the same time can easily expand to support hundreds of billions of objects.

[0037] In step S102 of the embodiment of the present invention, the object storage metadata in the read-only library is merged into the stable library at a set time interval, including: randomly selecting a set number of object storage metadata from the current stable library at a set time interval, and calculating a split point based on the object storage metadata; generating a metadata table and a metadata range record table based on the split point; wherein the metadata range record table records the storage range of the metadata table; based on the metadata table and the metadata range record table, merging the object storage metadata in the read-only library into the stable library.

[0038] In the embodiment of the present invention, the data in the read-only library is periodically merged into the stable library to reduce the reading consumption. The merging process is performed on two read-only data sets (that is, data can only be read from the read-only library and the stable library, and data cannot be written to the read-only library and the stable library), which reduces the difficulty of operation in the actual application process.

[0039] Based on step S102 of the above embodiment, in the embodiment of the present invention, based on the metadata table and the metadata range record table, the object storage metadata in the read-only library is merged into the stable library, including: traversing and reading the object storage metadata of the current stable library and the object storage metadata of the read-only library in sequence, and writing them into the metadata table; the metadata table and the metadata range record table constitute a new stable library; performing offline comparison and verification on the object storage metadata in the current stable library and the read-only library with the object storage metadata in the new stable library; after the verification passes, modifying the configuration pointing to the current stable library in the cluster configuration to point to the new stable library, and deleting the current stable library.

[0040] In the embodiment of the present invention, the data in the read-only database is periodically merged into the stable database to reduce the consumption of reading. The merging process is performed on two read-only data sets, which reduces the difficulty of operation in the actual application process.

[0041] The method in the embodiment of the present invention further includes: based on the object name, querying the object storage metadata corresponding to N object names; wherein N is a positive integer; and returning a query result set according to the object storage metadata corresponding to the object name. The embodiment of the present invention has a list query function, for example, querying the metadata corresponding to 1000 keys starting from a certain object name (Key).

[0042] The query of object storage metadata corresponding to N object names in the embodiment of the present invention includes: randomly reading N object storage metadata from a write library, a read-only library, and a stable library to obtain three read result sets; merging and traversing the three read result sets in turn according to the object names; after the merging and traversing is completed, if the amount of data in the query result set does not reach a set threshold, re-execute the process of merging and traversing the three read result sets starting from the last object name until the amount of data in the query result set reaches the set threshold.

[0043] For example, to query 1000 keys starting from a certain key, these 1000 keys may appear in the write library, read-only library and stable library. 1000 metadata are taken from each of these three databases, and then the returned data is determined according to the above business rules. This part may only have 500 items, so there will be a cycle to process until 1000 results are obtained. The embodiment of the present invention can solve the problem that the query becomes complicated due to the use of sub-library and sub-table, and the query process is efficient and the returned results are accurate.

[0044] Based on the above embodiments, the embodiments of the present invention sequentially merge and traverse three read result sets according to the object name, including: setting priorities, the priorities of the read result sets of the write library, the read-only library, and the stable library are successively reduced; sequentially merge and traverse the three read result sets according to the object name according to the priority: if an object name appears once in the three read result sets, then determine whether the object storage metadata corresponding to the object name carries a deletion mark, if it does, jump out, start merging and traversing with the next object name, if it does not carry, return the object storage metadata corresponding to the object name to the query result set; if an object name appears multiple times in the three read result sets, then read the object storage metadata corresponding to the object name from the read result set with the highest priority, and return it to the query result set. The embodiment of the present invention uses the method of setting priorities to query and filter in the three libraries, which can solve the problem of making the query complicated due to the use of sub-libraries and sub-tables, and the query process is efficient and the returned results are accurate.

[0045] Figure 2 Schematic diagram of basic modules of an apparatus for processing object storage metadata according to an embodiment of the present invention. Figure 2 As shown, an embodiment of the present invention provides a device 200 for processing object storage metadata, including:

[0046] The downgrade processing module 201 is used to: store the current object storage metadata in a write library, and when the write library is full or fails, downgrade the write library to a read-only library;

[0047] The merge processing module 202 is used to: merge the object storage metadata in the read-only library into the stable library according to a set time interval;

[0048] The standby processing module 203 is used to: clear the read-only library and use the cleared read-only library as a standby library; when the write library is full or fails, use the standby library as a new write library.

[0049] Because the embodiment of the present invention adopts the technical means of multiple database clusters cooperating with each other, it overcomes the technical problem of unstable performance of a single-machine database when the number of objects is large, thereby achieving the technical effect of stably processing metadata even when the number of objects is extremely large, and at the same time can easily expand to support hundreds of billions of objects.

[0050] In an embodiment of the present invention, the merge processing module 202 is further used to: randomly select a set number of object storage metadata from the current stable library at a set time interval, and calculate a split point based on the object storage metadata; generate a metadata table and a metadata range record table based on the split point; wherein the metadata range record table records the storage range of the metadata table; based on the metadata table and the metadata range record table, merge the object storage metadata in the read-only library into the stable library.

[0051] In the embodiment of the present invention, the data in the read-only library is periodically merged into the stable library to reduce the reading consumption. The merging process is performed on two read-only data sets (that is, data can only be read from the read-only library and the stable library, and data cannot be written to the read-only library and the stable library), which reduces the difficulty of operation in the actual application process.

[0052] In an embodiment of the present invention, the merge processing module is further used to: traverse and read the object storage metadata of the current stable repository and the object storage metadata of the read-only repository in sequence, and write them into the metadata table; the metadata table and the metadata range record table constitute a new stable repository; perform offline comparison and verification on the object storage metadata in the current stable repository and the read-only repository with the object storage metadata in the new stable repository; after the verification is passed, modify the configuration in the cluster configuration that points to the current stable repository to point to the new stable repository, and delete the current stable repository.

[0053] In the embodiment of the present invention, the data in the read-only database is periodically merged into the stable database to reduce the consumption of reading. The merging process is performed on two read-only data sets, which reduces the difficulty of operation in the actual application process.

[0054] In an embodiment of the present invention, the device further comprises a list query module, which is used to: query the object storage metadata corresponding to N object names based on the object name; wherein N is a positive integer; and return a query result set according to the object storage metadata corresponding to the object name. The embodiment of the present invention has a list query (List) function, for example, querying the metadata corresponding to 1000 keys starting from a certain object name (Key).

[0055] In an embodiment of the present invention, the list query module is also used to: randomly read N pieces of object storage metadata from the write library, the read-only library, and the stable library to obtain three read result sets; merge and traverse the three read result sets in turn according to the object name; after the merge and traverse is completed, if the amount of data in the query result set does not reach the set threshold, re-execute the process of merging and traversing the three read result sets starting from the last object name until the amount of data in the query result set reaches the set threshold. The embodiment of the present invention can solve the problem that the query becomes complicated due to the use of sub-libraries and sub-tables, and the query process is efficient and the returned results are accurate.

[0056] In an embodiment of the present invention, the list query module is also used to: randomly read N object storage metadata from the write library, the read-only library, and the stable library to obtain three read result sets; set the priority, and the priority of the read result sets of the write library, the read-only library, and the stable library is reduced in sequence; merge and traverse the three read result sets in sequence according to the object name according to the priority: if an object name appears once in the three read result sets, then determine whether the object storage metadata corresponding to the object name carries a deletion mark, if it does, jump out, start merging and traversing with the next object name, if it does not carry, return the object storage metadata corresponding to the object name to the query result set; if an object name appears multiple times in the three read result sets, then read the object storage metadata corresponding to the object name from the read result set with the highest priority and return it to the query result set. The embodiment of the present invention uses the method of setting priorities to query and filter in the three libraries, which can solve the problem that the query becomes complicated due to the use of sub-libraries and sub-tables, and the query process is efficient and the returned results are accurate.

[0057] Figure 3 Schematic diagram of the system architecture for managing object storage metadata according to an embodiment of the present invention. Figure 3As shown, the system for managing object storage metadata includes: a stable library, a read-only library, a write library, a spare library, a monitoring service, a merge service and a configuration center; the stable library consists of multiple database tables, which are used to: store historical stock data, the stable library only supports reading data, and the stable library also includes a metadata table and a metadata range record table; the write library consists of multiple database tables, which are used to: store the data currently being read and written, and can read data from the write library or write data to the write library, and the write library also includes a metadata table and a metadata range record table; the read-only library is switched from the write library when the write library fails or is full, and only supports reading data; the spare library is an empty database table cluster, which is used to switch to a new write library when the write library fails or is full; the monitoring service is used to: monitor whether the write library fails or is full; the configuration center is used to: provide consistent cluster configuration information to the database table; the merge service is used to: merge the read-only library into the stable library at a set time interval, and clear the read-only library after the merge is completed.

[0058] The basic operations of the system that manages object storage metadata are as follows:

[0059] 1. Write (Put), you can directly write to the writeable library (Writeable Meta).

[0060] 2. Delete: write data with a delete mark in the write library.

[0061] 3. Read (Get), read the data corresponding to the object name from the writeable library (Writeable Meta), read-only library (ReadOnly Meta), and stable library (Stable Meta) in turn, and make the following judgments:

[0062] If data without a delete mark is read, the loop is exited and the data is returned;

[0063] When data with a delete mark is read, the loop is exited and the object does not exist;

[0064] If no data is read, the returned object does not exist.

[0065] 4. List query (List)

[0066] Take out N data from each of the writeable meta, read-only meta, and stable meta.

[0067] Set the priority, the priority of write library, read-only library, and stable library decreases in sequence;

[0068] Traverse the three result sets above according to the priority and merge them as follows until one result set has been taken:

[0069] If a key appears only once, if the object storage metadata (i.e. metadata) corresponding to the key carries a deletion mark, it will jump out directly. If it does not carry a deletion mark, it will return to the metadata;

[0070] If a key appears multiple times, the metadata with the highest priority will prevail.

[0071] After the merging is completed, if the number of results is insufficient, the last key is used as the starting mark and the merging traversal process is re-executed. Figure 4 FIG. 1 is a schematic diagram of implementing the merging function in a system for processing object storage metadata according to an embodiment of the present invention. Figure 4 As shown in the figure, the data in the read-only database will be regularly merged into the stable database to reduce the consumption of reading. In this system, the merging process is performed on two read-only data sets. The process is as follows:

[0072] Determine whether the stable database needs to be split (when the number of objects is too large or the current system service capacity is insufficient, the stable database needs to be split). If splitting is required, randomly select some data and calculate the split point; based on the split point, generate a new metadata table (MetaN table) and a metadata range record table (Table Meta). Table Meta records the range of MetaN, forming a tree structure as a whole; 1. Traverse and read the data of the stable database and write it to the new MetaN; 2. Traverse and read the data of ReadOnly Meta and write it to MetaN (and delete the data with deletion mark); offline compare the object storage metadata in the current stable database, the read-only database participating in the merger, and the object storage metadata in the new stable database to ensure the integration of the logic before and after the merger; 3. Modify the cluster configuration, and modify the configuration pointing to the current stable database in the cluster configuration to point to the new stable database; delete the current stable database to complete the merger.

[0073] In actual operation, various new storage systems will continue to emerge. Based on the needs of architecture evolution, we will try to introduce new systems in the production environment. Generally, in this case, double writing will be required, but the verification of double-written data and rollback in case of failure will be a more troublesome task. This system provides a mechanism for double-write verification and rollback of architecture evolution. Figure 5 FIG. 1 is a schematic diagram of architecture evolution failure rollback in a system for processing object storage metadata according to an embodiment of the present invention. Figure 5As shown, under normal circumstances, new writes are made to the dual-write new storage device (New Store) and database (Mysql), with the new storage device being the main one. The write library is regularly changed to a read-only library, and the consistency of the data in the New Store and Mysql is compared. Once any problem occurs, the backup library is directly enabled, and all rollbacks will be made to the original database service status.

[0074] This system implements a highly available, large-scale, globally ordered KV cluster based on multiple Mysql clusters. Based on the idea of ​​LSM Tree, a large globally ordered KV is created on top of multiple databases. The write points of multiple clusters can be migrated arbitrarily, solving the problem of Mysql single point failure. The tree structure of the logic inside the logical cluster solves the problem of Mysql reading and writing. Dynamic data comparison is converted into static data comparison. The HBase, Tikv and other systems in the existing technology implement globally ordered KV, and the scale can be greatly expanded to meet the needs of large-scale global order. However, the use of these clusters will encounter the problem that a single cluster failure will cause the entire metadata management service to be unwritable. The write point of this system can be migrated at will to solve the problem of single cluster failure. At the same time, this system can also keep the basic architecture unchanged, and only use HBase, Tikv and other systems to replace Mysql to achieve higher availability.

[0075] The cluster data scale of this system is not limited to the service capacity of a single database, and can reach hundreds of billions of objects. Specifically:

[0076] The processing capability can be greatly expanded: the existing data is read-only data, and the reading capability can be improved by simply increasing the backup of the database; it has the ability to rebuild the existing data, and each access only requires access to one database, and the capacity can be increased by infinitely expanding the number of databases without causing a decrease in processing performance; the incremental data can be quickly made read-only, and only a very small incremental data cluster needs to be maintained to greatly reduce the access to incremental data.

[0077] Improved reliability: The existing data is read-only, and the availability of the existing data can be improved by simply increasing the number of slave libraries. When the incremental data cluster (write library) is unavailable, it can be quickly switched to the backup cluster (read-only library or standby library) without affecting the system's write.

[0078] With support for architectural evolution, the current database can be switched to a new system (or new storage device). Since the existing data is read-only, it is very simple to import the existing data into the new system, and when problems occur in the new system, it can be switched back to the original database (or current system) without risk. The incremental data cluster can be switched quickly, and the new system can be used as a writable cluster (write library). Only a small amount of data is stored in the read-only cluster (read-only library). In the event of a failure, the writable cluster can be quickly switched to the backup cluster of the old system.

[0079] Figure 6 An exemplary system architecture 600 is shown to which the method for processing object storage metadata or the apparatus for processing object storage metadata according to the embodiments of the present invention can be applied.

[0080] like Figure 6 As shown, system architecture 600 may include terminal devices 601, 602, 603, network 604 and server 605. Network 604 is used to provide a medium for communication links between terminal devices 601, 602, 603 and server 605. Network 604 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0081] The user can use the terminal devices 601, 602, 603 to interact with the server 605 through the network 604 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 601, 602, 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0082] The terminal devices 601 , 602 , and 603 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers, etc.

[0083] Server 605 may be a server that provides various services, such as a backend management server that provides support for shopping websites browsed by users using terminal devices 601, 602, and 603. The backend management server may analyze and process received data such as product information query requests, and feed back the processing results to the terminal device.

[0084] It should be noted that the method for processing object storage metadata provided in the embodiment of the present invention is generally executed by the server 605 , and accordingly, the device for processing object storage metadata is generally set in the server 605 .

[0085] It should be understood that Figure 6 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0086] According to an embodiment of the present invention, an electronic device and a readable storage medium are also provided.

[0087] The electronic device of an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method for storing metadata of a processing object provided by the present invention.

[0088] The computer-readable medium of the embodiment of the present invention stores a computer program, and when the program is executed by a processor, the method for processing object storage metadata provided by the present invention is implemented.

[0089] Reference below Figure 7 , which shows a schematic diagram of the structure of a computer system 700 of a terminal device suitable for implementing an embodiment of the present invention. Figure 7 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0090] like Figure 7 As shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage part 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0091] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that a computer program read therefrom is installed into the storage section 708 as needed.

[0092] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above-mentioned functions defined in the system of the present invention are executed.

[0093] It should be noted that the computer-readable medium shown in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0094] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0095] The modules involved in the embodiments of the present invention may be implemented in software or hardware. The modules described may also be set in a processor, for example, may be described as: a processor, including: a degradation processing module, a merging processing module, and a standby processing module. The names of these modules do not, in some cases, constitute limitations on the modules themselves. For example, the merging processing module may also be described as "a module for merging the object storage metadata in the read-only library into the stable library".

[0096] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiment; or may exist independently without being assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes: storing the current object storage metadata in a write library, and when the write library is full or fails, downgrading the write library to a read-only library; merging the object storage metadata in the read-only library into a stable library at a set time interval; clearing the read-only library, and using the cleared read-only library as a backup library, and when the write library is full or fails, using the backup library as a new write library.

[0097] It can be seen from the method for processing object storage metadata according to an embodiment of the present invention that, because a plurality of database clusters cooperate with each other, the technical problem of unstable performance of a single-machine database when the number of objects is large is overcome, thereby achieving the technical effect of stably processing metadata even when the number of objects is extremely large, and at the same time, achieving the purpose of easily expanding to support hundreds of billions of objects.

[0098] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for processing object storage metadata, characterized in that: include: The current object storage metadata is stored in a write library, and when the write library is full or fails, the write library is downgraded to a read-only library; At a set time interval, randomly select a set number of object storage metadata from the current stable library, and calculate the split point based on the object storage metadata; Generate a metadata table and a metadata range record table according to the split point; wherein the metadata range record table records the storage range of the metadata table; Based on the metadata table and the metadata range record table, merging the object storage metadata in the read-only library into the stable library to reduce reading consumption; The read-only library is emptied and used as a spare library. When the write library is full or fails, the spare library is used as a new write library.

2. The method according to claim 1, characterized in that Based on the metadata table and the metadata range record table, merging the object storage metadata in the read-only library into the stable library includes: The object storage metadata of the current stable repository and the object storage metadata of the read-only repository are traversed and read in sequence, and written into the metadata table; The metadata table and the metadata range record table constitute a new stable repository; Performing offline comparison and verification on the object storage metadata in the current stable repository and the read-only repository and the object storage metadata in the new stable repository; After the verification passes, modify the configuration in the cluster configuration that points to the current stable repository to point to the new stable repository, and delete the current stable repository.

3. The method according to claim 1, characterized in that The method further comprises: Based on the object name, query the object storage metadata corresponding to N object names; where N is a positive integer; Return a query result set based on the object storage metadata corresponding to the object name.

4. The method according to claim 3, characterized in that Query the object storage metadata corresponding to N object names, including: Randomly read N pieces of object storage metadata from the write library, read-only library, and stable library, and obtain three read result sets; Merge and traverse the three read result sets in sequence according to the object names; After the merge traversal is completed, if the amount of data in the query result set does not reach the set threshold, the process of merge traversal of the three read result sets is re-executed starting from the last object name until the amount of data in the query result set reaches the set threshold.

5. The method according to claim 4, characterized in that According to the object name, three read result sets are merged and traversed in turn, including: Set the priority, the priority of the read result set of the write library, read-only library, and stable library decreases in sequence; According to the object names, the three read result sets are merged and traversed in order of priority: If an object name appears once in three read result sets, determine whether the object storage metadata corresponding to the object name carries a deletion mark. If so, jump out and start merging and traversing with the next object name. If not, return the object storage metadata corresponding to the object name to the query result set. If an object name appears multiple times in total in three read result sets, the object storage metadata corresponding to the object name read from the read result set with the highest priority will be returned to the query result set.

6. A device for processing object storage metadata, characterized in that: include: The downgrade processing module is used to: store the current object storage metadata in a write library, and when the write library is full or fails, downgrade the write library to a read-only library; The merging processing module is used to: randomly select a set number of object storage metadata from the current stable library at a set time interval, and calculate a split point based on the object storage metadata; generate a metadata table and a metadata range record table based on the split point; wherein the metadata range record table records the storage range of the metadata table; based on the metadata table and the metadata range record table, merge the object storage metadata in the read-only library into the stable library to reduce the consumption of reading; The standby processing module is used to: clear the read-only library and use the cleared read-only library as a standby library; when the write library is full or fails, use the standby library as a new write library.

7. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Database management method and database system

    CN108319602A

  • Splitting and moving ranges in a distributed system

    CN109074362A