Data system, data management method and device, electronic device and storage medium
By constructing metadata with a multi-level tree structure in the database, combining the baseline data of the object storage layer and the incremental data of the database, the problem of low performance of the object storage data management system in the existing technology is solved, data sharing and efficient management are realized, and system performance is improved.
Patent Information
- Application Number
- CN202510052100.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The data management system used for object storage in the prior art has low performance and cannot effectively meet the needs of high scalability, high durability and high availability.
A data system is designed, including an object storage layer and a database connected to it. By constructing metadata of a multi-level tree structure in the database, storing the baseline data of the object storage layer and the incremental data of the database, data sharing and efficient management are achieved.
By reducing the use of local storage space in the database, data sharing between different databases is realized, the IO load during metadata management and maintenance and database restart is reduced, and the performance of the data system is improved.
Smart Images

Figure CN119474107B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of database technology, and in particular, to a data system, a data management method and device, an electronic device, and a storage medium. Background Art
[0002] Object storage has the characteristics of low cost (unit price is a fraction of cloud disk), payment based on actual usage (capacity and read and write request volume), high scalability, high durability, high availability, and shareability. Using object storage as shared storage can significantly reduce storage costs.
[0003] However, in the related art, the performance of the data system used for data management of object storage is relatively low. Summary of the invention
[0004] In view of this, one or more embodiments of the present specification provide a data system, a data management method and device, an electronic device, and a storage medium.
[0005] To achieve the above objectives, one or more embodiments of this specification provide the following technical solutions:
[0006] According to a first aspect of one or more embodiments of this specification, a data system is provided, the data system comprising:
[0007] The object storage layer stores the baseline data of the managed data;
[0008] At least one database is communicatively connected to the object storage layer, storing incremental data of the managed data and metadata of the managed data, wherein the metadata is in a multi-level tree structure with database metadata as a root node and data block metadata as leaf nodes;
[0009] Among them, the subordinate components contained in the upper-level component of the managed data form child nodes under the node corresponding to the upper-level component in the metadata; in the metadata, the nodes in any level except the leaf nodes are used to record the metadata of the corresponding components and the object address of the child nodes, and the nodes in the leaf nodes are used to record the metadata of the corresponding data blocks.
[0010] In a possible embodiment of the present specification, from the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata.
[0011] In a possible embodiment of the present specification, the data system includes different databases of the same tenant for reaching consensus on the tenant status in the metadata of the same tenant;
[0012] Different databases containing the same log stream in the data system are used to reach consensus on metadata of the same log stream;
[0013] The data system includes different databases of the same partition for reaching consensus on metadata of the same partition.
[0014] In a possible embodiment of the present specification, the data system includes forming a consensus protocol group between the same log streams in different databases, and the consensus protocol group is used to reach consensus on the metadata of the same tenant, the metadata of the same log stream, and the metadata of the same partition.
[0015] In a possible embodiment of the present specification, the partition metadata includes metadata of a baseline SSTable and metadata of an incremental SSTable, the data blocks corresponding to the data block metadata under the metadata of the baseline SSTable are data blocks in the object storage layer, and the data blocks corresponding to the data block metadata under the metadata of the incremental SSTable are data blocks local to the database.
[0016] In a possible embodiment of the present specification, the data system includes different databases in the same partition for periodically performing the following steps:
[0017] A consensus is reached on the incremental SSTable of the same partition;
[0018] Merge the baseline SSTable and incremental SSTable of the same partition to update the baseline SSTable;
[0019] The metadata corresponding to the incremental SSTable merged into the baseline SSTable is deleted from the metadata, and the metadata corresponding to the baseline SSTable and the metadata corresponding to the partition are updated.
[0020] In a possible embodiment of the present specification, the database is used to create a new object and write the updated metadata to be updated when updating the metadata to be updated in the partition metadata level within the metadata;
[0021] The database is used for updating the metadata to be updated in the object where the metadata to be updated is located when updating the metadata to be updated at any level except the partition metadata level in the metadata.
[0022] In a possible embodiment of the present specification, the database is used to load the metadata of the internal hierarchical levels of the metadata into the memory starting from the root node of the metadata when restarting, so that the database is partially restored to the state before shutdown and obtains data reading capability.
[0023] In a possible embodiment of the present specification, the database is used to load metadata of three levels, namely, database metadata, tenant metadata and log stream metadata, within the metadata into memory starting from the root node of the metadata when restarted, wherein the log stream metadata is used to record the metadata of the corresponding log stream and the object address of each partition metadata under the log stream metadata.
[0024] In a possible embodiment of the present specification, the database is used to load the metadata into the memory when restarting, so that the database is restored to the state before shutdown.
[0025] According to a second aspect of one or more embodiments of the present specification, a data management method is provided, which is applied to any database of a data system, wherein the database system includes an object storage layer and at least one database, wherein the object storage layer stores baseline data of managed data, and the database stores incremental data of the managed data;
[0026] The method comprises:
[0027] Based on the baseline data and the incremental data, metadata of the managed data is constructed in the database, and the metadata is in a multi-level tree structure with database metadata as the root node and data block metadata as the leaf nodes, wherein the subordinate components contained in the upper-level components of the managed data form child nodes under the nodes corresponding to the upper-level components in the metadata; in the metadata, nodes in any level except leaf nodes are used to record the meta-information of the corresponding components and the object address of the child nodes, and the nodes in the leaf nodes are used to record the meta-information of the corresponding data blocks.
[0028] According to a third aspect of one or more embodiments of the present specification, a data management device is provided, which is applied to any database of a data system, wherein the database system includes an object storage layer and at least one database, wherein the object storage layer stores baseline data of managed data, and the database stores incremental data of the managed data;
[0029] The device comprises:
[0030] A construction module is used to construct metadata of the managed data in the database based on the baseline data and the incremental data, wherein the metadata is in a multi-level tree structure with database metadata as the root node and data block metadata as the leaf nodes, wherein the subordinate components contained in the upper-level components of the managed data form child nodes under the nodes corresponding to the upper-level components in the metadata; in the metadata, nodes in any level except leaf nodes are used to record the metadata of the corresponding components and the object addresses of the child nodes, and the nodes in the leaf nodes are used to record the metadata of the corresponding data blocks.
[0031] According to a fourth aspect of one or more embodiments of this specification, a computer program product is proposed, comprising a computer program / instruction, which implements the steps of the method described in the second aspect when executed by a processor.
[0032] According to a fifth aspect of one or more embodiments of this specification, an electronic device is provided, including:
[0033] processor;
[0034] a memory for storing processor-executable instructions;
[0035] The processor implements the method as described in the second aspect by running the executable instructions.
[0036] According to a sixth aspect of one or more embodiments of the present specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in the second aspect are implemented.
[0037] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:
[0038] The data system provided in the embodiments of the present specification includes an object storage layer and at least one database that is communicatively connected to the object storage layer. The object storage layer stores baseline data of data to be managed, and the database locally stores only incremental data of the data to be managed. Since the space occupied by incremental data is much smaller than the space occupied by baseline data, the database locally does not need to provide too much storage space, and data sharing is achieved between different databases through the baseline data on the object storage layer; the metadata of the multi-level tree structure stored in the database can facilitate the database to access the data to be managed through the metadata, and the metadata of the multi-level tree structure is adapted to the characteristics of object storage, and there is no need to continuously append write operations to the pre-write log file according to the WAL mechanism, thereby reducing metadata management and maintenance, and the IO load during the database restart process, and improving the performance of the data system based on object storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a structural diagram of a data system provided by an exemplary embodiment.
[0040] Figure 2 It is a data architecture of data to be managed provided by an exemplary embodiment.
[0041] Figure 3 is a data architecture of metadata provided by an exemplary embodiment.
[0042] Figure 4 is a flow chart of a data management method provided by an exemplary embodiment.
[0043] Figure 5 It is a structural schematic diagram of a device provided by an exemplary embodiment.
[0044] Figure 6 is a block diagram of a data management device provided by an exemplary embodiment. DETAILED DESCRIPTION
[0045] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0046] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0047] First, some concepts involved in this specification are explained.
[0048] WAL: Write Ahead Log, a solution that records logs before writing data, used for state recovery and synchronization between replicas.
[0049] Redo_log: An implementation of the WAL mechanism that persists database status information through log records and is used to restore the database status when the system is shut down and restarted.
[0050] LSM-Tree: An efficient storage structure widely used in database systems that aims to optimize write performance and disk space utilization. It reduces the overhead of random writes by caching write operations in memory and periodically merging data to disk.
[0051] Object storage: Object storage is a data management architecture that treats data as objects rather than files or blocks to efficiently and flexibly store and retrieve massive amounts of unstructured data. Its "everything is an object" concept emphasizes that all data, whether it is documents, pictures, videos or logs, can be processed as objects, which helps improve access efficiency and scalability.
[0052] In the related art, the storage method of hard disk plus file system generally adopts WAL mechanism for metadata management, state recovery when the system restarts, etc., but WAL mechanism is not suitable for the storage method of object storage. Specifically, after the WAL mechanism is applied to object storage, a large amount of IO load will be generated during metadata maintenance and system restart, which will seriously affect the operation efficiency of the database and reduce the performance of the database. For example, in the WAL mechanism, when the database is restarted, the metadata change process files need to be read out in sequence, so as to play back in sequence to the state before the database is shut down, which will generate a large amount of IO load, which is not conducive to the restart rate of the object storage database.
[0053] Based on the above technical problems, at least one embodiment of this specification provides a data system, which can be implemented based on the LSM-Tree storage engine and is in the form of a distributed cluster, that is, it includes multiple databases.
[0054] Please refer to the attached Figure 1 , the data system includes an object storage layer and at least one database in communication connection with the object storage layer. The object storage layer stores baseline data of the managed data. The database stores incremental data of the managed data and metadata of the managed data, and the metadata is in a multi-level tree structure with database metadata as the root node and data block metadata as the leaf node; wherein the lower-level components contained in the upper-level components of the managed data form child nodes under the nodes corresponding to the upper-level components in the metadata; in the metadata, the nodes in any level except the leaf nodes are used to record the meta-information of the corresponding components and the object address of the child nodes, and the nodes in the leaf nodes are used to record the meta-information of the corresponding data blocks.
[0055] Each database exists as an instance or copy of the data system. The data system forms a storage-computing separation architecture based on object storage, and object storage itself has the characteristics of high reliability. Therefore, the baseline data stored in the object storage layer of the system not only avoids the occupation of local storage space of the database, but also realizes data sharing between different databases, and can also ensure the security and reliability of data. In addition, the data system realizes the effective management and maintenance of metadata of the data to be managed that is scattered in the object storage layer and locally through metadata in a multi-level tree structure.
[0056] Among them, in the metadata of the multi-level tree structure, the higher the level (that is, the level closer to the root node), the less metadata there is, and the updates are less frequent; the lower the level (that is, the level closer to the leaf node), the more metadata there is, and the updates are more frequent.
[0057] For example, the database is used to load the metadata into the memory when restarting, so that the database can be restored to the state before shutdown. The metadata management method in the related art, such as the WAL mechanism, will record each change of the metadata. When the database is restarted, the metadata change process file needs to be read out in sequence, so as to play back in order to the state before the database was shut down. Since the data system records the latest state of the metadata, when the database is restarted, it only needs to read out the metadata step by step to restore the state of the database before shutdown, avoiding the large amount of IO load caused by reading out the historical state logs one by one, and improving the efficiency of database restart and operation.
[0058] As a software system that stores data in a centralized manner, the database needs to be able to restart and support features such as disaster recovery and downtime recovery. To restore the state before the database was shut down, it is necessary to rely on metadata to restore the state. The following is a detailed introduction to the database restart process using an optional example.
[0059] When the database is restarted, the metadata can be pulled up step by step in the following way to restore the state before the database was shut down:
[0060] First, the database loads the database metadata from the fixed path of the database metadata into the memory, and obtains the object address of the metadata of each tenant contained in the database after operations such as deserialization.
[0061] Next, according to the object address of each tenant's metadata, the metadata of each tenant is loaded into the memory to restore the state of each tenant, and then the object address of the metadata of the log stream contained in the metadata of each tenant is read through the metadata of each tenant.
[0062] Next, for each tenant, the metadata of each log stream is loaded into the memory according to the object address of the metadata of each log stream it contains, so as to restore the state of the log stream and the consensus protocol group formed by the log stream, and then read the object address of the metadata of the partition contained in it through the metadata of each log stream.
[0063] Next, for each log stream, the metadata of each partition is loaded into the memory according to the object address of the metadata of each partition it contains to restore the state of the partition, and then the object address of the metadata of the SSTable contained in it is read through the metadata of each partition.
[0064] Finally, for each SSTable, the metadata of each data block contained in it is loaded into the memory according to the object address of the metadata of each data block to restore the address of the data block. At this point, the metadata and component status of all levels in the database have been restored, and all persistent data has been restored, that is, the database has been restored to the state before shutdown, and the restart is completed.
[0065] In this example, the hierarchical metadata management solution effectively reduces metadata IO during the restart process and reduces the load of IO operations in the database system. Furthermore, this metadata management method records the mapping relationship and final status between metadata. When restarting, you only need to read out level by level to restore the database status before shutdown, without having to read out the historical status logs one by one.
[0066] The data system provided in the embodiments of the present specification includes an object storage layer and at least one database in communication with the object storage layer. The object storage layer stores baseline data of data to be managed, and the database locally stores only incremental data of the data to be managed. Since the space occupied by incremental data is much smaller than the space occupied by baseline data, the database locally does not need to provide too much storage space, and data sharing is achieved between different databases through the baseline data on the object storage layer; the metadata of the multi-level tree structure stored in the database can facilitate the database to access the data to be managed through the metadata, and the metadata of the multi-level tree structure is adapted to the characteristics of object storage, and there is no need to continuously append write operations to the pre-write log file according to the WAL mechanism, thereby reducing the metadata management and maintenance, and the IO load during the database restart process, and improving the performance of the data system based on object storage.
[0067] The data system provided in this specification can fully utilize the advantages of the storage characteristics of object storage and avoid the disadvantages of object storage compared to traditional file systems by modifying the metadata of different components in situ through hierarchical metadata management. Moreover, from a global perspective, the hierarchical metadata management method fully utilizes the storage characteristics of object storage and avoids its disadvantages. It reduces the IO load of the database system during operation and restart recovery, and improves the overall operation efficiency of the database.
[0068] In one embodiment of the present disclosure, please refer to the attached Figure 2 The components of the managed data are database, tenant, log stream, partition, SSTable and data block. Figure 3 , from the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata.
[0069] Among them, the database metadata (Server Super Block) is used to maintain metadata related to the current database instance. It is the top-level metadata, that is, the root node of the tree structure, which records the database status and the object address of the next-level metadata, that is, the tenant metadata. It should be understood that the database contains data of multiple tenants, so under the database metadata is the metadata of the multiple tenants, and the database metadata and the metadata of the multiple tenants under it form a parent-child node relationship.
[0070] Among them, the tenant metadata (Tenant Super Block) is used to maintain metadata related to a tenant, recording the status of the corresponding tenant and the object address of the next-level metadata, namely the log stream metadata. It should be understood that the tenant's data contains multiple partitions, and these partitions are aggregated into at least one log stream. Therefore, under the tenant metadata is the metadata of the at least one log stream, and the tenant metadata and the log stream metadata under it form a parent-child node relationship.
[0071] Among them, the log stream data (Replication Group Meta) is used to maintain metadata related to the log stream, recording the state of the corresponding log stream and the object address of the next-level metadata, namely the partition metadata. As mentioned above, the log stream contains multiple partitions, so the metadata of the multiple partitions is under the log stream metadata, and the log stream metadata and the partition metadata under it form a parent-child node relationship.
[0072] Among them, the partition data (Partition Meta) is used to maintain the metadata related to the partition, recording the status of the corresponding partition and the object address of the next-level metadata, namely the SSTable metadata. It should be understood that the partition data is divided into multiple SSTables, such as a baseline SSTable that has completed the compaction and at least one incremental SSTable that has not been merged into the baseline SSTable. Therefore, under the partition metadata are the metadata of the multiple SSTables, and the partition metadata and the SSTable metadata under it form a parent-child node relationship.
[0073] Among them, SSTable metadata (SStable Meta) is used to maintain SSTable-related metadata, recording the state of the corresponding SSTable and the object address of the next-level metadata, namely, Block metadata. It should be understood that the data in SSTable is stored in multiple data blocks Block in the object storage layer or the local physical space of the database. Therefore, under the metadata of SSTable is the metadata of the multiple data blocks, and the metadata of SSTable and the metadata of the data blocks under it form a parent-child node relationship.
[0074] Among them, the block metadata (Block Object Meta) is used to maintain the metadata related to the data block, recording the status and object address of the corresponding data block. The block metadata is the lowest level of metadata, that is, the leaf node of the tree structure.
[0075] In this embodiment, the database metadata is stored as an object in a fixed path, and the database can read the database metadata from the fixed path, and then find the metadata of any data block based on the above tree structure to find any data block.
[0076] Optionally, based on the above embodiment, the data system contains different databases of the same tenant for reaching consensus on the tenant status in the metadata of the same tenant; the data system contains different databases of the same log stream for reaching consensus on the metadata of the same log stream; the data system contains different databases of the same partition for reaching consensus on the metadata of the same partition.
[0077] For example, the data system includes forming a consensus protocol group between the same log streams in different databases, and the consensus protocol group is used to complete the above-mentioned consensus operation in this example, that is, to reach consensus on the metadata of the same tenant, the metadata of the same log stream, and the metadata of the same partition.
[0078] It should be understood that databases are independent of each other and their metadata do not need to be synchronized. For different databases, the mapping relationship between tenants and log streams may be different, so this part of metadata does not need to be synchronized, that is, only the status needs to be synchronized between tenants. SSTables and data blocks are determined by the specific physical configuration and do not need to be synchronized between databases.
[0079] In this optional example, tenant status, log stream metadata, and partition metadata in tenant metadata are synchronized between different databases, thereby achieving data consistency between different replicas. For example, database 1 contains log stream 1 and log stream 2 of tenant 1, database 2 contains log stream 2 and log stream 3 of tenant 1, database 3 contains log stream 1 and log stream 3 of tenant 1, log stream 1 contains partition 1, partition 2, and partition 3, log stream 2 contains partition 4, partition 5, and partition 6, and log stream 3 contains partition 7, partition 8, and partition 9. In this example, the state of tenant 1 in the metadata stored on database 1, database 2, and database 3 can be made consistent, the metadata of log stream 2 in the metadata stored on database 1 and database 2 can be made consistent, the metadata of log stream 1 in the metadata stored on database 1 and database 3 can be made consistent, the metadata of log stream 3 in the metadata stored on database 2 and database 3 can be made consistent, the metadata of partition 4, partition 5, and partition 6 in the metadata stored on database 1 and database 2 can be made consistent, the metadata of partition 1, partition 2, and partition 3 in the metadata stored on database 1 and database 3 can be made consistent, and the metadata of partition 7, partition 8, and partition 9 in the metadata stored on database 2 and database 3 can be made consistent.
[0080] Optionally, the database is used to synchronously update the metadata when the managed data is updated.
[0081] For example, based on the above embodiment, the database creates a new non-partitioned table under a log stream of a tenant. Then, partition metadata can be created for the non-partitioned table, and the object address of the partition metadata can be recorded in the metadata of the log stream, and the relevant parts of the metadata of the log stream, such as the number of partitions, can be updated.
[0082] For another example, based on the above embodiment, the database updates the data in a certain partition and updates the related metadata.
[0083] Next, we will introduce the process of updating data in a partition of the data to be managed.
[0084] When writing data, the database will first insert it into the local memory table MemTable. When the MemTable reaches a certain amount of data, it will be marked as immutable to form an Immutable MemTable. The database will create a new MemTable in memory to receive data writes, and write the Immutable MemTable to the disk to generate a bottom-level SSTable in the local disk. The bottom-level SSTables formed by writing to the disk will gradually increase as the database runs. When the number of bottom-level SSTables formed by writing to the disk is large (for example, exceeding a certain threshold), the database will perform Minor Compaction on these bottom-level SSTables formed by writing to the disk to merge them into a neighboring bottom-level SSTable. The neighboring bottom-level SSTables formed by the merger will gradually increase as the database runs. When the number of neighboring bottom-level SSTables is large (for example, exceeding a certain threshold), the database will further perform Minor Compaction on these neighboring bottom-level SSTables to merge them into a higher-level SSTable. As the database runs, the number of SSTable levels will increase, gradually forming a complex multi-layer structure.
[0085] These SSTables local to the database can be called incremental SSTables of the partition, and the object storage layer also stores the baseline SSTable of the partition. That is, the partition metadata contains metadata of the baseline SSTable and metadata of the incremental SSTable, the data blocks corresponding to the data block metadata under the metadata of the baseline SSTable are the data blocks in the object storage layer, and the data blocks corresponding to the data block metadata under the metadata of the incremental SSTable are the data blocks local to the database.
[0086] The data system also periodically performs Major Compaction, which merges all SSTables in the partition, namely the incremental SSTables in the database and the baseline SSTables in the object storage layer.
[0087] Next, we will take the regularly executed Major Compaction as an example to introduce the update process of the managed data and related metadata.
[0088] The data system includes different databases in the same partition for periodically executing the following steps:
[0089] First, a consensus is reached on the incremental SSTables of the same partition. That is, different databases reach a consensus on all incremental SSTables in the local partition at the same read point, so that the incremental SSTables between different databases are consistent.
[0090] Next, the baseline SSTable and incremental SSTable of the same partition are merged to update the baseline SSTable. That is, Major Compaction is performed to merge the incremental SSTable of the local database into the baseline SSTable of the object storage layer. This step can be performed by the database where the leader copy of the partition is located.
[0091] Finally, the metadata corresponding to the incremental SSTable merged into the baseline SSTable is deleted from the metadata, and the metadata corresponding to the baseline SSTable and the metadata corresponding to the partition are updated.
[0092] In this optional example, the database updates its metadata as the data to be managed is updated, and the database periodically shares the incremental data generated locally to the object storage layer, thereby releasing local storage space in a timely manner, improving the utilization rate of local storage space, and reducing the requirements for local storage space.
[0093] For example, in the above optional example, when the database updates the metadata, it may create a new object and write the updated metadata to be updated when updating the metadata to be updated within the partition metadata level within the metadata; and when updating the metadata to be updated at any level except the partition metadata level within the metadata, the metadata to be updated is updated within the object where the metadata to be updated is located.
[0094] Among them, the metadata to be updated in the partition metadata level, that is, the partition metadata adopts a special update method, that is, the old metadata is not immediately overwritten with the new metadata, but the new metadata is written into the new object, so that the old and new versions of the metadata exist at the same time, so as to avoid the historical information of the partition being immediately erased, causing the unfinished data operation to be interrupted, thereby ensuring that the data operation is not affected by the metadata update. As for the old version of the partition metadata, it will be asynchronously recovered by the background GC county, that is, it will be deleted by the GC thread after the unfinished data operation.
[0095] In other words, when metadata at any level except the partition metadata level is deleted, it is deleted immediately, while when partition metadata is deleted, it is deleted asynchronously.
[0096] In some embodiments of the present disclosure, the database is used to load the metadata of the internal hierarchical levels of the metadata into the memory starting from the root node of the metadata when restarting, so that the database is partially restored to the state before shutdown and obtains data reading capability.
[0097] The metadata of the internal hierarchical level of metadata mentioned in this embodiment refers to metadata that enables the database to have the ability to restore the state before shutdown and the data reading ability, and enables databases to reach a consensus on the newly generated incremental data.
[0098] After the database is restarted, not all the data to be managed will be used immediately, so the database is partially restored to the state before the shutdown so that consensus can be processed for the newly generated incremental data, and when the unloaded metadata is needed, it can be accessed and loaded step by step through the loaded metadata. Loading partial metadata can improve the efficiency of database restart.
[0099] For the data corresponding to the unloaded metadata, since deleting this data also requires loading its metadata to determine its object address, the data corresponding to the unloaded metadata will not be deleted. In addition to the normal deletion method, the database also has a bottom-line verification measure to solve the problem of data corresponding to the unloaded metadata being leaked and deleted. The bottom-line verification also requires metadata to determine the object address of the data, so the data corresponding to the unloaded metadata will not be deleted during the bottom-line verification.
[0100] For data corresponding to unloaded metadata, when these data are accessed, metadata of each level corresponding to the accessed data can be loaded step by step through the already loaded higher-level metadata to determine the object address of the accessed data and further complete the access to the accessed data.
[0101] It can be seen that this embodiment provides a method of partially loading metadata to realize database restart, and loading the remaining metadata on demand during the subsequent operation of the database. This method not only speeds up the efficiency of database restart, but also enables successful access when accessing data, that is, it improves the efficiency of database restart without affecting database performance.
[0102] Optional, with Figure 3 Taking the example shown as an example, this embodiment can load the metadata of three levels, namely, database metadata, tenant metadata and log stream metadata, in the metadata into the memory starting from the root node of the metadata, wherein the log stream metadata is used to record the meta information of the corresponding log stream and the object address of each partition metadata under the log stream metadata. The meta information of the log stream may include the number of managed partitions, the managed partition identifiers, the key value range of each managed partition, etc.
[0103] Specifically, this optional example may include the following steps:
[0104] First, the database loads the database metadata from the fixed path of the database metadata into the memory, and obtains the object address of the metadata of each tenant contained in the database after operations such as deserialization.
[0105] Next, according to the object address of each tenant's metadata, the metadata of each tenant is loaded into the memory to restore the state of each tenant, and then the object address of the metadata of the log stream contained in the metadata of each tenant is read through the metadata of each tenant.
[0106] Finally, for each tenant, the metadata of each log stream is loaded into the memory according to the object address of the metadata of each log stream it contains, so as to restore the state of the log stream and the consensus protocol group formed by the log stream, and then read and record the object address of the metadata of the partition it contains through the metadata of each log stream.
[0107] Among them, the consensus protocol group formed by the log stream can achieve data consensus among different databases.
[0108] In this optional example, the object address of the partition's metadata recorded in memory can provide guidance when the database accesses data in the partition, that is, the object address of the partition metadata, the object address of the SSTable metadata, the object address of the data block metadata, and the object address of the data block are read in sequence until the data to be accessed is obtained. In other words, although the database has only loaded the log stream metadata and the metadata of the layers above it, the multi-level tree-like metadata provided by the data system enables the database to quickly find the data to be accessed based on the metadata that has been loaded when accessing specific data.
[0109] It should be understood that in the metadata of the multi-level tree structure, the metadata in the data level closer to the leaf node is larger, and the metadata in the data level closer to the root node increases exponentially. In this optional example, after the database is restarted, the metadata of the partition and its lower levels are not loaded, thereby reducing the delay of database restart, dispersing the delay in the access process of specific data, and solving the problem of low restart efficiency.
[0110] In this embodiment, when the database is restarted, only the highest level of local metadata needs to be loaded into the memory, so as to partially restore the database to the state before shutdown and obtain data reading capability, thereby completing the restart and providing external services. The restart efficiency is high and it is adapted to the characteristics of object storage. It avoids the large amount of IO load caused by reading out historical status logs one by one when the database is restarted in related technologies such as the WAL mechanism, thereby improving the efficiency of database restart and operation.
[0111] At least one embodiment of this specification provides a data management method, which can be applied to any database of a data system, wherein the database system includes an object storage layer and at least one database, wherein the object storage layer stores baseline data of the managed data, and the database stores incremental data of the managed data. Figure 4 , the method comprising:
[0112] In step S401, metadata of the managed data is constructed in the database based on the baseline data and the incremental data. The metadata is in a multi-level tree structure with database metadata as a root node and data block metadata as leaf nodes.
[0113] Among them, the subordinate components contained in the upper-level component of the managed data form child nodes under the node corresponding to the upper-level component in the metadata; in the metadata, the nodes in any level except the leaf nodes are used to record the metadata of the corresponding components and the object address of the child nodes, and the nodes in the leaf nodes are used to record the metadata of the corresponding data blocks.
[0114] Exemplarily, from the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata. Based on this, the method may also include:
[0115] The following steps are periodically performed for each partition included in the database:
[0116] Reach a consensus with other databases in the data system that contain the same partition on the incremental SSTable of the same partition;
[0117] Merge the incremental SSTable of the same partition into the baseline SSTable to update the baseline SSTable;
[0118] The metadata corresponding to the incremental SSTable merged into the baseline SSTable is deleted within the metadata, and the metadata corresponding to the baseline SSTable and the metadata corresponding to the same partition are updated.
[0119] Exemplarily, from the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata, and the data system includes a consensus protocol group formed between the same log streams in different databases. Based on this, the method may also include:
[0120] Based on the consensus protocol group, reaching a consensus with other databases in the data system that contain the same tenant on the tenant status in the metadata of the same tenant;
[0121] Based on the consensus protocol group, consensus is reached on metadata of the same log stream with other databases in the data system that contain the same log stream;
[0122] Based on the consensus protocol group, consensus is reached with other databases in the data system that contain the same partition on metadata of the same partition.
[0123] Exemplarily, the method may further include:
[0124] After the database is restarted, the metadata is loaded into the memory to restore the database to the state before shutdown.
[0125] Exemplarily, the method may further include:
[0126] When the database is restarted, the metadata of the internal hierarchical levels of the metadata are loaded into the memory starting from the root node of the metadata, so that the database is partially restored to the state before shutdown and obtains data reading capability.
[0127] For example, from the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata; then this example can start from the root node of the metadata to load the three levels of metadata, namely database metadata, tenant metadata and log stream metadata, into the memory when the database is restarted, so as to partially restore the database to the state before shutdown and obtain data reading capabilities, wherein the log stream metadata is used to record the metadata of the corresponding log stream and the object address of each partition metadata under the log stream metadata.
[0128] More details about the above steps have been introduced in detail in the embodiment of the data system, and will not be repeated here.
[0129] Figure 5 is a schematic structural diagram of a device provided by an exemplary embodiment. Figure 5At the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510, and may also include hardware required for other tasks. One or more embodiments of this specification may be implemented based on software, such as the processor 502 reading the corresponding computer program from the non-volatile memory 510 into the memory 508 and then running it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0130] Please refer to Figure 6 , the data management device can be used for Figure 5 The data management device may include:
[0131] Construction module 601 is used to construct the metadata of the managed data in the database based on the baseline data and the incremental data, and the metadata is in a multi-level tree structure with the database metadata as the root node and the data block metadata as the leaf nodes, wherein the subordinate components contained in the upper-level components of the managed data form child nodes under the nodes corresponding to the upper-level components in the metadata; in the metadata, the nodes in any level except the leaf nodes are used to record the metadata of the corresponding components and the object addresses of the child nodes, and the nodes in the leaf nodes are used to record the metadata of the corresponding data blocks.
[0132] In a possible embodiment of the present specification, from the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata;
[0133] The device comprises an updating module, which is used for:
[0134] The following steps are periodically performed for each partition included in the database:
[0135] Reach a consensus with other databases in the data system that contain the same partition on the incremental SSTable of the same partition;
[0136] Merge the incremental SSTable of the same partition into the baseline SSTable to update the baseline SSTable;
[0137] The metadata corresponding to the incremental SSTable merged into the baseline SSTable is deleted within the metadata, and the metadata corresponding to the baseline SSTable and the metadata corresponding to the same partition are updated.
[0138] In a possible embodiment of the present specification, from the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata, and the data system includes a consensus protocol group formed between the same log streams in different databases;
[0139] The device also includes a consensus module, which is used to:
[0140] Based on the consensus protocol group, reaching a consensus with other databases in the data system that contain the same tenant on the tenant status in the metadata of the same tenant;
[0141] Based on the consensus protocol group, consensus is reached on metadata of the same log stream with other databases in the data system that contain the same log stream;
[0142] Based on the consensus protocol group, consensus is reached with other databases in the data system that contain the same partition on metadata of the same partition.
[0143] In a possible embodiment of the present specification, the device further includes a restart module, configured to:
[0144] After the database is restarted, the metadata is loaded into the memory to restore the database to the state before shutdown.
[0145] In a possible embodiment of the present specification, the device includes a restart module, which is used to:
[0146] When the database is restarted, the metadata of the internal hierarchical levels of the metadata are loaded into the memory starting from the root node of the metadata, so that the database is partially restored to the state before shutdown and obtains data reading capability.
[0147] In a possible embodiment of the present specification, from the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata; the restart module is used to load the metadata of the internal hierarchical level of the metadata into the memory starting from the root node of the metadata, and is used to:
[0148] Starting from the root node of the metadata, the metadata of three levels, namely, database metadata, tenant metadata and log stream metadata, within the metadata are loaded into the memory, wherein the log stream metadata is used to record the metadata of the corresponding log stream and the object address of each partition metadata under the log stream metadata.
[0149] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver, a game console, a tablet computer, a wearable device or a combination of any of these devices.
[0150] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0151] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0152] Computer readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0153] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0154] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0155] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0156] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0157] It should be understood that although the terms first, second, third, etc. may be used to describe various information in one or more embodiments of this specification, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0158] The above description is merely a preferred embodiment of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.
Claims
1. A data system, comprising: The object storage layer stores the baseline data of the managed data; At least one database is communicatively connected to the object storage layer, storing incremental data of the managed data and metadata of the managed data, wherein the metadata is in a multi-level tree structure with database metadata as a root node and data block metadata as leaf nodes; The lower-level components contained in the upper-level component of the managed data form child nodes under the node corresponding to the upper-level component in the metadata; in the metadata, the nodes in any level except the leaf nodes are used to record the meta information of the corresponding components and the object address of the child nodes, and the nodes in the leaf nodes are used to record the meta information of the corresponding data blocks; From the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata; The database is used to create a new object and write the updated metadata to be updated when updating the metadata to be updated in the partition metadata level within the metadata; The database is used for updating the metadata to be updated in the object where the metadata to be updated is located when updating the metadata to be updated at any level except the partition metadata level in the metadata.
2. The data system according to claim 1, wherein different databases of the same tenant are included in the data system for reaching consensus on the tenant status in the metadata of the same tenant; Different databases containing the same log stream in the data system are used to reach consensus on metadata of the same log stream; The data system includes different databases of the same partition for reaching consensus on metadata of the same partition.
3. According to the data system of claim 2, the data system includes forming a consensus protocol group between the same log streams in different databases, and the consensus protocol group is used to reach consensus on the metadata of the same tenant, the metadata of the same log stream, and the metadata of the same partition.
4. According to the data system of claim 1, the partition metadata includes metadata of a baseline SSTable and metadata of an incremental SSTable, the data blocks corresponding to the data block metadata under the metadata of the baseline SSTable are data blocks in the object storage layer, and the data blocks corresponding to the data block metadata under the metadata of the incremental SSTable are data blocks local to the database.
5. The data system according to claim 4, wherein different databases in the same partition are used to periodically perform the following steps: A consensus is reached on the incremental SSTable of the same partition; Merge the baseline SSTable and incremental SSTable of the same partition to update the baseline SSTable; The metadata corresponding to the incremental SSTable merged into the baseline SSTable is deleted from the metadata, and the metadata corresponding to the baseline SSTable and the metadata corresponding to the partition are updated.
6. According to the data system of claim 1, the database is used to load the metadata of the internal hierarchical level of the metadata into the memory starting from the root node of the metadata when restarting, so that the database is partially restored to the state before shutdown and obtains data reading capability.
7. The data system according to claim 6, wherein the database is used to load metadata of three levels, namely, database metadata, tenant metadata and log stream metadata, in the metadata into the memory starting from the root node of the metadata when restarting, wherein: The log stream metadata is used to record the meta information of the corresponding log stream and the object address of each partition metadata under the log stream metadata.
8. The data system according to claim 1, wherein the database is used to load the metadata into the memory when restarting so that the database is restored to the state before shutdown.
9. A data management method, applied to any database of a data system, the data system comprising an object storage layer and at least one database, the object storage layer storing baseline data of managed data, the database storing incremental data of the managed data; The method comprises: The metadata of the managed data is constructed in the database based on the baseline data and the incremental data, wherein the metadata is in a multi-level tree structure with the database metadata as the root node and the data block metadata as the leaf node, wherein the lower-level components contained in the upper-level components of the managed data form child nodes under the nodes corresponding to the upper-level components in the metadata; in the metadata, the nodes in any level except the leaf nodes are used to record the meta information of the corresponding components and the object addresses of the child nodes, and the nodes in the leaf nodes are used to record the meta information of the corresponding data blocks; From the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata; The method further comprises: When updating the metadata to be updated in the partition metadata level within the metadata, creating a new object and writing the updated metadata to be updated; When updating metadata to be updated at any level except the partition metadata level in the metadata, the metadata to be updated is updated in the object where the metadata to be updated is located.
10. The data management method according to claim 9, the method comprising: The following steps are periodically performed for each partition included in the database: Reach a consensus with other databases in the data system that contain the same partition on the incremental SSTable of the same partition; Merge the incremental SSTable of the same partition into the baseline SSTable to update the baseline SSTable; The metadata corresponding to the incremental SSTable merged into the baseline SSTable is deleted within the metadata, and the metadata corresponding to the baseline SSTable and the metadata corresponding to the same partition are updated.
11. The data management method according to claim 9, wherein the data system includes forming a consensus protocol group between the same log streams in different databases; The method further comprises: Based on the consensus protocol group, reaching a consensus with other databases in the data system that contain the same tenant on the tenant status in the metadata of the same tenant; Based on the consensus protocol group, consensus is reached on metadata of the same log stream with other databases in the data system that contain the same log stream; Based on the consensus protocol group, consensus is reached with other databases in the data system that contain the same partition on metadata of the same partition.
12. The data management method according to claim 9, further comprising: When the database is restarted, the metadata is loaded into the memory to restore the database to the state before shutdown.
13. The data management method according to claim 9, further comprising: When the database is restarted, the metadata of the internal hierarchical levels of the metadata are loaded into the memory starting from the root node of the metadata, so that the database is partially restored to the state before shutdown and obtains data reading capability.
14. The data management method according to claim 13, wherein the step of loading the metadata of the internal hierarchical levels of the metadata into the memory starting from the root node of the metadata comprises: Starting from the root node of the metadata, the metadata of three levels, namely, database metadata, tenant metadata and log stream metadata, within the metadata are loaded into the memory, wherein the log stream metadata is used to record the metadata of the corresponding log stream and the object address of each partition metadata under the log stream metadata.
15. A data management device, applied to any database of a data system, the database system comprising an object storage layer and at least one database, the object storage layer storing baseline data of managed data, the database storing incremental data of the managed data; The device comprises: A construction module, used for constructing metadata of the managed data in the database based on the baseline data and the incremental data, wherein the metadata is in a multi-level tree structure with database metadata as the root node and data block metadata as the leaf node, wherein the subordinate components contained in the superior component of the managed data form child nodes under the nodes corresponding to the superior components in the metadata; in the metadata, nodes in any level except leaf nodes are used to record the metadata of the corresponding components and the object addresses of the child nodes, and the nodes in the leaf nodes are used to record the metadata of the corresponding data blocks; from the root node to the leaf node, the metadata includes database metadata, tenant metadata, log stream metadata, partition metadata, ordered string table SSTable metadata and data block metadata; The update module is used to create a new object and write the updated metadata to be updated when updating the metadata to be updated in the partition metadata level in the metadata, and to update the metadata to be updated in the object where the metadata to be updated is located when updating the metadata to be updated in any level except the partition metadata level in the metadata.
16. The data management device according to claim 15, comprising an update module, configured to: The following steps are periodically performed for each partition included in the database: Reach a consensus with other databases in the data system that contain the same partition on the incremental SSTable of the same partition; Merge the incremental SSTable of the same partition into the baseline SSTable to update the baseline SSTable; The metadata corresponding to the incremental SSTable merged into the baseline SSTable is deleted within the metadata, and the metadata corresponding to the baseline SSTable and the metadata corresponding to the same partition are updated.
17. The data management device according to claim 15, comprising a restart module, configured to: When the database is restarted, the metadata of the internal hierarchical levels of the metadata are loaded into the memory starting from the root node of the metadata, so that the database is partially restored to the state before shutdown and obtains data reading capability.
18. The data management device according to claim 17, wherein the restart module is used to, when loading the metadata of the internal hierarchical level of the metadata into the memory starting from the root node of the metadata,: Starting from the root node of the metadata, the metadata of three levels, namely, database metadata, tenant metadata and log stream metadata, in the metadata are loaded into the memory, wherein: The log stream metadata is used to record the meta information of the corresponding log stream and the object address of each partition metadata under the log stream metadata.
19. A computer program product comprising a computer program / instructions, which when executed by a processor implement the steps of the method according to any one of claims 9 to 14.
20. An electronic device, comprising: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 9 to 14 by running the executable instructions.
21. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 9 to 14.
Citation Information
Patent Citations
Implementation method and system of file system oriented to nonvolatile memory and medium
CN111221776A