Metadata storage method and apparatus, metadata query method and apparatus, and computer device and storage medium

By building a Hoffman coding tree and generating a target key, the problem that metadata storage cannot efficiently utilize storage space in the prior art is solved, and the effect of reducing data storage costs is achieved.

WO2025130630A1PCT designated stage expired Publication Date: 2025-06-26HANGZHOU OPENPIE TECH DEV CO LTD

Patent Information

Application Number
PCT/CN2024/137017
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-05
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing metadata storage methods cannot efficiently utilize storage space, resulting in high data storage costs.

Method used

By building a Hoffman encoding tree, the target key corresponding to each encoding path is generated, and the target key is associated with the corresponding related values ​​to store it on the target disk page, optimizing storage space utilization.

Benefits of technology

It effectively reduces the storage space required for each target key, solves the problem of inability to efficiently utilize the storage space, and achieves the effect of reducing data storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137017_26062025_PF_FP_ABST
    Figure CN2024137017_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A metadata storage method and apparatus, a metadata query method and apparatus, and a computer device and a storage medium. The metadata storage method comprises: on the basis of identity identification information in each metadata block, constructing a Huffman coding tree corresponding to each metadata block; generating a target key corresponding to each coding path in the Huffman coding tree, and determining a target disk page corresponding to the target key; and furthermore, using, as a correlation value, original data information corresponding to each coding path, and storing the target key and the corresponding correlation value in the target disk page in an associated manner.
Need to check novelty before this filing date? Find Prior Art

Description

Metadata storage and query method, device, computer equipment and storage medium

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 202311755927.1, filed on December 20, 2023, entitled “Metadata Storage and Query Method, Device, Computer Equipment and Storage Medium,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of information processing technology, and in particular to a metadata storage and query method, apparatus, computer equipment, and storage medium. Background Art

[0004] Metadata provides information about data attributes and characteristics, including its structure, format, and content. Metadata enables operations such as data query and data quality management. Furthermore, as critical data in a database, corruption can cause the database to cease service and become irrecoverable. Therefore, secure metadata storage is essential.

[0005] Current metadata storage methods determine the type of metadata to be stored, process it using different storage formats based on the type, generate new metadata, and store the new metadata in a pre-defined metadata storage system. However, this storage method fails to efficiently utilize storage space, resulting in high data storage costs.

[0006] There is currently no effective solution to the problem that related technologies cannot efficiently utilize storage space, resulting in high data storage costs. Summary of the Invention

[0007] According to various embodiments of the present application, a metadata storage and query method, apparatus, computer device, and storage medium are provided.

[0008] In a first aspect, this embodiment provides a metadata storage method, the method comprising:

[0009] Based on the identity information in each metadata block, construct a Huffman coding tree corresponding to each metadata block;

[0010] Generate a target key corresponding to each coding path in the Huffman coding tree, and determine a target disk page corresponding to the target key;

[0011] The original data information corresponding to each encoding path is used as a related value, and the target key and the corresponding related value are associated and stored in the target disk page.

[0012] In some embodiments, constructing a Huffman coding tree corresponding to each metadata block based on the identity information in each metadata block includes:

[0013] Obtaining the identity information in each metadata block; the identity information includes a metadata identifier, a database identifier, a view identifier, and a domain identifier;

[0014] Encoding the identity information in each metadata block to obtain the corresponding Huffman coding tree;

[0015] Wherein, each level of the Huffman coding tree corresponds to a different category of the identity identification information.

[0016] In some embodiments, generating a target key corresponding to each coding path in the Huffman coding tree includes:

[0017] Obtain each coding path in the Huffman coding tree;

[0018] Obtaining data feature information corresponding to the coding path; the data feature information includes version information and data area feature code;

[0019] The encoding of each node in the encoding path is determined, and the corresponding target key is generated based on the encoding of each node and the data feature information.

[0020] In some embodiments, after storing the target key and the corresponding associated value in association with each other in the target disk page, the method further includes:

[0021] When detecting that the target keys of the metadata blocks in the target disk page have the same encoding segment, generating a tag value corresponding to the same encoding segment; the tag value is associated with a reference address of the same encoding segment;

[0022] The same coding segment in the metadata block is updated to the marker value.

[0023] In some embodiments, after storing the target key and the corresponding related value in association with each other in the target disk page, the method further includes:

[0024] Upon receiving a new metadata block, generating a query key corresponding to the new metadata block;

[0025] Obtaining coding tree information corresponding to the query key, and encoding the query key based on the coding tree information to obtain a target key;

[0026] Using the original data information in the new metadata block as the relevant value;

[0027] Determine a data node corresponding to the target key, and associate the target key and the corresponding related value and store them in the data node.

[0028] In a second aspect, this embodiment provides a metadata query method, the method comprising:

[0029] Generate a corresponding query key according to the received user demand data, and obtain the coding tree information corresponding to the query key;

[0030] Encoding the query key based on the encoding tree information to obtain a target key;

[0031] A target disk page corresponding to the target key is determined, and a related value corresponding to the target key is obtained from the target disk page; the related value stores original data information in a metadata block.

[0032] In some embodiments, obtaining the coding tree information corresponding to the query key includes:

[0033] Retrieving the coding tree information corresponding to the query key in the coding tree cache;

[0034] When the coding tree information is not retrieved, the coding tree information is extracted from the metadata index node.

[0035] In a third aspect, a metadata storage device is provided in this embodiment, the device comprising: a construction module, a generation module, and a storage module;

[0036] The construction module is used to construct a Huffman coding tree corresponding to each metadata block based on the identity identification information in each metadata block;

[0037] The generating module is configured to generate a target key corresponding to each coding path in the Huffman coding tree, and determine a target disk page corresponding to the target key;

[0038] The storage module is used to use the original data information corresponding to each encoding path as a related value, and associate the target key with the corresponding related value and store it in the target disk page.

[0039] In a fourth aspect, a computer device is provided in this embodiment, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the metadata storage method described in the first aspect when executing the computer program.

[0040] In a fifth aspect, a storage medium is provided in this embodiment, on which a computer program is stored. When the program is executed by a processor, the metadata storage method described in the first aspect is implemented.

[0041] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.

[0043] FIG1 is a block diagram of the hardware structure of a terminal device of a metadata storage method provided in an embodiment of the present application.

[0044] FIG2 is a flowchart of a metadata storage method provided in an embodiment of the present application.

[0045] FIG3 is a schematic diagram of the structure of a Huffman coding tree provided in an embodiment of the present application.

[0046] FIG4 is a schematic diagram of the structure of sub-node inlining provided by an embodiment of the present application.

[0047] FIG5 is a flowchart of a metadata query method provided by an embodiment of the present application.

[0048] FIG6 is a flowchart of a metadata storage method provided by an optional embodiment of the present application.

[0049] FIG7 is a structural block diagram of a metadata query system provided in an embodiment of the present application.

[0050] FIG8 is a structural block diagram of a metadata storage device provided in an embodiment of the present application.

[0051] FIG9 is a structural block diagram of a metadata query device provided in an embodiment of the present application.

[0052] In the figure: 102, processor; 104, memory; 106, transmission device; 108, input and output device; 10, construction module; 20, generation module; 30, storage module; 40, acquisition module; 50, encoding module; 60, query module; 100, control module; 200, input module; 300, encoder; 400, coding tree cache; 500, metadata index node; 600, processor; 700, metadata storage node. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0054] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0055] Unless otherwise defined, the technical terms or scientific terms involved in this application should have the general meaning understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "an", "a", "the", "these" and the like in this application do not indicate quantitative restrictions, and they can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application refers to two or more. "And / or" describes the relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. Generally, the character " / " indicates that the related objects are in an "or" relationship. The terms "first," "second," "third," etc. used in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0056] The method embodiment provided in this embodiment can be executed in a terminal, a computer, or a similar computing device. For example, when running on a terminal, Figure 1 is a hardware structure block diagram of the terminal of the metadata storage method of this embodiment. As shown in Figure 1, the terminal may include one or more (only one is shown in Figure 1) processors 102 and a memory 104 for storing data, wherein the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above-mentioned terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned terminal. For example, the terminal may also include more or fewer components than those shown in Figure 1, or have a different configuration than that shown in Figure 1.

[0057] Memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the metadata storage method in this embodiment. Processor 102 executes the computer programs stored in memory 104 to execute various functional applications and data processing, thereby implementing the aforementioned method. Memory 104 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, memory 104 may further include memory remotely located from processor 102, which can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0058] The transmission device 106 is used to receive or send data via a network. The network may include a wireless network provided by the terminal's communications provider. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0059] In this embodiment, a metadata storage method is provided. FIG2 is a flow chart of the metadata storage method of this embodiment. As shown in FIG2 , the flow chart includes the following steps:

[0060] Step S210: constructing a Huffman coding tree corresponding to each metadata block based on the identity information in each metadata block.

[0061] It's important to note that in a key-value (KV) distributed storage system, each key corresponds to a unique value. When storing metadata in KV format, each metadata block is converted into corresponding KV data. The associated value is used to store the metadata block's raw data information, and the target key is the metadata feature value used to retrieve and query the corresponding data.

[0062] Specifically, the identity identification information in each metadata block is obtained, and the identity identification information includes four data areas, namely metadata identification, database identification, view identification and domain identification. The identity identification information in each metadata block is encoded to obtain the corresponding Huffman coding tree.

[0063] Among them, the Huffman coding tree is a coding tree structure used for data compression, and each level of the Huffman coding tree corresponds to a different category of identity information, that is, each level of the Huffman coding tree is only responsible for encoding the same category of identity information.

[0064] Step S220 , generating a target key corresponding to each coding path in the Huffman coding tree, and determining a target disk page corresponding to the target key.

[0065] Specifically, each level of the Huffman coding tree is traversed to obtain each coding path in the Huffman coding tree. For each coding path, corresponding data feature information is obtained, wherein the data feature information includes version information and data area feature code.

[0066] Furthermore, the code of each node in the coding path is extracted, and based on the code of each node, the associated version information and data area feature code, a corresponding target key is generated and a target disk page corresponding to the target key is determined.

[0067] It should be noted that this embodiment only performs encoding operations on the identity identification information in the metadata block, and the version information and data area characteristic coding of the metadata block remain unchanged in their original forms.

[0068] In step S230 , the original data information corresponding to each encoding path is used as a related value, and the target key and the corresponding related value are associated and stored in the target disk page.

[0069] Specifically, each encoding path corresponds to a different metadata block. Therefore, after generating a target key corresponding to the encoding path, the original data in the associated metadata block is used as the associated value. The target key and the corresponding associated value are then associated and stored on the target disk page, thus achieving ordered storage of the metadata blocks. Furthermore, when querying metadata, only the corresponding disk page location is required to complete the query, avoiding scanning all disk pages and improving the efficiency of subsequent data queries.

[0070] It's important to note that target disk pages correspond to data nodes, and all metadata is sharded according to the target key and stored across different data nodes to support load balancing within the database system. Furthermore, all leaf nodes in the Huffman coding tree store the corresponding data node location information, further improving metadata query performance.

[0071] Current metadata storage methods determine the type of metadata to be stored, process it using different storage formats based on the type, generate new metadata, and store the new metadata in a pre-defined metadata storage system. However, this storage method fails to efficiently utilize storage space, resulting in high data storage costs.

[0072] Compared with related technologies, this application performs Huffman encoding on each metadata block to be stored to obtain a corresponding Huffman encoding tree, and generates a target key corresponding to each encoding path in the Huffman encoding tree. Based on this, the original data information corresponding to each encoding path is used as the relevant value, and the target disk page corresponding to the target key is determined. The target key and the corresponding relevant value are associated and stored on the target disk page. By converting the data information constituting the target key into Huffman encoding, the storage space required for each target key is reduced, solving the problem of inefficient use of storage space, resulting in high data storage costs, and achieving reduced data storage costs.

[0073] In some embodiments, constructing a Huffman coding tree corresponding to each metadata block based on the identity information in each metadata block includes the following steps:

[0074] Step S211, obtaining the identity information in each metadata block; the identity information includes metadata identification, database identification, view identification and domain identification;

[0075] Step S212 , encoding the identity information in each metadata block to obtain a corresponding Huffman coding tree; wherein each level of the Huffman coding tree corresponds to a different type of identity information.

[0076] Specifically, each metadata block to be stored consists of multiple identification information, version information, and corresponding original data information. The identification information includes metadata identifiers, database identifiers, view identifiers, and domain identifiers, which are used to identify different metadata blocks, distinguish between different databases within the same database system, and distinguish between the views associated with the metadata block and the different domains to which it belongs.

[0077] Furthermore, each metadata block to be stored is scanned and its contained identity information is encoded to generate a corresponding Huffman coding tree. The metadata block encoding process is illustrated using Figure 3 as an example. First, the metadata identifier of the metadata block to be stored is obtained and encoded as A, which serves as the root node of the coding tree. The database identifier of the metadata block is obtained and encoded as a child node under node A. The view identifier and domain identifier of the metadata block are then obtained and encoded as child nodes under a node at the previous level.

[0078] Through this embodiment, identity identification information in each metadata block is obtained, and the identity identification information includes a metadata identifier, a database identifier, a view identifier, and a domain identifier. The identity identification information in each metadata block is encoded to obtain a corresponding Huffman coding tree, and each level of the Huffman coding tree corresponds to a different category of identity identification information. In this way, the Huffman coding tree is optimized, the number of levels of the Huffman coding tree is effectively reduced, and the encoding and decoding capabilities of the metadata block are improved.

[0079] In some embodiments, generating a target key corresponding to each coding path in a Huffman coding tree includes the following steps:

[0080] Step S221, obtaining each coding path in the Huffman coding tree;

[0081] Step S222: Acquire data feature information corresponding to the encoding path; the data feature information includes version information and data area feature code;

[0082] Step S223: Determine the code of each node in the coding path, and generate a corresponding target key based on the code and data feature information of each node.

[0083] Specifically, each level of the Huffman coding tree is traversed to obtain each coding path in the Huffman coding tree. As shown in Figure 3, "AB-AB-AAA", "AB-AB-AAB", and "AB-AB-AAC" are three different coding paths, and each code in each coding path corresponds to different identity information in the same metadata block.

[0084] Furthermore, data feature information corresponding to the encoding path is obtained, which includes the version information of the metadata block and the data area feature code. The version information includes information such as the version number or timestamp of the metadata block; and the data area feature code is used to verify the integrity and consistency of the metadata block.

[0085] Based on this, when generating the target key of each metadata block, the first data area in the current metadata block, namely the metadata identifier, is obtained, the root node corresponding to the metadata identifier is found from the Huffman coding tree, and the corresponding code is extracted from the root node; secondly, the second data area in the current metadata block, namely the database identifier, is obtained, the child nodes under the root node are traversed, the child node corresponding to the database identifier is determined, and the corresponding code is extracted from the child node; the same operation is performed on the third data area, namely the view identifier, and the fourth data area, namely the domain identifier, in turn to obtain the codes corresponding to the view identifier and the domain identifier, respectively.

[0086] After obtaining the code for each data region in the metadata block, the obtained codes are combined with the data feature information to regenerate the target key corresponding to the metadata block. For example, if the version information and data region feature code in the metadata block is 010OPENPIECLOUDDB..., the identifiers of the first to fourth data regions are 314, 00100000560000, 000000000787899, and 00000005768000, respectively. The codes corresponding to these identifiers are A, A, AA, and AAA, respectively. In this case, the regenerated target key is "AA-AA-AAA-010OPENPIECLOUDDB..."

[0087] It's important to note that each database system contains multiple databases, each database contains multiple views, each view contains multiple domains, and each domain stores a large number of metadata blocks. In related key-value pair storage methods, the target key of a metadata block needs to contain sufficient data to mark itself, resulting in a large amount of target key data. This embodiment utilizes an optimized Huffman coding tree to encode "314-00100000560000-000000000787899-00000005768000-010OPENPIECLOUDDB..." as "AA-AA-AAA-010OPENPIECLOUDDB..." for storage. This ensures secure metadata storage while reducing the storage space required for each target key.

[0088] Through this embodiment, each coding path in the Huffman coding tree is obtained, and the data feature information corresponding to the coding path is obtained, wherein the data feature information includes version information and data area feature coding, and then the coding of each node in the coding path is determined. Based on the coding and data feature information of each node, the corresponding target key is generated. In this way, by converting the data information used to generate the target key into a Huffman code with smaller storage space, the storage space required for each target key is reduced, thereby achieving efficient use of storage space and reducing data storage costs.

[0089] In some embodiments, after the target key and the corresponding associated value are stored in the target disk page, the following steps are further included:

[0090] When it is detected that the target keys of the metadata blocks in the target disk page have the same encoding segment, a tag value corresponding to the same encoding segment is generated; the tag value is associated with a reference address of the same encoding segment;

[0091] Update the same encoding segment in the metadata block to the tag value.

[0092] Specifically, in each target disk page storing metadata blocks, it is determined whether the target keys of the metadata blocks in the page have the same encoding segment. If it is detected that the target keys of the metadata blocks have the same encoding segment, a flag value corresponding to the same encoding segment is generated.

[0093] Exemplarily, when the metadata blocks stored in the target disk page include “314-A-AA-AAA-…”, “314-A-AA-AAA-…”, “314-A-AA-AAB-…” and “314-A-AB-AAA-…”, the same coding segment is “314-A-AA”, a tag value Ref corresponding to the coding segment is generated, and the same coding segment in the above metadata block is updated to the tag value Ref, the updated target disk page includes “314-A-AA-AAA-…”, “Ref-AAA-…”, “Ref-AAB-…” and “314-A-AB-AAA-…”.

[0094] It should be noted that when updating the metadata block, the first metadata block with the same coding segment in the target disk page remains unchanged in its original form, and the same coding segment in the metadata block is used as the data to be referenced, and the tag value is associated with the reference address of the same coding segment.

[0095] Through this embodiment, when it is detected that the target keys of each metadata block in the target disk page have the same coding segment, a tag value corresponding to the same coding segment is generated, the tag value is associated with the reference address of the same coding segment, and the same coding segment in the metadata block is updated to the tag value, thereby achieving continuous deduplication and compression of the stored data, and further reducing the space occupied by metadata storage.

[0096] In some embodiments, after the target key and the corresponding associated value are stored in the target disk page, the following steps are further included:

[0097] Step S241, upon receiving a new metadata block, generating a query key corresponding to the new metadata block;

[0098] Step S242: Obtain the coding tree information corresponding to the query key, and encode the query key based on the coding tree information to obtain the target key;

[0099] Step S243, using the original data information in the new metadata block as the relevant value;

[0100] Step S244 , determining the target disk page corresponding to the target key, and associating the target key and the corresponding related value and storing them in the target disk page.

[0101] Specifically, when a new metadata block needs to be inserted into a data node, a query key corresponding to the new metadata block is generated and transmitted to the encoder. The encoder then searches the encoding tree cache for the corresponding encoding tree information based on the received query key. If the cache hits, the encoding tree cache sends the required encoding tree information to the encoder; if the cache misses, the encoding tree cache obtains the required encoding tree information from the metadata index node and returns it to the encoder.

[0102] After obtaining the code tree information corresponding to the query key, the encoder encodes the query key based on the code tree information, generates a target key, and sends it to the processor. The code tree information is used to provide information such as the encoding format to generate a target key that is compatible with the local storage method.

[0103] Furthermore, using the original data information in the new metadata block as the relevant value, the processor obtains the location of the target disk page corresponding to the target key from the coding tree cache, and associates the target key and the corresponding relevant value and stores them on the target disk page, completing the insertion and storage of the metadata block. Furthermore, the operation principles for deleting a metadata block are the same as those for inserting a metadata block.

[0104] It is important to note that if the code tree information required by the new metadata block is not in the local code tree, the code tree expansion is performed. For example, based on the code tree information required by the new metadata block, a code corresponding to a database identifier needs to be added. In this case, a child node is added to the level of the database identifier in the relevant Huffman code tree to obtain the latest version of the code tree. At this time, the processor sends the latest version of the code tree to the metadata index node, which updates the code tree version. At the same time, the processor refreshes the code tree cache to achieve a rolling update.

[0105] Through this embodiment, when a new metadata block is received, a query key corresponding to the new metadata block is generated; the coding tree information corresponding to the query key is obtained, and the query key is encoded based on the coding tree information to obtain a target key; the original data information in the new metadata block is used as the relevant value; and then the target disk page corresponding to the target key is determined, and the target key and the corresponding relevant value are associated and stored in the target disk page, thereby realizing the insertion of the metadata block.

[0106] In some embodiments, in the Huffman coding tree, each child node under each node is associated with each other through an inline data pointer.

[0107] Specifically, in the Huffman coding tree, the levels corresponding to the database identifier, view identifier, and domain identifier each contain multiple child nodes. For the child nodes under the same node, inline data pointers are set to associate the child nodes.

[0108] For example, Figure 4 shows that nodes AA, AB, and AC are all child nodes of node A. Inline data pointers are used to associate nodes AA, AB, and AC. This allows queries to be performed between child nodes of the same node when traversing the Huffman coding tree, without backtracking to the parent node.

[0109] Through this embodiment, in the Huffman coding tree, each child node under each node is associated through an inline data pointer, thereby improving the search efficiency of data on the same layer and reducing the traversal overhead of the Huffman coding tree.

[0110] In this embodiment, a metadata query method is provided. FIG5 is a flow chart of the metadata query method of this embodiment. As shown in FIG5 , the flow chart includes the following steps:

[0111] Step S510: Generate a corresponding query key according to the received user demand data, and obtain the coding tree information corresponding to the query key;

[0112] Step S520: Encode the query key based on the encoding tree information to obtain a target key;

[0113] Step S530 , determining the target disk page corresponding to the target key, and obtaining the relevant value corresponding to the target key from the target disk page; the relevant value stores the original data information in the metadata block.

[0114] Specifically, when user demand data is received, a corresponding query key is generated according to the user demand data, the query key is sent to the encoder, and the encoder queries the coding tree information required by the query key.

[0115] After obtaining the coding tree information, the query key is encoded based on the coding tree information to obtain the target key. The coding tree information is used to provide information such as the encoding format to generate the target key that is compatible with the local storage method.

[0116] It's important to note that during metadata storage, to support load balancing in the database system, all metadata is sharded by target key and stored across different data nodes, with all data nodes placed in the metadata storage node. Based on this, the target data node and the corresponding target disk page location corresponding to the target key are retrieved from the coding tree cache. In the metadata storage node, the corresponding value corresponding to the target key is extracted from the target disk page. This value stores the original data information in the metadata block.

[0117] Not only that, if a tag value is detected for the target key during the metadata query process, the reference address associated with the tag value is obtained, and the compressed and stored encoding segment can be restored according to the reference address, thereby reducing storage space while ensuring accurate metadata query.

[0118] Through this embodiment, a corresponding query key is generated according to the received user demand data, and the coding tree information corresponding to the query key is obtained; the query key is encoded based on the coding tree information to obtain a target key; further, the target disk page corresponding to the target key is determined, and the relevant value corresponding to the target key is obtained from the target disk page, wherein the relevant value stores the original data information in the metadata block, so that when querying metadata, it is only necessary to obtain the corresponding disk page position to complete the query, avoiding scanning all disk pages, and improving the efficiency of metadata query.

[0119] In some embodiments, obtaining the encoding tree information corresponding to the query key includes the following steps:

[0120] Step S511, searching the coding tree information corresponding to the query key in the coding tree cache;

[0121] Step S512: When the coding tree information is not retrieved, extract the coding tree information from the metadata index node.

[0122] Specifically, when user demand data is received, a corresponding query key is generated according to the user demand data, the query key is sent to the encoder, and the encoder queries whether the coding tree buffer contains the coding tree information required by the query key.

[0123] If the cache hits, that is, the currently required coding tree information is retrieved in the coding tree cache, the coding tree cache sends the coding tree information required by the encoder to the encoder; if the cache misses, that is, the currently required coding tree information is not retrieved in the coding tree cache, the coding tree cache obtains the required coding tree information from the metadata index node and returns it to the encoder.

[0124] Through this embodiment, the coding tree information corresponding to the query key is retrieved in the coding tree cache. When the coding tree information is not retrieved, the coding tree information is extracted from the metadata index node, so that the target key that is compatible with the local metadata storage method can be encoded.

[0125] The present embodiment is described and illustrated below through optional embodiments.

[0126] FIG6 is a flowchart of a metadata storage method according to an alternative embodiment. As shown in FIG6 , the metadata storage method includes the following steps:

[0127] Step S610: Obtain the identity information in each metadata block; the identity information includes metadata identification, database identification, view identification, and domain identification;

[0128] Step S620: Encode the identity information in each metadata block to obtain a corresponding Huffman coding tree; wherein each level of the Huffman coding tree corresponds to a different type of identity information;

[0129] Step S630, obtaining each coding path in the Huffman coding tree;

[0130] Step S640: Acquire data feature information corresponding to the coding path; the data feature information includes version information and data area feature code;

[0131] Step S650, determining the code of each node in the coding path, and generating a corresponding target key based on the code and data feature information of each node;

[0132] In step S660, the original data information corresponding to each encoding path is used as a related value, and the target key is associated with the corresponding related value and stored.

[0133] Through this embodiment, identity identification information in each metadata block is obtained; the identity identification information includes a metadata identifier, a database identifier, a view identifier, and a domain identifier; the identity identification information in each metadata block is encoded to obtain a corresponding Huffman coding tree; wherein each level of the Huffman coding tree corresponds to a different category of identity identification information, effectively reducing the number of levels of the Huffman coding tree and improving the encoding and decoding capabilities of the metadata block.

[0134] Afterwards, each coding path in the Huffman coding tree is obtained, and the data feature information corresponding to the coding path is obtained, the data feature information including version information and data area feature coding; the coding of each node in the coding path is determined, and the corresponding target key is generated based on the coding and data feature information of each node; the original data information corresponding to each coding path is used as the relevant value, and the target key is associated with the corresponding relevant value and stored, thereby reducing the storage space required for each target key by converting the data information constituting the target key into Huffman coding, solving the problem of inefficient use of storage space, resulting in high data storage costs, and achieving reduced data storage costs.

[0135] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0136] In this embodiment, a metadata query system is also provided. FIG7 is a structural block diagram of the metadata query system of this embodiment. As shown in FIG7 , the system includes: a control module 100, an input module 200, an encoder 300, a coding tree cache 400, a metadata index node 500, a processor 600, and a metadata storage node 700;

[0137] When the input module 200 receives user demand data, the control module 100 controls it to generate a corresponding query key based on the user demand data, sends the query key to the encoder 300, and queries the coding tree cache 400 through the encoder 300 to see whether it contains the coding tree information required by the query key. If the cache hits, the coding tree cache 400 sends the coding tree information required by the encoder 300 to the encoder 300; if the cache misses, the coding tree cache 400 obtains the required coding tree information from the metadata index node 500 and returns it to the encoder 300.

[0138] After the encoder 300 obtains the coding tree information corresponding to the query key, it encodes the query key based on the coding tree information, generates a target key, and sends it to the processor 600. The coding tree information is used to provide information such as the coding format to generate a target key that is compatible with the local storage method.

[0139] Furthermore, processor 600 obtains the location of the target disk page corresponding to the target key from coding tree cache 400 and extracts the relevant value corresponding to the target key from the target disk page in metadata storage node 700. The relevant value stores the original data information in the metadata block. It should be noted that metadata storage node 700 is used to store all data nodes.

[0140] Through this embodiment, a corresponding query key is generated according to the received user demand data, and the coding tree information corresponding to the query key is obtained; the query key is encoded based on the coding tree information to obtain a target key; and the target disk page corresponding to the target key is determined, and the relevant value corresponding to the target key is obtained from the target disk page, where the relevant value stores the original data information in the metadata block, thereby improving the metadata query efficiency.

[0141] This embodiment also provides a metadata storage device for implementing the aforementioned embodiments and alternative implementations. Details already described will not be repeated. The terms "module," "unit," "subunit," etc., used below, may refer to a combination of software and / or hardware that implements a predetermined function. While the devices described in the following embodiments are preferably implemented using software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0142] FIG8 is a structural block diagram of the metadata storage device of this embodiment. As shown in FIG8 , the device includes: a construction module 10, a generation module 20, and a storage module 30;

[0143] A construction module 10 is configured to construct a Huffman coding tree corresponding to each metadata block based on the identity information in each metadata block;

[0144] A generating module 20 is used to generate a target key corresponding to each coding path in the Huffman coding tree and determine a target disk page corresponding to the target key;

[0145] The storage module 30 is configured to use the original data information corresponding to each encoding path as a related value, and associate the target key with the corresponding related value and store it in a target disk page.

[0146] Through the device provided by this embodiment, a Huffman coding tree corresponding to each metadata block is constructed based on the identity identification information in each metadata block; a target key corresponding to each coding path in the Huffman coding tree is generated, and a target disk page corresponding to the target key is determined; further, the original data information corresponding to each coding path is used as a related value, and the target key and the corresponding related value are associated and stored in the target disk page, which solves the problem of inefficient use of storage space and resulting in high data storage costs, thereby reducing data storage costs.

[0147] In some embodiments, based on Figure 8, the device also includes an encoding module for obtaining identity identification information in each metadata block; the identity identification information includes a metadata identifier, a database identifier, a view identifier, and a domain identifier; the identity identification information in each metadata block is encoded to obtain a corresponding Huffman coding tree; wherein each level of the Huffman coding tree corresponds to a different category of identity identification information.

[0148] In some of the embodiments, based on Figure 8, the device also includes a combination module for obtaining each coding path in the Huffman coding tree; obtaining data feature information corresponding to the coding path; the data feature information includes version information and data area feature coding; determining the coding of each node in the coding path, and generating a corresponding target key based on the coding and data feature information of each node.

[0149] In some of the embodiments, based on Figure 8, the device also includes an update module for generating a tag value corresponding to the same coding segment when it is detected that the target keys of each metadata block in the target disk page have the same coding segment; the tag value is associated with the reference address of the same coding segment; and the same coding segment in the metadata block is updated to the tag value.

[0150] In some of the embodiments, based on Figure 8, the device also includes a query module for generating a query key corresponding to a new metadata block when a new metadata block is received; obtaining the coding tree information corresponding to the query key, and encoding the query key based on the coding tree information to obtain a target key; using the original data information in the new metadata block as a related value; determining the data node corresponding to the target key, and associating the target key and the corresponding related value and storing them in the data node.

[0151] In this embodiment, a metadata query device is also provided. FIG9 is a structural block diagram of the metadata storage device of this embodiment. As shown in FIG9 , the device includes: an acquisition module 40, an encoding module 50, and a query module 60;

[0152] The acquisition module 40 is used to generate a corresponding query key according to the received user demand data and obtain the coding tree information corresponding to the query key;

[0153] An encoding module 50 is configured to encode the query key based on the encoding tree information to obtain a target key;

[0154] The query module 60 is used to determine the target disk page corresponding to the target key, and obtain the relevant value corresponding to the target key from the target disk page; the relevant value stores the original data information in the metadata block.

[0155] Through the device provided by this embodiment, a corresponding query key is generated according to the received user demand data, and the coding tree information corresponding to the query key is obtained; the query key is encoded based on the coding tree information to obtain a target key; further, the target disk page corresponding to the target key is determined, and the relevant value corresponding to the target key is obtained from the target disk page, wherein the relevant value stores the original data information in the metadata block, thereby improving the metadata query efficiency.

[0156] In some of the embodiments, based on FIG. 9 , the apparatus further includes a retrieval module for retrieving the coding tree information corresponding to the query key in the coding tree cache; and extracting the coding tree information from the metadata index node when the coding tree information is not retrieved.

[0157] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0158] This embodiment further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0159] Optionally, the computer device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0160] It should be noted that, for specific examples in this embodiment, reference may be made to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.

[0161] In addition, in conjunction with the metadata storage method provided in the above embodiments, a storage medium may be provided in this embodiment to implement the metadata storage method. The storage medium stores a computer program that, when executed by a processor, implements any of the metadata storage methods in the above embodiments.

[0162] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0163] Obviously, the accompanying drawings are merely examples or embodiments of the present application. A person skilled in the art can also apply the present application to other similar situations based on these drawings without inventive effort. Furthermore, it is understandable that, although the work involved in this development process may be complex and lengthy, certain design, manufacturing, or production changes based on the technical content disclosed in this application are merely routine technical means for a person skilled in the art and should not be considered to constitute a deficiency in the disclosure of the present application.

[0164] The term "embodiment" as used in this application refers to specific features, structures, or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily mean that the embodiment is the same, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is understood, either explicitly or implicitly, by those skilled in the art that the embodiments described in this application can be combined with other embodiments when there is no conflict.

[0165] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0166] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A metadata storage method, characterized in that: The method comprises: Based on the identity information in each metadata block, construct a Huffman coding tree corresponding to each metadata block; Generate a target key corresponding to each coding path in the Huffman coding tree, and determine a target disk page corresponding to the target key; The original data information corresponding to each encoding path is used as a related value, and the target key is associated with the corresponding related value and stored in the target disk page.

2. The metadata storage method according to claim 1, wherein: The step of constructing a Huffman coding tree corresponding to each metadata block based on the identity information in each metadata block includes: Acquire the identity information in each metadata block; the identity information includes a metadata identifier, a database identifier, a view identifier, and a domain identifier; Encoding the identity information in each metadata block to obtain the corresponding Huffman coding tree; Wherein, each level of the Huffman coding tree corresponds to different categories of the identity identification information.

3. The metadata storage method according to claim 1, wherein: The generating a target key corresponding to each coding path in the Huffman coding tree includes: Obtain each of the coding paths in the Huffman coding tree; Acquire data feature information corresponding to the coding path; the data feature information includes version information and data area feature code; The encoding of each node in the encoding path is determined, and based on the encoding of each node and the data feature information, the corresponding target key is generated.

4. The metadata storage method according to claim 1, wherein: After the target key and the corresponding related value are stored in association with each other in the target disk page, the method further includes: When it is detected that the target keys of the metadata blocks in the target disk page have the same coding segment, a tag value corresponding to the same coding segment is generated; the tag value is associated with a reference address of the same coding segment; The same coded segment in the metadata block is updated to the marker value.

5. The metadata storage method according to claim 1, wherein: After the target key and the corresponding related value are associated and stored in the target disk page, the method further includes: When a new metadata block is received, generating a query key corresponding to the new metadata block; Obtaining coding tree information corresponding to the query key, and encoding the query key based on the coding tree information to obtain a target key; Using the original data information in the new metadata block as the relevant value; The target disk page corresponding to the target key is determined, and the target key and the corresponding related value are associated and stored in the target disk page.

6. A metadata query method, characterized in that: The method comprises: Generate a corresponding query key according to the received user demand data, and obtain the coding tree information corresponding to the query key; Encode the query key based on the encoding tree information to obtain a target key; A target disk page corresponding to the target key is determined, and a related value corresponding to the target key is obtained from the target disk page; the related value stores original data information in a metadata block.

7. The metadata query method according to claim 6, wherein: The obtaining the coding tree information corresponding to the query key includes: Retrieving the coding tree information corresponding to the query key in the coding tree cache; When the coding tree information is not retrieved, the coding tree information is extracted from the metadata index node.

8. A metadata storage device, characterized in that: The device comprises: a construction module, a generation module and a storage module; The construction module is used to construct a Huffman coding tree corresponding to each metadata block based on the identity identification information in each metadata block; The generating module is used to generate a target key corresponding to each coding path in the Huffman coding tree, and determine a target disk page corresponding to the target key; The storage module is used to use the original data information corresponding to each encoding path as a related value, and associate the target key with the corresponding related value and store it in the target disk page.

9. A computer device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the steps of the metadata storage method according to any one of claims 1 to 5.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the metadata storage method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Real-time database data hotspot balancing method and device, equipment and medium

    CN112835896A

  • Mass space POI search method and system based on multi-factor constraint

    CN112948717A

  • Data processing method and device, electronic equipment and storage medium

    CN116932467A

  • Metadata storage and query method and device, computer equipment and storage medium

    CN117435776A

  • Data processor, data storage device, data processing method, data storage method and program

    JP2012164031A

Cited By

  • Interstellar molecular database construction method and device, interstellar molecular database query method and device and computer equipment

    CN120727160A