Metadata management method and data system

By introducing multiple metadata service nodes and one metadata management node in the computing cluster system, the problem of low metadata management efficiency is solved and more efficient data access is achieved.

CN120104570APending Publication Date: 2025-06-06CHENGDU HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311660581.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In computing cluster systems, the management efficiency of metadata is low, resulting in low data access efficiency.

Method used

By introducing multiple metadata service nodes and one metadata management node in the data system, the metadata management node is responsible for obtaining query requests, determining the target metadata service node, and forwarding the query request to the corresponding metadata service node to obtain and update the metadata.

Benefits of technology

This method reduces the load on a single node, improves the amount of metadata cache and processing efficiency, and thus improves the access efficiency of the data system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104570A_ABST
    Figure CN120104570A_ABST
Patent Text Reader

Abstract

The invention provides a metadata management method which is applied to a data system, the data system comprises a plurality of metadata service nodes and a metadata management node, each metadata service node is used for managing at least one metadata partition, and each metadata partition corresponds to part of metadata of data in the data system. The method comprises the steps that a metadata management node obtains a first query request, and a target metadata service node where a metadata partition corresponding to target data is located is determined; the metadata management node sends a second query request to the target metadata service node; and the first metadata service node in the target metadata service node obtains metadata corresponding to the target data in the metadata partition managed by the first metadata service node, and obtains the target data according to the metadata. An existing metadata storage and management method in a data system is low in efficiency, the method is high in metadata file access efficiency, and the load stability of the data system is good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of computing clusters, and more specifically, to a method and a data system for managing metadata in a computing cluster. Background Art

[0002] In computing cluster systems, such as big data clusters, high-performance computing clusters, or AI clusters, metadata plays a vital role in data access due to the large amount of data processed. Metadata refers to descriptive information about data, including the source, format, structure, and relationship of the data. It provides a basis for data management and analysis, enabling applications to manage and utilize data more efficiently. Therefore, how to manage the metadata of computing cluster systems and improve the access efficiency of data in computing cluster systems is a technical problem that needs to be solved urgently. Summary of the invention

[0003] The present application provides a metadata management method and a data system, in order to improve the access efficiency of the data system.

[0004] In a first aspect, an embodiment of the present application provides a metadata management method, which is applied to a data system, wherein the data system includes multiple metadata service nodes and metadata management nodes, each metadata service node is used to manage at least one metadata partition, and each metadata partition corresponds to part of the metadata of data in the data system. The method includes: the metadata management node obtains a first query request;

[0005] The metadata management node determines at least one target metadata service node where the metadata partition corresponding to the first target data is located based on data information of the first target data corresponding to the first query request; the metadata management node sends a second query request generated based on the first query request to the target metadata service nodes respectively; a first metadata service node among the at least one target metadata service node obtains metadata corresponding to the first target data in the metadata partition managed by the first metadata service node based on the second query request, and obtains the first target data based on the metadata.

[0006] In this embodiment, multiple metadata service nodes are responsible for processing transaction requests for target data, thereby reducing the load on a single node. Multiple metadata service nodes are jointly responsible for metadata management, and each metadata service node is only responsible for a portion of the partition range, which reduces the demand for single-node computing resources and improves the overall metadata cache and processing efficiency, thereby improving the access efficiency of the data system.

[0007] In a possible implementation manner of the first aspect, each metadata service node stores metadata corresponding to a managed metadata partition.

[0008] In a possible implementation manner of the first aspect, the metadata management node also serves as a metadata service node.

[0009] In a possible implementation of the first aspect, the metadata management node stores partition information of the metadata partition managed by each metadata service node, and the partition information includes data information corresponding to the metadata partition; the metadata management node determines at least one target metadata service node where the metadata partition corresponding to the first target data is located based on the data information of the first target data corresponding to the first query request, including: the metadata management node determines the target metadata service node and the metadata partition corresponding to the target metadata service node based on the data information of the first target data corresponding to the first query request.

[0010] The metadata management node is used to manage the correspondence between data and partitions.

[0011] In a possible implementation of the first aspect, the metadata corresponding to the metadata partition includes a partition file, and the partition file records information about a data file corresponding to the metadata partition. The first metadata service node obtains the first target data according to the metadata, including: the first metadata service node determines the data file including the first target data according to the data information of the first target data requested in the second query request and the information about the data file recorded in the partition file.

[0012] In a possible implementation of the first aspect, the metadata corresponding to the metadata partition includes multiple partition range lists, each partition range list records information of multiple partition files. The first metadata service node obtains the first target data according to the metadata, including: the first metadata service node determines the first target partition file according to the data information of the first target data requested in the second query request and the information of the partition file recorded in the partition range list, and determines the data file including the first target data according to the information of the data file recorded in the partition file.

[0013] In a possible implementation of the first aspect, the metadata corresponding to the metadata partition includes multiple partition definition lists and multiple snapshots, each partition definition list records information of multiple partition range lists, and each snapshot includes information of a partition definition list. The metadata management node sends a second query request generated according to the first query request to the target metadata service node, including: the metadata management node obtains the target partition definition list corresponding to the first target data according to the data information and snapshot of the first target data requested in the first query request; the metadata management node determines the information of the target partition range list according to the data information of the first target data requested in the first query request and the information of the partition range list recorded in the target partition definition list, wherein the second query request includes the information of the target partition range list.

[0014] In a possible implementation manner of the first aspect, the metadata service node obtains the target partition range list according to information of the target partition range list; and determines the first target partition file according to information of the partition files recorded in the target partition range list.

[0015] In a possible implementation of the first aspect, the method also includes: the metadata management node obtains a first update request; the metadata management node determines at least one target metadata service node where the metadata partition corresponding to the second target data is located based on data information of the second target data corresponding to the first update request; the metadata management node sends a second update request generated based on the first update request to the target metadata service nodes respectively; a second metadata service node among the at least one target metadata service node updates the second target data based on the second update request, and updates the metadata corresponding to the second target data in the metadata partition managed by the second metadata service node.

[0016] In a possible implementation of the first aspect, partition information of the metadata partition managed by each metadata service node is stored in the metadata management node, and the partition information includes data information corresponding to the metadata partition; the metadata management node determines at least one target metadata service node where the metadata partition corresponding to the second target data is located based on the data information of the second target data corresponding to the first update request, including: the metadata management node determines the target metadata service node and the metadata partition corresponding to the target metadata service node based on the data information of the second target data corresponding to the first update request.

[0017] In a possible implementation of the first aspect, the second metadata service node updates the second target data, and updates the metadata corresponding to the second target data in the metadata partition managed by the second metadata service node, including: the second metadata service node updates the data file according to the data information of the second target data; updates the second target partition file corresponding to the second target data according to the information of the updated data file; and updates the corresponding partition range list according to the information of the updated second target partition file.

[0018] In a possible implementation manner of the first aspect, the method further includes: the metadata management node updates the corresponding partition definition list and the corresponding snapshot according to the information of the updated partition range list.

[0019] Since the number of partition files and partition range lists is determined in advance, after executing multiple update requests, each metadata service node still only manages a certain number of metadata files. This improves the access efficiency of the file system compared to the method of creating a new list for each update request.

[0020] When executing an update request, multiple metadata service nodes update metadata at each level respectively, so that the status of the metadata is consistent with the data, which facilitates the subsequent execution of query requests.

[0021] In a second aspect, an embodiment of the present application provides a data system, comprising multiple metadata service nodes and metadata management nodes, each metadata service node being used to manage at least one metadata partition, each metadata partition corresponding to part of the metadata of the data in the data system; the metadata management node being used to obtain a first query request; the metadata management node being used to determine a target partition range based on the target partition, wherein the target partition is a partition of the target data, and the target partition range includes the target partition; the metadata management node being used to determine at least one target metadata service node where the metadata partition corresponding to the first target data is located based on data information of the first target data corresponding to the first query request; the target metadata service node being used to send a second query request generated based on the first query request to the target metadata service node respectively; a first metadata service node among the at least one target metadata service node being used to obtain metadata corresponding to the first target data in the metadata partition managed by the first metadata service node based on the second query request, and obtain the first target data based on the metadata.

[0022] In a possible implementation manner of the second aspect, each metadata service node stores metadata corresponding to a managed metadata partition.

[0023] In a possible implementation manner of the second aspect, the metadata management node also serves as a metadata service node.

[0024] In a possible implementation of the second aspect, the metadata management node stores partition information of the metadata partition managed by each metadata service node, and the partition information includes data information corresponding to the metadata partition; the metadata management node is used to determine the target metadata service node and the metadata partition corresponding to the target metadata service node based on the data information of the first target data corresponding to the first query request.

[0025] In a possible implementation of the second aspect, the metadata corresponding to the metadata partition includes a partition file, which records information about the data file corresponding to the metadata partition; the first metadata service node is used to determine the data file including the first target data based on the data information of the first target data requested in the second query request and the information about the data file recorded in the partition file.

[0026] In a possible implementation of the second aspect, the metadata corresponding to the metadata partition includes multiple partition range lists, each partition range list records information of multiple partition files; the first metadata service node is used to determine the first target partition file based on the data information of the first target data requested in the second query request and the information of the partition file recorded in the partition range list, and determine the data file including the first target data based on the information of the data file recorded in the partition file.

[0027] In a possible implementation of the second aspect, the metadata corresponding to the metadata partition includes multiple partition definition lists and multiple snapshots, each partition definition list records information of multiple partition range lists, and each snapshot includes information of a partition definition list. The metadata management node is used to obtain the target partition definition list corresponding to the first target data according to the data information and snapshot of the first target data requested in the first query request; the metadata management node is used to determine the information of the target partition range list according to the data information of the first target data requested in the first query request and the information of the partition range list recorded in the target partition definition list, wherein the second query request includes the information of the target partition range list.

[0028] In a possible implementation manner of the second aspect, the metadata service node obtains the target partition range list according to information of the target partition range list; and determines the first target partition file according to information of the partition files recorded in the target partition range list.

[0029] In a possible implementation of the second aspect, the metadata management node is used to obtain a first update request; the metadata management node is used to determine at least one target metadata service node where the metadata partition corresponding to the second target data is located based on data information of the second target data corresponding to the first update request; the metadata management node is used to send a second update request generated according to the first update request to the target metadata service nodes respectively; and a second metadata service node among the at least one target metadata service node is used to update the second target data according to the second update request, and update the metadata corresponding to the second target data in the metadata partition managed by the second metadata service node.

[0030] In a possible implementation of the second aspect, the metadata management node stores partition information of the metadata partition managed by each metadata service node, and the partition information includes data information corresponding to the metadata partition; the metadata management node is used to determine the target metadata service node and the metadata partition corresponding to the target metadata service node based on the data information of the second target data corresponding to the first update request.

[0031] In a possible implementation of the second aspect, the second metadata service node is used to: update the data file according to the data information of the second target data; update the second target partition file corresponding to the second target data according to the information of the updated data file; and update the corresponding partition range list according to the information of the updated second target partition file.

[0032] In a possible implementation manner of the second aspect, the metadata management node is further used to update the corresponding partition definition list and the corresponding snapshot according to the information of the updated partition range list.

[0033] This embodiment provides a metadata service composed of multiple metadata service nodes and a partition-based metadata structure, which has high access efficiency to metadata files and good load stability of the data system. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a structural diagram of a big data system in related technologies.

[0035] Figure 2 It is a big data system structure diagram of an embodiment of the present application.

[0036] Figure 3 It is a schematic flow chart of a metadata management method according to an embodiment of the present application.

[0037] Figure 4 It is a schematic diagram of the correspondence between a partition range and metadata in an embodiment of the present application.

[0038] Figure 5 It is a metadata structure diagram of an embodiment of the present application.

[0039] Figure 6 This is an update request execution flow chart of an embodiment of the present application.

[0040] Figure 7 It is a metadata structure diagram of an embodiment of the present application.

[0041] Figure 8 This is a query request execution flow chart of an embodiment of the present application.

[0042] Fig. 9 It is a schematic block diagram of a metadata management device according to an embodiment of the present application.

[0043] Fig.10 It is a schematic block diagram of a controller according to an embodiment of the present application. DETAILED DESCRIPTION

[0044] The present application will present various aspects, embodiments or features around a system including multiple devices, components, modules, etc. It should be understood and appreciated that each system may include additional devices, components, modules, etc., and / or may not include all devices, components, modules, etc. discussed in conjunction with the figures. In addition, combinations of these schemes may also be used.

[0045] In addition, in the embodiments of the present application, words such as "exemplary" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present concepts in a concrete way.

[0046] In the embodiments of the present application, "corresponding (corresponding, relevant)" and "corresponding (corresponding)" can sometimes be used interchangeably. It should be pointed out that when the distinction between them is not emphasized, the meanings they intend to express are consistent.

[0047] The business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. A person of ordinary skill in the art can appreciate that, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0048] The following is a brief introduction to the commonly used technical terms in this field.

[0049] Metadata is data that describes data. It is information about the characteristics, structure, and context of data, rather than the actual data content. Metadata can include the attributes, format, date, author, permissions, connection relationships, and other related information of data objects to help users understand the purpose, source, quality, business rules, etc. of the data. Due to the above properties, metadata is often used to manage the structure of data systems.

[0050] Splits represent a reference to a part of a data file in a data system. It contains the metadata information of the file, the reference method and area, etc. Therefore, a slice can be regarded as the metadata of a part of a file.

[0051] A metadata snapshot is a metadata capture and storage of a data set at a specific point in time. It records the state, attributes, and configuration information of the data so that the data state at that specific point in time can be restored when needed. Metadata snapshots are usually used in backup and recovery operations, as well as in areas such as system monitoring and version control. In this embodiment, the metadata snapshot is referred to as a snapshot.

[0052] Distributed File Storage System: In a distributed file storage system, files are divided into multiple blocks or partitions and distributedly stored on multiple nodes. These nodes can be physical servers, virtual machines, or containers, etc. By storing file blocks on multiple nodes, the distributed file storage system achieves redundant replication and load balancing of data, thereby improving data reliability and performance. Common distributed file storage systems include HDFS, Google File System, Ceph, GlusterFS, and Lustre, etc. In each embodiment of the present application, the distributed file storage system is referred to as a file system.

[0053] Data files refer to data files in a data system, which are stored in a distributed file storage system. Common data files are columnar storage format files, also known as data tables. Specific file formats include ORC files or Parquet files. Data files generally store multiple records, and these records can be stored in a specific order based on the characteristics of the record fields or attributes.

[0054] Partition refers to a method of dividing data according to a certain logic, and can also be regarded as a logical view of data. When querying a partitioned data set, you can scan only the partitions that meet the conditions according to the query logic, instead of scanning the entire data set, thereby improving query efficiency.

[0055] The Consistent Hashing Algorithm, often referred to as the Hash Ring, is a data partitioning and routing strategy for distributed systems. The specific working principle is as follows: The Hash Ring calculates the hash value of the data through a hash function as the partition identifier of the data, where the hash value is mapped to a fixed-size ring hash space. Then, the identifiers of each physical node and the virtual node it manages are hashed through the same hash function and used as partition identifiers. When new data needs to be stored, the hash value of the data is calculated and the nearest virtual node clockwise on the hash ring is found, and the new data is stored by the virtual node. Since the virtual nodes are evenly distributed on the hash ring, the data can be relatively evenly distributed on each physical node. When adding or deleting a physical node, only the affected part of the data needs to be remapped without causing a large amount of data migration to the entire system.

[0056] The data system of the embodiment of the present application can be a big data system, a high-performance computing system, or an AI system, but the present application is not limited to this. The following uses a big data system as an example to illustrate the metadata management method and data system provided by the present application. When updating or querying target data in a big data system, it is usually necessary to first obtain the corresponding metadata to determine the various attributes of the data and storage location information. Figure 1 The traditional big data system shown includes a coordination node, a computing node, and a distributed file storage system. The coordination node includes a metadata service for managing metadata of the big data system.

[0057] In this technology, if a single point failure occurs in the coordination node, the metadata will be unavailable and need to be reloaded after the failure is resolved. During the failure, the business applications of the entire big data system are affected. In addition, the memory of the coordination node is limited and cannot cache a large number of metadata files; the number of CPU cores of the coordination node is limited, and when too many metadata files are loaded, a large amount of CPU computing resources are consumed. These problems cause the business applications of the entire big data system to have a high load and take a long time, sometimes causing blockage or data unavailability, affecting the access efficiency of the big data system.

[0058] In view of this, an embodiment of the present application provides a method for metadata management, in which multiple metadata service nodes are jointly responsible for metadata management, thereby reducing the bottleneck impact of single-node metadata management and improving the access efficiency of the big data system.

[0059] Figure 2 Schematic diagram of the big data system of the present application embodiment. Figure 2 As shown, the big data system includes a coordination node, multiple computing nodes and a distributed file storage system. Specifically, an instance of a metadata service node can be run on the computing node.

[0060] The distributed file storage system includes multiple file storage nodes, each of which stores multiple data files. It can be considered that all data stored in the big data system constitutes a total data file, and these data files in the disk are part of the total data file, so these data files can be called data. This embodiment manages the metadata of the data by metadata partitioning. Metadata partitioning is a logical view used to classify metadata according to features and is independent of the specific form of the file storage system. Unless otherwise specified, metadata partitions are referred to as partitions.

[0061] For example, in one embodiment, the data in the big data system is divided into four fields, namely order number, order date, user name and city, then the metadata can be partitioned according to the order date field, for example, the metadata of January 2021 corresponds to partition P101, and the order date of all metadata in partition P101 is January 2021, and the metadata of October 2022 corresponds to partition P210, and the order date of all metadata in partition P210 is October 2022. In this case, the name of the order date field in the metadata is called the partition key, and the name of the order date field is called the partition value, for example, the partition value of the metadata is February 2023. The correspondence between the partition key and the partition is recorded in the partition table.

[0062] Metadata can also be partitioned according to "order date, city" as the partition key. For example, the partition value of one set of metadata is "September 2022, Nanjing", and the partition key of another set of metadata is "December 2021, New York". Compared with the partitioning method using one field, the partitioning method using two fields can divide the metadata into more partitions, and each partition contains less metadata.

[0063] When a user sends a transaction request to a big data system, the transaction request includes data information of the target data, including a partition key and a partition value of metadata of the target data, and may also include a snapshot identifier of the target data.

[0064] In one embodiment, the big data system includes multiple metadata service nodes and metadata management nodes, each metadata service node is used to manage at least one metadata partition, and each metadata partition corresponds to part of the metadata of the data in the big data system. In one embodiment, the partitions can be divided into multiple partition ranges according to their numbers, such as partitions P1 to P100 forming partition range L1, and partitions P101 to P200 forming partition range L2. Multiple metadata service nodes respectively manage metadata of data in different partition ranges, wherein a metadata service node can manage metadata of data in one or more partition ranges. In addition, multiple metadata service nodes can elect a metadata service node as a metadata management node, and it should be understood that the metadata management node is also used as the metadata service node.

[0065] Figure 3 is a schematic flow chart of a metadata management method provided in an embodiment of the present application. The method can be applied to Figure 2 Specifically, the big data system performs the following Figure 3 Steps shown.

[0066] S110: The metadata management node obtains a first transaction request.

[0067] Specifically, the metadata management node obtains a first transaction request from the coordination node, and the first transaction request may be an update request or a query request for target data.

[0068] S120: The metadata management node determines, according to data information of the target data corresponding to the first transaction request, at least one target metadata service node where the metadata partition corresponding to the target data is located.

[0069] The metadata management node stores the partition range information managed by each metadata service node. The partition range information includes the partition information of the metadata partition. The partition information includes the data information corresponding to the metadata partition.

[0070] Specifically, the metadata management node can determine the target metadata service node and the metadata partition corresponding to the target metadata service node according to the data information of the target data corresponding to the first transaction request. Specifically, a partition table can be created in the big data system, the partition field can be specified, and the mapping relationship between the partition value and the partition can be specified. In other words, the partition table can be managed by the metadata service node, and the metadata management node is used to manage the corresponding relationship between data and partitions.

[0071] In one possible embodiment, partitions P1 to P100 are made up of partition range L1, and partitions P101 to P200 are made up of partition range L2, and the partition range information includes the corresponding relationship between such partitions and partition ranges. The partition information includes data such as the partition key, all partition values, and the latest snapshot identifier of the partition. The metadata management node requests the data information of the target data corresponding to the first transaction. According to the data information, it can be obtained that the modification time of part of the target data is December 2020; according to the partition table, the partition value "December 2020" corresponds to partition P12, and the modification time of another part of the data is May 2022, corresponding to partition P205, that is, the target data corresponds to partition ranges L1 and L3. Assuming that at this time, metadata service node N1 manages partition ranges L1 to L2, metadata service node N2 manages partition ranges L3 to L6, and metadata service node N3 manages partition range L7, then the target metadata service nodes are metadata service nodes N1 and N2. The way in which the metadata service node and the partition range are established will be further described in subsequent embodiments.

[0072] S130, the metadata management node sends a second transaction request generated according to the first transaction request to the target metadata service node respectively.

[0073] Specifically, the target metadata service node is the metadata service node corresponding to the target partition range. The second transaction request includes partial information of the first transaction request. For example, the target data includes data of partition range L1 and partition range L5, metadata service node N1 manages partition ranges L1 and L2, and metadata service node N3 manages partition ranges L5 and L6, then the target metadata service nodes are metadata service nodes N1 and N3.

[0074] S140: The target metadata service node processes the second transaction request according to the metadata corresponding to the target data.

[0075] Each metadata service node manages one or more partition ranges. Specifically, the metadata service node is responsible for reading and updating data files or metadata files within these partition ranges; and managing the physical location information of the partition ranges or partitions, such as the IP address, port number, disk location, etc. of the corresponding server.

[0076] Each metadata service node processes the second transaction request according to the metadata of the data in the corresponding partition range. For example, the first metadata service node among at least one target metadata service node obtains the metadata corresponding to the target data in the metadata partition managed by the first metadata service node according to the second transaction request, and obtains the target data according to the metadata; or, the second metadata service node among at least one target metadata service node updates the target data according to the second transaction request, and updates the metadata corresponding to the target data in the metadata partition managed by the second metadata service node. The specific query or update process will be further described in subsequent embodiments.

[0077] In this embodiment, multiple metadata service nodes are responsible for processing transaction requests for target data, thereby reducing the load on a single node. Multiple metadata service nodes are jointly responsible for metadata management, and each metadata service node is only responsible for a portion of the partition range, which reduces the demand for single-node computing resources and improves the overall metadata cache and processing efficiency, thereby improving the access efficiency of the data system.

[0078] The following is an example of how the metadata service node manages the partition range. In one embodiment, the metadata management node can calculate the hash values ​​of multiple metadata service nodes and the hash values ​​of multiple partition ranges respectively through a hash function, and set the same hash value interval as the number of metadata service nodes according to the hash values ​​of the multiple metadata service nodes. Each metadata service node corresponds to a hash value interval. The metadata service node manages the partition range whose hash value is located in the hash value interval corresponding to the metadata service node.

[0079] In one embodiment, Figure 4As shown, the metadata management node is responsible for managing a hash ring. The multiple metadata service nodes of this embodiment are distributed on the hash ring as storage nodes of the hash ring, and the partition ranges are data nodes of the hash ring. For the convenience of explanation, this embodiment numbers the multiple metadata service nodes in the order of clockwise arrangement on the hash ring, and each node manages one or more partition ranges. For example, the partition range P301~P400 is called the partition range L4, the partition range P401~P500 is called the partition range L5, the partition range P501~P600 is called the partition range L6, and so on.

[0080] In this embodiment, there are 3 storage nodes on the hash ring, namely metadata service nodes N4, N6 and N8, and their hash values ​​are 400, 600 and 800 respectively; in addition, there are 8 data nodes on the hash ring, namely partition ranges L1 to L8, and their hash values ​​are 50, 150, 250, 350, 450, 550, 650 and 750 respectively. The metadata management node is set with the same hash value interval as the number of metadata service nodes, where metadata service node N4 corresponds to the hash value interval 0 to 400, metadata service node N6 corresponds to the hash value interval 401 to 600, and metadata service node N8 corresponds to the hash value interval 601 to 800. Figure 4 As shown in the figure, each metadata service node and each partition range are arranged in order on the hash ring according to their own hash values. Starting from any metadata service node, move forward in a counterclockwise direction on the hash ring until you encounter other metadata service nodes. The range you pass through is the hash value interval corresponding to the arbitrary metadata service node.

[0081] In this embodiment, the metadata service node manages the partition range of the hash value interval corresponding to the metadata service node. Figure 4 As shown, the metadata service node to which each partition range belongs is the first metadata service node encountered by the partition range in the clockwise direction on the hash ring. Specifically, metadata service node N4 manages partition ranges L1, L2, L3, and L4, metadata service node N6 manages partition ranges L5 and L6, and metadata service node N8 manages partition ranges L7 and L8. When metadata service node N6 fails, partition ranges L5 and L6 will be assigned to metadata service node N8 for management, as shown in FIG. Figure 4 As shown by the dotted arrows of the partition ranges L5 and L6 in FIG, at this time, the metadata service node N4 manages the partition ranges L1 to L4, and the metadata service node N8 manages the partition ranges L5 to L8. In other words, the partition ranges remain unchanged, and each metadata service node can manage one or more partition ranges.

[0082] By managing the correspondence between metadata service nodes and partition ranges through a hash ring, the metadata cache capacity of multiple metadata service nodes can be evenly distributed. When a single node fails, the load of multiple metadata service nodes can be evenly redistributed, thereby improving the access efficiency of the data system.

[0083] In one embodiment of the present application, the metadata of the data in the big data system may include a partition file, a partition range list, a partition definition list, and a snapshot. The structure is as follows: Figure 5 shown.

[0084] First, the metadata corresponding to the metadata partition includes a partition file, which records the information of the data file corresponding to the metadata partition. Specifically, the information of the data file may include the file path, file size, modification time, partition identifier and snapshot identifier of the data file.

[0085] The file path in the big data system may include: the IP address and port of the server where the metadata service node is located, the local path of the data file on the server disk, and the identifier of the server in the server cluster, which is not limited in this embodiment.

[0086] This embodiment only involves the management of various metadata, and does not limit the storage location. For example, partition files can be stored in file storage nodes or metadata service nodes, which is not limited in this embodiment.

[0087] Furthermore, the metadata corresponding to the metadata partition includes a plurality of partition range lists, each of which records information of a plurality of partition files, wherein the information of the partition file may include a snapshot identifier and a file path of the partition file.

[0088] Similarly, the partition range list may be stored in a file storage node or a metadata service node, which is not limited in this embodiment. That is, each metadata service node may store metadata corresponding to the managed metadata partition.

[0089] The metadata corresponding to the metadata partition includes multiple partition definition lists and multiple snapshots, each partition definition list records information of multiple partition range lists, and each snapshot includes information of a partition definition list. Specifically, the newly generated partition definition list records information of all partition range lists under the current snapshot.

[0090] In the related art, the metadata service manages metadata in the following way: each update request creates one or more data files, a manifest and a manifest list, the manifest list records the information of the manifest, and the manifest records the information of the data file. Accordingly, each time a query request is executed, multiple manifests need to be frequently read and written. For example, when executing query transaction 938, the metadata service needs to locate and read 938 files from manifests 1 to 938; when executing query transaction 10429, the metadata service needs to locate and read 10429 files from manifests 1 to 10429. The access efficiency of the metadata service to the file system continues to decrease.

[0091] In this embodiment, since the number of partition files and partition range lists is determined in advance, after executing multiple update requests, each metadata service node still only manages a determined number of metadata files, thereby improving the access efficiency of the file system.

[0092] The following uses a specific update request and a specific query request as examples to illustrate how the metadata management node and the metadata service node manage metadata at each layer.

[0093] First combine Figure 5 and Figure 6 The update request process S200 is introduced. It should be understood that the update request can be a request to add new data or a request to modify existing data.

[0094] S210: The metadata management node obtains a first update request.

[0095] The first update request is usually received from the coordination node, and includes a new data file and its data information. The new data file is the target data, and the data information may include a partition key, a partition value, and a snapshot identifier of the metadata of the target data.

[0096] S220: The metadata management node sends a second update request to the target metadata service node.

[0097] The metadata management node generates a second update request based on the partition information of the metadata partition managed by each metadata service node, the data information of the target data requested in the first update request, and the information of the partition range list recorded in the partition definition list. The second update request includes the target data and the target partition, that is, the data information of the target data requested in the second update request includes the target data itself and the target partition. In one embodiment, the partition key of the metadata of a new data file is the modification time column, and the partition value is March 29. According to the partition table, the data file corresponds to partition P329, and P329 belongs to P301~P400, that is, the partition range L4, so the metadata service node corresponding to the data file manages the partition range L4, such as metadata service node N6 or N8. The metadata management node sends a second update request to the metadata service node that manages the partition range L4, and the target data of the second update request is the data file with the partition identifier P329.

[0098] In another embodiment, the second update request updates multiple new data files, and the metadata management node sends the second update request to the target metadata service node. For example, the target data includes three data files with target partitions P258, P113, and P759, and P258 belongs to P201-P300, that is, the partition range L3. In the second update request sent by the metadata management node to the metadata service node corresponding to the partition range L3, the target data only includes the data files with the target partition P258, and does not include the data files with the target partitions P113 or P759, thereby improving the transmission efficiency.

[0099] S230: The target metadata service node processes the second update request according to the target data.

[0100] The metadata service node updates the data file according to the target data in the metadata partition it manages and the data information of the target data in the second update request. Figure 5 Taking the data file D4-62 shown in as an example, data file D4-62 is created in partition P4 during transaction 62.

[0101] Furthermore, the metadata service node updates the target partition file corresponding to the target data according to the information of the updated data file. In this embodiment, the target metadata service node finds the partition file P4 according to the records in the partition range list L1, and appends one or more pieces of information in the partition file P4 for the data file D4-62. If the partition does not contain a data file and there is no partition file, a partition file is created and the information of the new data file is recorded, otherwise the information of the new data file is directly appended to the partition file. It should be understood that the data file D4-62 in this embodiment can be one data file or multiple data files, and their metadata and the data contained therein are different. In addition, the specific method of finding the partition file P4 is further described in subsequent embodiments.

[0102] Furthermore, the metadata service node updates the corresponding partition range list according to the information of the updated target partition file. Figure 5 Taking the data file D219-57 shown in as an example, the current snapshot identifier is 57. This update request writes the data file in partition P219 for the first time. The existing records in the partition range list L3 are the information of partition files P201~P218 and partition files P220~P300. Then, the information of partition file P219 is appended to the partition range list L3, where the snapshot identifier in the information of partition file P219 is 57.

[0103] The target metadata service node can also update the information of existing partition files in the partition range list. When the partition file appends the information of the data file, the number of records, file size, modification time and other information of the partition file will change, and the metadata service node updates the information of these partition files in the partition range list.

[0104] S240, the metadata management node creates a new partition definition list and records information of all partition range lists after the update operation.

[0105] The metadata management node records the information of all current partition range lists in the partition definition list. For example, the current snapshot identifier is 57, and the data files updated by transaction 57 are distributed in partitions P323 and P678. Specifically, partition definition list 57 is created. Since partitions P323 and P678 belong to partition ranges L4 and L7 respectively, and partition range list L8 also exists, the information of partition range lists L4, L7, and L8 is appended to partition definition list 57.

[0106] S250: The metadata management node creates a new snapshot and records information about a new partition definition list.

[0107] Specifically, the metadata management node creates a corresponding partition definition list and a corresponding snapshot according to the information of the updated partition range list.

[0108] When executing an update request, multiple metadata service nodes update metadata at each level respectively, so that the state of the metadata is consistent with the state of the data, which facilitates the subsequent execution of query requests.

[0109] Combine the following Figure 7 and Figure 8 The query request process S300 is introduced.

[0110] S310: The metadata management node obtains a first query request and determines a historical partition definition list corresponding to the target data according to a snapshot identifier.

[0111] The first query request is usually a request sent by a user, including data information of the target data and query conditions of the target data, and the data information includes a partition key, a partition value, and a snapshot identifier. The metadata service node determines the historical partition definition list based on the data information of the target data requested in the first query request and the records in the snapshot, and obtains the location and content of the historical partition definition list.

[0112] The query conditions can be a time range, a specific user name, a behavioral feature composed of multiple fields, etc., which are related to the specific business.

[0113] This embodiment takes the example of an application querying a user name whose modification time of snapshot 67 is after 2021. Figure 7 As shown, the metadata service node locates snapshot 67 according to the snapshot identifier 67 in the request, and then obtains the location and content of the historical partition definition list 67 according to the partition definition list information contained in snapshot 67.

[0114] S320: The metadata management node obtains the historical partition range corresponding to the target data according to the historical partition definition list, and sends a second query request to the target metadata service node.

[0115] The metadata management node determines the target partition range list information based on the data information of the target data requested in the first query request and the information of the partition range list recorded in the target partition definition list, and then sends the generated second query request to the corresponding metadata service node. The second query request includes the snapshot identifier, the target partition, the query condition of the target data, and the information of the target partition range list, that is, the data information of the target data requested in the second query request includes the target data itself, the snapshot identifier, the target partition, the query condition of the target data, and the information of the target partition range list.

[0116] In this embodiment, the partition definition list with snapshot identifier 67 records information of partition range lists L1 and L6, and does not record information of other partition range lists, that is, data files of historical snapshot 67 are all distributed in partition ranges L1 and L6.

[0117] Assuming that the current metadata service node N1 manages the partition range L1 and the metadata service node N6 manages the partition range L6, the metadata management node sends a second query request to the metadata service nodes N1 and N6.

[0118] S330, the metadata service node determines a target partition range list corresponding to the target data.

[0119] The metadata service node determines the target partition range list according to the information of the target partition range list in the second query request. In this embodiment, metadata service node N1 manages partition ranges L1 and L2, the target partition range is partition range L1, and the subsequent steps of metadata service node N1 will be performed in partition range L1. Metadata service node N6 manages partition ranges L4 to L6, the target partition range is partition range L6, and the subsequent steps of metadata service node N6 will be performed in partition range L6.

[0120] S340: The metadata service node determines a target partition file corresponding to the target data according to the target partition range list and the snapshot identifier.

[0121] The metadata service node determines the target partition file based on the information of the partition file recorded in the target partition range list. Figure 7 As shown, for the metadata service node N6, the target partitions are P556, P557 and P590, the partition range list L6 is read, and the partition files of these three partitions are located according to the file names in the recorded information of the partition files P556, P557 and P590.

[0122] like Figure 7 As shown, partition file P556 contains information updated at snapshot 39 and information updated at snapshot 44, and partition file P590 contains information updated at snapshot 5 and information updated at snapshot 6, and they all point to the data contained in snapshot 67. Partition file P557 only contains information updated at snapshot 154, so there is no need to locate the data file corresponding to partition file P557.

[0123] S350, the metadata service node determines the target data according to the target partition file, the snapshot identifier and the query condition.

[0124] The metadata service node N6 determines the data file including the target data according to the information of the data file recorded in the partition file. Figure 7As shown, the metadata of the data files with snapshot identifiers of 1 to 67 are filtered out according to the partition files P556 and P590, and it is found that four files D556-39, D556-44, D590-5 and D590-6 represent all the data in the user name column after the completion of transaction 67. The query condition of this embodiment is "user names with modification time after 2021". According to the file name in the information of the data in the partition files P556 and P590, the files corresponding to D556-39, D556-44, D590-5 and D590-6 are found in the storage system, and read respectively, and then the user name data contained in these four files are combined into a general table with N rows of data, and the target data is M rows of user name data after 2021, where M≤N.

[0125] It should be understood that other target metadata service nodes, such as metadata service node N1 responsible for managing partition range L1, also perform similar operations in parallel to obtain target data of all partitions.

[0126] S360, the metadata service node generates slices based on the target data and sends them to the metadata management node.

[0127] The specific form of the target data is a data file. The slice may include the information and index range of the data file. In this embodiment, four slices are generated to represent the data corresponding to the above-mentioned N rows of username data in each data file. Specifically, assuming that the N rows of username data in D556-44 include the 12th, 56th, 328th and 492nd rows of data of D556-44, the slice of D556-44 at least includes the information of the data file D556-44 and the two information of "the 12th, 56th, 328th and 492nd rows".

[0128] S370, the metadata management node sends the slice to the coordination node.

[0129] It should be understood that after each query request is executed, the coordination node will record the correspondence between the query request statement and the slice in the cache. When the query request is executed multiple times using the same query statement, the coordination node will directly read the corresponding slice from the cache.

[0130] S380, the coordination node reads the target data according to the slice.

[0131] Specifically, the coordination node reads target data from the distributed file storage system.

[0132] When executing a query request, multiple metadata service nodes determine the location of the target data based on the metadata, reducing unnecessary scanning in the file system and improving the access efficiency of the file system.

[0133] To summarize, metadata management is jointly completed by multiple metadata service nodes, and a single functional module does not need to cache all metadata files to ensure the load stability of the data system.

[0134] It should be understood that the related processes in the above query process and update process can refer to each other. For example, the process of searching for partition files in the update process can refer to the query process.

[0135] Fig. 9 It is a schematic block diagram of a metadata management device according to an embodiment of the present application. Fig. 9 The device 600 shown can be used to execute the metadata management method of the embodiment of the present application.

[0136] like Fig. 9 As shown, the device includes an acquisition unit 610, a processing unit 620 and an output unit 630. In multiple steps of any embodiment of the present application, the above three units need to be used simultaneously. For example, the output unit is required to store new data files and update partition files; when obtaining information of data files from data files, it is necessary to obtain part of the data in the data file through the acquisition unit, and it is necessary to execute a certain algorithm through the processing unit to extract valid information from these data.

[0137] The term "unit" here can be implemented in software and / or hardware form and is not specifically limited.

[0138] For example, a "unit" can be a software program, a hardware circuit, or a combination of the two that implements the above functions, and can include code running on a computing instance. Exemplarily, the following takes a processing unit as an example to introduce the implementation of the processing unit. Similarly, the implementation of the acquisition unit and the output unit can refer to the implementation of the processing unit.

[0139] Specifically, the processing unit may include codes running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with close geographical locations. Generally, a region may include multiple AZs.

[0140] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0141] As an example of a hardware functional unit, the processing unit may include at least one computing device, such as a server, etc. Alternatively, the processing unit may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0142] The multiple computing devices included in the processing unit can be distributed in the same region or in different regions. The multiple computing devices included in the processing unit can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the processing unit can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0143] Therefore, the modules of each example described in the embodiments of the present application can be implemented with electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present application.

[0144] The present application also provides a controller 700. Fig.10As shown, the controller 700 includes: a processor 704 and a communication interface 708. Further, the controller 700 may also include a bus 702 and a memory 706. It should be understood that the bus 702 and the memory 706 are optional. The processor 704, the memory 706 and the communication interface 708 communicate through the bus 702. Exemplarily, the controller 700 can be a computing device or a device in a computing device for implementing the method of the embodiment of the present application. The controller 700 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the controller 700.

[0145] The bus 702 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.10 The bus 704 is represented by only one line, but it does not mean that there is only one bus or one type of bus. The bus 704 may include a path for transmitting information between various components of the controller 700 (eg, the memory 706, the processor 704, and the communication interface 708).

[0146] The processor 704 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0147] The memory 706 may include a volatile memory, such as a random access memory (RAM). The memory 706 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0148] The memory 706 stores executable program codes, and the processor 704 executes the executable program codes to respectively implement the functions of the aforementioned acquisition unit and processing unit, thereby implementing the method for distributed data storage. That is, the memory 706 stores instructions for executing the method for distributed data storage.

[0149] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the controller 700 and other devices or communication networks (eg, multiple servers). The communication interface may also be referred to as an interface circuit.

[0150] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes a controller 700 and multiple computing devices. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop, a laptop, or a smart phone. The multiple computing devices can be the multiple servers mentioned above.

[0151] In a possible implementation, the controller 700 is a computing device among the multiple computing devices or is a device in the computing device for implementing the above method.

[0152] In another possible implementation, the controller 700 is another computing device other than the multiple computing devices or is a device in another computing device for implementing the above method.

[0153] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method in the present application embodiment.

[0154] The present application embodiment also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the method in the embodiment of the present application, or instruct the computing device to execute the method in the embodiment of the present application.

[0155] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0156] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0157] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0158] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0159] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0160] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0161] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0162] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A method for metadata management, Features: Applied to a data system, the data system includes a plurality of metadata service nodes and a metadata management node, each metadata service node is used to manage at least one metadata partition, each metadata partition corresponds to partial metadata of data in the data system; The method comprises: The metadata management node obtains a first query request; The metadata management node determines, according to data information of the first target data corresponding to the first query request, at least one target metadata service node where the metadata partition corresponding to the first target data is located; The metadata management node sends a second query request generated according to the first query request to the target metadata service node respectively; The first metadata service node among the at least one target metadata service node obtains metadata corresponding to the first target data in the metadata partition managed by the first metadata service node according to the second query request, and obtains the first target data according to the metadata.

2. The method for metadata management according to claim 1, Features: Each metadata service node stores metadata corresponding to the managed metadata partition.

3. The method for metadata management according to claim 1 or 2, Features: The metadata management node also serves as the metadata service node.

4. The method for metadata management according to any one of claims 1 to 3, Features: The metadata management node stores partition information of the metadata partition managed by each metadata service node, and the partition information includes data information corresponding to the metadata partition; The metadata management node determines, according to data information of the first target data corresponding to the first query request, at least one target metadata service node where the metadata partition corresponding to the first target data is located, including: The metadata management node determines the target metadata service node and the metadata partition corresponding to the target metadata service node according to data information of the first target data corresponding to the first query request.

5. The method for metadata management according to any one of claims 1 to 4, Features: The metadata corresponding to the metadata partition includes a partition file, and the partition file records information of a data file corresponding to the metadata partition; The first metadata service node obtains the first target data according to the metadata, including: The first metadata service node determines the data file including the first target data according to the data information of the first target data requested in the second query request and the information of the data file recorded in the partition file.

6. The method for metadata management according to claim 5, Features: The metadata corresponding to the metadata partition includes a plurality of partition range lists, each partition range list records information of a plurality of the partition files; The first metadata service node obtains the first target data according to the metadata, including: The first metadata service node determines the first target partition file based on the data information of the first target data requested in the second query request and the information of the partition file recorded in the partition range list, and determines the data file including the first target data based on the information of the data file recorded in the partition file.

7. The method for metadata management according to claim 6, Features: The metadata corresponding to the metadata partition includes multiple partition definition lists and multiple snapshots, each partition definition list records information of multiple partition range lists, and each snapshot includes information of one partition definition list; The metadata management node sends a second query request generated according to the first query request to the target metadata service node respectively, including: The metadata management node obtains a target partition definition list corresponding to the first target data according to the data information of the first target data requested in the first query request and the snapshot; The metadata management node determines the information of the target partition range list based on the data information of the first target data requested in the first query request and the information of the partition range list recorded in the target partition definition list, wherein the second query request includes the information of the target partition range list.

8. The method for metadata management according to claim 7, Features: The metadata service node obtains the target partition range list according to information of the target partition range list; and determines the first target partition file according to information of the partition files recorded in the target partition range list.

9. The method for metadata management according to any one of claims 1 to 8, It is characterized in that The method further comprises: The metadata management node obtains a first update request; The metadata management node determines, according to data information of the second target data corresponding to the first update request, at least one target metadata service node where the metadata partition corresponding to the second target data is located; The metadata management node sends a second update request generated according to the first update request to the target metadata service node respectively; A second metadata service node among the at least one target metadata service node updates the second target data according to the second update request, and updates the metadata corresponding to the second target data in the metadata partition managed by the second metadata service node.

10. The method for metadata management according to claim 9, Features: The metadata management node stores partition information of the metadata partition managed by each metadata service node, and the partition information includes data information corresponding to the metadata partition; The metadata management node determines, according to data information of the second target data corresponding to the first update request, at least one target metadata service node where the metadata partition corresponding to the second target data is located, including: The metadata management node determines the target metadata service node and the metadata partition corresponding to the target metadata service node according to data information of the second target data corresponding to the first update request.

11. The method for metadata management according to claim 10, Features: The second metadata service node updates the second target data, and updates metadata corresponding to the second target data in the metadata partition managed by the second metadata service node, including: The second metadata service node updates the data file according to the data information of the second target data; Update the second target partition file corresponding to the second target data according to the updated information of the data file; Update the corresponding partition range list according to the updated information of the second target partition file.

12. The method for metadata management according to claim 11, It is characterized in that The method further comprises: The metadata management node updates the corresponding partition definition list and the corresponding snapshot according to the updated information of the partition range list.

13. A data system, Features: The data system includes a plurality of metadata service nodes and metadata management nodes, each metadata service node is used to manage at least one metadata partition, and each metadata partition corresponds to partial metadata of the data in the data system; The metadata management node is used to obtain a first query request; The metadata management node is used to determine a target partition range according to a target partition, wherein the target partition is a partition of the target data, and the target partition range includes the target partition; The metadata management node is used to determine at least one target metadata service node where the metadata partition corresponding to the first target data is located according to data information of the first target data corresponding to the first query request; The target metadata service node is used to send a second query request generated according to the first query request to the target metadata service node respectively; The first metadata service node among the at least one target metadata service node is used to obtain metadata corresponding to the first target data in the metadata partition managed by the first metadata service node according to the second query request, and obtain the first target data according to the metadata.

14. The data system according to claim 13, Features: Each metadata service node stores metadata corresponding to the managed metadata partition.

15. A data system according to claim 13 or 14, Features: The metadata management node also serves as the metadata service node.

16. A data system according to any one of claims 13 to 15, Features: The metadata management node stores partition information of the metadata partition managed by each metadata service node, and the partition information includes data information corresponding to the metadata partition; The metadata management node is used to determine the target metadata service node and the metadata partition corresponding to the target metadata service node according to data information of the first target data corresponding to the first query request.

17. A data system according to any one of claims 13 to 16, Features: The metadata corresponding to the metadata partition includes a partition file, and the partition file records information of a data file corresponding to the metadata partition; The first metadata service node is used to determine the data file including the first target data according to the data information of the first target data requested in the second query request and the information of the data file recorded in the partition file.

18. The data system according to claim 17, Features: The metadata corresponding to the metadata partition includes a plurality of partition range lists, each partition range list records information of a plurality of the partition files; The first metadata service node is used to determine the first target partition file based on the data information of the first target data requested in the second query request and the information of the partition file recorded in the partition range list, and determine the data file including the first target data based on the information of the data file recorded in the partition file.

19. The data system according to claim 18, Features: The metadata corresponding to the metadata partition includes multiple partition definition lists and multiple snapshots, each partition definition list records information of multiple partition range lists, and each snapshot includes information of one partition definition list; The metadata management node is used to obtain a target partition definition list corresponding to the first target data according to the data information of the first target data requested in the first query request and the snapshot; The metadata management node is used to determine the information of the target partition range list based on the data information of the first target data requested in the first query request and the information of the partition range list recorded in the target partition definition list, wherein the second query request includes the information of the target partition range list.

20. The data system according to claim 19, Features: The metadata service node obtains the target partition range list according to information of the target partition range list; and determines the first target partition file according to information of the partition files recorded in the target partition range list.

21. A data system according to any one of claims 13 to 20, It is characterized in that Also includes: The metadata management node is used to obtain a first update request; The metadata management node is used to determine at least one target metadata service node where the metadata partition corresponding to the second target data is located according to data information of the second target data corresponding to the first update request; The metadata management node is used to send a second update request generated according to the first update request to the target metadata service node respectively; The second metadata service node among the at least one target metadata service node is used to update the second target data according to the second update request, and to update the metadata corresponding to the second target data in the metadata partition managed by the second metadata service node.

22. The data system according to claim 21, Features: The metadata management node stores partition information of the metadata partition managed by each metadata service node, and the partition information includes data information corresponding to the metadata partition; The metadata management node is used to determine the target metadata service node and the metadata partition corresponding to the target metadata service node according to data information of the second target data corresponding to the first update request.

23. The data system according to claim 22, It is characterized in that The second metadata service node is used for: updating a data file according to the data information of the second target data; Update the second target partition file corresponding to the second target data according to the updated information of the data file; Update the corresponding partition range list according to the updated information of the second target partition file.

24. The data system according to claim 23, It is characterized in that The metadata management node is further used to update the corresponding definition list and the corresponding snapshot according to the updated information of the partition range list.