Metadata processing method and device, electronic equipment and readable storage medium

By obtaining the load status information of the metadata tree and identifying and migrating the overloaded metadata subtree, the problem of load imbalance in the metadata server cluster is solved, and more efficient load balancing and performance improvement is achieved.

CN120216158APending Publication Date: 2025-06-27BEIJING DIDI INFINITY TECH & DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311828190.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The load imbalance of different metadata servers in the metadata server cluster leads to degradation of metadata service performance.

Method used

By obtaining the load status information of each metadata tree, the overloaded metadata tree is determined, and its subtrees are generated, and migrating to other servers to reduce the load and improve the load balancing of the cluster.

Benefits of technology

It effectively reduces the load of the metadata tree, improves the load balancing of the server cluster, and improves the performance of metadata services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216158A_ABST
    Figure CN120216158A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a metadata processing method and device, electronic equipment and a readable storage medium. According to the embodiment of the invention, a plurality of pieces of load state information of each metadata tree in a server cluster within a preset time period are obtained, the metadata trees of which the plurality of pieces of load state information meet a preset load condition are determined as target metadata trees, and at least one metadata sub-tree corresponding to the target metadata trees is determined, the load state information of the metadata tree comprises storage resource consumption and / or operation request parameters of the metadata tree. Therefore, in the embodiment of the invention, whether the metadata tree is overloaded or not is determined according to the load state information of each metadata tree, and the metadata subtree of the metadata tree is generated and migrated to other servers when the metadata tree is overloaded, so that the load of the overloaded metadata tree is reduced, and the load balance of the server cluster is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and more particularly, to a method, apparatus, electronic device, and readable storage medium for metadata processing. Background Art

[0002] In order to store and access data, data files are stored in a data storage system, and the data storage system stores and manages the metadata of different data files through a metadata server cluster to provide metadata services. The continuous growth of the number of data files makes the number of metadata also increase continuously. However, the growth of the scale of the metadata trees of different metadata servers in the metadata server cluster and the access volume are different, which causes the phenomenon of load imbalance in the metadata server cluster. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a method, apparatus, electronic device, and readable storage medium for metadata processing, which determine whether a metadata tree is overloaded according to the load status information of each metadata tree, and generate a metadata subtree of the metadata tree and migrate it to other servers when the metadata tree is overloaded, so as to reduce the load of the metadata tree and improve the load balance of the server cluster.

[0004] In a first aspect, an embodiment of the present invention provides a method for metadata processing, the method comprising:

[0005] Obtaining a plurality of load status information of each metadata tree in the server cluster within a predetermined time period, the load status information including at least one of the storage resource consumption amount and operation request parameters of the corresponding metadata tree;

[0006] Determining the metadata trees whose plurality of load status information meet a predetermined load condition as target metadata trees;

[0007] Determining at least one metadata subtree corresponding to the target metadata tree.

[0008] Optionally, the determining the metadata trees whose plurality of load status information meet a predetermined load condition as target metadata trees includes:

[0009] For each metadata tree, determining a corresponding first load status according to the corresponding plurality of load status information;

[0010] In response to the first load status indicating that the corresponding metadata tree is overloaded, determining the metadata tree as the target metadata tree.

[0011] Optionally, the for each metadata tree, determining a corresponding first load status according to the corresponding plurality of load status information includes:

[0012] For each of the metadata trees, determine a plurality of first load scores within the predetermined duration according to the corresponding plurality of load status information.

[0013] For each of the metadata trees, in response to a first number of the first load scores among the plurality of first load scores satisfying a first condition, determine that the corresponding first load status is overloaded.

[0014] Optionally, the load status information includes the storage resource consumption of the corresponding metadata tree, and the first load score is determined according to the storage resource consumption of the metadata tree.

[0015] Optionally, the load status information includes the operation request parameters of the corresponding metadata tree, and the first load score is determined according to the operation request parameters of the metadata tree and the corresponding operation request scores.

[0016] Optionally, the load status information includes the storage resource consumption and operation request parameters of the corresponding metadata tree, and the first load score is determined according to a first score and a second score. The first score is determined according to the storage resource consumption of the metadata tree, and the second score is determined according to the operation request parameters of the metadata tree and the corresponding operation request scores.

[0017] Optionally, the determining of at least one metadata subtree corresponding to the target metadata tree includes:

[0018] Determine the second load status of each server in the server cluster according to the plurality of load status information corresponding to each of the target metadata trees;

[0019] Determine at least one of the metadata subtrees corresponding to the target metadata tree and the target server corresponding to each of the metadata subtrees according to each of the second load statuses, so as to migrate each of the metadata subtrees to the corresponding target server.

[0020] Optionally, the determining of the second load status of each server in the server cluster according to the plurality of load status information corresponding to each of the target metadata trees includes:

[0021] For each of the servers, determine a plurality of second load scores within the predetermined duration according to the plurality of load status information of the corresponding metadata trees;

[0022] For each of the servers, in response to a second number of the second load scores among the plurality of second load scores satisfying a second condition, determine that the corresponding second load status is overloaded.

[0023] Optionally, the second condition is determined according to the average value of the first load scores of the metadata trees corresponding to the server and the minimum balance parameter of the metadata tree.

[0024] Optionally, the determining of at least one metadata subtree corresponding to the target metadata tree and the target server corresponding to each metadata subtree according to each of the second load states includes:

[0025] Determining the servers with the first load state characterized as not overloaded as candidate servers;

[0026] Determining the expected migration amount of the target metadata tree and the expected reception amount of the candidate servers;

[0027] Determining at least one metadata subtree corresponding to the target metadata tree and the target server corresponding to each metadata subtree according to each of the expected migration amounts and each of the expected reception amounts.

[0028] Optionally, the determining of the expected migration amount of the target metadata tree and the expected reception amount of the candidate servers includes:

[0029] Obtaining the first balance parameter of each candidate server and the second balance parameter of each target metadata tree;

[0030] Determining the corresponding expected reception amount according to the first load score and the first balance parameter of each candidate server;

[0031] Determining the corresponding expected migration amount according to the second load score and the second balance parameter of each target metadata tree.

[0032] Optionally, the determining of at least one metadata subtree corresponding to the target metadata tree and the target server corresponding to each metadata subtree according to each of the expected migration amounts and each of the expected reception amounts includes:

[0033] Matching the target metadata tree and the candidate servers according to the expected migration amount and the expected reception amount;

[0034] Determining at least one of the candidate servers matched with the target metadata tree as the corresponding target server;

[0035] Determining the metadata subtree corresponding to each target server according to each of the expected reception amounts.

[0036] Optionally, the operation request parameter is the operation request parameter of the corresponding metadata tree under each operation request category;

[0037] The storage resource consumption is determined according to the amount of metadata written to the metadata tree and the amount of metadata deleted from the metadata tree.

[0038] In a second aspect, an embodiment of the present invention provides a metadata processing device, which includes:

[0039] An information acquisition unit, configured to acquire multiple load status information of each metadata tree in the server cluster within a predetermined time period, where the load status information includes at least one of the storage resource consumption of the corresponding metadata tree and the operation request parameter;

[0040] A tree determination unit, configured to determine the metadata tree whose multiple load status information meets a predetermined load condition as a target metadata tree;

[0041] A subtree determination unit, configured to determine at least one metadata subtree corresponding to the target metadata tree.

[0042] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of the first aspects.

[0043] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method according to any one of the first aspects is implemented.

[0044] In the embodiments of the present invention, multiple load status information of each metadata tree in the server cluster is acquired within a predetermined time period, the metadata tree whose multiple load status information meets a predetermined load condition is determined as a target metadata tree, and at least one metadata subtree corresponding to the target metadata tree is determined. Among them, the load status information of the metadata tree includes the storage resource consumption of the metadata tree and / or the operation request parameter. Thus, in the embodiments of the present invention, it is determined whether the metadata tree is overloaded according to the load status information of each metadata tree, and when the metadata tree is overloaded, a metadata subtree of the metadata tree is generated and migrated to other servers to reduce the load of the overloaded metadata tree and improve the load balance of the server cluster. Description of the Drawings

[0045] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features and advantages of the present invention will become clearer. In the drawings:

[0046] Figure 1 It is a schematic diagram of a data storage system according to an embodiment of the present invention;

[0047] Figure 2It is a schematic diagram of the metadata server cluster according to an embodiment of the present invention;

[0048] Figure 3 It is a flowchart of a metadata processing method according to an embodiment of the present invention;

[0049] Figure 4 It is a schematic diagram of determining the first load status of a metadata tree according to multiple first load scores of the metadata tree within a predetermined time period in an embodiment of the present invention;

[0050] Figure 5 It is a schematic diagram of the target metadata tree according to an embodiment of the present invention;

[0051] Figure 6 It is a schematic diagram of the metadata processing apparatus according to an embodiment of the present invention;

[0052] Figure 7 It is a schematic diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners

[0053] The following describes the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these detail parts. In order to avoid obscuring the essence of the present application, well-known methods, processes, procedures, elements and circuits are not described in detail.

[0054] In addition, those of ordinary skill in the art should understand that the drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.

[0055] Unless the context clearly requires otherwise, words such as "including" and "comprising" in the entire application document should be interpreted in an inclusive sense rather than an exclusive or exhaustive sense; that is, the meaning of "including but not limited to".

[0056] In the description of the present application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0057] This embodiment provides a data storage system, which provides a unified data storage entry for various computing applications in the data lake ecosystem and integrates various storage protocols to support storage semantic fusion of various storage services such as Hadoop storage services (such as storage services based on HDFS distributed file system and HBase distributed NoSQL database), S3 (S3 Simple Storage Service), K8S CSI storage service, and Posix (Portable Operating System Interface of UNIX). It is applied to multiple data ecosystems and application scenarios, improving the performance of data access and reducing the storage space required for data applications.

[0058] Figure 1 It is a schematic diagram of the data storage system according to an embodiment of the present invention. As Figure 1 shown, the data storage system 20 in this embodiment is connected to the service layer 10 and the storage layer 30.

[0059] In this embodiment, the data storage system 20 performs semantic fusion on various storage services so that each service layer 10 can call an appropriate storage interface to operate on data files or the data therein (such as read, write, delete, etc.). Among them, the data storage system 20 includes a storage interface service module 21 with multiple storage interfaces and a metadata service module 22. Optionally, the storage interface service module 21 of the data storage system 20 may include interface services such as Posix file interface, HDFS SDK, CSI, S3 interface, S2 interface, and image processing interface. It should be understood that this embodiment is not limited to the above storage interfaces, and other storage interfaces for implementing storage services can also be integrated into this embodiment.

[0060] In this embodiment, regardless of which storage protocol (such as S3 or file system, etc.), the data file consists of two parts: metadata and data. Among them, the metadata is stored in the metadata service, and the data is stored in the corresponding data storage space (such as GIFT DFS storage system, S3 storage system, OSS storage system, COS storage system, etc.).

[0061] Furthermore, in this embodiment, the metadata service module 22 is configured to store and manage the metadata of the file. Among them, the metadata of the file is stored in the form of a metadata tree or a metadata subtree.

[0062] In this embodiment, the metadata service module 22 is implemented by a metadata server cluster. The metadata server (Meta Data Server, MDS) in the metadata server cluster is a node in the metadata server cluster. To facilitate the client to access the metadata of various files in the data storage system, each metadata server stores the corresponding metadata tree or metadata subtree in the local storage module or the remote storage module. As the number of files continues to grow, the scale of the metadata tree or metadata subtree also continues to grow. Moreover, there are differences in the load status information such as the access volume between the metadata trees or metadata subtrees, which results in differences in the loads of different metadata servers in the metadata server cluster, thus causing the phenomenon of load imbalance.

[0063] Figure 2 is a schematic diagram of the metadata server cluster according to an embodiment of the present invention. As Figure 2 shown, the metadata service module 22 includes two metadata server groups. One group is server 221, server 222, and server 223, and the other group is server 224, server 225, and server 226. Among them, server 221 and server 224 are respectively the primary metadata servers (i.e., the primary nodes) in their respective groups. Server 222 and server 223 are two secondary metadata servers (i.e., secondary nodes) of server 221, and server 225 and server 226 are two secondary metadata servers of server 224. Among them, server 221, server 222, and server 223 store and manage the same metadata tree and the metadata in the metadata tree. Server 224, server 225, and server 226 store and manage the same metadata tree and the metadata in the metadata tree. Further, server 221 and server 224 are configured to read / write the metadata tree and the metadata in the metadata tree, and server 222, server 223, server 225, and server 226 are responsible for reading the metadata tree and the metadata in the metadata tree. It is easy to understand that Figure 2 the structure of the metadata server cluster shown and the number of metadata servers are merely illustrative.

[0064] To solve the above problems, the data storage system according to the embodiment of the present invention obtains at least one of the storage resource consumption amount and operation request parameters of each metadata tree by collecting the heartbeats reported by the metadata trees of each metadata server, determines the target metadata tree with a larger load from each metadata tree according to at least one of the storage resource consumption amount and operation request parameters of each metadata tree, and determines the load status of each metadata server in the server cluster. Furthermore, at least one metadata subtree corresponding to the target metadata tree and the target server corresponding to the metadata subtree are determined according to the load status of each metadata server, so as to migrate each metadata subtree to the corresponding target server, thereby reducing the load of the overloaded metadata tree and improving the load balance of the server cluster.

[0065] The following is illustrated by method embodiments. Figure 3 It is a flowchart of a metadata processing method according to an embodiment of the present invention. As Figure 3 shown, the method of this embodiment includes the following steps:

[0066] Step S100, obtain multiple load status information of each metadata tree in the server cluster within a predetermined time period.

[0067] In this embodiment, the data storage system maintains the load balance of the metadata server cluster through the management server (Master Server). Each metadata server in the metadata server cluster will report its heartbeat to the management server according to a predetermined cycle, and the load status information of each metadata tree stored and managed by itself is carried in the heartbeat.

[0068] The load status information of the metadata tree is used to reflect the current load situation of the metadata tree, and may specifically include at least one of the storage resource consumption amount of the metadata tree and the operation request parameters. Among them, the storage resource consumption amount characterizes the size of the storage space occupied by the metadata tree, and can be specifically determined according to the data amount of the metadata written into the metadata tree and the data amount of the metadata deleted from the metadata; the operation request parameters include the request quantities of read / write / delete and other operation requests received by the metadata tree, and can specifically be the queries-per-second (QPS) of the operation request, that is, the number of operation requests per 1 second, and may also include the access latency of a single read / write / delete and other operation requests received by the metadata tree.

[0069] The larger the storage resource consumption amount of the metadata tree, the larger the data amount of the metadata in the metadata tree, resulting in a longer time for the metadata server to search for the corresponding metadata in response to the client operation request, and thus a longer time for returning the result. The operation requests of the client to the metadata server are usually stored in the request queue and processed in the order of being stored in the request queue. Therefore, the larger the operation request parameters of the metadata tree, the longer the response time of the metadata server to the operation request, resulting in a longer time for returning the result. It is easy to understand that the load status information may also include other information, such as the usage rate of the central processing unit (CPU) of the metadata tree, etc., and this embodiment does not make specific limitations.

[0070] The write operation of the client to the metadata will be written into a predetermined database (such as RocksDB) in the form of key-value pairs. Therefore, the management server can determine the storage resource consumption of the metadata tree according to the data volume of the values written into / removed from the metadata tree. The key-value pairs written into the metadata tree will increase the storage resource consumption of the metadata tree, while the key-value pairs removed from the metadata tree will reduce the storage resource consumption of the metadata tree. Therefore, for each metadata tree, the management server can calculate the difference between the total data volume of the key-value pairs written into the metadata tree and the total data volume of the key-value pairs removed from the metadata tree to obtain the storage resource consumption of the metadata tree.

[0071] Step S200, determine multiple metadata trees whose load status information meets the predetermined load conditions as target metadata trees.

[0072] In this step, the management server will determine the metadata trees with larger loads as target metadata trees according to the multiple load status information of each metadata tree within a predetermined time period. Specifically, for each metadata tree, the management server will determine the first load status of the metadata tree according to the multiple load status information of the metadata tree, and when the first load status of the metadata tree indicates that the metadata tree is overloaded, the metadata tree will be determined as the target metadata tree.

[0073] Furthermore, the management server can determine the first load score corresponding to each load status information according to the multiple load status information of each metadata tree within a predetermined time period. If there are a first number of first load scores among the multiple first load scores within the predetermined time period that meet the first condition, the management server will determine the first load status of the metadata tree as overloaded.

[0074] In an optional implementation manner, the priorities of the storage resource consumption and the operation request parameters can be set in advance. The priority indicates the importance of the storage resource consumption and the operation request parameters. The management server can determine the first load score of the server according to the load status information with a higher priority. If there are servers with the same first load score, the management server can further update the first load score according to the load status information with a lower priority.

[0075] Taking the example that the priority of the storage resource consumption is higher than the priority of the operation request parameters, the management server will determine the first load score of the metadata tree according to the storage resource consumption of the metadata tree. If there are multiple metadata trees with the same storage resource consumption, the management server will determine the first load score of the metadata tree according to the storage resource consumption and the operation request parameters of the metadata tree at the same time.

[0076] In another alternative implementation, the management server can determine the first load score of the metadata tree only based on the storage resource consumption of the metadata tree. For example, if the storage resource consumption of metadata tree 1 is 1 GB, the management server can convert the measurement unit of this resource consumption into bytes, obtaining that the storage resource consumption of metadata tree 1 is 1073741824 bytes, and determine the first load score of metadata tree 1 as 1073741824. For another example, if the storage resource consumption of metadata tree 1 is 1 GB, the management server can convert the measurement unit of this resource consumption into bytes, obtaining that the storage resource consumption of metadata tree 1 is 1073741824 bytes, and multiply the storage resource consumption of metadata tree 1 by 0.0000001 (i.e., the predetermined coefficient) and take the integer part to obtain that the first load score of metadata tree 1 is 107 points. For yet another example, if the resource storage consumption of metadata tree 1 is 1 GB, the management server can directly determine this resource consumption as the first load score of metadata tree 1, that is, the first load score of metadata tree 1 is 1.

[0077] In another alternative implementation, the management server can determine the first load score of the metadata tree only based on the operation request parameters of the metadata tree. Specifically, the management server can determine based on the operation request parameters of the metadata tree and the corresponding operation request scores.

[0078] Different operation requests sent by the client to the data storage system will cause different degrees of load on the metadata server. Therefore, in this embodiment, different operation request categories correspond to different operation request scores. For example, for operation request categories such as StatFS, Open, Close, Access, GetAttr, Resolve, ReadLink, GetXattr, ListXattr, and RemoveXAttr, the corresponding operation request scores can be set to 1 point; for operation request categories such as Lookup, Read, NextSlice, and NextINode, the corresponding operation request scores can be set to 2 points; for operation request categories such as Write, Create, MKnod, Mkdir, Rename, SetAttr, Rmdir, Unlink, Truncate, Fallocate, Flock, SetLlk, Link, Symlink, AppendFile, CommitCompact, NewSession, and SessionHeartbeat, the corresponding operation request scores can be set to 6 points; for operation request categories such as ReadDir and CommitAppend, the corresponding operation request scores can be set to 10 points.

[0079] For each metadata tree, after determining the product of the number of operation requests generated by each operation request category and / or the access latency caused by the operation requests and the operation request score corresponding to the operation request category, the management server may determine the sum of the above products as the first load score corresponding to the metadata tree.

[0080] In another alternative implementation, the management server may determine the first load score of the metadata tree according to the storage resource consumption of the metadata tree and the operation request parameters. Specifically, for each metadata tree, the management server may determine the first score according to the storage resource consumption of the metadata tree and determine the second score according to the operation request parameters, so as to determine the first load score of the metadata tree according to the first score and the second score. Among them, the determination methods of the first score and the second score are similar to the method of determining the first load score in the above alternative implementation, and will not be elaborated here. Optionally, the first score and the second score may be set with the same or different weights respectively, and the management server may calculate the weighted sum of the first score and the second score of the metadata tree to determine the first load score of the metadata tree.

[0081] Furthermore, the weights of the first score and the second score may be determined according to the load conditions caused by the storage resource consumption and the operation request parameters on the metadata server cluster in actual applications. Larger weights are set for the load status information with a greater impact on the load, such as 0.6, and smaller weights are set for the load status information with a smaller impact on the load, such as 0.4.

[0082] To prevent large fluctuations in load in a short period from having a greater impact on the load status evaluation of the metadata tree, after determining the first load scores of the same metadata tree at different times within a predetermined duration, the management server will determine the metadata tree as the target metadata tree when there are a first number of first load scores among these multiple first load scores that meet the first condition.

[0083] In this embodiment, both the first number and the first condition may be determined according to the actual load conditions of the metadata server cluster. For example, if the metadata tree has a high update frequency and the server cluster needs to perform load balancing every 60 minutes, the predetermined duration may be set to 30 minutes, 45 minutes, etc.; if the number of first load scores in the predetermined duration is 30, the first number may be set to 10, 15, etc.; when the first load score is not higher than 10000000 points, the metadata tree is not overloaded, so the first condition may be set to the first load score not being higher than 10000000 points.

[0084] Optionally, the first condition may be set to the first load score ≥ the predetermined load parameter and the load score ≥ the minimum value of all load scores, where the load score of the metadata tree is the median of the first load scores in the sliding window sequence.

[0085] The frequency at which the server reports heartbeats is usually fixed. Therefore, within the same duration, the number of heartbeats reported by the server is the same. In an alternative implementation, for any metadata tree, when the management server determines that the first load score of this metadata tree meets the first condition, this first load score can be added to a sliding window queue of a predetermined length, and it is determined that the consecutive (n - 1) first load scores after this first load score of this metadata tree are added to the sliding window queue. Then, these n first load scores are counted one by one from the 1st to the nth. If the i-th (1 ≤ i ≤ n) first load score meets the first condition, the management server increments the hot degree value of this metadata tree by 1 and sets the elimination value (Counter) of this server to 0; if the i-th first load score does not meet the first condition, the management server decrements the hot degree value of this metadata tree by 1 and increments the elimination value of this metadata tree by 1. When the elimination value of this metadata tree is not less than the minimum elimination parameter, that is, when there are a first number of first load scores among the multiple first load scores corresponding to this metadata tree that do not meet the first condition, the management server can determine that the first load state of this metadata tree is not overloaded; if the hot degree value of this metadata tree is not less than the minimum hot degree parameter and the load score is not less than the minimum value of all load scores, the management server can determine that the first load state of this metadata tree is overloaded. Here, n is the number of load status information of this server obtained within a predetermined duration.

[0086] Figure 4 It is a schematic diagram for determining the first load state of a metadata tree according to multiple first load scores of the metadata tree within a predetermined duration in an embodiment of the present invention. Figure 4 The first load scores shown are the first load states determined according to the load status information reported by the same metadata tree at different times within a predetermined duration. As Figure 4 As shown, the management server sequentially stores the n first load scores of metadata tree 1, that is, the first load score 1 - the first load score n, into the sliding window queue 40. Among them, the window 41 in the sliding window queue 40 stores the first load score 1, the window 42 stores the first load score 2,..., and the window 4n stores the first load score n. When the management server determines that the first load score 1 does not meet the first condition, it eliminates the first load score 1. At this time, the window 41 is emptied and changes from the first window in the sliding window queue 40 to the last window in the sliding window queue 40. The management server will obtain the (n + 1)-th first load score of metadata tree 1, that is, the first load score (n + 1), and store the first load score (n + 1) in the window 41 to continue determining the first load state of metadata tree 1.

[0087] Step S300: Determine at least one metadata subtree corresponding to the target metadata tree.

[0088] In this embodiment, the management server determines the second load status of each server in the server cluster according to the multiple load status information corresponding to each metadata tree, and then determines the target status corresponding to each metadata subtree and the target server corresponding to each metadata subtree according to the second load status of each server, so as to migrate each metadata subtree to the corresponding target server. In this embodiment, the target server can be an existing metadata server in the metadata server cluster or a newly added server in the metadata server cluster, and this embodiment does not make specific limitations.

[0089] In an alternative implementation, for each server, the management server determines multiple second load scores of the server within a predetermined time period according to the load status information of each metadata tree corresponding to the server, and when there are a second number of second load scores among the multiple second load scores that meet the second condition, determines the second load status of the server as overloaded.

[0090] Specifically, the management server can determine the sum of the first load scores of each metadata tree corresponding to the same server at the same moment as the second load score of the server at that moment. The management server can also determine the second load status of the server by counting the multiple second load scores corresponding to each server in the sliding window sequence. Specifically, it can refer to the determination method of the first load status in step S200, which will not be elaborated here.

[0091] Optionally, for each server, the management server determines the second condition according to the average value of the first load scores of each metadata tree corresponding to the server and the minimum balance parameter of the metadata tree. The minimum balance parameter is the minimum value at which the metadata tree can maintain an equilibrium state. Specifically, the management server can set the second condition as the second load score ≤ the average value of the second load scores of each metadata tree * (1 + the minimum balance parameter). Or, similar to the first condition and the first number, the second condition and the second number can also be determined according to the actual load situation of the metadata server cluster.

[0092] After determining the second load status of each server, the management server determines the candidate servers whose second load status indicates not overloaded, then determines the expected migration amount of the target metadata tree and the expected reception amount of the candidate servers, and further determines at least one metadata subtree corresponding to the target metadata tree and the target server corresponding to each metadata subtree according to the expected migration amount of each target metadata tree and the expected reception amount of the candidate servers, so as to migrate each metadata subtree to the corresponding target server and improve the load balance of the metadata server cluster.

[0093] Specifically, the management server can obtain the first balance parameter of each target metadata tree and the second balance parameter of each candidate server, and determine the expected migration amount of the target metadata tree according to the first load score and the first balance parameter of each target metadata tree, and determine the expected reception amount of the candidate server according to the second load score and the second balance parameter of each candidate server. Among them, the first balance parameter is the load status score corresponding to when the target metadata tree reaches the balanced state, specifically, it can be the maximum load status score, and the second balance parameter is the load status score corresponding to when the candidate server reaches the balanced state, specifically, it can be the maximum load status score.

[0094] The management server can determine the expected migration amount of the target metadata tree according to the difference between the first load score and the first balance parameter of each target metadata tree. For a candidate server including the target metadata tree, the management server can determine the expected reception amount of the candidate server according to the expected migration amounts of the target metadata trees corresponding to the candidate server and the difference between the second load score and the second balance parameter of the candidate server; for a candidate server not including the target metadata tree, the management server can determine the expected reception amount of the candidate server according to the difference between the second load score and the second balance parameter of the candidate server.

[0095] For example, the target metadata trees include metadata tree T1 and metadata tree T2, and the candidate servers include server MDS1, server MDS2, and server MDS3, where server MDS2 includes metadata tree T1. The first load scores of metadata tree T1 and metadata tree T2 are the storage resource consumption amounts of the corresponding target metadata trees. Among them, the first load score of metadata tree T1 is 580 (GB), and the first load score of metadata tree T2 is 620 (GB). The second load scores of server MDS1, server MDS2, and server MDS3 are the sum of the storage resource consumption amounts of the corresponding metadata trees. Among them, the second load score of server MDS1 is 1860 (GB), the second load score of server MDS2 is 2010 (GB), and the second load score of server MDS3 is 1980 (GB). The second balance parameter of the candidate server is 2048 (GB), and the first balance parameter of the target metadata tree is 500 GB.

[0096] The management server determines that the expected migration amount of metadata tree T1 is 80 (GB) based on the difference between the first load score of metadata tree T1 and the first balancing parameter, and the expected migration amount of metadata tree T2 is 120 (GB) based on the difference between the first load score of metadata tree T2 and the first balancing parameter. The expected reception amount of server MDS1 is determined to be 188 (GB) according to the difference between the second load score of server MDS1 and the second balancing parameter. The expected reception amount of server MDS2 is determined to be 2048 - (2010 - 80) = 118 (GB) based on the expected migration amount of metadata tree T1 and the difference between the second load score of server MDS2 and the second balancing parameter. The expected reception amount of server MDS3 is determined to be 68 (GB) according to the difference between the second load score of server MDS3 and the second balancing parameter.

[0097] After determining the expected migration amounts corresponding to the target metadata trees and the expected reception amounts corresponding to the candidate servers, the management server matches the target metadata trees and the candidate servers according to the expected migration amounts and the expected reception amounts, and then determines at least one candidate server that matches the target metadata tree as the target server corresponding to the metadata subtree of the target metadata tree. Then, the metadata subtree corresponding to each target server is determined according to the expected reception amount of each target server.

[0098] Figure 5 It is a schematic diagram of the target metadata tree in an embodiment of the present invention. Figure 5 The expected migration amount of the metadata tree 50 shown is 100 (GB). The management server matches the expected migration amount of the metadata tree 50 with the expected migration amounts of the candidate servers in the metadata server cluster, and determines that the candidate servers that match the metadata tree 50 are server S1 and server S2 respectively. The expected reception amount of server S1 is 60 (GB), and the expected reception amount of server S2 is 40 (GB). The management server determines that the load score of the metadata subtree t1 with node 51 as the root node is 60 (GB) and the load score of the metadata subtree t2 with node 52 as the root node is 40 (GB) by calculating the load scores of the metadata nodes in the metadata tree 50. Therefore, it can be determined that server S1 is the target server of the metadata subtree t1 and server S2 is the target server of the metadata subtree t2.

[0099] In this embodiment, the management server can also determine at least one metadata subtree corresponding to the target metadata tree and the target server corresponding to each metadata subtree according to the second load status of each server in other ways, which is not limited in this embodiment.

[0100] For example, a maximum migration amount of the metadata subtree is preset in advance, a metadata subtree corresponding to the target metadata tree is determined according to the maximum migration amount, and the target server corresponding to the metadata subtree is determined as the candidate server with the lowest second load score; or a migration quantity (i.e., the quantity of metadata subtrees expected to be migrated) is preset in advance, and the quantity of metadata subtrees is determined according to the ratio of the expected migration amount of the target metadata tree to the number of migrations, and then the lowest second load score of the migration quantity of candidate servers are respectively determined as the target servers corresponding to the respective metadata subtrees.

[0101] In an embodiment of the present invention, multiple load status information of each metadata tree in the server cluster within a predetermined time period is obtained, and the metadata tree whose multiple load status information meets a predetermined load condition is determined as the target metadata tree, and at least one metadata subtree corresponding to the target metadata tree is determined, where the load status information of the metadata tree includes the storage resource consumption amount of the metadata tree and / or operation request parameters. Thus, in the embodiment of the present invention, it is determined whether the metadata tree is overloaded according to the load status information of each metadata tree, and when the metadata tree is overloaded, a metadata subtree of the metadata tree is generated and migrated to other servers to reduce the load of the overloaded metadata tree and improve the load balance of the server cluster. Operation request parameters

[0102] Figure 6 is a schematic diagram of the metadata processing device according to an embodiment of the present invention. As Figure 6 shown, the metadata processing device 6 according to an embodiment of the present invention includes an information acquisition unit 61, a tree determination unit 62, and a subtree determination unit 63.

[0103] Among them, the information acquisition unit 61 is configured to obtain multiple load status information of each metadata tree in the server cluster within a predetermined time period, and the load status information includes at least one of the storage resource consumption amount of the corresponding metadata tree and operation request parameters. The tree determination unit 62 is configured to determine the metadata tree whose multiple load status information meets a predetermined load condition as the target metadata tree. The subtree determination unit 63 is configured to determine at least one metadata subtree corresponding to the target metadata tree.

[0104] In an optional implementation manner, the tree determination unit 62 includes a first status determination subunit and a tree determination subunit.

[0105] Among them, the first status determination subunit is configured to, for each metadata tree, determine a corresponding first load status according to the corresponding multiple load status information. The tree determination subunit is configured to, in response to the first load status indicating that the corresponding metadata tree is overloaded, determine the metadata tree as the target metadata tree.

[0106] In an alternative implementation, the first status determination subunit includes a first score determination module and a first status determination module.

[0107] Among them, the first score determination module is used to determine a plurality of first load scores within the predetermined time length for each of the metadata trees according to the corresponding plurality of load status information. The first status determination module is used to determine that the corresponding first load status is overloaded for each of the metadata trees in response to the existence of a first number of the first load scores among the plurality of first load scores satisfying the first condition.

[0108] In an alternative implementation, the load status information includes the storage resource consumption of the corresponding metadata tree, and the first load score is determined according to the storage resource consumption of the metadata tree.

[0109] In an alternative implementation, the load status information includes the operation request parameters of the corresponding metadata tree, and the first load score is determined according to the operation request parameters of the metadata tree and the corresponding operation request score.

[0110] In an alternative implementation, the load status information includes the storage resource consumption and operation request parameters of the corresponding metadata tree, and the first load score is determined according to a first score and a second score. The first score is determined according to the storage resource consumption of the metadata tree, and the second score is determined according to the operation request parameters of the metadata tree and the corresponding operation request score.

[0111] In an alternative implementation, the subtree determination unit 63 includes a second status determination subunit and a server determination subunit.

[0112] Among them, the second status determination subunit is used to determine the second load status of each server in the server cluster according to the corresponding plurality of load status information of each metadata tree. The server determination subunit is used to determine at least one metadata subtree corresponding to the target metadata tree and the target server corresponding to each metadata subtree according to each second load status, so as to migrate each metadata subtree to the corresponding target server.

[0113] In an alternative implementation, the second status determination subunit includes a second score determination module and a second status determination module.

[0114] Among them, the second scoring determination module is used to determine, for each of the servers, a plurality of second load scores within the predetermined time period according to the plurality of load status information of the corresponding metadata trees. The second status determination module is used to determine, for each of the servers, that the corresponding second load status is overloaded in response to the existence of a second number of the second load scores among the plurality of second load scores satisfying a second condition.

[0115] In an alternative implementation, the second condition is determined according to the average value of the first load scores of the metadata trees corresponding to the server and the minimum balance parameter of the metadata tree.

[0116] In an alternative implementation, the server determination subunit includes a candidate server determination module, a quantity determination module, and a target server determination module.

[0117] Among them, the candidate server determination module is used to determine the servers whose first load status indicates no overload as candidate servers. The quantity determination module is used to determine the expected migration quantity of the target metadata tree and the expected reception quantity of the candidate servers. The target server determination module is used to determine at least one metadata subtree corresponding to the target metadata tree and the target servers corresponding to the metadata subtrees according to the respective expected migration quantities and the respective expected reception quantities.

[0118] In an alternative implementation, the quantity determination module includes a parameter acquisition sub-module, a migration quantity determination sub-module, and a reception quantity determination sub-module.

[0119] Among them, the parameter acquisition sub-module is used to acquire the first balance parameter of each of the target metadata trees and the second balance parameter of each of the candidate servers. The migration quantity determination sub-module is used to determine the corresponding expected migration quantity according to the first load score of each of the target metadata trees and the first balance parameter. The reception quantity determination sub-module is used to determine the corresponding expected reception quantity according to the second load score of each of the candidate servers and the second balance parameter.

[0120] In an alternative implementation, the target server determination module includes a matching sub-module, a target server determination sub-module, and a subtree determination sub-module.

[0121] Among them, the matching sub-module is used to match the target metadata tree and the candidate servers according to the expected migration quantity and the expected reception quantity. The target server determination sub-module is used to determine at least one of the candidate servers that matches the target metadata tree as the corresponding target server. The subtree determination sub-module is used to determine the metadata subtrees corresponding to the target servers according to the respective expected reception quantities.

[0122] In an alternative implementation, the operation request parameter is the operation request parameter of the corresponding metadata tree under each operation request category;

[0123] The storage resource consumption is determined according to the amount of data of the metadata written into the metadata tree and the amount of data of the metadata deleted from the metadata tree.

[0124] Embodiments of the present invention obtain multiple load status information of each metadata tree in a server cluster within a predetermined time period, determine the metadata tree whose multiple load status information meets a predetermined load condition as the target metadata tree, and determine at least one metadata subtree corresponding to the target metadata tree. Among them, the load status information of the metadata tree includes the storage resource consumption of the metadata tree and / or the operation request parameter. Thus, in the embodiments of the present invention, it is determined whether the metadata tree is overloaded according to the load status information of each metadata tree, and when the metadata tree is overloaded, a metadata subtree of the metadata tree is generated and migrated to other servers to reduce the load of the overloaded metadata tree and improve the load balancing of the server cluster.

[0125] Figure 7 is a schematic diagram of an electronic device according to an embodiment of the present invention. As Figure 7 shown, the electronic device 7 is a general data processing device, which includes a general computer hardware structure, and at least includes a processor 701 and a memory 702. The processor 701 and the memory 702 are connected through a bus 703. The memory 702 is suitable for storing instructions or programs executable by the processor 701. The processor 701 can be an independent microprocessor or a set of one or more microprocessors. Thus, the processor 701 processes data and controls other devices by executing the instructions stored in the memory 702 to implement the method flow of the embodiments of the present invention as described above. The bus 703 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to a display controller 704, a display device, and an input / output (I / O) device 705. The input / output (I / O) device 705 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices well known in the art. Typically, the input / output (I / O) device 705 is connected to the system through an input / output (I / O) controller 706.

[0126] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device (equipment), or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be implemented as a computer program product on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This application is described with reference to the flowcharts of methods, apparatuses (devices), and computer program products according to embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.

[0128] These computer program instructions can be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the process Figure 1 the functions specified in one or more of the processes.

[0129] These computer program instructions can also be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the Figure 1 functions specified in one or more of the processes.

[0130] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, and the computer-readable program is used for a computer to execute the above-mentioned partial or all method embodiments.

[0131] That is, those skilled in the art can understand that all or part of the steps of implementing the methods in the above embodiments can be completed by specifying relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0132] The foregoing are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A metadata processing method, characterized in that The method includes: Obtaining multiple load status information of each metadata tree in the server cluster within a predetermined time period, where the load status information includes at least one of the storage resource consumption amount and operation request parameters of the corresponding metadata tree; Determining the metadata tree for which the multiple load status information meets the predetermined load condition as the target metadata tree; Determining at least one metadata subtree corresponding to the target metadata tree.

2. The method according to claim 1, wherein The determining the metadata tree for which the multiple load status information meets the predetermined load condition as the target metadata tree includes: For each metadata tree, determining a corresponding first load status according to the corresponding multiple load status information; In response to the first load status indicating that the corresponding metadata tree is overloaded, determining the metadata tree as the target metadata tree.

3. The method according to claim 2, characterized in that, The for each metadata tree, determining a corresponding first load status according to the corresponding multiple load status information includes: For each metadata tree, determining multiple first load scores within the predetermined time period according to the corresponding multiple load status information; For each metadata tree, in response to there being a first number of the first load scores among the multiple first load scores that meet the first condition, determining the corresponding first load status as overloaded.

4. The method according to claim 3, characterized in that, The load status information includes the storage resource consumption amount of the corresponding metadata tree, and the first load score is determined according to the storage resource consumption amount of the metadata tree.

5. The method according to claim 3, characterized in that The load status information includes the operation request parameters of the corresponding metadata tree, and the first load score is determined according to the operation request parameters of the metadata tree and the corresponding operation request score.

6. The method according to claim 3, wherein The load status information includes the storage resource consumption amount and operation request parameters of the corresponding metadata tree, and the first load score is determined according to a first score and a second score. The first score is determined according to the storage resource consumption amount of the metadata tree, and the second score is determined according to the operation request parameters of the metadata tree and the corresponding operation request score.

7. The method according to claim 1, characterized in that, The determining at least one metadata subtree corresponding to the target metadata tree includes: Determining the second load status of each server in the server cluster according to the multiple load status information corresponding to each metadata tree; Determining at least one of the metadata subtrees corresponding to the target metadata tree and the target server corresponding to each metadata subtree according to each second load status, so as to migrate each metadata subtree to the corresponding target server.

8. The method according to claim 7, characterized in that The determining the second load status of each server in the server cluster according to the multiple load status information corresponding to each metadata tree includes: For each server, determining multiple second load scores within the predetermined time period according to the multiple load status information of the corresponding metadata trees; For each server, in response to there being a second number of the second load scores among the multiple second load scores that meet the second condition, determining the corresponding second load status as overloaded.

9. The method according to claim 8, wherein The second condition is determined according to the average value of the first load scores of the metadata trees corresponding to the server and the minimum balance parameter of the metadata tree.

10. The method according to claim 7, characterized in that The determining of at least one metadata subtree corresponding to the target metadata tree and the target server corresponding to each metadata subtree according to each of the second load states includes: Determining the servers with the second load state characterized as not overloaded as candidate servers; Determining the expected migration amount of the target metadata tree and the expected reception amount of the candidate servers; Determining at least one metadata subtree corresponding to the target metadata tree and the target server corresponding to each metadata subtree according to each of the expected migration amounts and each of the expected reception amounts.

11. The method according to claim 10, wherein The determining of the expected migration amount of the target metadata tree and the expected reception amount of the candidate servers includes: Obtaining the first balance parameter of each target metadata tree and the second balance parameter of each candidate server; Determining the corresponding expected migration amount according to the first load score and the first balance parameter of each target metadata tree; Determining the corresponding expected reception amount according to the second load score and the second balance parameter of each candidate server.

12. The method according to claim 10, wherein The determining of at least one metadata subtree corresponding to the target metadata tree and the target server corresponding to each metadata subtree according to each of the expected migration amounts and each of the expected reception amounts includes: Matching the target metadata tree and the candidate servers according to the expected migration amount and the expected reception amount; Determining at least one of the candidate servers matched with the target metadata tree as the corresponding target server; Determining the metadata subtree corresponding to each target server according to each of the expected reception amounts.

13. The method according to claim 1, characterized in that, The operation request parameter is the operation request parameter of the corresponding metadata tree under each operation request category; The storage resource consumption is determined according to the data volume of the metadata written into the metadata tree and the data volume of the metadata deleted from the metadata tree.

14. A metadata processing device, characterized in that, The device includes: An information acquisition unit, configured to acquire multiple load state information of each metadata tree in a server cluster within a predetermined time period, where the load state information includes at least one of the storage resource consumption and the operation request parameter of the corresponding metadata tree; A tree determination unit, configured to determine the metadata trees whose multiple load state information meets a predetermined load condition as target metadata trees; A subtree determination unit, configured to determine at least one metadata subtree corresponding to the target metadata tree.

15. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1-13.

16. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-13 is implemented.