Fragment Allocation Method, Device, Electronic Device and Computer Readable Medium

By dynamically determining the number of shards and reasonably allocating shards, the storage system performance decline and cost increase caused by creating indexes and shards based on experience is solved, and the storage system performance improvement and node load balancing are achieved, and the user experience is improved.

CN116204273BActive Publication Date: 2025-07-08MULTIPOINT LIFE (WUHAN) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310098948.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-30
Publication Date
2025-07-08
Estimated Expiration
2043-01-30

AI Technical Summary

Technical Problem

In the prior art, creating a fixed number of indexes and shards based on experience leads to degradation in storage system performance, excessive server load and increased cluster consumption costs, and hotspot sharding problems are prone to occur during high concurrent access, resulting in unbalanced node load and reduced system performance.

Method used

By obtaining the total log data and the number of nodes in the distributed storage cluster, dynamically determine the average log data, shard log data, total shard number and index number per day, filter out available nodes that meet the conditions, and reasonably allocate shards to regulate node load.

Benefits of technology

It improves the performance of the storage system, saves resource costs, regulates the load of cluster nodes, improves user experience, solves the performance decline and cost increase caused by improper shards, and balances the load of node clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116204273B_ABST
    Figure CN116204273B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a shard allocation method, apparatus, electronic device, and computer-readable medium. A specific implementation of the method includes: obtaining the total amount of log data; determining the average daily log data volume; obtaining the shard log data volume; determining the total number of shards created per day according to the average daily log data volume and the shard log data volume; obtaining the number of nodes in the distributed storage cluster; determining the number of indexes according to the total number of shards and the number of nodes; determining the number of shards corresponding to each index according to the total number of shards and the number of indexes; screening out node information that meets the first preset condition from the node information cluster to obtain an available node information cluster; and allocating the shards corresponding to the number of shards to the available node cluster corresponding to the available node information cluster. This implementation can improve the performance of the storage system, save resource costs, adjust the load of cluster nodes, and improve the user experience by creating indexes and allocating reasonable numbers of shards.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and particularly to a shard allocation method, apparatus, electronic device, and computer-readable medium. Background Art

[0002] To store a large amount of data, it is necessary to create indexes in a storage system and reasonably shard the created indexes. The amount allocated to each index is an important factor affecting the performance of the storage system and the user experience. For creating indexes and setting the number of shards, the commonly adopted method is: based on one's own experience, create a specified threshold number of indexes, and allocate a fixed number of shards to each index.

[0003] However, the inventors have found that when the above method is used to allocate the number of shards, the following technical problems often exist:

[0004] First, since users create a fixed number of indexes and shards based on their own experience, it is easy to cause too many or too few shards, which in turn leads to a decline in the performance of the storage system, too high a load on the server, and an increase in the cluster consumption cost.

[0005] Second, when allocating a fixed number of shards, in the case of high-concurrency access, the problem of hot shards is likely to occur. When hot shards are concentrated on certain nodes, it causes too high a load on the nodes and uneven load on the node cluster, reducing the system performance and stability of the node cluster.

[0006] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention

[0007] This summary of the present disclosure is used to introduce concepts in a brief form, and these concepts will be described in detail in the following detailed implementation section. This summary of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.

[0008] Some embodiments of the present disclosure propose a shard allocation method, apparatus, electronic device, and computer-readable medium to solve one or more of the technical problems mentioned in the above background art section.

[0009] In a first aspect, some embodiments of the present disclosure provide a shard allocation method, including: obtaining the total amount of log data, where the total amount of log data is the total amount of log data within a preset number of days; determining the average daily log data volume according to the total amount of log data; obtaining the shard log data volume of each shard storing log data; determining the total number of shards created per day according to the average daily log data volume and the shard log data volume; obtaining the number of nodes in the distributed storage cluster; determining the number of indexes according to the total number of shards and the number of nodes; determining the number of shards corresponding to each index according to the total number of shards and the number of indexes; screening out node information that meets the first preset condition from the node information cluster as available node information to obtain an available node information cluster, where the node information cluster is a node information cluster corresponding to the distributed storage cluster; allocating the shards corresponding to the number of shards to the available node cluster corresponding to the available node information cluster to adjust the load of the available node cluster corresponding to the available node information cluster.

[0010] In a second aspect, some embodiments of the present disclosure provide a shard allocation device, including: a first obtaining unit configured to obtain the total amount of log data, where the total amount of log data is the total amount of log data within a preset number of days; a first determining unit configured to determine the average daily log data volume according to the total amount of log data; a second obtaining unit configured to obtain the shard log data volume of each shard storing log data; a second determining unit configured to determine the total number of shards created per day according to the average daily log data volume and the shard log data volume; a third obtaining unit configured to obtain the number of nodes in the distributed storage cluster; a third determining unit configured to determine the number of indexes according to the total number of shards and the number of nodes; a fourth determining unit configured to determine the number of shards corresponding to each index according to the total number of shards and the number of indexes; a screening unit configured to screen out node information that meets the first preset condition from the node information cluster as available node information to obtain an available node information cluster, where the node information cluster is a node information cluster corresponding to the distributed storage cluster; an allocation unit configured to allocate the shards corresponding to the number of shards to the available node cluster corresponding to the available node information cluster to adjust the load of the available node cluster corresponding to the available node information cluster.

[0011] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the first aspect.

[0012] Fourthly, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in any implementation manner of the first aspect.

[0013] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The shard allocation method of some embodiments of the present disclosure can improve the performance of the storage system, save resource costs, and adjust the load of cluster nodes by creating indexes and allocating a reasonable number of shards, thereby improving the user experience. Specifically, the reasons for the decline in the performance of the relevant storage system, the high server load, and the increase in cluster consumption costs are as follows: Since users create a fixed number of indexes and shards based on their own experience, it is easy to cause too many or too few shards, resulting in a decline in the performance of the storage system, an excessive server load, and an increase in cluster consumption costs. Based on this, the shard allocation method of some embodiments of the present disclosure can, first, obtain the total amount of log data, where the total amount of log data is the total amount of log data within a preset number of days. According to the total amount of log data, determine the average daily log data volume. Here, the obtained average daily log data volume is used to determine the total number of shards created per day subsequently. Second, obtain the shard log data volume of each shard for storing log data. Here, the obtained shard log data volume is used to determine the total number of shards created per day subsequently. Third, according to the average daily log data volume and the shard log data volume, determine the total number of shards created per day. Here, create the total number of shards according to the actual situation, avoiding users creating a fixed number of shards based on their own experience and causing waste of system resources. Then, obtain the number of nodes in the distributed storage cluster. Here, the obtained number of nodes is used to determine the number of indexes subsequently. Subsequently, according to the total number of shards and the number of nodes, determine the number of indexes. Here, determine the number of indexes according to the actual node cluster situation, avoiding users creating a fixed number of indexes based on their own experience and causing waste of system resources and a decline in node performance. Then, according to the total number of shards and the number of indexes, determine the number of shards corresponding to each index. Here, obtaining a reasonable number of shards corresponding to the index can improve the stability of the node cluster, reduce the waste of resources in the node cluster, and improve the search speed of the node cluster, thereby improving the user experience. Finally, screen out the node information that meets the first preset condition from the node information cluster as the available node information to obtain the available node information cluster, where the node information cluster is the node information cluster corresponding to the distributed storage cluster. Here, the available node cluster corresponding to the obtained available node information cluster is a node with relatively high performance that can also be allocated shards, avoiding allocating shards to nodes with poor performance, thereby adjusting the performance of the nodes, reducing the server load, and the consumption cost of the node cluster. Allocate the shards corresponding to the above-mentioned number of shards to the available node cluster corresponding to the available node information cluster to adjust the load of the available node cluster corresponding to the available node information cluster. Thus, it can be seen that the present disclosure comprehensively considers the performance of the node cluster and the total amount of log service data to obtain reasonable numbers of indexes and shards.This shard allocation method can improve the performance of the storage system, save resource costs, adjust the load of cluster nodes, and enhance the user experience by creating indexes and allocating a reasonable number of shards. Description of the Drawings

[0014] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.

[0015] Figure 1 is a flowchart of some embodiments of the shard allocation method according to the present disclosure;

[0016] Figure 2 is a schematic structural diagram of some embodiments of the shard allocation device according to the present disclosure;

[0017] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Description of the Embodiments

[0018] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0019] In addition, it should be noted that only parts related to the relevant invention are shown in the drawings for the sake of convenience of description. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0020] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules, or units.

[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0024] Figure 1 Flow 100 of some embodiments of a shard allocation method according to the present disclosure is shown. The shard allocation method includes the following steps:

[0025] Step 101, obtain the total amount of log data.

[0026] In some embodiments, the execution subject of the above shard allocation method may obtain the total amount of log data from each network server through a wired connection method or a wireless connection method, where the total amount of log data is the total amount of log data within a preset number of days. For example, the preset number of days may be one week. The above log data may be data that records changes in data in a distributed storage cluster.

[0027] Step 102, determine the average daily log data volume according to the total amount of log data.

[0028] In some embodiments, the above execution subject may determine the average daily log data volume according to the total amount of log data.

[0029] As an example, the above execution subject may determine the ratio of the total amount of log data divided by the preset number of days as the average daily log data volume.

[0030] Step 103, obtain the shard log data volume of each shard storing log data.

[0031] In some embodiments, the above execution subject may obtain the shard log data volume of each shard storing log data through a wired connection method or a wireless connection method. Among them, the above shard log data volume may be the data volume of shard log data determined by the performance of a computer. The above shard log data volume may be the log data volume that each shard can store. For example, the above shard log data volume may be 20 - 50 GB.

[0032] Step 104, determine the total number of shards created per day according to the average daily log data volume and the shard log data volume.

[0033] In some embodiments, the above execution subject may determine the total number of shards created per day according to the above average daily log data volume and the above shard log data volume.

[0034] As an example, the above-mentioned execution entity may first determine the product of the above-mentioned average daily log data volume and the first weight to obtain a first product. Secondly, determine the product of the above-mentioned sharded log data volume and the second weight to obtain a second product. Then, determine the sum of the first product and the second product to obtain a sum value. Among them, the sum of the first weight and the second weight is 1. The values of the first weight and the second weight can be determined according to specific actual situations. Then, determine the ratio of the sum value to the above-mentioned sharded log data volume as the total number of shards created per day.

[0035] In some optional implementation manners of some embodiments, the above-mentioned determining the total number of shards created per day according to the above-mentioned average daily log data volume and the above-mentioned sharded log data volume may include the following steps:

[0036] In the first step, determine the sum of the above-mentioned average daily log data volume and the above-mentioned sharded log data volume as the total data volume.

[0037] In the second step, determine the ratio of the above-mentioned total data volume to the above-mentioned sharded log data volume as the above-mentioned total number of shards created per day.

[0038] Step 105, obtain the number of nodes in the distributed storage cluster.

[0039] In some embodiments, the above-mentioned execution entity may obtain the number of nodes in the distributed storage cluster. For example, the above-mentioned execution entity may obtain the number of nodes in the distributed storage cluster by calling the interface of the distributed storage cluster. The above-mentioned distributed storage cluster may be a computer cluster used to store log data.

[0040] Step 106, determine the number of indexes according to the total number of shards and the number of nodes.

[0041] In some embodiments, the above-mentioned execution entity may determine the number of indexes according to the above-mentioned total number of shards and the above-mentioned number of nodes.

[0042] As an example, the above-mentioned execution entity may first determine the product value of the product of the above-mentioned number of nodes and the predetermined index ratio. Then, determine the sum of the product value multiplied by the third weight and the sum of the above-mentioned total number of shards multiplied by the fourth weight. Among them, the sum of the third weight and the fourth weight is 1. The corresponding values of the third weight and the fourth weight can be determined according to the actual situation. Finally, determine the ratio of the sum to the product value as the number of indexes.

[0043] In some optional implementation manners of some embodiments, the above-mentioned determining the number of indexes according to the above-mentioned total number of shards and the above-mentioned number of nodes may include the following steps:

[0044] In the first step, determine the first index number by multiplying the above-mentioned number of nodes by the predetermined index ratio. Among them, the above-mentioned predetermined index ratio can be a value determined according to the computer performance of the above-mentioned distributed storage cluster. For example, the above-mentioned predetermined index ratio can be 2.

[0045] In the second step, determine the second index number by adding the above-mentioned first index number to the above-mentioned total number of shards.

[0046] In the third step, determine the above-mentioned index number by taking the ratio of the above-mentioned second index number to the above-mentioned first index number.

[0047] Step 107: Determine the number of shards corresponding to each index according to the total number of shards and the index number.

[0048] In some embodiments, the above-mentioned execution entity can determine the number of shards corresponding to each index according to the above-mentioned total number of shards and the above-mentioned index number. Among them, the above-mentioned index can be a logical storage structure for storing log data.

[0049] As an example, the above-mentioned execution entity can determine the ratio of the above-mentioned total number of shards to the above-mentioned index number as the number of shards corresponding to each index.

[0050] In some optional implementation manners of some embodiments, the above-mentioned determining the number of shards corresponding to each index according to the above-mentioned total number of shards and the above-mentioned index number may include the following steps:

[0051] In the first step, determine the ratio shard number by taking the ratio of the above-mentioned total number of shards to the above-mentioned index number.

[0052] In the second step, determine the number of shards corresponding to each index by adding the above-mentioned ratio shard number to a predetermined threshold. In practice, the above-mentioned predetermined threshold can be 1.

[0053] Step 108: Screen out the node information that meets the first preset condition from the node information cluster as the available node information to obtain the available node information cluster.

[0054] In some embodiments, the above-mentioned execution entity may screen out node information that meets the first preset condition from the node information cluster as available node information, obtaining an available node information cluster. Among them, the above-mentioned node information cluster is the node information cluster corresponding to the above-mentioned distributed storage cluster. The node information in the above-mentioned node information cluster may represent the information of available nodes. The above-mentioned first preset condition may be a condition that the disk usage rate is less than or equal to the preset disk usage rate, and the number of allocated shards is less than or equal to the preset node allocation shard number threshold. In practice, the above-mentioned preset disk usage rate may be 0.90. The above-mentioned preset node allocation shard number threshold may be the maximum number of shards that can be allocated to each node determined according to the actual performance of each node. For example, the above-mentioned preset node allocation shard number threshold may be 20.

[0055] As an example, the above-mentioned execution entity may perform the following determination steps for each node in the above-mentioned node cluster:

[0056] First step, obtain the disk usage rate of the above-mentioned node and the number of allocated shards of the above-mentioned node. Among them, the above-mentioned execution entity may obtain the disk usage rate of the above-mentioned node and the number of allocated shards of the above-mentioned node by controlling a monitoring tool. For example, the above-mentioned monitoring tool may be kibana.

[0057] Second step, in response to determining that the above-mentioned disk usage rate is less than or equal to the preset disk usage rate, and the above-mentioned number of allocated shards is less than or equal to the preset node allocation shard number, determine the above-mentioned node as an available node. Among them, the above-mentioned preset disk usage rate may be 0.90. The above-mentioned preset node allocation shard number may be the maximum number of shards that can be allocated to each node.

[0058] Step 109, allocate the shards corresponding to the shard number to the available node cluster corresponding to the available node information cluster to adjust the load of the available node cluster corresponding to the available node information cluster.

[0059] In some embodiments, the above-mentioned execution entity may allocate the shards corresponding to the above-mentioned shard number to the available node cluster corresponding to the above-mentioned available node information cluster to adjust the load of the available node cluster corresponding to the above-mentioned available node information cluster.

[0060] As an example, the above-mentioned execution entity may, by means of request shunting, allocate the shards corresponding to the above-mentioned shard number to each available node in the available node cluster to adjust the load of the available node cluster. Among them, the above-mentioned request shunting may be to disperse user requests to each node through a routing rule or a distribution algorithm.

[0061] In some alternative implementation manners of some embodiments, the step of allocating the shards corresponding to the number of shards to the available node cluster corresponding to the available node information cluster may include the following steps:

[0062] First, perform a performance evaluation on each available node in the available node cluster to generate performance evaluation values and obtain a set of performance evaluation values.

[0063] As an example, the execution entity may use the linear weighted method to construct a performance evaluation formula, perform a performance evaluation on each available node in the available node cluster, and obtain a set of performance evaluation values. Among them, the performance evaluation formula may be a formula composed of the sum of the first weight coefficient multiplied by the average load of the available node, the second weight coefficient multiplied by the number of shards of the available node, and the third weight coefficient multiplied by the disk usage rate of the available node. The sum of the first weight coefficient, the second weight coefficient, and the third weight coefficient is 1. The values of the first weight coefficient, the second weight coefficient, and the third weight coefficient can be obtained by using statistical analysis or expert consultation for the actual distributed storage cluster environment. The execution entity can obtain the average load, the number of shards, and the disk usage rate of the available node by controlling the monitoring tool of the distributed storage cluster. For example, the monitoring tool may be kibana.

[0064] Second, sort the set of performance evaluation values to obtain a sequence of performance evaluation values. The sorting may be in ascending order.

[0065] Third, according to the order of the sequence of performance evaluation values, allocate the shards to each available node in the available node cluster to adjust the load of the available node cluster.

[0066] As an example, the execution entity may first allocate the first shard to the available node corresponding to the first evaluation value in the sequence of performance evaluation values, allocate the second shard to the available node corresponding to the second evaluation value in the sequence of performance evaluation values, until the number of shards is allocated. Secondly, when the available nodes corresponding to the sequence of performance evaluation values are allocated, start allocating from the first available node again. Finally, determine whether the number of shards already allocated to the available node being allocated plus the number being allocated exceeds the preset number of shards. When it exceeds the preset number of shards, no shard is allocated to the available node being allocated. Among them, the preset number of shards per node may be the maximum number of shards that can be allocated to each node for shards of the same index.

[0067] Optionally, after step 109, the execution entity may further perform the following steps:

[0068] First step, screen out the available node information that meets the second preset condition from the above available node information cluster as the load node information to be adjusted, and obtain the load node information cluster to be adjusted. Among them, the above second preset condition can be the condition that the number of hot shards in the hot shard cluster on the available nodes corresponding to the available node information cluster is greater than the preset hot shard number threshold. The above preset hot shard number threshold can be a shard number threshold obtained through statistical analysis or expert consultation for the actual distributed storage cluster environment. The above hot shard can be the shard where the hot data is located. The hot data can be the data that receives a large number of access requests within a preset time. The preset time can be 1 minute. The load node information to be adjusted in the load node information cluster to be adjusted can represent the information of the load node to be adjusted. For example, the above load node information to be adjusted can be the node number of the load node to be adjusted. The above load node to be adjusted can be the node where the number of hot shards is greater than or equal to the preset hot shard number threshold.

[0069] Second step, based on the above load node information cluster to be adjusted and the above available node information cluster, perform load adjustment on the load node cluster corresponding to the above load node information cluster to be adjusted.

[0070] As an example, the above execution entity can first select the hot shards in the above load node cluster to be adjusted in the shard cluster corresponding to the hot shards in the available node cluster corresponding to the available node information cluster by using the random walk algorithm, and obtain the selectable secondary shards. Then, set the status of the hot shards in the load node to be adjusted to unavailable. Finally, determine the above selectable secondary shards as the primary shards to adjust the load of the load node cluster to be adjusted.

[0071] In some optional implementation manners of some embodiments, the above screening out the available node information that meets the second preset condition from the above available node information cluster as the load node information to be adjusted, and obtaining the load node information cluster to be adjusted may include the following steps:

[0072] First step, for each available node in the above available node cluster, perform the following adding steps:

[0073] Sub-step 1, determine the number of hot shards in the hot shard cluster in the above available node. Among them, the hot shards in the above hot shard cluster are the shards that receive a preset number of access requests within a preset time. The preset time can be 1 minute. The above preset number can be a number obtained through statistical analysis or expert consultation for the actual distributed storage cluster environment.

[0074] Sub-step 2: In response to determining that the number of hot shards is greater than or equal to the preset hot shard number threshold, add the available node information corresponding to the available nodes to the preset load-adjustable node information cluster to obtain the added preset load-adjustable node information cluster, which is used as the load-adjustable node information cluster. The preset hot shard number threshold can be a shard number threshold obtained through statistical analysis or expert consultation for the actual distributed storage cluster environment. The preset load-adjustable node information cluster can be a node cluster with the number of hot shards greater than the preset hot shard number threshold.

[0075] In some alternative implementation manners of some embodiments, the load adjustment of the load-adjustable node cluster corresponding to the load-adjustable node information cluster based on the load-adjustable node information cluster and the available node information cluster may include the following steps:

[0076] First step: For each load-adjustable node in the load-adjustable node cluster, perform the following adjustment steps:

[0077] Sub-step 1: Determine the number of times the adjustment step has been executed.

[0078] Sub-step 2: Perform a load assessment on each available node in the available node cluster to generate a load assessment value, obtaining a set of load assessment values.

[0079] As an example, the above execution entity can use the linear weighting method to construct a load evaluation formula, evaluate the load of each available node in the above available node cluster, and obtain a set of load evaluation values. Among them, the above load evaluation formula can be obtained through the following steps: First, determine the first weight coefficient and the input / output utilization rate of the available node as the first load evaluation value. Second, determine the second weight coefficient and the network bandwidth utilization rate of the available node as the second load evaluation value. Then, determine the third weight coefficient and the CPU (central processing unit) utilization rate of the available node as the third load evaluation value. Then, determine the fourth weight coefficient and the memory utilization rate of the available node as the fourth load evaluation value. Finally, determine the sum of the first load evaluation value, the second load evaluation value, the third load evaluation value, and the fourth load evaluation value as the load evaluation formula. The sum of the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient is 1. The values of the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient can be obtained by using statistical analysis or expert consultation for the actual distributed storage cluster environment. The above execution entity can obtain the input / output utilization rate, network bandwidth utilization rate, CPU utilization rate, and memory utilization rate of the node through the monitoring tool that controls the distributed storage cluster. For example, the above monitoring tool can be kibana.

[0080] Sub-step 3, screen out at least one load evaluation value that meets the third preset condition from the above set of load evaluation values to obtain a target set of load evaluation values. Among them, the above third preset condition can be the condition that the load evaluation value is less than the preset load evaluation threshold. The preset load evaluation threshold is a threshold obtained by using statistical analysis or expert consultation for the actual distributed storage cluster environment.

[0081] Sub-step 4, sort the above target set of load evaluation values to obtain a target sequence of load evaluation values. The above sorting can be a forward sorting.

[0082] Sub-step 5, determine the available node cluster corresponding to the above target sequence of load evaluation values as the target sequence of load nodes.

[0083] Sub-step 6, according to the order of the above target sequence of load nodes, send the hot shards in the hot shard cluster of the above load node to be adjusted to the target load node. The above hot shard can be one hot shard.

[0084] As an example, the above-mentioned execution entity may first mark the status of a hot shard in the to-be-adjusted load node as read-only. Secondly, establish a data transmission channel between the to-be-adjusted load node and the target load node. Thirdly, use the transmission channel to transfer the data in the hot shard to the target load node. Then, in response to determining that the target load node has completed receiving the data of the hot shard, update the node number where the hot shard is located. Finally, delete the hot shard from the to-be-adjusted load node.

[0085] Sub-step 7: Update the number of to-be-adjusted hot shards corresponding to the to-be-adjusted load node and the number of target hot shards corresponding to the target load node to obtain the updated number of to-be-adjusted hot shards and the updated number of target hot shards.

[0086] Sub-step 8: Update the to-be-adjusted load node cluster according to the updated number of to-be-adjusted hot shards, the updated number of target hot shards, and the preset hot shard number threshold.

[0087] As an example, the above-mentioned execution entity may, in response to determining that the updated number of to-be-adjusted hot shards is less than the preset hot shard number threshold, delete the to-be-adjusted load node corresponding to the updated number of to-be-adjusted hot shards from the to-be-adjusted load node cluster. In response to determining that the updated number of target hot shards is greater than or equal to the preset hot shard number threshold, add the target load node corresponding to the updated number of target hot shards to the to-be-adjusted load node cluster.

[0088] Sub-step 9: Update the available node information cluster according to the first preset condition, the to-be-adjusted load node information corresponding to the updated number of to-be-adjusted hot shards, and the target load node information corresponding to the updated number of target hot shards.

[0089] As an example, both the to-be-adjusted load node and the target load node are available nodes. In response to determining that the to-be-adjusted load node does not meet the first preset condition, delete the to-be-adjusted load node information corresponding to the updated number of to-be-adjusted hot shards from the available node information cluster. In response to determining that the target load node does not meet the first preset condition, delete the target load node corresponding to the updated number of target hot shards from the available node cluster.

[0090] Sub-step 10: End the above adjustment step in response to determining that the updated to-be-adjusted load node list is empty or the above-mentioned executed times is equal to the preset executed times. Wherein, the preset executed times may be the maximum number of times that the adjustment step can loop. For example, the preset executed times may be one-half of the above-mentioned number of nodes.

[0091] In the second step, in response to determining that the list of load nodes to be adjusted is not empty and the above-mentioned number of executed times is less than the preset number of executed times, determine the updated load node cluster to be adjusted as the load node cluster to be adjusted, and determine the updated available node cluster as the available node cluster, continue to execute the above adjustment steps, and increase the number of executed times of the first load node to be adjusted by 1.

[0092] The above technical solution and its related content, as an inventive point of the embodiments of the present disclosure, solve the second technical problem mentioned in the background art: "When allocating a fixed number of shards, in the case of high-concurrency access, hot shard problems are likely to occur. When hot shards are concentrated on certain nodes, it causes excessive load on the nodes and uneven load on the node cluster, reducing the system performance and stability of the node cluster." The factors that lead to excessive load on the nodes and uneven load on the node cluster, reducing the system performance and stability of the node cluster are usually as follows: allocating a fixed number of shards, in the case of high-concurrency access, hot shard problems occur, and hot shards are concentrated on certain nodes. If the above factors are solved, the effect of adjusting the load of available nodes, improving the systematicness and stability of the available node cluster, and making the available node cluster load balanced can be achieved. To achieve this effect, based on the above-mentioned node information cluster with load to be adjusted and the above-mentioned available node information cluster, load adjustment is performed on the node cluster with load to be adjusted corresponding to the above-mentioned node information cluster with load to be adjusted, which may include the following steps: For each node with load to be adjusted in the above-mentioned node cluster with load to be adjusted, perform the following adjustment steps: First, determine the number of times the adjustment step has been executed. Perform a load assessment on each available node in the above-mentioned available node cluster to generate a load assessment value, and obtain a set of load assessment values. Here, performing a load assessment on each node is beneficial for grasping the load conditions of available nodes and determining whether hot shard migration is required subsequently. Second, screen out at least one load assessment value that meets the third preset condition from the above-mentioned set of load assessment values to obtain a target set of load assessment values. Sort the above-mentioned target set of load assessment values to obtain a target sequence of load assessment values. Determine the available node cluster corresponding to the above-mentioned target sequence of load assessment values as the target sequence of load nodes. Here, the obtained target sequence of load nodes is used for subsequent hot shard migration to adjust the load of the node cluster with load to be adjusted. Third, according to the order of the above-mentioned target sequence of load nodes, send the hot shards in the hot shard cluster of the above-mentioned node with load to be adjusted to the target load nodes. Here, according to the order of the target sequence of load nodes, migrate one hot shard in the node with load to be adjusted to the target load node with the smallest load assessment, so as to achieve the load of the first node with load to be adjusted, in order to achieve the load balance of the node cluster with load to be adjusted. Then, update the number of hot shards to be adjusted corresponding to the above-mentioned node with load to be adjusted and the number of target hot shards corresponding to the above-mentioned target load node to obtain the updated number of hot shards to be adjusted and the updated number of target hot shards. According to the above-mentioned updated number of hot shards to be adjusted, the above-mentioned updated number of target hot shards, and the above-mentioned preset threshold of the number of hot shards, update the above-mentioned node cluster with load to be adjusted.Update the available node information cluster according to the above first preset condition, the to-be-adjusted load node information corresponding to the updated to-be-adjusted hot shard number, and the target load node information corresponding to the updated target hot shard number. Here, update the to-be-adjusted load node cluster and the available node cluster for subsequent determination of whether to continue the adjustment step. Finally, in response to determining that the updated to-be-adjusted load node list is empty, or the above-mentioned executed times are equal to the preset executed times, end the above adjustment step. In response to determining that the to-be-adjusted load node list is not empty and the above-mentioned executed times are less than the preset executed times, determine the updated to-be-adjusted load node cluster as the to-be-adjusted load node cluster, and determine the updated available node cluster as the available node cluster, continue to execute the above adjustment step, and increase the executed times of the first to-be-adjusted load node by 1. Thus, cyclically detect the number of hot shards in the hot shard cluster of the to-be-adjusted load nodes in the to-be-adjusted load node cluster, perform hot shard migration on the to-be-adjusted load nodes, balance the load of the to-be-adjusted load node cluster, so as to adjust the load of the to-be-adjusted load node cluster and improve the system performance and stability of the node cluster.

[0093] The above embodiments of the present disclosure have the following beneficial effects: The shard allocation method of some embodiments of the present disclosure can improve the performance of the storage system, save resource costs, and adjust the load of cluster nodes by creating indexes and allocating a reasonable number of shards, thereby improving the user experience. Specifically, the reasons for the degradation of the performance of the relevant storage system, the high load of the server, and the increase in the cost consumed by the cluster are as follows: Since users create a fixed number of indexes and shards based on their own experience, it is easy to cause too many or too few shards, which in turn leads to a decrease in the performance of the storage system, an excessive load on the server, and an increase in the cost consumed by the cluster. Based on this, the shard allocation method of some embodiments of the present disclosure can first, obtain the total amount of log data, where the total amount of log data is the total amount of log data within a preset number of days. According to the total amount of log data, determine the average daily log data volume. Here, the obtained average daily log data volume is used to determine the total number of shards created per day subsequently. Secondly, obtain the shard log data volume of each shard storing log data. Here, the obtained shard log data volume is used to determine the total number of shards created per day subsequently. Thirdly, according to the average daily log data volume and the shard log data volume, determine the total number of shards created per day. Here, the obtained total number of shards is used to determine the number of indexes subsequently. Then, obtain the number of nodes in the distributed storage cluster. Here, the obtained number of nodes is used to determine the number of indexes subsequently. Subsequently, according to the total number of shards and the number of nodes, determine the number of indexes. Here, the obtained number of indexes is used to determine the number of shards subsequently. Then, according to the total number of shards and the number of indexes, determine the number of shards corresponding to each index. Here, the obtained number of shards is used by the user to adjust the load of the nodes subsequently. Finally, screen out the node information that meets the first preset condition from the node information cluster as the available node information to obtain an available node information cluster, where the node information cluster is the node information cluster corresponding to the distributed storage cluster. Here, the node information cluster is screened to obtain an available node information cluster, so as to adjust the load of the available node cluster corresponding to the available node information cluster subsequently. Allocate the shards corresponding to the above number of shards to the available node cluster corresponding to the available node information cluster to adjust the load of the available node cluster corresponding to the available node information cluster. Thus, it can be seen that the present disclosure comprehensively considers the performance of the node cluster and the total amount of log service data to obtain reasonable numbers of indexes and shards. The shard allocation method can improve the performance of the storage system, save resource costs, adjust the load of cluster nodes, and improve the user experience by creating indexes and allocating a reasonable number of shards.

[0094] Further referring to Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a shard allocation device, and these device embodiments are related to Figure 1Corresponding to the method embodiments shown, the shard allocation device can be specifically applied to various electronic devices.

[0095] As Figure 2 shown, a shard allocation device 200 includes: a first acquisition unit 201, a first determination unit 202, a second acquisition unit 203, a second determination unit 204, a third acquisition unit 205, a third determination unit 206, a fourth determination unit 207, a screening unit 208, and an allocation unit 209. Among them, the first acquisition unit 201 is configured to: acquire the total amount of log data, where the total amount of log data is the total amount of log data within a preset number of days. The first determination unit 202 is configured to: determine the average daily log data volume according to the total amount of log data. The second acquisition unit 203 is configured to: acquire the shard log data volume of each shard storing log data. The second determination unit 204 is configured to: determine the total number of shards created per day according to the average daily log data volume and the shard log data volume. The third acquisition unit 205 is configured to: acquire the number of nodes in the distributed storage cluster. The third determination unit 206 is configured to: determine the number of indexes according to the total number of shards and the number of nodes. The fourth determination unit 207 is configured to: determine the number of shards corresponding to each index according to the total number of shards and the number of indexes. The screening unit 208 is configured to: screen out node information that meets the first preset condition from the node information cluster as available node information to obtain an available node information cluster, where the node information cluster is a node information cluster corresponding to the distributed storage cluster. The allocation unit 209 is configured to: allocate the shards corresponding to the number of shards to the available node cluster corresponding to the available node information cluster to adjust the load of the available node cluster corresponding to the available node information cluster.

[0096] It can be understood that the various units described in the shard allocation device 200 correspond to the respective steps in the method described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the shard allocation device 200 and the units included therein, and will not be elaborated here.

[0097] Next, with reference to Figure 3 , which shows a schematic structural diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0098] As Figure 3As shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0099] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 an electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 3 Each block shown in may represent a device or, as needed, multiple devices.

[0100] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are executed.

[0101] It should be noted that, in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0102] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LAN”), wide area networks (“WAN”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0103] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain the total amount of log data, where the total amount of log data is the total amount of log data within a preset number of days; determine the average daily log data amount according to the total amount of log data; obtain the shard log data amount of each shard storing log data; determine the total number of shards created per day according to the average daily log data amount and the shard log data amount; obtain the number of nodes in the distributed storage cluster; determine the number of indexes according to the total number of shards and the number of nodes; determine the number of shards corresponding to each index according to the total number of shards and the number of indexes; screen out node information that meets the first preset condition from the node information cluster as available node information to obtain an available node information cluster, where the node information cluster is a node information cluster corresponding to the distributed storage cluster; allocate the shards corresponding to the number of shards to the available node cluster corresponding to the available node information cluster to adjust the load of the available node cluster corresponding to the available node information cluster.

[0104] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0106] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a first acquisition unit, a first determination unit, a second acquisition unit, a second determination unit, a third acquisition unit, a third determination unit, a fourth determination unit, a screening unit, and an allocation unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring the total amount of log data".

[0107] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0108] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A shard allocation method, comprising: Obtaining the total amount of log data, where the total amount of log data is the total amount of log data within a preset number of days; Determining the average daily log data volume according to the total amount of log data; Obtaining the shard log data volume of each shard storing log data; Determining the total number of shards created per day according to the average daily log data volume and the shard log data volume; Obtaining the number of nodes in the distributed storage cluster; Determining the number of indexes according to the total number of shards and the number of nodes, where determining the number of indexes according to the total number of shards and the number of nodes includes: Determining the first number of indexes as the product of the number of nodes and a predetermined index ratio; Determining the second number of indexes as the sum of the first number of indexes and the total number of shards; Determining the number of indexes as the ratio of the second number of indexes to the first number of indexes; Determining the number of shards corresponding to each index according to the total number of shards and the number of indexes; Filtering out node information that meets the first preset condition from the node information cluster as available node information to obtain an available node information cluster, where the node information cluster is a node information cluster corresponding to the distributed storage cluster; Allocating the shards corresponding to the number of shards to the available node cluster corresponding to the available node information cluster to adjust the load of the available node cluster corresponding to the available node information cluster.

2. The method according to claim 1, wherein The method further includes: Filtering out available node information that meets the second preset condition from the available node information cluster as node information with load to be adjusted to obtain a node information cluster with load to be adjusted; Based on the node information cluster with load to be adjusted and the available node information cluster, performing load adjustment on the node cluster with load to be adjusted corresponding to the node information cluster with load to be adjusted.

3. The method according to claim 1, wherein, The determining the total number of shards created per day according to the average daily log data volume and the shard log data volume includes: Determining the total data volume as the sum of the average daily log data volume and the shard log data volume; Determining the total number of shards created per day as the ratio of the total data volume to the shard log data volume.

4. The method according to claim 1, wherein The determining the number of shards corresponding to each index according to the total number of shards and the number of indexes includes: Determining the ratio shard number as the ratio of the total number of shards to the number of indexes; Determining the number of shards corresponding to each index as the sum of the ratio shard number and a predetermined threshold.

5. The method according to claim 1, wherein The allocating the shards corresponding to the number of shards to the available node cluster corresponding to the available node information cluster includes: Performing performance evaluation on each available node in the available node cluster to generate performance evaluation values to obtain a set of performance evaluation values; Sorting the set of performance evaluation values to obtain a sequence of performance evaluation values; Allocating shards to each available node in the available node cluster according to the order of the sequence of performance evaluation values to adjust the load of the available node cluster.

6. The method according to claim 2, wherein Screening out the available node information that meets the second preset condition from the available node information cluster as the node information of the load to be adjusted, and obtaining the node information cluster of the load to be adjusted, including: For each available node in the available node cluster, perform the following addition steps: Determine the number of hot shards in the hot shard cluster of the available node, where the hot shards in the hot shard cluster are the shards that receive a preset number of access requests within a preset time; In response to determining that the number of hot shards is greater than or equal to the preset hot shard number threshold, add the available node information corresponding to the available node to the preset node information cluster of the load to be adjusted to obtain the added preset node information cluster of the load to be adjusted as the node information cluster of the load to be adjusted.

7. A shard allocation device, including: A first acquisition unit configured to acquire the total amount of log data, where the total amount of log data is the total amount of log data within a preset number of days; A first determination unit configured to determine the average daily log data volume according to the total amount of log data; A second acquisition unit configured to acquire the shard log data volume of each shard storing log data; A second determination unit configured to determine the total number of shards created per day according to the average daily log data volume and the shard log data volume; A third acquisition unit configured to acquire the number of nodes in the distributed storage cluster; A third determination unit configured to determine the number of indexes according to the total number of shards and the number of nodes, where determining the number of indexes according to the total number of shards and the number of nodes includes: determining the product of the number of nodes and a predetermined index ratio as the first number of indexes; determining the sum of the first number of indexes and the total number of shards as the second number of indexes; determining the ratio of the second number of indexes to the first number of indexes as the number of indexes; A fourth determination unit configured to determine the number of shards corresponding to each index according to the total number of shards and the number of indexes; A screening unit configured to screen out the node information that meets the first preset condition from the node information cluster as the available node information, and obtain the available node information cluster, where the node information cluster is the node information cluster corresponding to the distributed storage cluster; An allocation unit configured to allocate the shards corresponding to the number of shards to the available node cluster corresponding to the available node information cluster to adjust the load of the available node cluster corresponding to the available node information cluster.

8. An electronic device, including: One or more processors; A storage device on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, wherein, The computer program implements the method according to any one of claims 1-6 when executed by the processor.

Citation Information

Patent Citations

  • Method and device for assessing distributed cluster index fragmenting and electronic equipment

    CN108897858A

  • An Elasticsearch index fragment optimization method

    CN109582758A