Parallel Pond Sampling Dynamic Consistent Hashing Partition Processing Method and System
Through parallel pond sampling and dynamically consistent hash partitioning strategy, node processing speed is calculated and virtual nodes is set, which solves the load imbalance problem of MapReduce framework in heterogeneous environments, and improves overall operating efficiency and node utilization.
Patent Information
- Application Number
- CN202111628827.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-28
AI Technical Summary
The MapReduce framework has unbalanced load on Reduce nodes in heterogeneous environments, resulting in problems of low overall operation efficiency and low node utilization.
The parallel pond sampling algorithm is used to sample data, the node processing speed is calculated using the heartbeat mechanism, and data allocation is performed through the dynamic consistent hash partitioning strategy, and virtual nodes are set to improve the data volume of nodes with fast node processing speed.
It improves the overall operation efficiency of the MapReduce framework in a heterogeneous environment and the utilization rate of Reduce nodes, and solves the problem of load imbalance of Reduce nodes.
Smart Images

Figure CN114327893B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed parallel computing, and particularly to a parallel pond sampling dynamic consistent hash partitioning processing method and system. Background Art
[0002] The heterogeneity problem of the MapReduce framework is an important issue affecting the framework performance. In a heterogeneous environment, due to problems such as the computing power of Reduce nodes and network latency, the idea of balanced allocation is not applicable to heterogeneous environments. In response to such problems, many scholars at home and abroad have proposed various strategies and methods, which can be summarized as using partitioning strategies, task scheduling algorithms, and other aspects for processing.
[0003] The different performances of each node in MapReduce in a heterogeneous environment result in a reduction in the framework efficiency. The most direct factor is that the heterogeneity problem of nodes is not considered when allocating data. Therefore, a custom data partitioning strategy can be used according to the node performance to solve this problem. Sum et al. considered the diversity of cluster performance (such as CPU write speed and disk write speed, etc.) in terms of dynamically loading data in a heterogeneous cluster and designed a heterogeneous load-aware partitioning function PPF to allocate data by sensing the load status of Reduce nodes. Hanif et al. extended the Hadoop partitioning scheme by adopting an efficient method for allocating Key values, and proposed a node data-aware partitioning algorithm to find the best data allocation position and fairly allocate data to the Reduce side. Maleki et al. proposed a two-stage Map and Reduce task scheduler TMaR, and used a dynamic partitioning binder in the Reduce stage to reduce task transmission when shuffling part of the data, so as to reduce the traffic caused by shuffling. Venkatesh et al. designed a three-layer traffic-aware clustering algorithm to change the original data partitioning, and added an aggregator in the process to reduce the traffic of multiple Map task merges, ultimately shortening the task response time and reducing network traffic. Ma Qingyun et al. believed that there is a performance bottleneck in heterogeneous bandwidth. In this case, according to the uplink and downlink bandwidth of nodes and the size of the initial data volume, the optimal data allocation ratio of each node is calculated, and a performance transmission model between nodes is established, and a bandwidth-based data partitioning method is designed based on the model. Ning et al. proposed a streaming image partitioning process according to the characteristics of image data. Based on modeling the computing power and network environment of heterogeneous nodes, a new adaptive flow graph partitioning function is proposed to minimize the processing duration of graphic jobs.
[0004] The allocation of cluster resources in a heterogeneous environment is determined by task scheduling. Therefore, an efficient task scheduling algorithm is also the key to improving the performance of the MapReduce framework. Dang T et al. established a trust-based framework to handle tasks, assigned each task a certain trust level, and then represented the task scheduling problem of MapReduce as a maximum weighted matching problem of a bipartite graph, which aims to maximize the total trust value of all possible assignments according to the different trust requirements of different tasks. Naik, Wang et al. both formulated task scheduling strategies based on the data locality principle, and input data blocks to the selected nodes according to node performance and the number of available slots respectively. Liao Bin et al. proposed an ItemBased algorithm based on distributed caching to address the problem of resource utilization decline caused by a large number of redundant operations between jobs with unified algorithms. The algorithm processes the I / O data between multiple tasks using caching, breaks the independence defect between jobs, and reduces the waiting latency between Map and Reduce tasks. Chen C et al. proposed a new task scheduling strategy using a bipartite graph model for some MapReduce tasks with deadline requirements, transforming the deadline-constrained problem into a minimum weighted bipartite matching problem of a graph to obtain the optimal solution to this problem. In addition, some scholars have achieved good results by applying various swarm intelligence algorithms to the MapReduce framework for task scheduling.
[0005] In summary, when solving the problem of node heterogeneity in the MapReduce framework, whether using data partitioning strategies or task scheduling algorithms, factors such as the data processing capabilities of nodes and data locality are considered. However, due to the high time complexity of task scheduling algorithms and the fact that traditional data partitioning strategies often use dynamic allocation to solve heterogeneity problems, the initial allocation results directly affect the overall performance of the framework. Therefore, overall performance needs to be considered. The default Hash partitioning strategy in the MapReduce framework in a heterogeneous environment results in uneven processing times of data at Reduce nodes and reduces the utilization rate of each Reduce node. Traditional MapReduce partitioning strategies also have problems such as load imbalance, slow speed, and resource consumption in a heterogeneous environment. Summary of the Invention
[0006] The purpose of the present invention is to provide a parallel reservoir sampling dynamic consistent hash partitioning processing method and system to solve the problem of load balancing of Reduce nodes in a heterogeneous environment and improve the overall operation efficiency of the MapReduce framework and the utilization rate of each Reduce node in a heterogeneous environment.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A parallel pond sampling dynamic consistent hashing partitioning processing method, comprising:
[0009] Performing parallel data sampling using a parallel pond sampling algorithm to obtain sampled data;
[0010] Calculating the processing speed of each node using a heartbeat mechanism according to the sampled data;
[0011] Performing data allocation using a dynamic consistent hashing partitioning strategy according to the processing speed of each node, and allocating the data to be processed to the corresponding Reduce nodes for data processing.
[0012] Optionally, the performing parallel data sampling using a parallel pond sampling algorithm to obtain sampled data specifically includes:
[0013] Setting the sampling rate and the maximum number of sampling partitions of the parallel pond sampling algorithm;
[0014] Reading the partition information splits array through an InputFormat component;
[0015] Determining the number of sampling partitions according to the maximum number of sampling partitions and the total number of partitions in the splits array;
[0016] Calculating the number of samples required for each partition according to the sampling rate and the number of sampling partitions;
[0017] Performing parallel sampling in each partition at the map side according to the number of samples required for each partition to generate sampled data.
[0018] Optionally, the calculating the processing speed of each node using a heartbeat mechanism according to the sampled data specifically includes:
[0019] Marking the sampled data to generate marked data;
[0020] Allocating the marked data to the reduc nodes through the default shuffle partitioning method;
[0021] Obtaining the marked data volume and the reduce processing time of the marked data processed by each reduc node through a heartbeat mechanism;
[0022] Calculating the processing speed of each node according to the marked data volume and the reduce processing time.
[0023] Optionally, the performing data allocation using a dynamic consistent hashing partitioning strategy according to the processing speed of each node, and allocating the data to be processed to the corresponding Reduce nodes for data processing specifically includes:
[0024] Construct a circular space of the consistent hashing algorithm and place the real nodes on the circular space;
[0025] Set virtual nodes for the real nodes according to the processing speed of each node;
[0026] Obtain the data to be processed;
[0027] When the data to be processed enters the circular space, calculate the position of the data to be processed in the circular space;
[0028] According to the position of the data to be processed in the circular space, distribute the data to be processed to the corresponding Reduce nodes for data processing in a clockwise traversal manner.
[0029] Optionally, the setting of virtual nodes for the real nodes according to the processing speed of each node specifically includes:
[0030] Calculate the greatest common divisor of all speeds according to the processing speed of each node;
[0031] Calculate the number of virtual nodes that each real node needs to set on the circular space according to the greatest common divisor;
[0032] Set the virtual nodes of each real node on the circular space according to the number of virtual nodes.
[0033] A parallel pond sampling dynamic consistent hashing partition processing system, including:
[0034] A parallel data sampling module, which is used to perform parallel data sampling by using the parallel pond sampling algorithm to obtain sampled data;
[0035] A processing speed calculation module, which is used to calculate the processing speed of each node by using the heartbeat mechanism according to the sampled data;
[0036] A dynamic consistent hashing partition module, which is used to perform data allocation by using the dynamic consistent hashing partition strategy according to the processing speed of each node, and distribute the data to be processed to the corresponding Reduce nodes for data processing.
[0037] Optionally, the parallel data sampling module specifically includes:
[0038] Sampling parameter setting corresponding, which is used to set the sampling rate and the maximum number of sampling partitions of the parallel pond sampling algorithm;
[0039] Partition information reading corresponding, which is used to read the partition information splits array through the InputFormat component;
[0040] A sampling partition number calculation unit, configured to determine the sampling partition number according to the maximum sampling partition number and the total number of partitions in the splits array;
[0041] A partition sampling number calculation unit, configured to calculate the number of samples required for each partition according to the sampling rate and the sampling partition number;
[0042] A parallel sampling unit, configured to perform parallel sampling in each partition at the map side according to the number of samples required for each partition, and generate sampled data.
[0043] Optionally, the processing speed calculation module specifically includes:
[0044] A data marking unit, configured to mark the sampled data to generate marked data;
[0045] A marked data distribution unit, configured to distribute the marked data to the reduc nodes through a default shuffle partition method;
[0046] A speed parameter acquisition unit, configured to obtain the marked data volume processed by each reduc node and the reduce processing time through a heartbeat mechanism;
[0047] A processing speed calculation unit, configured to calculate the processing speed of each node according to the marked data volume and the reduce processing time.
[0048] Optionally, the dynamic consistent hash partition module specifically includes:
[0049] A circular space construction unit, configured to construct a circular space of the consistent hash algorithm and place real nodes on the circular space;
[0050] A virtual node setting unit, configured to perform virtual node setting on the real nodes according to the processing speed of each node;
[0051] A to-be-processed data acquisition unit, configured to acquire to-be-processed data;
[0052] A data position calculation unit, configured to calculate the position of the to-be-processed data in the circular space when the to-be-processed data enters the circular space;
[0053] A data distribution unit, configured to distribute the to-be-processed data to corresponding Reduce nodes for data processing in a clockwise traversal manner according to the position of the to-be-processed data in the circular space.
[0054] Optionally, the virtual node setting unit specifically includes:
[0055] The greatest common divisor calculation subunit is used to calculate the greatest common divisor of all speeds according to the processing speeds of each node;
[0056] The virtual node number calculation subunit is used to calculate the number of virtual nodes that each real node needs to set on the circular space according to the greatest common divisor;
[0057] The virtual node setting subunit is used to set the virtual nodes of each real node on the circular space according to the number of virtual nodes.
[0058] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:
[0059] The present invention provides a parallel pond sampling dynamic consistent hash partitioning processing method and system. The method includes: performing parallel data sampling by using a parallel pond sampling algorithm to obtain sampled data; calculating the processing speed of each node by using a heartbeat mechanism according to the sampled data; and performing data allocation by using a dynamic consistent hash partitioning strategy according to the processing speeds of each node, and allocating the data to be processed to the corresponding Reduce nodes for data processing. The present invention proposes a two-stage partitioning strategy for the heterogeneity problem in the MapReduce framework. In the first stage, the parallel pond sampling algorithm is used to sample the data and calculate the processing speeds of each node. In the second stage, the dynamic consistent hash partitioning strategy is used for data allocation. Virtual nodes are set for the consistent hash partitioning algorithm according to the processing speeds of the nodes, so that the nodes with faster speeds process more data, thereby improving the overall operation efficiency of the MapReduce framework in a heterogeneous environment and the utilization rate of each Reduce node, and solving the load balancing problem of the Reduce nodes in a heterogeneous environment. Description of the Drawings
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0061] Figure 1 It is a flowchart of a parallel pond sampling dynamic consistent hash partitioning processing method of the present invention;
[0062] Figure 2 It is a schematic diagram of the inventive concept of a parallel pond sampling dynamic consistent hash partitioning processing method of the present invention;
[0063] Figure 3Schematic diagram of the parallel pond sampling process of a parallel pond sampling dynamic consistent hashing partitioning method of the present invention;
[0064] Figure 4 Schematic diagram of the dynamic consistent hashing partitioning process of a parallel pond sampling dynamic consistent hashing partitioning method of the present invention;
[0065] Figure 5 Schematic diagram of the sampling time comparison provided by an embodiment of the present invention;
[0066] Figure 6 Schematic diagram of the comparison of the overall running time of tasks provided by an embodiment of the present invention. Detailed implementation manners
[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0068] The purpose of the present invention is to provide a parallel pond sampling dynamic consistent hashing partitioning method and system to solve the problem of load balancing of Reduce nodes in a heterogeneous environment and improve the overall running efficiency of the MapReduce framework and the utilization rate of each Reduce node in a heterogeneous environment.
[0069] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0070] Figure 1 Flowchart of a parallel pond sampling dynamic consistent hashing partitioning method of the present invention, Figure 2 Schematic diagram of the inventive concept of a parallel pond sampling dynamic consistent hashing partitioning method of the present invention. Refer to Figure 2 , the inventive concept of a parallel pond sampling dynamic consistent hashing partitioning method of the present invention is based on a two-stage partitioning strategy, namely a parallel pond sampling consistent hashing partitioning strategy (PSDC). This strategy performs parallel data sampling using the pond sampling algorithm in the first stage and then estimates the running efficiency of each node using the heartbeat mechanism. In the second stage, virtual nodes are dynamically set for each real node by combining the running efficiency of each node with the consistent hashing algorithm to implement a new partitioning strategy.
[0071] Refer to Figure 1 and Figure 2, a parallel pond sampling dynamic consistent hashing partition processing method of the present invention specifically includes the following steps:
[0072] Step 101: Use the parallel pond sampling algorithm to perform parallel data sampling to obtain the sampled data.
[0073] The parallel pond sampling consistent hashing partition strategy (PSDC) adopted by the method of the present invention is generally processed in two stages. In the first stage, parallel pond sampling is performed for data sampling, and in the second stage, dynamic consistent hashing partition allocation is performed. In the first stage, the original sampling strategy is changed from being operated by the client side to using the map side parallel calculation of the mapReduce computing framework. At the same time, the original random sampler algorithm is changed to the pond sampling idea, which can reduce the time complexity of some intermediate calculations.
[0074] Figure 3 It is a schematic diagram of the parallel pond sampling process of a parallel pond sampling dynamic consistent hashing partition processing method of the present invention. Refer to Figure 3 , in step 101, the parallel pond sampling algorithm is used to perform parallel data sampling to obtain the sampled data, which specifically includes:
[0075] Step 1.1: Set the sampling rate and the maximum number of sampling partitions of the parallel pond sampling algorithm.
[0076] Set the sampling rate samprobablity and the maximum number of sampling partitions maxSplits of the parallel pond sampling algorithm to control the specific amount of sampled data. Among them, the sampling rate samprobablity represents the proportion of the sampled data in the total number, and the maximum number of sampling partitions maxSplits represents the number of slices sampled when the data enters the Map stage for data slicing.
[0077] Step 1.2: Read the partition information splits array through the InputFormat component.
[0078] Next, read the entire slice information splits array through the InputFormat component. The splits array represents all the slice information obtained after slicing in the Map stage, and a new sampled data set samples is created to store the sampled data.
[0079] Step 1.3: Determine the number of sampling partitions according to the maximum number of sampling partitions and the total number of partitions in the splits array.
[0080] After determining the maximum number of sampling partitions maxSplits and the array of split information splits, the number of sampling partitions splitsToSample can be obtained to determine how many partitions need to be sampled for subsequent sampling. The specific process of obtaining the number of sampling partitions splitsToSample is as follows: Select the smaller number between the maximum number of sampling partitions and the total number of partitions to prevent the set maximum number of sampling partitions maxSplits from exceeding the original number of splits and causing errors in the program. The calculation formula for the number of sampling partitions splitsToSample is as follows:
[0081] splitsToSample = Math.min(maxSplits, splits.length)
[0082] where splits.length represents the length of the splits array, that is, the total number of partitions; Math.min() represents returning the smaller value in ().
[0083] Step 1.4: Calculate the number of samples required for each partition according to the sampling rate and the number of sampling partitions.
[0084] After determining the number of partitions splitsToSample that need to be sampled, randomly select splitsToSample partitions for sampling. First, it is necessary to calculate the number of samples samplesPerSplit required for each partition to facilitate subsequent sampling in each partition. The amount of data to be sampled for each partition is the total amount of overall data * the sampling rate set above / the number of partitions to be sampled. The calculation formula is as follows:
[0085] samplesPerSplit = dataSum * samprobablity / splitsToSample
[0086] where samplesPerSplit represents the number of samples required for each partition; dataSum represents the total amount of overall data.
[0087] Step 1.5: Perform parallel sampling in each partition on the map side according to the number of samples required for each partition to generate the sampled data.
[0088] To ensure the complete randomness of the sampling algorithm, the idea of the reservoir sampling algorithm is incorporated into the hadoop sampler, and parallel sampling is performed in each partition on the map side according to the number of samples required for each partition samplesPerSplit to generate the sampled data. See Figure 3First, take out the first samplesPerSplit data in the partition and store them in the data set samples. Next, judge each subsequent element and generate a random number from 1 to row (row represents the number of current data rows). If this random number is less than samplesPerSplit, replace this element. When all elements in the partition are traversed, sampling is completed and the sampled data is obtained.
[0089] In the first stage of the present invention, the default sampler of Hadoop is changed from being operated by the original client to being operated by the mapreduce computing framework, and data sampling is performed on the map side through the idea of parallel computing of the framework, thereby enhancing the sampling efficiency. In addition, the first stage of the present invention uses a parallel pond sampling algorithm to solve the problems that the default shard sampler of Hadoop has poor randomness, the random sampler has a slow sampling speed, relatively more operations, and wastes a lot of time.
[0090] Step 102: Calculate the processing speed of each node using the heartbeat mechanism based on the sampled data.
[0091] The step 102 marks the sampled data, performs common shuffle partitioning and reduce function processing, and distributes the marked data to the reduce node. Through the heartbeat mechanism of Yarn, a resource management control platform in MapReduce 2.x, the data that needs to be calculated from the client input is allocated resources. Next, the child node regularly sends the load status of each node to the ResourceManager through the heartbeat mechanism of Yarn, including the occupancy rate of each node, the amount of data processed, and the processing time. Finally, the approximate computing speed of each node is obtained based on the amount of data processed and the processing time, which is convenient for the consistent hash algorithm to set the virtual node.
[0092] Unlike MapReduce 1.x, MapReduce 2.x proposed Yarn to decouple the new-era JobTracker function. From then on, Yarn is only responsible for resource scheduling, which is equivalent to a distributed operating system platform, while computing programs such as MapReduce are equivalent to applications running on the operating system.
[0093] Therefore, the step 102 calculates the processing speed of each node using the heartbeat mechanism according to the sampled data, and specifically includes:
[0094] Step 2.1: labeling the sampled data to generate labeled data;
[0095] Step 2.2: distribute the marked data to the reduce node through the default shuffle partitioning method;
[0096] Step 2.3: Obtain the marked data volume and the reduce processing time of the marked post-data processed by each of the reduc nodes through a heartbeat mechanism;
[0097] Step 2.4: Calculate the processing speed of each node according to the marked data volume and the reduce processing time.
[0098] Mark the sampled data, allocate it to the reduc nodes through the default shuffle partitioning method, and then obtain the data volume and processing time processed by each node through the heartbeat mechanism of Hadoop, so as to obtain the node operation speed.
[0099] Step 103: According to the processing speed of each node, adopt a dynamic consistent hashing partitioning strategy for data allocation, and allocate the data to be processed to the corresponding Reduce nodes for data processing.
[0100] Figure 4 It is a schematic diagram of the dynamic consistent hashing partitioning process of a parallel pond sampling dynamic consistent hashing partitioning processing method of the present invention. Refer to Figure 2 and Figure 4 , enter the second stage of the parallel pond sampling dynamic consistent hashing partitioning strategy, refer to Figure 4 First, construct a circular space; and based on the node calculation speed obtained in the first stage, use the Euclidean algorithm to find the number of virtual nodes to be set, and then dynamically set the virtual nodes after hash processing. Adopt the consistent hashing algorithm for data allocation. When the data enters the allocation partition stage, calculate the position where this data falls into the circular space through the consistent hashing algorithm, and then find the first virtual node / real node on the circular space through the clockwise traversal method for allocation.
[0101] Therefore, the step 103 adopts a dynamic consistent hashing partitioning strategy for data allocation according to the processing speed of each node, and allocates the data to be processed to the corresponding Reduce nodes for data processing, specifically including:
[0102] Step 3.1: Construct a circular space of the consistent hashing algorithm and place the real nodes on the circular space.
[0103] Enter the second stage of the two-stage partitioning strategy. First, construct a circular space of the consistent hashing algorithm, and then place the real nodes on the circular space. To ensure that the nodes with higher operation speed process more data, it is necessary to set virtual nodes for the real nodes.
[0104] Step 3.2: Perform virtual node setting on the real nodes according to the processing speed of each node.
[0105] To ensure that nodes with faster computing speeds receive more data, more virtual nodes need to be set for them. The step 3.2 performs virtual node setting on the real nodes according to the processing speed of each node, which specifically includes:
[0106] Step 3.2.1: Calculate the greatest common divisor of all speeds according to the processing speed of each node;
[0107] First, round the computing speeds (processing speeds) of each node, and then use the Euclidean algorithm to find the greatest common divisor Vcmax of all speeds.
[0108] Step 3.2.2: Calculate the number of virtual nodes that each real node needs to set on the circular space according to the greatest common divisor.
[0109] Calculate the number of virtual nodes that each node needs to set on the circular space according to the greatest common divisor Vcmax:
[0110] VMNum[i] = V[i] / Vcmax - 1
[0111] where VMNum[i] represents the number of virtual nodes that the i-th node needs to set on the circular space; V[i] represents the running speed (processing speed) of node i.
[0112] Step 3.2.3: Set the virtual nodes of each real node on the circular space according to the number of virtual nodes.
[0113] Set the virtual nodes of each real node on the circular space according to the number of virtual nodes VMNum[i]. At this time, the sum Vnums of all virtual nodes and real nodes existing on the ring is as follows:
[0114]
[0115] where n is the number of Reduce nodes.
[0116] Step 3.3: Obtain the data to be processed.
[0117] Step 3.4: When the data to be processed enters the circular space, calculate the position of the data to be processed in the circular space.
[0118] Set a new partitioning function to change the original default hash partitioning to consistent hash partitioning. When the data to be processed enters the circular space, calculate the position of the data in the circular space:
[0119] (key.hashCode() & 2^32 - 1) % Vnums
[0120] Where key represents the key value of the element, hashcode represents the hash value obtained based on key, and Vnums represents the sum of all nodes.
[0121] Step 3.5: According to the position of the data to be processed in the circular space, the data to be processed is allocated to the corresponding Reduce nodes for data processing in a clockwise traversal manner.
[0122] See Figure 4 , after the position of the data in the circular space has been calculated, starting from this position, traverse the circular space clockwise like a clock one by one to find the first node in the circular space. If the node is a real node, it is the machine position where the data needs to be allocated. If it is a virtual node, map it to a real node.
[0123] Finally, through this partitioning strategy, the data is allocated to each Reduce node, achieving a load balancing effect in the processing time of the Reduce nodes, and finally completing the parallel pond sampling dynamic consistent hash partitioning strategy of the present invention.
[0124] Figure 5 This is a schematic diagram of the sampling time comparison provided by the embodiments of the present invention. Refer to Figure 5 for a sampling comparison experiment to verify the sampling efficiency. As Figure 5 shown, the parallel pond sampling method adopted by the method of the present invention significantly reduces the sampling time and improves the sampling efficiency compared with the default random sampling and the default sharding sampling.
[0125] Figure 6 This is a schematic diagram of the overall running time comparison of tasks provided by the embodiments of the present invention. Refer to Figure 6 for an algorithm comparison experiment to verify the efficiency of the parallel pond sampling dynamic consistent hash partitioning strategy. As Figure 6 shown, the parallel pond sampling consistent hash partitioning strategy (i.e., the PSDC partitioning strategy) adopted by the method of the present invention reduces the overall running time and improves the overall running efficiency of the MapReduce framework in a heterogeneous environment compared with the DTHE partitioning strategy and the Hadoop default partitioning strategy, regardless of the data volume.
[0126] In view of the problem of the deteriorated performance of the mapReduce framework due to the different computing performances of each node, the present invention proposes a parallel pond sampling dynamic consistent hash partitioning processing method, which is based on a two-stage data partitioning strategy, namely the parallel pond sampling consistent hash partitioning strategy (PSDC). The method of the present invention performs parallel data sampling by using the idea of the pond sampling algorithm in the first stage, and then estimates the operating efficiency of each node by using the heartbeat mechanism. In the second stage, virtual nodes are dynamically set for each real node by combining the operating efficiency of each node with the consistent hash algorithm to implement a new partitioning strategy. Finally, it can be seen from the experimental comparison that the method of the present invention speeds up the overall running speed of MapReduce in a heterogeneous cluster and improves the utilization rate of nodes.
[0127] Based on a parallel pond sampling dynamic consistent hash partitioning processing method provided by the present invention, the present invention also provides a parallel pond sampling dynamic consistent hash partitioning processing system, which includes:
[0128] A parallel data sampling module, which is used to perform parallel data sampling by using a parallel pond sampling algorithm to obtain sampled data;
[0129] A processing speed calculation module, which is used to calculate the processing speed of each node by using the heartbeat mechanism according to the sampled data;
[0130] A dynamic consistent hash partitioning module, which is used to perform data allocation by using a dynamic consistent hash partitioning strategy according to the processing speed of each node, and allocate the data to be processed to the corresponding Reduce nodes for data processing.
[0131] Among them, the parallel data sampling module specifically includes:
[0132] Sampling parameter setting corresponding, which is used to set the sampling rate and the maximum number of sampling partitions of the parallel pond sampling algorithm;
[0133] Partition information reading corresponding, which is used to read the partition information splits array through the InputFormat component;
[0134] A sampling partition number calculation unit, which is used to determine the sampling partition number according to the maximum number of sampling partitions and the total number of partitions in the splits array;
[0135] A partition sampling number calculation unit, which is used to calculate the number of samples required for each partition according to the sampling rate and the sampling partition number;
[0136] A parallel sampling unit, which is used to perform parallel sampling in each partition at the map end according to the number of samples required for each partition to generate sampled data.
[0137] The processing speed calculation module specifically includes:
[0138] A data marking unit, configured to mark the sampled data to generate marked data;
[0139] A marked data distribution unit, configured to distribute the marked data to the reduc nodes through the default shuffle partitioning method;
[0140] A speed parameter acquisition unit, configured to obtain the marked data volume and the reduce processing time of the marked data processed by each reduc node through a heartbeat mechanism;
[0141] A processing speed calculation unit, configured to calculate the processing speed of each node according to the marked data volume and the reduce processing time.
[0142] The dynamic consistent hash partitioning module specifically includes:
[0143] A circular space construction unit, configured to construct a circular space of the consistent hash algorithm and place real nodes on the circular space;
[0144] A virtual node setting unit, configured to perform virtual node setting on the real nodes according to the processing speed of each node;
[0145] A to-be-processed data acquisition unit, configured to acquire to-be-processed data;
[0146] A data position calculation unit, configured to calculate the position of the to-be-processed data in the circular space when the to-be-processed data enters the circular space;
[0147] A data distribution unit, configured to distribute the to-be-processed data to the corresponding Reduce nodes for data processing in a clockwise traversal manner according to the position of the to-be-processed data in the circular space.
[0148] The virtual node setting unit specifically includes:
[0149] A greatest common divisor calculation subunit, configured to calculate the greatest common divisor of all speeds according to the processing speed of each node;
[0150] A virtual node quantity calculation subunit, configured to calculate the number of virtual nodes that each real node needs to set on the circular space according to the greatest common divisor;
[0151] A virtual node setting subunit, configured to set the virtual nodes of each real node on the circular space according to the number of virtual nodes.
[0152] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method section.
[0153] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A parallel pond sampling dynamic consistency hash partitioning processing method, characterized in that, Including: Using the parallel reservoir sampling algorithm to perform parallel data sampling to obtain the sampled data; The step of using the parallel reservoir sampling algorithm to perform parallel data sampling to obtain the sampled data specifically includes: Setting the sampling rate and the maximum number of sampling partitions of the parallel reservoir sampling algorithm; Reading the partition information splits array through the InputFormat component; Determining the number of sampling partitions according to the maximum number of sampling partitions and the total number of partitions in the splits array; Calculating the number of samples required for each partition according to the sampling rate and the number of sampling partitions; Performing parallel sampling in each partition at the map side according to the number of samples required for each partition to generate the sampled data; Setting the sampling rate samprobablity and the maximum number of sampling partitions maxSplits of the parallel reservoir sampling algorithm to control the specific amount of sampled data; where the sampling rate samprobablity represents the proportion of the sampled data in the total number, and the maximum number of sampling partitions maxSplits represents the number of slices sampled after the data is sliced in the Map stage; Reading the entire splits array of slice information through the InputFormat component, where the splits array represents all the slice information obtained after slicing in the Map stage, and creating a sampled data set samples to store the sampled data; The calculation formula for the number of sampling partitions splitsToSample is as follows: splitsToSample = Math.min(maxSplits, splits.length) where splits.length represents the length of the splits array, that is, the total number of partitions; Math.min() represents returning the smaller value in (); The calculation formula for the number of samples required for each partition is as follows: samplesPerSplit = dataSum * samprobablity / splitsToSample where samplesPerSplit represents the number of samples required for each partition; dataSum represents the total amount of the overall data; First, take out the first samplesPerSplit data in the partition and store them in the data set samples, and then judge each subsequent element, generate a random number from 1 to row, where row represents the current data row number. If this random number is less than samplesPerSplit, then replace this element; when all the elements in the partition are traversed, the sampling is completed to obtain the sampled data; Calculating the processing speed of each node using the heartbeat mechanism according to the sampled data; Performing data allocation using the dynamic consistent hashing partition strategy according to the processing speed of each node, and allocating the data to be processed to the corresponding Reduce nodes for data processing.
2. The method according to claim 1, characterized in that, The step of calculating the processing speed of each node using the heartbeat mechanism according to the sampled data specifically includes: Marking the sampled data to generate the marked data; Allocate the marked data to the Reduc nodes through the default shuffle partitioning method; Obtain the marked data volume and reduce processing time of the marked data processed by each Reduc node through the heartbeat mechanism; Calculate the processing speed of each node according to the marked data volume and the reduce processing time.
3. The method according to claim 2, wherein According to the processing speed of each node, adopt the dynamic consistent hashing partitioning strategy for data allocation, and allocate the data to be processed to the corresponding Reduce nodes for data processing. Specifically, it includes: Construct a circular space of the consistent hashing algorithm and place the real nodes on the circular space; Perform virtual node settings on the real nodes according to the processing speed of each node; Obtain the data to be processed; When the data to be processed enters the circular space, calculate the position of the data to be processed in the circular space; According to the position of the data to be processed in the circular space, allocate the data to be processed to the corresponding Reduce nodes for data processing in a clockwise traversal manner.
4. The method according to claim 3, characterized in that The virtual node setting of the real nodes according to the processing speed of each node specifically includes: Calculate the greatest common divisor of all speeds according to the processing speed of each node; Calculate the number of virtual nodes that each real node needs to set on the circular space according to the greatest common divisor; Set the virtual nodes of each real node on the circular space according to the number of virtual nodes.
5. A parallel pond sampling dynamic consistent hash partitioning processing system, characterized in that, It includes: A parallel data sampling module for performing parallel data sampling using the parallel reservoir sampling algorithm to obtain the sampled data; The parallel data sampling module specifically includes: Sampling parameter setting corresponding for setting the sampling rate and the maximum number of sampling partitions of the parallel reservoir sampling algorithm; Partition information reading corresponding for reading the partition information splits array through the InputFormat component; Sampling partition number calculation unit for determining the sampling partition number according to the maximum number of sampling partitions and the total number of partitions in the splits array; Partition sampling number calculation unit for calculating the number of samples required for each partition according to the sampling rate and the sampling partition number; Parallel sampling unit for performing parallel sampling in each partition at the map side according to the number of samples required for each partition to generate the sampled data; Set the sampling rate samprobablity and the maximum number of sampling partitions maxSplits of the parallel reservoir sampling algorithm to control the specific amount of sampled data; where the sampling rate samprobablity represents the proportion of the sampled data in the total, and the maximum number of sampling partitions maxSplits represents the number of slices sampled after the data is sliced in the Map stage; Read all the split information splits array through the InputFormat component. The splits array represents all the split information obtained after the data is sliced in the Map stage, and create a sampled data set samples to store the sampled data; The calculation formula for the number of sampling partitions splitsToSample is as follows: splitsToSample = Math.min(maxSplits, splits.length) where splits.length represents the length of the splits array, i.e., the total number of partitions; Math.min() represents returning the smaller value in (). The calculation formula for the number of samples to be taken for each partition is as follows: samplesPerSplit = dataSum * samprobablity / splitsToSample where samplesPerSplit represents the number of samples to be taken for each partition; dataSum represents the total amount of overall data; First, take the first samplesPerSplit data in the partition and store it in the data set samples. Next, judge each subsequent element, generate a random number from 1 to row, where row represents the current data row number. If this random number is less than samplesPerSplit, then replace this element; when all elements in the partition have been traversed, the sampling is completed, and the sampled data is obtained; A processing speed calculation module, which is used to calculate the processing speed of each node by using the heartbeat mechanism according to the sampled data; A dynamic consistent hash partitioning module, which is used to perform data allocation by adopting a dynamic consistent hash partitioning strategy according to the processing speed of each node, and allocate the data to be processed to the corresponding Reduce nodes for data processing.
6. The system according to claim 5, wherein The processing speed calculation module specifically includes: A data marking unit, which is used to mark the sampled data to generate marked data; A marked data allocation unit, which is used to allocate the marked data to the reduc nodes by the default shuffle partitioning method; A speed parameter acquisition unit, which is used to obtain the marked data volume and the reduce processing time of the marked data processed by each reduc node through the heartbeat mechanism; A processing speed calculation unit, which is used to calculate the processing speed of each node according to the marked data volume and the reduce processing time.
7. The system according to claim 6, wherein The dynamic consistent hash partitioning module specifically includes: A circular space construction unit, which is used to construct a circular space of the consistent hash algorithm and place the real nodes on the circular space; A virtual node setting unit, which is used to perform virtual node setting on the real nodes according to the processing speed of each node; A data to be processed acquisition unit, which is used to acquire the data to be processed; A data position calculation unit, which is used to calculate the position of the data to be processed in the circular space when the data to be processed enters the circular space; A data allocation unit, which is used to allocate the data to be processed to the corresponding Reduce nodes for data processing in a clockwise traversal manner according to the position of the data to be processed in the circular space.
8. The system according to claim 7, wherein The virtual node setting unit specifically includes: The greatest common divisor calculation subunit is used to calculate the greatest common divisor of all speeds according to the processing speed of each node; The virtual node number calculation subunit is used to calculate the number of virtual nodes that each real node needs to set on the circular space according to the greatest common divisor; The virtual node setting subunit is used to set the virtual nodes of each real node on the circular space according to the number of virtual nodes.
Citation Information
Patent Citations
PaaS platform load balancing method based on consistency hash strategy
CN107979646A