HIVE grouping operation performance optimization and scheduling method based on key column partitioning
Through the key column partitioning method, the partitioning strategy and resource allocation of Hive packet operations are dynamically adjusted, which solves the problems of data skew and low resource utilization, and achieves more efficient computing performance and more accurate results.
Patent Information
- Application Number
- CN202510144047.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-10
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] The present invention relates to data processing technology, and in particular to a HIVE grouping operation performance optimization and scheduling method based on key column partitioning. Background Art
[0002] Hive is a data warehouse tool based on Hadoop and is widely used in the field of big data analysis. In Hive, group operation is a common and important operation used to aggregate and analyze large-scale data. As the scale of data continues to grow, how to improve the performance of Hive group operation has become an urgent problem to be solved.
[0003] The traditional Hive grouping operation method mainly relies on the MapReduce framework, which improves computing efficiency by distributing data to multiple nodes for parallel processing. However, this method still has some limitations when processing large-scale data. First, the data partitioning strategy is often fixed and cannot be dynamically adjusted according to the actual data distribution characteristics, which easily leads to data skew problems and affects computing performance. Secondly, resource allocation and task scheduling lack flexibility, making it difficult to fully utilize cluster resources, resulting in resource waste. Finally, for complex grouping operations, existing methods are difficult to effectively handle the correlation between data, affecting the accuracy of the calculation results.
[0004] In order to solve these problems, it is necessary to develop a more intelligent and efficient Hive grouping operation optimization method. This method should be able to adaptively adjust the partitioning strategy according to the data characteristics to achieve a more balanced data distribution. At the same time, it is also necessary to consider the resource requirements and data relevance of the computing tasks and adopt a more optimized scheduling algorithm to improve the cluster resource utilization. In addition, it should also have dynamic load balancing capabilities and be able to monitor and adjust the load status of computing nodes in real time to ensure the efficiency and stability of the entire distributed computing process. Summary of the invention
[0005] The embodiment of the present invention provides a HIVE grouping operation performance optimization and scheduling method based on key column partitioning, which can solve the problems in the prior art.
[0006] According to a first aspect of the embodiments of the present invention,
[0007] Provides HIVE grouping operation performance optimization and scheduling methods based on key column partitioning, including:
[0008] Receive a HIVE computing task submitted by a user, obtain data table information and group operation key column information in the HIVE computing task, build a key column distribution histogram based on the data table information and the group operation key column information, calculate the distribution density and data skewness of the key column data according to the key column distribution histogram, determine a key column partitioning strategy based on the distribution density and data skewness of the key column data, the key column partitioning strategy includes a data partition size, a partition number, and a sampling rate, perform partition preprocessing on the data of the HIVE computing task according to the key column partitioning strategy, and generate an initial data partition set;
[0009] Perform resource evaluation on each data partition in the initial data partition set, generate a resource demand vector based on the data volume, computational complexity and data correlation of each data partition, input the resource demand vector into a pre-trained load prediction model, obtain the predicted execution time and resource occupancy rate of each data partition, construct a cost matrix of the data partition according to the predicted execution time and resource occupancy rate, optimize and reorganize the initial data partition set based on the cost matrix using a dynamic programming algorithm, and generate a target data partition set with the optimal computational cost;
[0010] A plurality of computing nodes that meet computing requirements are selected from a preset computing node resource pool to build a distributed computing cluster. Based on the cost matrix of each data partition in the target data partition set, a minimum spanning tree algorithm is used to calculate the optimal data transmission path between the plurality of computing nodes. The data partitions in the target data partition set are allocated to corresponding computing nodes according to the optimal data transmission path. A grouping operation subtask is started at each computing node. The load status of each computing node is monitored in real time. When a load imbalance is detected, the allocation scheme of the data partitions is dynamically adjusted based on the cost matrix and the optimal data transmission path to ensure that the computing load of the distributed computing cluster is in a dynamically balanced state.
[0011] Constructing a key column distribution histogram based on the data table information and the group operation key column information, calculating the distribution density and data skewness of the key column data according to the key column distribution histogram, and determining the key column partitioning strategy based on the distribution density and data skewness of the key column data includes:
[0012] Perform a full table scan on the data table to obtain key column value frequency statistics information, construct the key column values in the key column value frequency statistics information into a key column value set, construct the frequencies in the key column value frequency statistics information into a frequency set, construct a key column distribution histogram based on the value interval of the key column value set, divide the value interval into multiple sub-intervals, and calculate the frequency density of each sub-interval;
[0013] Calculating data skewness and standard deviation based on the key column distribution histogram, calculating the standard deviation according to the difference between the frequency density and the average density of the multiple sub-intervals, and calculating the Gini coefficient as the data skewness based on the frequency density of the multiple sub-intervals;
[0014] The optimal number of partitions is calculated based on the total amount of data and the data inclination, and the subintervals among the multiple subintervals whose frequency density is greater than the sum of the average density and the standard deviation are taken as high-density areas, and the subintervals among the multiple subintervals whose frequency density is less than or equal to the sum of the average density and the standard deviation are taken as low-density areas, and the high-density areas are divided into a first partition width by fine-grained division, and the low-density areas are divided into a second partition width by coarse-grained division;
[0015] Monitor the data distribution changes of the high-density area and the low-density area in real time, calculate the density difference between two adjacent time points, and when the density difference is greater than a preset density threshold, or the density ratio of adjacent areas is greater than the preset density threshold, or the change in the data inclination is greater than the preset inclination threshold, trigger partition adjustment, re-divide the high-density area and the low-density area according to the first partition width and the second partition width, and generate a new partition scheme.
[0016] Performing resource evaluation on each data partition in the initial data partition set, generating a resource demand vector based on the data volume, computational complexity and data association of each data partition, inputting the resource demand vector into a pre-trained load prediction model, and obtaining the predicted execution time and resource occupancy rate of each data partition includes:
[0017] Constructing a data volume feature subvector according to the data size, number of records and number of columns of the data partition, constructing a computational complexity feature subvector according to the operation operator complexity coefficient, number of data connection operations, number of aggregation functions and number of grouping conditions of the data partition, constructing a data relevance feature subvector according to the input dependent data volume, output dependent data volume and data dependency relationship complexity coefficient of the data partition, and connecting the data volume feature subvector, the computational complexity feature subvector and the data relevance feature subvector to form a unified feature vector;
[0018] Building a load prediction model based on a multi-layer perceptron model, inputting the unified feature vector into an input layer of the multi-layer perceptron model, performing nonlinear transformation processing on the unified feature vector in multiple hidden layers of the multi-layer perceptron model, inputting output results of the multiple hidden layers into an output layer of the multi-layer perceptron model, and generating predicted execution time and predicted resource occupancy rate;
[0019] A training data set is constructed based on historical execution records, data enhancement processing is performed on the training data set, the number of samples in the training data set is expanded by means of characteristic value perturbation, load scenario combination and abnormal sample injection, the multi-layer perceptron model is trained by using the training data set, the mean square error of the predicted execution time and the predicted resource occupancy rate is calculated, and a weighted sum of the mean square error and a regularization parameter is constructed as a loss function;
[0020] Obtain the actual execution time and actual resource occupancy rate of the data partition, calculate a first deviation rate between the predicted execution time and the actual execution time, calculate a second deviation rate between the predicted resource occupancy rate and the actual resource occupancy rate, calculate a dynamic correction coefficient based on the first deviation rate and the second deviation rate, take the product of the predicted execution time and the corresponding dynamic correction coefficient as the final predicted execution time, and take the product of the predicted resource occupancy rate and the corresponding dynamic correction coefficient as the final predicted resource occupancy rate.
[0021] Constructing a cost matrix of data partitions according to the predicted execution time and resource occupancy rate, optimizing and reorganizing the initial data partition set using a dynamic programming algorithm based on the cost matrix, and generating a target data partition set with the optimal computational cost includes:
[0022] Obtaining the predicted execution time and resource occupancy rate of the data partition, constructing a cost function of a single data partition according to the ratio of the predicted execution time to the maximum predicted execution time, the ratio of the resource occupancy rate to the maximum resource occupancy rate, and the ratio of the data migration cost to the maximum data migration cost, and calculating the single body cost of the data partition based on the cost function;
[0023] Calculate the data dependency, data similarity and load balancing factor between two data partitions, take the weighted sum of the data dependency, data similarity and load balancing factor as the additional cost of merging the two data partitions, take the sum of the single cost of the two data partitions and the additional cost as the merging cost of the two data partitions, and construct a cost matrix based on the merging cost;
[0024] A state transfer equation is constructed based on the cost matrix, and the minimum cost of reorganizing the previous data partition into the target data partition is calculated by the state transfer equation. During the reorganization process, the size of the data partition is limited to ensure that the size of the reorganized data partition does not exceed a preset upper limit, the load of the data partition is balanced to ensure that the deviation between the load of the reorganized data partition and the average load does not exceed a preset load threshold, and the locality of the data partition is constrained to ensure that the locality of the reorganized data partition is not lower than a preset lower limit;
[0025] The load balance, data locality and overall cost of the data partition reorganization scheme are calculated, and the weighted sum of the load balance, data locality and overall cost is used as the score of the reorganization scheme. The score change rate of two adjacent iterations is calculated. When the score change rate is less than a preset convergence threshold, the weight parameters in the cost function are adaptively adjusted based on the load balance change rate, resource utilization change rate and migration cost change rate, and the data partition reorganization process is re-executed using the adjusted cost function.
[0026] Selecting multiple computing nodes that meet computing requirements from a preset computing node resource pool to build a distributed computing cluster, and using a minimum spanning tree algorithm to calculate an optimal data transmission path between the multiple computing nodes based on a cost matrix of each data partition in the target data partition set includes:
[0027] Obtain resource information of the computing node, the resource information including CPU usage, memory capacity, network bandwidth and storage capacity, construct a resource vector of the computing node based on the resource information, calculate the ratio of the resource amount of each dimension in the resource vector to the resource demand of the data partition, and use the weighted sum of the ratios as the resource fitness of the computing node;
[0028] Determine the location affinity of the computing node and the data partition, determine the matching degree of the computing node and the data partition based on the resource fitness, the location affinity and the resource consumption cost, calculate the total load demand of the data partition, determine the number of nodes of the distributed computing cluster according to the ratio of the total load demand to the average node capacity, and select multiple computing nodes with the highest matching degree from a preset computing node resource pool to construct the distributed computing cluster;
[0029] Calculate the data size transmitted between computing nodes in the distributed computing cluster, obtain the available bandwidth, network delay and network congestion between the computing nodes, calculate the transmission cost between the computing nodes based on the data size, the available bandwidth, the network delay and the network congestion, and construct a transmission cost matrix;
[0030] Constructing a minimum spanning tree based on the transmission cost matrix, calculating the data flow on the path between any two computing nodes in the minimum spanning tree, ensuring that the data flow does not exceed the bandwidth capacity on the path, and calculating the transmission delay of the path between any two computing nodes in the minimum spanning tree, ensuring that the transmission delay does not exceed a preset delay threshold;
[0031] Calculate the path cost of the data transmission path in the minimum spanning tree, the ratio of the transmission delay to the preset delay threshold, and the ratio of the available bandwidth to the required bandwidth, and calculate the path quality evaluation value based on the path cost, the ratio of the transmission delay to the preset delay threshold, and the ratio of the available bandwidth to the required bandwidth. When the path quality evaluation value is less than the preset quality threshold, reconstruct, split or merge the data transmission path according to the network congestion level and load conditions.
[0032] The method includes allocating data partitions in the target data partition set to corresponding computing nodes according to the optimal data transmission path, starting a grouping operation subtask at each computing node, monitoring the load status of each computing node in real time, and dynamically adjusting the allocation scheme of the data partitions based on the cost matrix and the optimal data transmission path when load imbalance is detected. The method includes:
[0033] Calculating the data location correlation between the data partition and the computing node, the computing power of the computing node and the current load level of the computing node, calculating the affinity between the data partition and the computing node based on the data location correlation, the computing power and the current load level, and allocating the data partition to the corresponding computing node according to the affinity;
[0034] Obtain the data volume of the data partition and the processing capacity of the computing node, determine the task segmentation granularity based on the data ratio of the data volume to the processing capacity, calculate the ratio of the operator load to the preset load threshold, round up the data ratio as the operator parallelism, and the operator parallelism does not exceed the preset maximum parallelism;
[0035] Monitor the CPU usage, memory usage, input / output usage, and network bandwidth usage of the computing nodes in real time, calculate the standard deviation of the computing node load and the average load of all computing nodes, use the ratio of the standard deviation to the average load as the load balancing degree, and trigger a load warning when the computing node load exceeds the load balancing degree;
[0036] Calculate the migration cost of the data partition between the source computing node and the target computing node, the migration cost includes the ratio of the data partition size to the bandwidth between the nodes, the migration preparation overhead and the sum of the service interruption compensation, and calculate the weighted difference between the load balancing improvement and the migration cost as the migration benefit;
[0037] When the migration benefit is greater than the preset benefit threshold, data partition migration is performed; when the data partition size is greater than the preset split threshold, data partition splitting is performed; when the sum of the sizes of adjacent data partitions is less than the preset merge threshold, data partition merging is performed; the weight coefficient of the migration cost is adjusted based on changes in system performance; when the system performance improves, the load balancing adjustment period is reduced; when the system performance decreases, the load balancing adjustment period is increased.
[0038] A second aspect of an embodiment of the present invention provides a HIVE grouping operation performance optimization and scheduling system based on key column partitioning, including:
[0039] The first unit is used to receive a HIVE computing task submitted by a user, obtain data table information and group operation key column information in the HIVE computing task, construct a key column distribution histogram based on the data table information and the group operation key column information, calculate the distribution density and data skewness of the key column data according to the key column distribution histogram, determine a key column partitioning strategy based on the distribution density and data skewness of the key column data, the key column partitioning strategy includes a data partition size, a partition number, and a sampling rate, and perform partition preprocessing on the data of the HIVE computing task according to the key column partitioning strategy to generate an initial data partition set;
[0040] The second unit is used to perform resource evaluation on each data partition in the initial data partition set, generate a resource demand vector based on the data volume, computational complexity and data correlation of each data partition, input the resource demand vector into a pre-trained load prediction model, obtain the predicted execution time and resource occupancy rate of each data partition, construct a cost matrix of the data partition according to the predicted execution time and resource occupancy rate, and optimize and reorganize the initial data partition set using a dynamic programming algorithm based on the cost matrix to generate a target data partition set with the optimal computational cost;
[0041] The third unit is used to select multiple computing nodes that meet the computing requirements from a preset computing node resource pool to build a distributed computing cluster, and based on the cost matrix of each data partition in the target data partition set, use a minimum spanning tree algorithm to calculate the optimal data transmission path between the multiple computing nodes, and allocate the data partitions in the target data partition set to the corresponding computing nodes according to the optimal data transmission path, start a group operation subtask at each of the computing nodes, monitor the load status of each computing node in real time, and when a load imbalance is detected, dynamically adjust the data partition allocation plan based on the cost matrix and the optimal data transmission path to ensure that the computing load of the distributed computing cluster is in a dynamically balanced state.
[0042] A third aspect of the embodiments of the present invention
[0043] An electronic device is provided, comprising:
[0044] processor;
[0045] a memory for storing processor-executable instructions;
[0046] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0047] A fourth aspect of the embodiments of the present invention is:
[0048] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0049] The beneficial effects of this application are as follows:
[0050] The present invention constructs a key column distribution histogram and calculates data distribution characteristics, formulates a targeted key column partitioning strategy, realizes efficient preprocessing of HIVE grouping operation data, and effectively improves the balance and computing efficiency of data partitioning.
[0051] Through resource evaluation and load prediction models, the present invention can accurately estimate the execution cost of each data partition, and use dynamic programming algorithms to optimize and reorganize to obtain a data partitioning scheme with the optimal computational cost, thereby significantly improving the overall performance of HIVE grouping operations.
[0052] The present invention adopts the minimum spanning tree algorithm to optimize the data transmission path and monitors the load status in real time for dynamic adjustment. While ensuring the load balancing of the distributed computing cluster, it minimizes the data transmission overhead and realizes the efficient scheduling and execution of HIVE group computing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a flow chart of a method for optimizing and scheduling HIVE grouping operation performance based on key column partitioning according to an embodiment of the present invention;
[0054] Figure 2 The structure diagram of the HIVE group operation performance optimization and scheduling system based on key column partitioning according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0056] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0057] Figure 1FIG. 1 is a flow chart of a method for optimizing and scheduling HIVE grouping operation performance based on key column partitioning according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0058] S101. Receive a HIVE computing task submitted by a user, obtain data table information and group operation key column information in the HIVE computing task, construct a key column distribution histogram based on the data table information and the group operation key column information, calculate the distribution density and data inclination of the key column data according to the key column distribution histogram, determine a key column partitioning strategy based on the distribution density and data inclination of the key column data, the key column partitioning strategy includes a data partition size, a partition number, and a sampling rate, perform partition preprocessing on the data of the HIVE computing task according to the key column partitioning strategy, and generate an initial data partition set;
[0059] S102. Perform resource evaluation on each data partition in the initial data partition set, generate a resource demand vector based on the data volume, computational complexity and data correlation of each data partition, input the resource demand vector into a pre-trained load prediction model, obtain the predicted execution time and resource occupancy rate of each data partition, construct a cost matrix of the data partition according to the predicted execution time and resource occupancy rate, optimize and reorganize the initial data partition set based on the cost matrix using a dynamic programming algorithm, and generate a target data partition set with the optimal computational cost;
[0060] S103. Select multiple computing nodes that meet the computing requirements from a preset computing node resource pool to build a distributed computing cluster, and based on the cost matrix of each data partition in the target data partition set, use a minimum spanning tree algorithm to calculate the optimal data transmission path between the multiple computing nodes, and allocate the data partitions in the target data partition set to the corresponding computing nodes according to the optimal data transmission path, start a group operation subtask at each of the computing nodes, monitor the load status of each computing node in real time, and when a load imbalance is detected, dynamically adjust the data partition allocation plan based on the cost matrix and the optimal data transmission path to ensure that the computing load of the distributed computing cluster is in a dynamically balanced state.
[0061] In an optional implementation, constructing a key column distribution histogram based on the data table information and the group operation key column information, calculating the distribution density and data skewness of the key column data according to the key column distribution histogram, and determining the key column partitioning strategy based on the distribution density and data skewness of the key column data includes:
[0062] Perform a full table scan on the data table to obtain key column value frequency statistics information, construct the key column values in the key column value frequency statistics information into a key column value set, construct the frequencies in the key column value frequency statistics information into a frequency set, construct a key column distribution histogram based on the value interval of the key column value set, divide the value interval into multiple sub-intervals, and calculate the frequency density of each sub-interval;
[0063] Calculating data skewness and standard deviation based on the key column distribution histogram, calculating the standard deviation according to the difference between the frequency density and the average density of the multiple sub-intervals, and calculating the Gini coefficient as the data skewness based on the frequency density of the multiple sub-intervals;
[0064] The optimal number of partitions is calculated based on the total amount of data and the data inclination, and the subintervals among the multiple subintervals whose frequency density is greater than the sum of the average density and the standard deviation are taken as high-density areas, and the subintervals among the multiple subintervals whose frequency density is less than or equal to the sum of the average density and the standard deviation are taken as low-density areas, and the high-density areas are divided into a first partition width by fine-grained division, and the low-density areas are divided into a second partition width by coarse-grained division;
[0065] Monitor the data distribution changes of the high-density area and the low-density area in real time, calculate the density difference between two adjacent time points, and when the density difference is greater than a preset density threshold, or the density ratio of adjacent areas is greater than the preset density threshold, or the change in the data inclination is greater than the preset inclination threshold, trigger partition adjustment, re-divide the high-density area and the low-density area according to the first partition width and the second partition width, and generate a new partition scheme.
[0066] This embodiment provides an adaptive partitioning strategy method based on data distribution characteristics. The method first performs a full table scan on a data table to obtain frequency statistics of key column values. A key column value set and a frequency set are constructed based on the statistics, and a key column distribution histogram is constructed according to the value interval of the key column value. The value interval is divided into multiple sub-intervals, and the frequency density of each sub-interval is calculated.
[0067] Next, the data skewness and standard deviation are calculated based on the key column distribution histogram. Specifically, the standard deviation is calculated based on the difference between the frequency density of each subinterval and the average density, and the Gini coefficient is calculated based on the subinterval frequency density as the data skewness. For example, assuming there are 5 subintervals, the frequency densities are 10, 20, 15, 30, and 25, and the average density is 20. Calculate the sum of the squares of the differences between each subinterval and the average density, divide it by the number of subintervals, and take the square root to get the standard deviation. The calculation of the Gini coefficient is based on the cumulative distribution of the frequency density.
[0068] Then, the optimal number of partitions is calculated based on the total amount of data and the data tilt. The sub-intervals with frequency density greater than the sum of the average density and the standard deviation are defined as high-density areas, and the others are low-density areas. The high-density area is divided into the first partition width using fine-grained division, and the low-density area is divided into the second partition width using coarse-grained division. For example, the high-density area can be divided into 1 / 4 of the atomic interval size, and the low-density area can be divided into 2 times the atomic interval size.
[0069] Finally, the data distribution changes of high-density and low-density areas are monitored in real time. The density difference between two adjacent time points is calculated. When the difference is greater than the preset density threshold (such as 20%), or the density ratio of adjacent areas is greater than the preset density threshold (such as 2), or the change in data inclination is greater than the preset inclination threshold (such as 0.1), the partition adjustment is triggered. The high-density area and the low-density area are re-divided according to the first partition width and the second partition width to generate a new partition scheme.
[0070] The specific implementation process of this method is as follows:
[0071] First, perform a full table scan on the data table and count the frequency of occurrence of each key column value. For example, for a table containing user IDs and transaction amounts, you can count the number of times each user ID appears. Store the statistical results in the form of key-value pairs, with the key being the user ID and the value being the number of occurrences.
[0072] Next, construct the key column values in the statistical results into a key column value set, and construct the frequencies into a frequency set. Determine the value range of the key column value, for example, the range of user ID is 1-1000000. Divide the interval into several sub-intervals, such as 100 sub-intervals, each of which contains 10,000 IDs.
[0073] Then, traverse the key column value set and frequency set, map each key column value to the corresponding sub-interval, and accumulate the frequency of the sub-interval. Finally, the total frequency of each sub-interval is obtained, which is the frequency density of the sub-interval. Visualize these frequency density data to get the key column distribution histogram.
[0074] Based on the histogram data, the average frequency density of all subintervals is calculated. The frequency density of each subinterval is subtracted from the average value, and the sum of the squares of the differences is taken, divided by the number of subintervals, and then squared to obtain the standard deviation. At the same time, the Gini coefficient is calculated based on the subinterval frequency density as a data tilt indicator.
[0075] According to the total amount of data and the inclination, the optimal number of partitions is calculated using an empirical formula. The sub-intervals with frequency density greater than the sum of the average density and the standard deviation are marked as high-density areas, and the others are low-density areas. A finer granularity is used for high-density areas, such as 1 / 4 of the atomic interval size; a coarser granularity is used for low-density areas, such as twice the atomic interval size.
[0076] Finally, continuously monitor changes in data distribution. Rescan the data table regularly and update frequency statistics. Calculate the change in density of each sub-interval between the old and new statistical results. When the density change, the ratio of density of adjacent areas, or the overall slope change exceeds the preset threshold, trigger the partition adjustment. Regenerate the partition scheme based on the granularity of the high and low density areas.
[0077] Beneficial effects:
[0078] This method can adaptively adjust the partitioning strategy according to the data distribution characteristics, effectively solve the data skew problem, and improve query performance. By real-time monitoring of data distribution changes and timely adjusting the partitioning scheme, the timeliness and accuracy of the partitioning strategy are guaranteed.
[0079] This method adopts hierarchical partition granularity to perform fine-grained division on high-density areas and coarse-grained division on low-density areas, thus reducing the number of partitions and system overhead while ensuring load balancing.
[0080] This method constructs a key column distribution histogram based on statistical information, and quantifies data distribution characteristics by calculating indicators such as standard deviation and Gini coefficient, providing a reliable basis for the formulation of partitioning strategies and improving the scientificity and rationality of the partitioning scheme.
[0081] In an optional implementation, performing resource evaluation on each data partition in the initial data partition set, generating a resource demand vector based on the data volume, computational complexity, and data association of each data partition, inputting the resource demand vector into a pre-trained load prediction model, and obtaining the predicted execution time and resource occupancy rate of each data partition includes:
[0082] Constructing a data volume feature subvector according to the data size, number of records and number of columns of the data partition, constructing a computational complexity feature subvector according to the operation operator complexity coefficient, number of data connection operations, number of aggregation functions and number of grouping conditions of the data partition, constructing a data relevance feature subvector according to the input dependent data volume, output dependent data volume and data dependency relationship complexity coefficient of the data partition, and connecting the data volume feature subvector, the computational complexity feature subvector and the data relevance feature subvector to form a unified feature vector;
[0083] Building a load prediction model based on a multi-layer perceptron model, inputting the unified feature vector into an input layer of the multi-layer perceptron model, performing nonlinear transformation processing on the unified feature vector in multiple hidden layers of the multi-layer perceptron model, inputting output results of the multiple hidden layers into an output layer of the multi-layer perceptron model, and generating predicted execution time and predicted resource occupancy rate;
[0084] A training data set is constructed based on historical execution records, data enhancement processing is performed on the training data set, the number of samples in the training data set is expanded by means of characteristic value perturbation, load scenario combination and abnormal sample injection, the multi-layer perceptron model is trained by using the training data set, the mean square error of the predicted execution time and the predicted resource occupancy rate is calculated, and a weighted sum of the mean square error and a regularization parameter is constructed as a loss function;
[0085] Obtain the actual execution time and actual resource occupancy rate of the data partition, calculate a first deviation rate between the predicted execution time and the actual execution time, calculate a second deviation rate between the predicted resource occupancy rate and the actual resource occupancy rate, calculate a dynamic correction coefficient based on the first deviation rate and the second deviation rate, take the product of the predicted execution time and the corresponding dynamic correction coefficient as the final predicted execution time, and take the product of the predicted resource occupancy rate and the corresponding dynamic correction coefficient as the final predicted resource occupancy rate.
[0086] This embodiment provides a method for performing resource evaluation and load prediction on data partitions. First, resource evaluation is performed on each data partition in the initial data partition set, and a resource demand vector is generated based on the data volume, computational complexity, and data association. Then, the resource demand vector is input into a pre-trained load prediction model to obtain the predicted execution time and resource occupancy rate of each data partition.
[0087] Specifically, the resource assessment process includes the following steps:
[0088] First, construct the data volume feature subvector. This is done based on the data size, number of records, and number of columns of the data partition. For example, for a data partition containing 1 million records, 50 columns, and a total size of 10GB, a data volume feature subvector of [1000000, 50, 10] can be constructed.
[0089] Secondly, construct the computational complexity feature subvector. This is done based on the operator complexity coefficient of the data partition, the number of data join operations, the number of aggregation functions, and the number of grouping conditions. For example, for a query that contains 2 join operations, 3 aggregation functions, and 1 grouping condition, you can construct a computational complexity feature subvector of [2, 3, 1], and then multiply it by an operator complexity coefficient (such as 1.5) to get [3, 4.5, 1.5].
[0090] Then, construct the data correlation feature subvector. This is done based on the input dependency data volume, output dependency data volume, and data dependency complexity coefficient of the data partition. For example, for a data partition with an input dependency data volume of 5GB and an output dependency data volume of 2GB, you can construct a data correlation feature subvector of [5, 2], and then multiply it by a data dependency complexity coefficient (such as 1.2) to get [6, 2.4].
[0091] Finally, the above three feature sub-vectors are connected to form a unified feature vector, such as [1000000, 50, 10, 3,4.5, 1.5, 6, 2.4]. This unified feature vector is the resource demand vector of the data partition.
[0092] Next, a load prediction model is constructed. This embodiment adopts a multi-layer perceptron model as the basic structure of the load prediction model. Specifically, the multi-layer perceptron model includes an input layer, multiple hidden layers, and an output layer. The number of neurons in the input layer is the same as the dimension of the unified feature vector, and the output layer includes two neurons, corresponding to the predicted execution time and the predicted resource occupancy rate, respectively.
[0093] In the model training phase, we first build a training data set based on historical execution records. In order to enhance the generalization ability of the model, we perform data enhancement on the training data set. Specifically, we:
[0094] 1. Eigenvalue perturbation: Perform a small random perturbation on the values in the original eigenvector, such as adding or subtracting no more than 10% random noise on the basis of the original value.
[0095] 2. Load scenario combination: Combine historical data under different load scenarios to generate new samples. For example, mix scene data with high CPU load and high memory load.
[0096] 3. Abnormal sample injection: Artificially construct samples of some extreme cases, such as extremely large data volumes or extremely complex queries, and add them to the training set.
[0097] Through the above method, the number and diversity of samples in the training data set can be significantly expanded.
[0098] In the process of model training, the mean square error is used as the main component of the loss function, and the regularization term is introduced to prevent overfitting. Specifically, the loss function can be expressed as the weighted sum of the mean square error of the predicted execution time and the predicted resource occupancy and the regularization parameter.
[0099] After the training is completed, the unified feature vector is input into the input layer of the multi-layer perceptron model. After nonlinear transformation processing of multiple hidden layers, the predicted execution time and predicted resource occupancy are generated in the output layer.
[0100] In order to further improve the prediction accuracy, this implementation introduces a dynamic correction mechanism. The specific steps are as follows:
[0101] First, obtain the actual execution time and actual resource usage of the data partition. This can be achieved by executing the data partition in the actual operating environment and recording relevant indicators.
[0102] Then, the first deviation rate between the predicted execution time and the actual execution time, and the second deviation rate between the predicted resource occupancy rate and the actual resource occupancy rate are calculated. For example, if the predicted execution time is 100 seconds and the actual execution time is 90 seconds, the first deviation rate is (100-90) / 90 ≈ 11.11%.
[0103] Next, the dynamic correction coefficient is calculated based on the first deviation rate and the second deviation rate. A weighted average method can be used, such as giving a weight of 70% to the execution time deviation and a weight of 30% to the resource occupancy rate deviation. Assuming the second deviation rate is 5%, the dynamic correction coefficient can be calculated as: 1 - (11.11% * 0.7 + 5% * 0.3) ≈ 0.9167.
[0104] Finally, multiply the predicted execution time by the corresponding dynamic correction coefficient to get the final predicted execution time. Similarly, multiply the predicted resource occupancy by the corresponding dynamic correction coefficient to get the final predicted resource occupancy. In the above example, the final predicted execution time is 100 * 0.9167 = 91.67 seconds, which is closer to the actual execution time.
[0105] Through the above steps, accurate resource evaluation and load prediction of data partitions can be achieved, providing an important basis for subsequent task scheduling and resource allocation.
[0106] Beneficial effects:
[0107] This implementation method constructs a multi-dimensional feature vector, comprehensively considers factors such as data volume, computational complexity, and data correlation, and can more accurately characterize the resource demand characteristics of data partitions, providing comprehensive input information for load prediction.
[0108] The multi-layer perceptron model is used as the load prediction model, combined with data enhancement and regularization techniques, which effectively improves the generalization ability and prediction accuracy of the model. Through nonlinear transformation, it can capture the complex relationship between features and adapt to different types of data processing tasks.
[0109] The introduction of a dynamic correction mechanism can effectively reduce the prediction error and improve the accuracy and adaptability of load prediction by adjusting the prediction results through real-time feedback. This adaptive method enables the prediction model to be continuously optimized as the actual execution environment changes, maintaining long-term effectiveness.
[0110] In an optional implementation, a cost matrix of data partitions is constructed according to the predicted execution time and resource occupancy rate, and the initial data partition set is optimized and reorganized using a dynamic programming algorithm based on the cost matrix to generate a target data partition set with the optimal computational cost, including:
[0111] Obtaining the predicted execution time and resource occupancy rate of the data partition, constructing a cost function of a single data partition according to the ratio of the predicted execution time to the maximum predicted execution time, the ratio of the resource occupancy rate to the maximum resource occupancy rate, and the ratio of the data migration cost to the maximum data migration cost, and calculating the single body cost of the data partition based on the cost function;
[0112] Calculate the data dependency, data similarity and load balancing factor between two data partitions, take the weighted sum of the data dependency, data similarity and load balancing factor as the additional cost of merging the two data partitions, take the sum of the single cost of the two data partitions and the additional cost as the merging cost of the two data partitions, and construct a cost matrix based on the merging cost;
[0113] A state transfer equation is constructed based on the cost matrix, and the minimum cost of reorganizing the previous data partition into the target data partition is calculated by the state transfer equation. During the reorganization process, the size of the data partition is limited to ensure that the size of the reorganized data partition does not exceed a preset upper limit, the load of the data partition is balanced to ensure that the deviation between the load of the reorganized data partition and the average load does not exceed a preset load threshold, and the locality of the data partition is constrained to ensure that the locality of the reorganized data partition is not lower than a preset lower limit;
[0114] The load balance, data locality and overall cost of the data partition reorganization scheme are calculated, and the weighted sum of the load balance, data locality and overall cost is used as the score of the reorganization scheme. The score change rate of two adjacent iterations is calculated. When the score change rate is less than a preset convergence threshold, the weight parameters in the cost function are adaptively adjusted based on the load balance change rate, resource utilization change rate and migration cost change rate, and the data partition reorganization process is re-executed using the adjusted cost function.
[0115] This embodiment provides a data partition optimization method based on dynamic programming. The method first constructs a cost matrix for data partitions, and then uses a dynamic programming algorithm to optimize and reorganize the initial data partition set to generate a target data partition set with the optimal computational cost.
[0116] First, obtain the predicted execution time and resource usage of the data partition. For each data partition, predict its execution time and resource usage through historical data analysis or performance model. For example, the predicted execution time of a data partition is 100ms, the CPU usage is 60%, and the memory usage is 40%.
[0117] Then, construct the cost function of a single data partition. This function takes into account three factors: execution time, resource usage, and data migration cost. Specifically, calculate the ratio of the predicted execution time to the maximum execution time of all partitions, the ratio of resource usage to the maximum usage, and the ratio of data migration cost to the maximum migration cost. The weighted sum of these three ratios is used as the cost of a single data partition. For example, if the execution time weight is set to 0.5, the resource usage weight is set to 0.3, and the migration cost weight is set to 0.2, the cost value of a single partition can be obtained.
[0118] Next, calculate the merge cost between the two data partitions. First, evaluate the degree of data dependency, such as calculating the frequency of data interaction between the two partitions by analyzing the data access pattern. Then calculate the data similarity, such as by comparing the cosine similarity of the data feature vectors. Then calculate the load balancing factor, such as by comparing the resource usage of the two partitions. The weighted sum of these three indicators is used as the additional cost, and added to the single cost of the two partitions to get the merge cost. For example, if the single costs of the two partitions are 0.6 and 0.7 respectively, and the additional cost is 0.2, then the merge cost is 1.5.
[0119] Based on the above calculation results, a cost matrix is constructed. Each element in the matrix represents the merge cost of two corresponding data partitions. For example, if there are 4 initial data partitions, a 4x4 cost matrix can be obtained.
[0120] Then, a state transition equation is constructed based on the cost matrix. This equation describes the process of reorganizing the previous data partition into the target data partition. In each step of the transfer, the cost of merging different partition combinations is calculated, and the combination with the lowest cost is selected. At the same time, constraints are imposed on the reorganization process: 1) The partition size is limited to not exceed the preset upper limit, such as 1TB; 2) The deviation between the partition load and the average load is ensured not to exceed the preset threshold, such as 20%; 3) The partition locality is ensured not to be lower than the preset lower limit, such as 0.8. Through iterative calculation, the optimal reorganization plan is obtained.
[0121] During the reorganization process, the score of each iteration is calculated. The score consists of three parts: load balance, data locality, and overall cost. Load balance can be measured by calculating the variance of resource usage of each partition. Data locality can be measured by calculating the data correlation within the partition. The overall cost is the sum of all partition costs. The three indicators are normalized and weighted to get the score. For example, if the three weights are set to 0.4, 0.3, and 0.3 respectively, the score of each iteration can be calculated.
[0122] Calculate the rate of change of the score between two adjacent iterations. When the rate of change is less than the preset convergence threshold (such as 0.01), it means that the optimization process has stabilized. At this time, based on the changes in load balancing, resource utilization, and migration cost, the weight parameters in the cost function are adaptively adjusted. For example, if the load balancing degree does not change significantly, its weight can be appropriately reduced; if the migration cost changes significantly, its weight can be appropriately increased. After adjustment, re-execute the data partition reorganization process until the final convergence.
[0123] Through the above steps, the target data partition set with the optimal computational cost can be obtained. This method can effectively balance execution efficiency, resource utilization and data migration cost, and realize dynamic optimization of data partitions.
[0124] Beneficial effects:
[0125] This method can more comprehensively evaluate the advantages and disadvantages of data partitioning schemes by constructing cost functions and cost matrices that comprehensively consider multiple factors, thereby obtaining better reorganization results. Compared with traditional methods, it can better balance multiple goals such as computing efficiency, resource utilization, and data locality.
[0126] Compared with heuristic methods such as greedy algorithms, the dynamic programming algorithm is used for optimization and reorganization, which can obtain the global optimal solution and avoid falling into the local optimal solution. At the same time, by setting multiple constraints, the practicality and feasibility of the reorganization results are guaranteed.
[0127] The introduction of adaptive weight adjustment mechanism enables the optimization process to dynamically adjust the importance of each indicator according to the actual situation, improving the robustness and adaptability of the algorithm. This method can better cope with complex and changeable practical application scenarios.
[0128] In an optional implementation, multiple computing nodes that meet computing requirements are selected from a preset computing node resource pool to build a distributed computing cluster, and based on the cost matrix of each data partition in the target data partition set, a minimum spanning tree algorithm is used to calculate the optimal data transmission path between the multiple computing nodes, including:
[0129] Obtain resource information of the computing node, the resource information including CPU usage, memory capacity, network bandwidth and storage capacity, construct a resource vector of the computing node based on the resource information, calculate the ratio of the resource amount of each dimension in the resource vector to the resource demand of the data partition, and use the weighted sum of the ratios as the resource fitness of the computing node;
[0130] Determine the location affinity of the computing node and the data partition, determine the matching degree of the computing node and the data partition based on the resource fitness, the location affinity and the resource consumption cost, calculate the total load demand of the data partition, determine the number of nodes of the distributed computing cluster according to the ratio of the total load demand to the average node capacity, and select multiple computing nodes with the highest matching degree from a preset computing node resource pool to construct the distributed computing cluster;
[0131] Calculate the data size transmitted between computing nodes in the distributed computing cluster, obtain the available bandwidth, network delay and network congestion between the computing nodes, calculate the transmission cost between the computing nodes based on the data size, the available bandwidth, the network delay and the network congestion, and construct a transmission cost matrix;
[0132] Constructing a minimum spanning tree based on the transmission cost matrix, calculating the data flow on the path between any two computing nodes in the minimum spanning tree, ensuring that the data flow does not exceed the bandwidth capacity on the path, and calculating the transmission delay of the path between any two computing nodes in the minimum spanning tree, ensuring that the transmission delay does not exceed a preset delay threshold;
[0133] Calculate the path cost of the data transmission path in the minimum spanning tree, the ratio of the transmission delay to the preset delay threshold, and the ratio of the available bandwidth to the required bandwidth, and calculate the path quality evaluation value based on the path cost, the ratio of the transmission delay to the preset delay threshold, and the ratio of the available bandwidth to the required bandwidth. When the path quality evaluation value is less than the preset quality threshold, reconstruct, split or merge the data transmission path according to the network congestion level and load conditions.
[0134] Select multiple computing nodes that meet the computing requirements from the preset computing node resource pool to build a distributed computing cluster. Based on the cost matrix of each data partition in the target data partition set, the specific implementation method of using the minimum spanning tree algorithm to calculate the optimal data transmission path between multiple computing nodes is as follows:
[0135] First, obtain the resource information of the computing node, including the CPU usage, memory capacity, network bandwidth, and storage capacity. Based on this resource information, construct the resource vector of the computing node. For example, the resource vector of a computing node can be expressed as [CPU usage 80%, memory capacity 16GB, network bandwidth 100Mbps, storage capacity 1TB]. Then calculate the ratio of the resource amount of each dimension in the resource vector to the resource demand of the data partition. Assuming that the resource demand of a data partition is [CPU usage 50%, memory 8GB, network bandwidth 50Mbps, storage 500GB], the ratios of each dimension are [1.6, 2, 2, 2]. The weighted sum of these ratios is used as the resource fitness of the computing node. The weights can be set according to actual needs, for example, [0.3, 0.3, 0.2, 0.2], then the resource fitness of the node for the data partition is 1.6*0.3+2*0.3+2*0.2+2*0.2=1.88.
[0136] Next, determine the location affinity between the computing node and the data partition. This can be calculated based on the physical location distance between the node and the data partition. The closer the distance, the higher the affinity. For example, the node affinity in the same rack can be set to 1, the node affinity in the same data center can be set to 0.8, and the node affinity in different data centers can be set to 0.5. Then determine the matching degree between the computing node and the data partition based on resource fitness, location affinity, and resource consumption cost. The matching degree can be expressed as the weighted sum of these three factors, and the weights can be set according to actual needs. For example, if the weights are [0.4, 0.3, 0.3], the resource fitness of a node for a data partition is 1.88, the location affinity is 0.8, and the resource consumption cost is 0.7 (normalized value), then the matching degree is 1.88*0.4+0.8*0.3+0.7*0.3=1.23.
[0137] The total load demand of the data partition is counted, and the number of nodes in the distributed computing cluster is determined based on the ratio of the total load demand to the average node capacity. For example, if the total load demand is 1,000 units and each node can bear 100 units of load on average, 10 nodes are required. Select multiple computing nodes with the highest matching degree from the preset computing node resource pool to build a distributed computing cluster.
[0138] Calculate the size of data transmitted between computing nodes in a distributed computing cluster, and obtain the available bandwidth, network delay, and network congestion between computing nodes. Based on this information, calculate the transmission cost between computing nodes and construct a transmission cost matrix. For example, if 100GB of data is transmitted between nodes A and B, the available bandwidth is 1Gbps, the network delay is 10ms, and the network congestion is 0.8, then the transmission cost can be expressed as 100*8 / 1*0.8+10=640.01. Calculate the transmission cost for all node pairs and obtain the transmission cost matrix.
[0139] Construct a minimum spanning tree based on the transmission cost matrix. This can be implemented using the Kruskal algorithm or the Prim algorithm. Calculate the data flow on the path between any two computing nodes in the minimum spanning tree to ensure that the data flow does not exceed the bandwidth capacity on the path. Calculate the transmission delay of the path between any two computing nodes in the minimum spanning tree to ensure that the transmission delay does not exceed a preset delay threshold, such as 100ms.
[0140] Calculate the path cost of the data transmission path in the minimum spanning tree, the ratio of the transmission delay to the preset delay threshold, and the ratio of the available bandwidth to the required bandwidth. Calculate the path quality evaluation value based on these indicators. For example, if the path cost is 1000, the transmission delay is 80ms, the preset delay threshold is 100ms, the available bandwidth is 2Gbps, and the required bandwidth is 1Gbps, then the path quality evaluation value can be expressed as 1000*0.4+80 / 100*0.3+2 / 1*0.3=424. When the path quality evaluation value is less than the preset quality threshold, such as 500, the data transmission path is reconstructed, split or merged according to the network congestion and load conditions. You can consider choosing other alternative paths, or splitting the transmission task of large data volume into multiple small tasks for parallel transmission.
[0141] Beneficial effects:
[0142] This method comprehensively considers the resource fitness, location affinity and resource consumption cost of computing nodes, selects the most suitable nodes to build a distributed computing cluster, and improves resource utilization efficiency and computing performance.
[0143] The minimum spanning tree algorithm is used to calculate the optimal data transmission path, and the efficiency and reliability of data transmission are guaranteed through path quality evaluation and dynamic adjustment.
[0144] Through fine-grained resource management and transmission path optimization, this method can effectively cope with complex network topologies and dynamic load changes in large-scale distributed computing environments, improving the scalability and robustness of the system.
[0145] In an optional implementation, the data partitions in the target data partition set are allocated to corresponding computing nodes according to the optimal data transmission path, a grouping operation subtask is started at each computing node, the load status of each computing node is monitored in real time, and when a load imbalance is detected, the allocation scheme of the data partitions is dynamically adjusted based on the cost matrix and the optimal data transmission path, including:
[0146] Calculating the data location correlation between the data partition and the computing node, the computing power of the computing node and the current load level of the computing node, calculating the affinity between the data partition and the computing node based on the data location correlation, the computing power and the current load level, and allocating the data partition to the corresponding computing node according to the affinity;
[0147] Obtain the data volume of the data partition and the processing capacity of the computing node, determine the task segmentation granularity based on the data ratio of the data volume to the processing capacity, calculate the ratio of the operator load to the preset load threshold, round up the data ratio as the operator parallelism, and the operator parallelism does not exceed the preset maximum parallelism;
[0148] Monitor the CPU usage, memory usage, input / output usage, and network bandwidth usage of the computing nodes in real time, calculate the standard deviation of the computing node load and the average load of all computing nodes, use the ratio of the standard deviation to the average load as the load balancing degree, and trigger a load warning when the computing node load exceeds the load balancing degree;
[0149] Calculate the migration cost of the data partition between the source computing node and the target computing node, the migration cost includes the ratio of the data partition size to the bandwidth between the nodes, the migration preparation overhead and the sum of the service interruption compensation, and calculate the weighted difference between the load balancing improvement and the migration cost as the migration benefit;
[0150] When the migration benefit is greater than the preset benefit threshold, data partition migration is performed; when the data partition size is greater than the preset split threshold, data partition splitting is performed; when the sum of the sizes of adjacent data partitions is less than the preset merge threshold, data partition merging is performed; the weight coefficient of the migration cost is adjusted based on changes in system performance; when the system performance improves, the load balancing adjustment period is reduced; when the system performance decreases, the load balancing adjustment period is increased.
[0151] This embodiment provides a data partition allocation and dynamic load balancing method based on an optimal data transmission path. The method first allocates data partitions in a target data partition set to corresponding computing nodes according to the optimal data transmission path, then starts a group operation subtask on each computing node, and monitors the load status of each computing node in real time. When load imbalance is detected, the allocation scheme of the data partition is dynamically adjusted based on the cost matrix and the optimal data transmission path.
[0152] Specifically, the method comprises the following steps:
[0153] First, calculate the data location correlation between the data partition and the computing node, the computing power of the computing node, and the current load level of the computing node. The data location correlation can be measured by calculating the network distance between the data partition and the computing node. The closer the distance, the higher the correlation. The computing power can be measured by hardware indicators such as the number of CPU cores and memory size of the node. The current load level can be measured by indicators such as the CPU usage and memory usage of the node.
[0154] Then, the affinity of the data partitions to the computing nodes is calculated based on the above three factors. The affinity can be calculated by weighted summation, giving higher weights to data location relevance to prioritize data locality. For example, the weight of data location relevance can be set to 0.5, the weight of computing power can be set to 0.3, and the weight of current load level can be set to 0.2.
[0155] Next, the data partitions are assigned to the corresponding computing nodes according to the calculated affinity. A greedy algorithm can be used to select the node with the highest affinity to assign the data partition each time until all data partitions are assigned.
[0156] After the allocation is completed, the data volume of each data partition and the processing capacity of the corresponding computing node are obtained. The data volume can be measured by the partition size, and the processing capacity can be measured by the CPU frequency of the node. The task segmentation granularity is determined based on the ratio of data volume to processing capacity. For example, if the data volume is 100GB and the processing capacity is 10GB / s, the data ratio is 10, and the task can be divided into 10 subtasks.
[0157] Then calculate the ratio of the operator load to the preset load threshold. The load can be measured by CPU usage, and the preset load threshold can be set to 80%. Round up the data ratio as the operator parallelism, but do not exceed the preset maximum parallelism. For example, if the data ratio is 10.5, the parallelism is 11, but if the preset maximum parallelism is 8, the final parallelism is 8.
[0158] Next, monitor the CPU usage, memory usage, input / output usage, and network bandwidth usage of the computing nodes in real time. You can collect data for these indicators every 10 seconds. Calculate the average of these indicators for all computing nodes, and then calculate the standard deviation of each node from the average. The ratio of the standard deviation to the average is used as the load balance degree. For example, if the CPU usage of a node is 90%, the average is 60%, and the standard deviation is 30%, the load balance degree is 0.5.
[0159] When the load of the computing node exceeds the load balance degree, a load warning is triggered. A threshold can be set, for example, when the load balance degree is greater than 0.3, a warning is triggered. After the warning is triggered, the cost and benefit of data partition migration need to be calculated.
[0160] Calculate the migration cost of the data partition between the source compute node and the target compute node. The migration cost includes the ratio of the data partition size to the inter-node bandwidth, the migration preparation overhead, and the sum of the service interruption compensation. For example, if the data partition size is 10GB and the inter-node bandwidth is 1GB / s, the transfer time is 10 seconds. The migration preparation overhead may take 5 seconds, and the service interruption compensation may take 2 seconds, so the total migration cost is 17 seconds.
[0161] Calculate the weighted difference between the load balance improvement and the migration cost as the migration benefit. You can set the weight of the load balance improvement to 0.7 and the weight of the migration cost to 0.3. For example, if the load balance improvement is 0.2 and the migration cost is 17 seconds, the migration benefit is 0.2 * 0.7 - 17 * 0.3 = -3.7.
[0162] When the migration benefit is greater than the preset benefit threshold, data partition migration is performed. The benefit threshold can be set to 0, that is, migration is performed as long as the benefit is positive. When performing migration, it is necessary to suspend the relevant tasks, transfer the data from the source node to the target node, and then resume the task execution on the target node.
[0163] When the size of a data partition is larger than the preset split threshold, the data partition is split. The split threshold can be set to 100GB. When splitting, the data partition is evenly divided into two sub-partitions, which are respectively distributed to different nodes. When the sum of the sizes of adjacent data partitions is smaller than the preset merge threshold, the data partition is merged. The merge threshold can be set to 50GB. When merging, the data of two adjacent partitions are merged to one node.
[0164] Finally, adjust the weight coefficient of the migration cost based on the change in system performance. Performance can be measured by monitoring the throughput of the system. When system performance improves, it means that the current strategy is effective and the load balancing adjustment cycle can be reduced, for example, from 60 seconds to 30 seconds. When system performance decreases, it means that there may be problems with the current strategy and the load balancing adjustment cycle needs to be increased, for example, from 60 seconds to 90 seconds, to reduce the overhead caused by frequent adjustments.
[0165] Beneficial effects:
[0166] This method realizes initial data allocation by calculating the affinity between data partitions and computing nodes, fully considering the data location relevance, computing power and current load level, and can effectively improve data locality and computing resource utilization.
[0167] This method can effectively balance the load of each node and improve the overall performance and stability of the system by monitoring various indicators of computing nodes in real time, calculating the load balance, and dynamically adjusting the data partition allocation based on the migration cost and benefit.
[0168] This method can adaptively optimize according to changes in system performance by dynamically adjusting the parameters of the load balancing strategy, such as the migration cost weight and adjustment period, thereby reducing the adjustment overhead while ensuring the load balancing effect, and improving the scalability and robustness of the system.
[0169] Figure 2 FIG. 1 is a schematic diagram of the structure of a HIVE grouping operation performance optimization and scheduling system based on key column partitioning according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0170] The first unit is used to receive a HIVE computing task submitted by a user, obtain data table information and group operation key column information in the HIVE computing task, construct a key column distribution histogram based on the data table information and the group operation key column information, calculate the distribution density and data skewness of the key column data according to the key column distribution histogram, determine a key column partitioning strategy based on the distribution density and data skewness of the key column data, the key column partitioning strategy includes a data partition size, a partition number, and a sampling rate, and perform partition preprocessing on the data of the HIVE computing task according to the key column partitioning strategy to generate an initial data partition set;
[0171] The second unit is used to perform resource evaluation on each data partition in the initial data partition set, generate a resource demand vector based on the data volume, computational complexity and data correlation of each data partition, input the resource demand vector into a pre-trained load prediction model, obtain the predicted execution time and resource occupancy rate of each data partition, construct a cost matrix of the data partition according to the predicted execution time and resource occupancy rate, and optimize and reorganize the initial data partition set using a dynamic programming algorithm based on the cost matrix to generate a target data partition set with the optimal computational cost;
[0172] The third unit is used to select multiple computing nodes that meet the computing requirements from a preset computing node resource pool to build a distributed computing cluster, and based on the cost matrix of each data partition in the target data partition set, use a minimum spanning tree algorithm to calculate the optimal data transmission path between the multiple computing nodes, and allocate the data partitions in the target data partition set to the corresponding computing nodes according to the optimal data transmission path, start a group operation subtask at each of the computing nodes, monitor the load status of each computing node in real time, and when a load imbalance is detected, dynamically adjust the data partition allocation plan based on the cost matrix and the optimal data transmission path to ensure that the computing load of the distributed computing cluster is in a dynamically balanced state.
[0173] According to a third aspect of the embodiments of the present invention,
[0174] An electronic device is provided, comprising:
[0175] processor;
[0176] a memory for storing processor-executable instructions;
[0177] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0178] A fourth aspect of the embodiments of the present invention is:
[0179] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0180] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. HIVE grouping operation performance optimization and scheduling method based on key column partitioning, characterized in that: include: Receive a HIVE computing task submitted by a user, obtain data table information and group operation key column information in the HIVE computing task, build a key column distribution histogram based on the data table information and the group operation key column information, calculate the distribution density and data skewness of the key column data according to the key column distribution histogram, determine a key column partitioning strategy based on the distribution density and data skewness of the key column data, the key column partitioning strategy includes a data partition size, a partition number, and a sampling rate, perform partition preprocessing on the data of the HIVE computing task according to the key column partitioning strategy, and generate an initial data partition set; Perform resource evaluation on each data partition in the initial data partition set, generate a resource demand vector based on the data volume, computational complexity and data correlation of each data partition, input the resource demand vector into a pre-trained load prediction model, obtain the predicted execution time and resource occupancy rate of each data partition, construct a cost matrix of the data partition according to the predicted execution time and resource occupancy rate, optimize and reorganize the initial data partition set based on the cost matrix using a dynamic programming algorithm, and generate a target data partition set with the optimal computational cost; A plurality of computing nodes that meet computing requirements are selected from a preset computing node resource pool to build a distributed computing cluster. Based on the cost matrix of each data partition in the target data partition set, a minimum spanning tree algorithm is used to calculate the optimal data transmission path between the plurality of computing nodes. The data partitions in the target data partition set are allocated to corresponding computing nodes according to the optimal data transmission path. A grouping operation subtask is started at each computing node. The load status of each computing node is monitored in real time. When a load imbalance is detected, the allocation scheme of the data partitions is dynamically adjusted based on the cost matrix and the optimal data transmission path to ensure that the computing load of the distributed computing cluster is in a dynamically balanced state.
2. The method according to claim 1, characterized in that Constructing a key column distribution histogram based on the data table information and the group operation key column information, calculating the distribution density and data skewness of the key column data according to the key column distribution histogram, and determining the key column partitioning strategy based on the distribution density and data skewness of the key column data includes: Perform a full table scan on the data table to obtain key column value frequency statistics information, construct the key column values in the key column value frequency statistics information into a key column value set, construct the frequencies in the key column value frequency statistics information into a frequency set, construct a key column distribution histogram based on the value interval of the key column value set, divide the value interval into multiple sub-intervals, and calculate the frequency density of each sub-interval; Calculating data skewness and standard deviation based on the key column distribution histogram, calculating the standard deviation according to the difference between the frequency density and the average density of the multiple sub-intervals, and calculating the Gini coefficient as the data skewness based on the frequency density of the multiple sub-intervals; The optimal number of partitions is calculated based on the total amount of data and the data inclination, and the subintervals among the multiple subintervals whose frequency density is greater than the sum of the average density and the standard deviation are taken as high-density areas, and the subintervals among the multiple subintervals whose frequency density is less than or equal to the sum of the average density and the standard deviation are taken as low-density areas, and the high-density areas are divided into a first partition width by fine-grained division, and the low-density areas are divided into a second partition width by coarse-grained division; Monitor the data distribution changes of the high-density area and the low-density area in real time, calculate the density difference between two adjacent time points, and when the density difference is greater than a preset density threshold, or the density ratio of adjacent areas is greater than the preset density threshold, or the change in the data inclination is greater than the preset inclination threshold, trigger partition adjustment, re-divide the high-density area and the low-density area according to the first partition width and the second partition width, and generate a new partition scheme.
3. The method according to claim 1, characterized in that Performing resource evaluation on each data partition in the initial data partition set, generating a resource demand vector based on the data volume, computational complexity and data association of each data partition, inputting the resource demand vector into a pre-trained load prediction model, and obtaining the predicted execution time and resource occupancy rate of each data partition includes: Constructing a data volume feature subvector according to the data size, number of records and number of columns of the data partition, constructing a computational complexity feature subvector according to the operation operator complexity coefficient, number of data connection operations, number of aggregation functions and number of grouping conditions of the data partition, constructing a data relevance feature subvector according to the input dependent data volume, output dependent data volume and data dependency relationship complexity coefficient of the data partition, and connecting the data volume feature subvector, the computational complexity feature subvector and the data relevance feature subvector to form a unified feature vector; Building a load prediction model based on a multi-layer perceptron model, inputting the unified feature vector into an input layer of the multi-layer perceptron model, performing nonlinear transformation processing on the unified feature vector in multiple hidden layers of the multi-layer perceptron model, inputting output results of the multiple hidden layers into an output layer of the multi-layer perceptron model, and generating predicted execution time and predicted resource occupancy rate; A training data set is constructed based on historical execution records, data enhancement processing is performed on the training data set, the number of samples in the training data set is expanded by a feature value perturbation method, a load scenario combination method, and an abnormal sample injection method, the multilayer perceptron model is trained using the training data set, the mean square error of the predicted execution time and the predicted resource occupancy rate is calculated, and a weighted sum of the mean square error and a regularization parameter is constructed as a loss function; Obtain the actual execution time and actual resource occupancy rate of the data partition, calculate a first deviation rate between the predicted execution time and the actual execution time, calculate a second deviation rate between the predicted resource occupancy rate and the actual resource occupancy rate, calculate a dynamic correction coefficient based on the first deviation rate and the second deviation rate, take the product of the predicted execution time and the corresponding dynamic correction coefficient as the final predicted execution time, and take the product of the predicted resource occupancy rate and the corresponding dynamic correction coefficient as the final predicted resource occupancy rate.
4. The method according to claim 1, characterized in that: Constructing a cost matrix of data partitions according to the predicted execution time and resource occupancy rate, optimizing and reorganizing the initial data partition set using a dynamic programming algorithm based on the cost matrix, and generating a target data partition set with the optimal computational cost includes: Obtaining the predicted execution time and resource occupancy rate of the data partition, constructing a cost function of a single data partition according to the ratio of the predicted execution time to the maximum predicted execution time, the ratio of the resource occupancy rate to the maximum resource occupancy rate, and the ratio of the data migration cost to the maximum data migration cost, and calculating the single body cost of the data partition based on the cost function; Calculate the data dependency, data similarity and load balancing factor between two data partitions, take the weighted sum of the data dependency, data similarity and load balancing factor as the additional cost of merging the two data partitions, take the sum of the single cost of the two data partitions and the additional cost as the merging cost of the two data partitions, and construct a cost matrix based on the merging cost; A state transfer equation is constructed based on the cost matrix, and the minimum cost of reorganizing the previous data partition into the target data partition is calculated by the state transfer equation. During the reorganization process, the size of the data partition is limited to ensure that the size of the reorganized data partition does not exceed a preset upper limit, the load of the data partition is balanced to ensure that the deviation between the load of the reorganized data partition and the average load does not exceed a preset load threshold, and the locality of the data partition is constrained to ensure that the locality of the reorganized data partition is not lower than a preset lower limit; The load balance, data locality and overall cost of the data partition reorganization scheme are calculated, and the weighted sum of the load balance, data locality and overall cost is used as the score of the reorganization scheme. The score change rate of two adjacent iterations is calculated. When the score change rate is less than a preset convergence threshold, the weight parameters in the cost function are adaptively adjusted based on the load balance change rate, resource utilization change rate and migration cost change rate, and the data partition reorganization process is re-executed using the adjusted cost function.
5. The method according to claim 1, characterized in that Selecting multiple computing nodes that meet computing requirements from a preset computing node resource pool to build a distributed computing cluster, and using a minimum spanning tree algorithm to calculate an optimal data transmission path between the multiple computing nodes based on a cost matrix of each data partition in the target data partition set includes: Obtain resource information of the computing node, the resource information including CPU usage, memory capacity, network bandwidth and storage capacity, construct a resource vector of the computing node based on the resource information, calculate the ratio of the resource amount of each dimension in the resource vector to the resource demand of the data partition, and use the weighted sum of the ratios as the resource fitness of the computing node; Determine the location affinity of the computing node and the data partition, determine the matching degree of the computing node and the data partition based on the resource fitness, the location affinity and the resource consumption cost, calculate the total load demand of the data partition, determine the number of nodes of the distributed computing cluster according to the ratio of the total load demand to the average node capacity, and select multiple computing nodes with the highest matching degree from a preset computing node resource pool to construct the distributed computing cluster; Calculate the data size transmitted between computing nodes in the distributed computing cluster, obtain the available bandwidth, network delay and network congestion between the computing nodes, calculate the transmission cost between the computing nodes based on the data size, the available bandwidth, the network delay and the network congestion, and construct a transmission cost matrix; Constructing a minimum spanning tree based on the transmission cost matrix, calculating the data flow on the path between any two computing nodes in the minimum spanning tree, ensuring that the data flow does not exceed the bandwidth capacity on the path, and calculating the transmission delay of the path between any two computing nodes in the minimum spanning tree, ensuring that the transmission delay does not exceed a preset delay threshold; Calculate the path cost of the data transmission path in the minimum spanning tree, the ratio of the transmission delay to the preset delay threshold, and the ratio of the available bandwidth to the required bandwidth, and calculate the path quality evaluation value based on the path cost, the ratio of the transmission delay to the preset delay threshold, and the ratio of the available bandwidth to the required bandwidth. When the path quality evaluation value is less than the preset quality threshold, reconstruct, split or merge the data transmission path according to the network congestion level and load conditions.
6. The method according to claim 1, characterized in that The method includes allocating data partitions in the target data partition set to corresponding computing nodes according to the optimal data transmission path, starting a grouping operation subtask at each computing node, monitoring the load status of each computing node in real time, and dynamically adjusting the allocation scheme of the data partitions based on the cost matrix and the optimal data transmission path when load imbalance is detected. The method includes: Calculating the data location correlation between the data partition and the computing node, the computing power of the computing node and the current load level of the computing node, calculating the affinity between the data partition and the computing node based on the data location correlation, the computing power and the current load level, and allocating the data partition to the corresponding computing node according to the affinity; Obtain the data volume of the data partition and the processing capacity of the computing node, determine the task segmentation granularity based on the data ratio of the data volume to the processing capacity, calculate the ratio of the operator load to the preset load threshold, round up the data ratio as the operator parallelism, and the operator parallelism does not exceed the preset maximum parallelism; Monitor the CPU usage, memory usage, input / output usage, and network bandwidth usage of the computing nodes in real time, calculate the standard deviation of the computing node load and the average load of all computing nodes, use the ratio of the standard deviation to the average load as the load balancing degree, and trigger a load warning when the computing node load exceeds the load balancing degree; Calculate the migration cost of the data partition between the source computing node and the target computing node, the migration cost includes the ratio of the data partition size to the bandwidth between the nodes, the migration preparation overhead and the sum of the service interruption compensation, and calculate the weighted difference between the load balancing improvement and the migration cost as the migration benefit; When the migration benefit is greater than the preset benefit threshold, data partition migration is performed; when the data partition size is greater than the preset split threshold, data partition splitting is performed; when the sum of the sizes of adjacent data partitions is less than the preset merge threshold, data partition merging is performed; the weight coefficient of the migration cost is adjusted based on changes in system performance; when the system performance improves, the load balancing adjustment period is reduced; when the system performance decreases, the load balancing adjustment period is increased.
7. A HIVE grouping operation performance optimization and scheduling system based on key column partitioning, used to implement the method as described in any one of claims 1 to 6, characterized in that: include: The first unit is used to receive a HIVE computing task submitted by a user, obtain data table information and group operation key column information in the HIVE computing task, construct a key column distribution histogram based on the data table information and the group operation key column information, calculate the distribution density and data skewness of the key column data according to the key column distribution histogram, determine a key column partitioning strategy based on the distribution density and data skewness of the key column data, the key column partitioning strategy includes a data partition size, a partition number, and a sampling rate, and perform partition preprocessing on the data of the HIVE computing task according to the key column partitioning strategy to generate an initial data partition set; The second unit is used to perform resource evaluation on each data partition in the initial data partition set, generate a resource demand vector based on the data volume, computational complexity and data correlation of each data partition, input the resource demand vector into a pre-trained load prediction model, obtain the predicted execution time and resource occupancy rate of each data partition, construct a cost matrix of the data partition according to the predicted execution time and resource occupancy rate, and optimize and reorganize the initial data partition set using a dynamic programming algorithm based on the cost matrix to generate a target data partition set with the optimal computational cost; The third unit is used to select multiple computing nodes that meet the computing requirements from a preset computing node resource pool to build a distributed computing cluster, and based on the cost matrix of each data partition in the target data partition set, use a minimum spanning tree algorithm to calculate the optimal data transmission path between the multiple computing nodes, and allocate the data partitions in the target data partition set to the corresponding computing nodes according to the optimal data transmission path, start a group operation subtask at each of the computing nodes, monitor the load status of each computing node in real time, and when a load imbalance is detected, dynamically adjust the data partition allocation plan based on the cost matrix and the optimal data transmission path to ensure that the computing load of the distributed computing cluster is in a dynamically balanced state.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Multi-tenant server-free platform resource management method and system
CN118467186A
Self-adaptive task scheduling execution unit management method and system
CN119376903A