Real-time partitioning method and device for relieving data skew problem

By optimizing the sampling and partitioning strategies for tuples in distributed computing tasks, the data skew problem was solved, task load balancing and efficient utilization of computing kernels were achieved, and system execution time was reduced.

CN120892181APending Publication Date: 2025-11-04709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510923499.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

In distributed parallel computing, the uneven distribution of data can lead to data skew, causing some nodes in the cluster to process a much larger amount of data than other nodes, which affects the overall computing efficiency and task execution performance of the nodes.

Method used

By sampling the pairs of data to be processed in the distributed computing task, the optimal number of partitions is determined based on the sampling results and the number of computing cores. Different data partitioning strategies are applied to high-frequency and low-frequency keys respectively. The weighted round-robin algorithm and hash method are used for partitioning, and the number of partitions is adaptively adjusted to match the number of computing cores.

Benefits of technology

It achieves task load balancing, improves the utilization of computing kernels, reduces system execution time, and reduces performance degradation caused by uneven data distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892181A_ABST
    Figure CN120892181A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed computing, in particular to a real-time partitioning method and device for relieving a data skew problem. The method comprises the following steps: sampling a to-be-processed two-tuple in a distributed computing task to obtain a sampling result; obtaining the number of calculation kernels in the distributed calculation task; obtaining an optimal partition number in the distributed computing task based on the sampling result and the number of the computing kernels; and dividing the sampling result into a high-frequency key and a low-frequency key, and respectively executing different data partitioning strategies on the high-frequency key and the low-frequency key based on the optimal partitioning number. The invention can improve the utilization rate of system resources and reduce the execution time of the system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of distributed computing, in particular to a real-time partitioning method and device for relieving data skew problem. BACKGROUND

[0002] With the advent of the information age, the amount of data is growing explosively, which poses a serious challenge to the processing speed of computers. To meet this challenge, various distributed parallel computing frameworks (such as Spark) have emerged. Distributed parallel computing framework divides large-scale tasks into multiple small tasks and executes them in parallel on multiple nodes in the cluster, which significantly improves the data processing efficiency. However, in real-world data sets, due to the differences in the properties and characteristics of the data itself, the data distribution often presents significant unevenness. This unevenness is further amplified in distributed parallel computing, which easily leads to data skew problem. Data skew problem specifically manifests as the amount of data processed by some nodes in the cluster is much larger than that of other nodes, leading to uneven load of each node, thereby affecting the overall node computing efficiency and task execution performance.

[0003] In view of this, overcoming the defects of the prior art is a problem urgently to be solved in the technical field. SUMMARY

[0004] In view of the above defects or improvement needs of the prior art, the present application provides a real-time partitioning method and device for relieving data skew problem, which can balance the task load, reduce the execution time of the system, and improve the utilization rate of system resources.

[0005] The embodiment of the present application adopts the following technical scheme: In a first aspect, the present application provides a real-time partitioning method for relieving data skew problem, specifically: sampling the to-be-processed two-tuple in the distributed computing task to obtain a sampling result; obtaining the number of computing kernels in the distributed computing task; based on the sampling result and the number of computing kernels, obtaining the optimal partition number in the distributed computing task; dividing the sampling result into high-frequency keys and low-frequency keys, and executing different data partitioning strategies on the high-frequency keys and the low-frequency keys based on the optimal partition number.

[0006] Preferably, the sampling of the to-be-processed two-tuple in the distributed computing task to obtain a sampling result comprises: dividing the to-be-processed two-tuple into different partitions; sampling the two-tuple in each partition to obtain a sampling result corresponding to each partition; The sampling results corresponding to all the partitions are obtained by integrating the sampling results corresponding to each partition.

[0007] Preferably, the sampling of the tuples in each partition to obtain the sampling result corresponding to each partition comprises: The tuples in a partition are divided according to a sampling interval to obtain a plurality of intermediate tuples, the intermediate tuples comprising a plurality of keys and a frequency corresponding to each key; The probability that a key in the intermediate tuples is sampled and retained is calculated; When the probability is greater than a first random value, the key and the corresponding frequency are taken as a sampling result, and the calculation is completed for all keys in all intermediate tuples to obtain the sampling result corresponding to each partition.

[0008] Preferably, before the sampling of the to-be-processed tuples in the distributed computing task to obtain the sampling result, the method further comprises: Data modeling is performed on the original tuples to obtain a skew degree and a load balancing level of the original tuples; When the skew degree is greater than the load balancing level, the to-be-processed tuples in the distributed computing task are sampled to obtain the sampling result.

[0009] Preferably, the obtaining of the optimal number of partitions in the distributed computing task based on the sampling result and the number of computing cores comprises: The total number of tuples in the sampling result is obtained; Based on the total number of tuples and the number of computing cores, the number of tuples that need to be calculated by a single core is obtained; The maximum value of the frequency in the sampling result is obtained; When the maximum value of the frequency is greater than the number of tuples that need to be calculated by a single core, the optimal number of partitions in the distributed computing task is obtained based on the number of tuples.

[0010] Preferably, the dividing of the sampling result into high-frequency keys and low-frequency keys and the execution of different data partitioning strategies on the high-frequency keys and the low-frequency keys based on the optimal number of partitions comprises: A partition information table is set, the partition information table comprising at least one partition ID and a partition weight corresponding to each partition ID; For each high-frequency key, the following steps are performed: The target partition ID corresponding to the maximum partition weight in the partition information table is obtained, the partition corresponding to the target partition ID is taken as the partition corresponding to the high-frequency key, and the partition information table is updated until the partitions corresponding to all high-frequency keys are obtained.

[0011] Preferably, the sampling results are divided into high-frequency keys and low-frequency keys, different data partition strategies are respectively performed on the high-frequency keys and the low-frequency keys based on the optimal number of partitions, and the different data partition strategies comprise the following steps: obtaining a partition information table, wherein the partition information table is a latest partition information table generated after the high-frequency key partition strategy is completed, and the partition information table comprises at least one partition ID, a partition weight corresponding to each partition ID, and a partition availability corresponding to each partition ID; for each low-frequency key, the following steps are performed: calculating a first hash value corresponding to the low-frequency key, and taking the first hash value as a partition ID; obtaining a partition weight corresponding to the partition ID; comparing the partition weight with a threshold value to obtain a partition corresponding to the low-frequency key.

[0012] Preferably, the threshold value comprises a lower threshold value, an upper threshold value and a full load threshold value, and the comparison of the partition weight with the threshold value to obtain the partition corresponding to the low-frequency key comprises the following steps: when the partition weight is greater than the lower threshold value and less than the upper threshold value, a partition corresponding to the partition ID is taken as the partition corresponding to the low-frequency key, and the partition information table is updated until the partitions corresponding to all low-frequency keys are obtained; when the partition weight is less than the lower threshold value or greater than the upper threshold value, a second hash value corresponding to the low-frequency key is recalculated to obtain the partition corresponding to the low-frequency key; when the partition weight is less than the full load threshold value, the partition availability is modified, and the hash value corresponding to the low-frequency key is recalculated to obtain the partition corresponding to the low-frequency key.

[0013] In a second aspect, the present application provides a real-time partition device for relieving data skew problems, specifically comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to execute the real-time partition method for relieving data skew problems in the first aspect.

[0014] In a third aspect, the present application further provides a non-volatile computer storage medium, wherein the computer storage medium stores computer executable instructions, and the computer executable instructions are executed by one or more processors to complete the method provided by the method in the first aspect.

[0015] Compared with the prior art, the application has the beneficial effects that: the application adjusts the number of partitions adaptively, matches the number of partitions with the number of computing cores, finds the balance between the number of computing cores and the number of to-be-processed binary tuples, ensures that each computing core can be evenly distributed to an appropriate amount of binary tuples, and ensures that each computing core in each batch of tasks has a corresponding task to execute, which can not only improve the utilization rate of computing cores, but also avoid performance degradation caused by uneven data distribution; based on the number of partitions, different partition strategies are performed on different frequency keys, which can improve the utilization rate of system resources and reduce the execution time of the system.

[0016] Further, the weighted round-robin algorithm is used for partitioning high-frequency keys, which greatly reduces the huge pressure on memory caused by polling of multiple key types, and reduces the time complexity; for low-frequency keys, a hash-based method is used for partitioning, which can exhibit better random distribution, and the hash has the characteristics of not being prone to collision, and can further achieve partition balance. The two partition methods for high-frequency keys and low-frequency keys are applied to the data partition process, and the two methods work together to balance the task load and reduce the execution time of the system. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments of the application. Obviously, the drawings described below are only some of the embodiments of the application, and other drawings can also be obtained from these drawings by those of ordinary skill in the art without any creative effort.

[0018] Figure 1 is a flowchart of a real-time partition method for relieving data skew provided by an embodiment of the application; Figure 2 is a flowchart of a to-be-processed binary tuple data modeling method provided by an embodiment of the application; Figure 3 is a flowchart of a method for implementing step 101 in Figure 1 is a flowchart of a method for implementing step 101 in Figure 4 is a flowchart of a method for implementing step 102-step 103 in Figure 1 is a flowchart of a method for implementing step 102-step 103 in Figure 5 is a flowchart of a method for implementing step 104 in Figure 1 is a flowchart of a method for implementing step 104 in Figure 6 is a flowchart of a method for implementing step 104 in Figure 1 is a flowchart of a method for implementing step 104 in Figure 7 is a flow diagram of a real-time partitioning method for relieving data skew problem provided by an embodiment of the present application; Figure 8 is a structural diagram of a real-time partitioning device for relieving data skew problem provided by an embodiment of the present application; In the drawings, reference numerals: 21: processor; 22: memory. DETAILED DESCRIPTION

[0019] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0020] Unless otherwise required by context, the term "comprises" in the specification and claims is to be construed as open-ended, i.e. as "comprising but not limited to". In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiment", "example", "specific example" or "some examples" are intended to mean that a particular feature, structure, material or characteristic is included in at least one embodiment or example of the disclosure. The illustrative representations of the above terms do not necessarily mean the same embodiment or example. In addition, the specific features, structures, materials or characteristics described can be included in any one or more embodiments or examples in any appropriate manner, i.e. although they are carried in the embodiments or examples of the above terms due to the order of appearance and location, they are not limited to being carried in combination by one embodiment or example.

[0021] In the description of the present application, the terms "first", "second" are only used for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features limited by "first", "second" can be explicitly or implicitly included in one or more features. In the description of the embodiments of the present disclosure, unless otherwise specified, the meaning of "multiple" is two or more. In addition, for example, in the description, the same type of nouns can also be described as two independent individuals by adding "A", "B" at the end, in which case the features limited by "A", "B" are only used for the purpose of distinguishing the same type of individual description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features.

[0022] In the description of the application, the expression "A and / or B" (wherein A and B represent a specific feature) is used in the sense that (A) alone, (B) alone or (A and B) in combination can be considered as the specific feature.

[0023] As used herein, "about," "approximately," or "substantially" with reference to a particular recited quantity, characteristic, or other dimension means that the quantity is the recited value, and means the quantity can vary from the recited value by an acceptable range of variation (e.g., as determined by one of ordinary skill in the art considering the measurement in question and the error in measuring the particular quantity (i.e., the limitations of the measurement system)).

[0024] Furthermore, the various features of the embodiments of the present application described herein can be combined with each other, as long as they do not contradict each other.

[0025] Example 1: In a distributed parallel computing framework, each Job is split into multiple Stages to maintain the dependency relationship of Resilient Distributed Datasets (RDD). The Shuffle process is used to complete the redistribution of data between Stages. There are many distributed computing frameworks for processing large amounts of data, and the present embodiment takes Spark as an example for illustration. Spark includes three key stages: Map stage, Shuffle stage and Reduce stage. The Map stage divides a large amount of data into multiple small portions of data, a Map task processes one small portion of data, and the Map task processes the small portion of data into the form of a pair, each pair containing multiple keys and values corresponding to each key; the Shuffle stage puts pairs with the same key into the same partition, at this time, the Map task writes the partitioned pairs into a file and generates an index file corresponding to the file, and the index file stores the partition ID corresponding to each key; the Reduce stage pulls the corresponding pairs for calculation and processing through the partition ID, and a Reduce task processes pairs in a partition. Since the number of pairs corresponding to each key is different, when the number of pairs of a certain key exceeds the capacity of a single partition, the pairs of the key will be scattered to multiple partitions. If the number of pairs in each partition is very different, it will cause the workloads of each Reduce task to be different, and thus cause some tasks to take too long to execute, slowing down the completion time of the entire job. Since the system must wait for all tasks in the application to complete before submitting new tasks, data skew will increase the overall processing time of the application and cause waste of idle waiting of some computing resources in the cluster. The data skew problem occurs between the Map task write and the Reduce task read process, and the Shuffle stage needs to be optimized, and on the basis of considering the computing power of the computing kernel in the cluster, the balanced division of pairs in the Map is realized to avoid the occurrence of slow tasks.

[0026] As shown in Figure 1 Embodiment 1 of the present application provides a real-time partitioning method for alleviating the data skew problem, which specifically includes the following steps: Step 101: Sampling the to-be-processed pairs in the distributed computing task to obtain a sampling result.

[0027] In one embodiment, the to-be-processed pairs include multiple keys and the frequency corresponding to each key, and the sampling result includes multiple keys and the frequency corresponding to each key, wherein a key and the frequency corresponding to the key in the sampling result form a pair.

[0028] The to-be-processed two-tuple is a partition result obtained after the Map task generates a default partition strategy based on Spark, each partition result includes multiple keys and the frequency corresponding to each key, and a sampling result is obtained based on the frequency corresponding to the key in the partition result, wherein the sampling result contains the key key (i.e., high-frequency key and low-frequency key) causing the data skew problem and the frequency of each key key.

[0029] In this embodiment, the key key causing the data skew problem in the default partition strategy is identified, and the partition strategy is optimized according to the sampling result.

[0030] In order to make the sampling result stable in the case of data skew, the dynamic weighting negative sampling algorithm is used in this embodiment, which adds the consideration of high-frequency key in the negative sampling algorithm to ensure that high-frequency key has a greater probability of being detected.

[0031] Step 102: Obtain the number of computing kernels in the distributed computing task.

[0032] In one embodiment, the Reduce task is essentially to process the two-tuple in the partition using the computing kernel. However, the number of two-tuples that each computing kernel can effectively process in unit time is limited. When the number of two-tuples in a certain partition far exceeds that in other partitions, the Reduce task corresponding to the partition will become a bottleneck, causing the overall load imbalance, i.e., the data skew problem. One of the root causes of the data skew problem is that the partition strategy is unreasonable, which causes the partition containing too many two-tuples to be unable to be effectively processed in parallel. Therefore, when optimizing the partition strategy, the key is to find the balance between the number of computing kernels and the number of two-tuples to be processed, to ensure that each computing kernel can be evenly distributed to an appropriate number of two-tuples. In this way, not only the utilization rate of the computing kernel can be improved, but also the performance degradation caused by uneven data distribution can be avoided.

[0033] Step 103: Obtain the optimal number of partitions in the distributed computing task based on the sampling result and the number of computing kernels.

[0034] In one embodiment, the main cause of data skew is uneven distribution of high-frequency keys. To alleviate the problem of data skew, the tuples of high-frequency keys are divided into multiple partitions, and multiple computing kernels are used to process in parallel, thereby reducing the data load of a single partition. The sampling result can more accurately reflect the actual distribution of keys, and can also reflect the distribution of key high-frequency keys that cause data skew. In combination with the number of computing kernels in the cluster and the sampling result, the number of partitions is adaptively adjusted, the tuples in the sampling result are uniformly divided into multiple batches, and each batch of tasks can be one-to-one corresponding to a computing kernel. This method not only helps to achieve load balancing between tasks, but also effectively improves the utilization rate of system kernels.

[0035] Step 104: dividing the sampling result into high-frequency keys and low-frequency keys, and performing different data partitioning strategies on the high-frequency keys and the low-frequency keys based on the optimal number of partitions.

[0036] In this embodiment, the key keys are divided into high-frequency keys and low-frequency keys according to the frequency of the key keys, the high-frequency keys are re-partitioned based on the weighted round robin algorithm, and the low-frequency keys are re-partitioned based on the hash method, so that each partition is load balanced and the execution time of the system is reduced.

[0037] In one embodiment, the sampling result is divided into high-frequency keys and low-frequency keys, and the high-frequency keys and the low-frequency keys are respectively enlarged according to the sampling rate to obtain the distribution of high-frequency keys and the distribution of low-frequency keys in the to-be-processed tuples. The high-frequency key partitioning strategy is performed on the high-frequency keys, and the low-frequency key partitioning strategy is performed on the low-frequency keys. The optimal number of partitions is the total number of partitions in the high-frequency key partitioning strategy and the low-frequency key partitioning strategy, and the sampling rate is the ratio of the number of tuples in the sampling result to the number of to-be-processed tuples.

[0038] In this embodiment, the number of partitions is adaptively adjusted to match the number of tuples, and the balance between the number of computing kernels and the number of to-be-processed tuples is found, so that each computing kernel can be evenly distributed to an appropriate number of tuples, and each computing kernel in each batch of tasks has a corresponding task to execute, thereby improving the utilization rate of computing kernels. Based on the above number of partitions, different partitioning strategies are performed on keys of different frequencies to balance the task load and reduce the execution time of the system.

[0039] Before sampling the to-be-processed binary tuples, the number of binary tuples received by the Reduce task can be counted by modeling the original binary tuple data to preliminarily determine whether there is a problem of unbalanced load between Reduce tasks, i.e., a data skew problem, wherein the original binary tuple includes a plurality of keys and values corresponding to each key, and the original binary tuple is original data before partitioning by a default partitioning strategy of Spark.

[0040] In one embodiment, in order to more clearly introduce the process of modeling the original binary tuple data, it is first declared that the key variables in Table 1.

[0041] Table 1 Variable declaration of skew model

[0042] As shown in Table 1, the key variables in the embodiment 1 include: Figure 2 Step 201: Model the original binary tuple to obtain the skew degree and load balancing level of the original binary tuple.

[0043] In one embodiment, the original binary tuple is the binary tuple received by the Reduce task from the Map task , and the number of binary tuples received by the Reduce task from the Map task can be calculated based on the original binary tuple . The calculation formula of the total number of all binary tuples transmitted between the Map task and the Reduce task is as follows:

[0044] . The number of binary tuples received by each Reduce task , under different data distribution characteristics, is represented as the skew degree , the number of binary tuples received by each Reduce task , is represented as the skew degree , and the number of binary tuples received by the Reduce task from the Map task is represented as the skew degree . The calculation formula of the number of binary tuples received by the Reduce task from the Map task is as follows:

[0045] . According to the above calculation, the average number of binary tuples received by each Reduce task An ideal Reduce task , the load balancing optimization result of , can be calculated by the following formula:

[0046] wherein, the skewness is , the number of binary tuples received by each Reduce task , the average number of binary tuples received by each Reduce task The greater the difference, the more serious the skewness of the original binary tuples, and the skewness of the original binary tuples can be calculated by the following formula:

[0047] The standard deviation commonly used in probability statistics is used to analyze the dispersion degree of the binary tuples in the task load, and the skewness is , the load balancing level of each Reduce task , wherein the load balancing level of each Reduce task , can be defined as:

[0048] Step 202: When the skewness is greater than the load balancing level, sampling the to-be-processed binary tuples in the distributed computing task to obtain a sampling result.

[0049] In one embodiment, when the skewness of the original binary tuples - is greater than the load balancing level of the Reduce task , it means that the Reduce task has data skew, and the load of the Reduce task exceeds the average level. At this time, the to-be-processed binary tuples in the distributed computing task are sampled to obtain a sampling result. In order to quantify the effect of load balancing between Reduce tasks , the embodiment defines as the skewness of the task load, representing the skewness between partitioned data, and the calculation method is: .

[0050] Wherein, the better the load balancing effect of the partition, that is, the smaller the value of , the smaller the value of . Therefore, the final effect of data skew optimization is to minimize the value of . ​​​​

[0051] When the data skew occurs, a sampling result is obtained based on the frequency of the key in the to-be-processed two-tuple, the sampling result is analyzed, the key causing the data skew in the default partition strategy is identified, and the partition strategy is optimized according to the sampling result. Since the keys with different frequencies have different calculation weights in the sampling, in order to make each key have an equal opportunity to be selected in the sampling and make the sampling result accurately describe the distribution of the to-be-processed two-tuple as much as possible, the embodiment adopts a negative sampling algorithm based on dynamic weighting to sample the to-be-processed two-tuple.

[0052] In one embodiment, in order to more clearly introduce the sampling process of the to-be-processed two-tuple, it is first declared that the sampling parameters in Table 2 are used.

[0053] Table 2 Sampling parameters used by the sampling method

[0054] As shown in Table 2, the sampling parameters used by the sampling method are as follows: Figure 3 The method for implementing step 101 in the embodiment 1 specifically includes the following steps: Figure 1 Step 301: Dividing the to-be-processed two-tuple into different partitions. In one embodiment, the to-be-processed two-tuple is divided into different partitions according to the Spark default partition strategy.

[0055] Step 302: Sampling the two-tuple in each partition to obtain a sampling result corresponding to each partition.

[0056] In one embodiment, the two-tuple in a partition is divided according to a sampling interval to obtain a plurality of intermediate two-tuples, the intermediate two-tuples include a plurality of keys and a frequency corresponding to each key; a probability that a key in the intermediate two-tuple is sampled and retained is calculated; when the probability is greater than a first random value, the key and the corresponding frequency are taken as a sampling result, and until all keys in all intermediate two-tuples are calculated, a sampling result corresponding to each partition is obtained. The embodiment adopts a negative sampling algorithm based on dynamic weighting to sample the to-be-processed two-tuple. The negative sampling algorithm is specifically as follows: the sampling process is performed in each partition at the same time, for a partition i, first, the sampling interval in the partition i is initialized

[0057] , the two-tuple in a partition is divided according to the sampling interval , a plurality of intermediate two-tuples are obtained, the intermediate two-tuples include a plurality of keys and a frequency corresponding to each key; a probability that a key in the intermediate two-tuple is sampled and retained is calculated; when the probability is greater than a first random value, the key and the corresponding frequency are taken as a sampling result, and until all keys in all intermediate two-tuples are calculated, a sampling result corresponding to each partition is obtained. Divide the data into multiple intermediate tuples. For each key in an intermediate tuple, determine whether it is a high-frequency key. If it is a high-frequency key, add a dynamic weight to the key and calculate the probability of the key being sampled and retained based on the dynamic weight. If it is a low-frequency key, set the weight of the key to 0 and calculate the probability of the key being sampled and retained based on the weight.

[0058] If the probability of a key being sampled and retained is greater than the first random value, then the key and its corresponding frequency are considered as a sampling result, and the key and its corresponding frequency are written into the sampling result table. And update the sampling results table. Partition Number of pairs not included in the sampling and partitions The number of pairs that still need to be sampled Based on partitions Number of pairs not included in the sampling and partitions The number of pairs that still need to be sampled Update the sampling interval in partition i. This ensures that the sampling intervals are consistent across partitions, giving each pair of pairs to be sampled a chance of being selected. After one sampling cycle, if there are still unsampled pairs and the required sample size has not yet been reached (i.e., partitions are excluded from the sampling), the remaining pairs are considered unsampled. Number of pairs to be sampled If no pairs are found to be sampled, sampling continues; if no pairs are found to be sampled or a sufficient number of pairs have been collected, sampling is complete. Finally, the sampling results for each partition are aggregated on the master node to obtain the sampling results for all partitions.

[0059] In one embodiment, the sampling interval in partition i The calculation formula is: Among them, random variables Values ​​range from 0 to 1, sampling interval Based on partition Number of binary pairs not included in the sampling and partitions The number of pairs that still need to be sampled The changes are influenced by the fact that each pair of pairs to be processed may be sampled. This represents the probability that a key will be sampled and retained, or the probability that a specific key will be sampled and retained. The calculation formula is as follows: .

[0060] in, Represents a specific key Total frequency of occurrence It is the sum of the frequencies of all keys. When a specific key... Frequency in partitions When it is greater than the average frequency, it indicates that the current It can be preliminarily determined to be a high-frequency key. At this point, in 0 and Choose a random number between As Additional weights, so that It is easier to pass the sampling; when When the frequency is less than or equal to the average frequency, it indicates that the current frequency is less than or equal to the average frequency. It can be preliminarily determined to be a low-frequency event. At this point, the [event name] should be [determined]. weights Set to zero.

[0061] When calculated corresponding Then, a first random value less than 1 will be obtained. Compare the first random value with... If the first random value is less than Then accept ,Will As a sampling result, and The corresponding frequencies are written into the sampling results table. And update the sampling results table. Partition Number of binary pairs not included in the sampling and partitions The number of pairs that still need to be sampled Otherwise, I will not accept it. Continue sampling. When evaluating each key in an intermediate tuple, since the random number is uncertain, some low-frequency keys may also be sampled and accepted. This setting takes into account the possibility of both high-frequency and low-frequency keys being sampled and accepted.

[0062] Step 303: Combine the sampling results corresponding to each partition to obtain the sampling results corresponding to all partitions.

[0063] In one embodiment, step 302 is performed on each partition to obtain the sampling results for each partition. The sampling results for each partition are then integrated into a table to obtain the sampling results for all partitions. Although the sampling process consumes some of the program's runtime, the sampling results are of great significance for guiding the number of partitions and the repartitioning of tuples.

[0064] In the Spark framework, the number of partitions directly affects the parallelism of the Reduce task. The system will start the corresponding number of Reduce tasks according to the number of partitions, and too many partitions will cause multiple batch thread switching overhead, and too few partitions will cause idle computing kernels, so it is necessary to find an optimal number of partitions to ensure efficient use of computing kernels and avoid unnecessary performance loss.

[0065] In one embodiment, in order to more clearly introduce the process of obtaining the optimal number of partitions, it is first declared that the parameters in Table 3 are as follows.

[0066] Table 3 Description of variables used in task statistics

[0067] As shown in Table 1, the method for implementing steps 102-103 in the embodiment 1 provided by the present embodiment comprises the following steps: Figure 4 Figure 1 The method for implementing steps 102-103 in the embodiment 1 provided by the present embodiment comprises the following steps: Step 401: Obtain the total number of two-tuples in the sampling result.

[0068] Step 402: Based on the total number of two-tuples and the number of computing kernels, obtain the number of two-tuples that need to be calculated by a single kernel.

[0069] In one embodiment, based on the statistical data in the above table, first, the two-tuples can be evenly divided to each computing kernel for execution in an ideal case, based on the total number of two-tuples in the sampling result and the number of computing kernels CN of each node in the cluster, the number of two-tuples that need to be calculated by a single kernel is obtained as follows: Here, the data result in the sampling data is used to approximate the overall state of the two-tuples to be processed.

[0070] Step 403: Obtain the maximum frequency value in the sampling result.

[0071] In one embodiment, the frequency record of the key in the sampling result is counted. Let represent the frequency value of a specific key, represent the maximum frequency value in the sampling result.

[0072] Step 404: When the maximum frequency value is greater than the number of two-tuples that need to be calculated by a single kernel, based on the number of two-tuples, obtain the optimal number of partitions in the distributed computing task.

[0073] In one embodiment, when the maximum frequency value is greater than the number of two-tuples that need to be calculated by a single kernel , the method for obtaining the optimal number of partitions in the distributed computing task comprises the following steps:​ and the greatest common divisor of , at this time the kernel computing resources can be fully utilized, but when and are prime numbers, and the greatest common divisor of is too small to be realistic, at this time a threshold is set to represent the minimum value of . Set and the maximum of the greatest common divisor of as the optimal number of binary tuples for single kernel computing , based on the optimal number of binary tuples for single kernel computing and the total number of binary tuples in the sampling result to get the optimal number of partitions .

[0074] In an embodiment, when , it means that the binary tuple data corresponding to the key exceeds the average level, and if the hash-based strategy (the default partition strategy in Spark) is directly partitioned, it will inevitably cause data skew to occur. At this time, the binary tuple needs to be divided into multiple small batches for batch execution, which increases the concurrency of the system while reducing the execution time of the task. When data is calculated in batches, in order to fully utilize kernel computing resources, the ideal division is to achieve the integral division of high-frequency key data, that is, to find the greatest common divisor between the data of and .

[0075] Define GCD as the greatest common divisor of and , represent the optimal number of binary tuples for single kernel computing, the calculation formula is: .

[0076] When and are prime numbers, set a threshold to represent the minimum value of , the calculation formula is: .

[0077] At this time, the optimal number of binary tuples for single kernel computing the calculation formula is: .

[0078] The optimal number of partitions The calculation formula is:

[0079] After determining the optimal number of partitions, a specific partitioning strategy needs to be designed to ensure that tuples are evenly distributed across all Reduce tasks, reducing system execution time. Since high-frequency keys have a large data volume, without proper control, some partitions may run for extended periods, slowing down the entire job. Conversely, low-frequency keys have a smaller data volume; processing them separately would be inefficient and wasteful of resources. Therefore, different partitioning strategies need to be applied to high-frequency and low-frequency keys.

[0080] First, for the pairs in the sampling results obtained in step 303, according to the Pareto principle, the keys in the pairs with the highest frequency (n%) are classified as high-frequency keys, and the keys in the pairs with the lowest frequency (n%) are classified as low-frequency keys, where n is a positive integer from 0 to 100.

[0081] In one embodiment, to more clearly explain the High Frequency Key Partition (HFKP) strategy, the following key variables are declared first.

[0082] (1) This is the optimal number of partitions calculated in step 404.

[0083] (2) Record the statistical values ​​of high-frequency keys in the sampling results of all partitions. High-frequency keys refer to the high-frequency keys obtained after partitioning using the Pareto principle. The formula for calculating FH is: .

[0084] in, Indicates partition Distribution of mid-to-high frequency keys This refers to the number of partitions in the Map task. It is the frequency of the high-frequency key appearing in the sampling results.

[0085] (3) Record the frequency distribution of high-frequency keys in the tuples to be processed. Based on The statistical values ​​of high-frequency keys obtained after Pareto principle partitioning, combined with the sampling rate ratio, can approximately estimate the frequency distribution of high-frequency keys in the pairs to be processed. After obtaining the frequency distribution of high-frequency keys, subsequent partitioning is based on this frequency distribution. The calculation formula is:

[0086] (4) This is a partition information table, which includes the partition ID, the partition weight corresponding to each partition ID, and the partition availability corresponding to each partition ID. The partition weight is the basis for data partitioning. Record the partition weight of partition j. Partition availability indicates whether partition j is available. The initial value of partition availability is 1.

[0087] (5) Store tuples (key, ),in, This is the partition ID corresponding to the key.

[0088] like Figure 5 As shown, this embodiment provides... Figure 1 The implementation method of step 104 specifically includes the following steps: Step 501: Set up a partition information table, which includes at least one partition ID and the partition weight corresponding to each partition ID.

[0089] In one embodiment, first according to step 404 Initialize with sampling rate ratio Value: During initialization, the partition weights corresponding to each partition are... The same and equal to the average number of pairs processed per partition. Partition weight This reflects the load level of the partition during runtime, and the partition weight. A higher value indicates a lower current load level for the partition, meaning the partition can accommodate more high-frequency keys. When assigning a tuple to a partition, the partition weight gradually decreases, thus reducing the number of tuples assigned to that partition and balancing the load across all partitions.

[0090] Step 502: For each high-frequency key, perform the following steps: Obtain the target partition ID corresponding to the largest partition weight in the partition information table, use the partition corresponding to the target partition ID as the partition corresponding to the high-frequency key, and update the partition information table until all partitions corresponding to the high-frequency keys are obtained.

[0091] In one embodiment, for each high-frequency key, the partition information table is first looked up. The target partition ID corresponding to the largest partition weight in the middle, that is, finding The partition information of the available partition corresponding to the largest partition weight in the middle. Match the high-frequency key with the partition information, and assign the high-frequency key to the partition corresponding to the target partition ID, and then (key, Adding binary form to In this process, a partition record is created in the RP (Real Resource Plane), serving as the result of this partitioning. After assigning the high-frequency keys to the partitions corresponding to the target partition IDs, the process is updated. Partition weight corresponding to the target partition ID Partition weights The update method is as follows: , When polling for each high-frequency key, the high-frequency key can be sorted according to partition weight and repartitioned into the corresponding partition ID. This enables a partitioning strategy based on high-frequency keys during the shuffle phase.

[0092] In one embodiment, to more clearly explain the Low Frequency Key Partition (LFKP) strategy, the following key variables are declared first.

[0093] (1) Record the statistical values ​​of low-frequency keys in the sampling results of all partitions. The low-frequency keys are those obtained after partitioning using the Pareto principle. The formula for calculating LH is: .

[0094] in, Indicates partition Distribution of low- to medium-frequency keys This is the number of partitions in the Map task. It represents the frequency of low-frequency keys appearing in the sampling results.

[0095] (2) Record the frequency distribution of low-frequency keys in the tuples to be processed. Based on The statistical values ​​of low-frequency keys obtained after Pareto principle partitioning, combined with the sampling rate ratio (i.e., the sampling rate ratio in HFKP), can approximately estimate the frequency distribution of low-frequency keys in the tuples to be processed. After obtaining the frequency distribution of low-frequency keys, subsequent partitioning is based on this frequency distribution. The calculation formula is: .

[0096] (3) The latest partition information table after inheriting the HFKP strategy .

[0097] (4) Update and in the following way: , .

[0098] (5) stores the key-value pair (key, value). When the partition PW obtained by Murmurhash for a low-frequency key cannot meet the space size required by the key, a pseudo-random probing rehashing method is used to obtain a new partition ID for the conflict key.

[0099] (6) records the number of partitions available for matching in the current . When a mismatched partition occurs, a pseudo-random probing rehashing method is used to obtain a new partition ID , which is calculated as .

[0100] As shown in Figure 6 , the method for implementing step 104 in the embodiment provided by the present application specifically includes the following steps: Figure 1 Step 601: Obtain a partition information table, wherein the partition information table is the latest partition information table generated after the high-frequency key partition strategy is completed, and the partition information table includes at least one partition ID, a partition weight corresponding to each partition ID, and a partition availability corresponding to each partition ID. Step 602: For each low-frequency key, the following steps are performed:

[0101] Calculate a first hash value corresponding to the low-frequency key, and use the first hash value as a partition ID. In one embodiment, a Murmurhash method is used to calculate a first hash value corresponding to the low-frequency key, and the first hash value is used as a partition ID.

[0102] Step 603: Obtain a partition weight corresponding to the partition ID.

[0103] In one embodiment, based on the above-mentioned partition information table

[0104] , the partition weight corresponding to the partition ID is found. Step 604: Compare the partition weight with a threshold value to obtain a partition corresponding to the low-frequency key.

[0105]

[0106] ​​In one embodiment, the threshold value comprises a lower threshold value, an upper threshold value and a full load threshold value, when the partition weight value is greater than the lower threshold value and less than the upper threshold value, the partition corresponding to the partition ID is taken as the partition corresponding to the low-frequency key, and the partition information table is updated (i.e., updated by the method in (4)), until all the partitions corresponding to the low-frequency key are obtained; When the partition weight value is less than the lower threshold value, it means that the space of the partition is too small to store the data corresponding to the key; when the partition weight value is greater than the upper threshold value, it means that the space of the partition is too large, and storing the data corresponding to the key will waste space, at this time, a new partition should be found, i.e., the second hash value of the low-frequency key is recalculated (i.e., the pseudo-random probe rehashing method in (6) is used), to obtain the partition corresponding to the low-frequency key; when the partition weight value is less than the full load threshold value, the partition availability is modified, the number of available partitions is updated, and the hash value of the low-frequency key is recalculated (i.e., the pseudo-random probe rehashing method in (6) is used), to obtain the partition corresponding to the low-frequency key.

[0107] As shown in Figure 7 , it is a flow chart of the real-time partition method provided by the embodiment for alleviating the data skew problem.

[0108] Specifically: in Spark, a large amount of data generated after a series of RDD conversion operations on the data read from the original data source (such as HDFS or a database) is divided into multiple Stages (i.e., Stage_0~Stage_j in Figure 7 ). For each Stage, according to the default partition strategy of Spark, the data in the Stage is further divided into multiple partitions (i.e., partition 0~partition m-1 in Figure 7 ), a sampling operation is performed on each partition in parallel to obtain a sampling result, the sampling result includes the distribution of keys in all partitions; and the granularity is adaptively adjusted based on the sampling result and the computing resources (the number of computing cores in step 102), the granularity can also be referred to as the number of partitions, to obtain the optimal number of partitions; the sampling result is divided into high-frequency keys and low-frequency keys, the high-frequency partition strategy is performed on the high-frequency keys, and the low-frequency partition strategy is performed on the low-frequency keys. Wherein, when a key is divided into the corresponding partition in the execution of the high-frequency partition strategy, the data corresponding to the key is written into a file and an index file (i.e., the RP file in the foregoing, which stores the keys and the partition ID corresponding to each key) is generated, and then each Reduce task pulls the data in the partition by the partition ID for computing and processing, and the specific content of the execution of the low-frequency partition strategy is similar to that of the execution of the high-frequency partition strategy, which will not be described here.

[0109] Embodiment 2 On the basis of the real-time partitioning method for relieving data skew problem provided in the foregoing embodiments, the application further provides a device for real-time partitioning for relieving data skew problem, which can be used to implement the method. Figure 8 As shown in FIG. 1, which is a schematic diagram of the device architecture of an embodiment of the application. The device for real-time partitioning for relieving data skew problem in this embodiment comprises one or more processors 21 and a memory 22. In this embodiment, the processor 21 is taken as an example. Figure 8 The processor 21 and the memory 22 can be connected through a bus or other means. In this embodiment, the connection through a bus is taken as an example.

[0110] Figure 8 The memory 22, as a non-volatile computer readable storage medium for real-time partitioning for relieving data skew problem, can be used to store non-volatile software programs and non-volatile computer executable programs, such as the real-time partitioning method for relieving data skew problem in the foregoing embodiments. The processor 21 executes various functional applications and data processing of the device for real-time partitioning for relieving data skew problem by running the non-volatile software programs, instructions and modules stored in the memory 22, that is, implements the real-time partitioning method for relieving data skew problem in the foregoing embodiments.

[0111] The memory 22 can include a high-speed random access memory and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 22 can include a memory remotely arranged relative to the processor 21, and these remote memories can be connected to the processor 21 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0112] The program instructions / modules are stored in the memory 22, and when executed by the one or more processors 21, the real-time partitioning method for relieving data skew problem in the foregoing embodiments is executed, for example, the various steps described above are executed.

[0113] The application further provides a non-volatile computer storage medium, which stores computer executable instructions. When executed by one or more processors, for example, the processor 21 in FIG. 1, the one or more processors can execute the real-time partitioning method for relieving data skew problem in the foregoing embodiments, for example, execute the various steps described above. Figures 1-6

[0114] The application further provides a non-volatile computer storage medium, which stores computer executable instructions. When executed by one or more processors, for example, the processor 21 in FIG. 1, the one or more processors can execute the real-time partitioning method for relieving data skew problem in the foregoing embodiments, for example, execute the various steps described above. Figure 8 Figures 1-6

[0115] ​​​It is worth noting that the information interaction and execution process between the modules and units in the above-mentioned device and system are based on the same concept as the processing method embodiment of the present invention. For details, please refer to the description in the method embodiment of the present invention, and will not be repeated here.

[0116] Those skilled in the art will understand that all or part of the steps in the various methods of the embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0117] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A real-time partitioning method for mitigating data skew issues, characterized in that, The method comprises: sampling the to-be-processed tuples in the distributed computing task to obtain a sampling result; obtaining the number of computing kernels in the distributed computing task; based on the sampling result and the number of computing kernels, obtaining the optimal number of partitions in the distributed computing task; dividing the sampling result into high-frequency keys and low-frequency keys, and performing different data partition strategies on the high-frequency keys and the low-frequency keys based on the optimal number of partitions.

2. The real-time partitioning method for mitigating data skew problem according to claim 1, wherein, The method further comprises: dividing the to-be-processed tuples into different partitions; sampling the tuples in each partition to obtain a sampling result corresponding to each partition; integrating the sampling result corresponding to each partition to obtain a sampling result corresponding to all partitions.

3. The real-time partitioning method for mitigating data skew problem according to claim 2, wherein, The method further comprises: dividing the tuples in a partition according to a sampling interval to obtain a plurality of intermediate tuples, the intermediate tuples comprising a plurality of keys and a frequency corresponding to each key; calculating the probability of a key in the intermediate tuples being sampled and retained; when the probability is greater than a first random value, taking the key and the corresponding frequency as a sampling result, and repeating the calculation until all keys in all intermediate tuples are calculated to obtain a sampling result corresponding to each partition.

4. The real-time partitioning method for mitigating data skew problem according to claim 1, wherein, The method further comprises: performing data modeling on the original tuples to obtain the skew degree and the load balancing level of the original tuples; when the skew degree is greater than the load balancing level, sampling the to-be-processed tuples in the distributed computing task to obtain a sampling result.

5. The real-time partitioning method for mitigating data skew problem according to claim 1, wherein, The method further comprises: obtaining the total number of tuples in the sampling result; based on the total number of tuples and the number of computing kernels, obtaining the number of tuples that need to be calculated by a single kernel; obtaining the maximum frequency value in the sampling result; when the maximum frequency value is greater than the number of tuples that need to be calculated by a single kernel, based on the number of tuples, obtaining the optimal number of partitions in the distributed computing task.

6. The real-time partitioning method for mitigating data skew problem according to claim 1, wherein, The method further comprises: setting a partition information table, the partition information table comprising at least one partition ID and a partition weight corresponding to each partition ID; for each high-frequency key, performing the following steps: obtaining a target partition ID corresponding to the maximum partition weight in the partition information table, taking the partition corresponding to the target partition ID as the partition corresponding to the high-frequency key, and updating the partition information table until the partitions corresponding to all high-frequency keys are obtained.

7. The real-time partitioning method for mitigating data skew problem according to claim 1, wherein, The method further comprises: obtaining a partition information table, wherein the partition information table is the latest partition information table generated after the high-frequency key partition strategy is completed, and the partition information table comprises at least one partition ID, a partition weight corresponding to each partition ID, and a partition availability corresponding to each partition ID; for each low-frequency key, performing the following steps: calculating a first hash value corresponding to the low-frequency key, and taking the first hash value as a partition ID; obtaining a partition weight corresponding to the partition ID; comparing the partition weight with a threshold value to obtain a partition corresponding to the low-frequency key.

8. The real-time partitioning method for mitigating data skew problem according to claim 7, wherein, The threshold value comprises a lower threshold value, an upper threshold value, and a full-load threshold value, and the method further comprises: when the partition weight is greater than the lower threshold value and less than the upper threshold value, taking the partition corresponding to the partition ID as the partition corresponding to the low-frequency key, and updating the partition information table until the partitions corresponding to all low-frequency keys are obtained; when the partition weight is less than the lower threshold value or greater than the upper threshold value, recalculating the second hash value corresponding to the low-frequency key to obtain the partition corresponding to the low-frequency key; when the partition weight is less than the full load threshold value, modifying the partition availability, and recalculating the hash value corresponding to the low-frequency key to obtain the partition corresponding to the low-frequency key.

9. A real-time partitioning apparatus for mitigating data skew issues, the apparatus comprising: a data skew detection module configured to detect data skew in a data stream; and a data skew mitigation module configured to mitigate the data skew in the data stream. comprise: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to execute the real-time partitioning method for relieving data skew problems in any one of claims 1-8.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor, so that the computer executes the real-time partitioning method for relieving data skew problems in any one of claims 1-8. The computer program is executed by the processor, so that the computer executes the real-time partitioning method for relieving data skew problems in any one of claims 1-8.