Data skew processing method and device, electronic equipment and storage medium

By real-time reception and analysis of statistical results on the data consumer side, the data skew problem existing in Spark Streaming is automatically processed, which solves the inefficiency problem caused by relying on manual processing in the existing technology, and improves the performance and stability of Spark Streaming.

CN120045603APending Publication Date: 2025-05-27BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510068268.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, data skew is a common and difficult problem in Spark Streaming, which leads to the processing time of some nodes being too long and may even have abnormalities such as memory overflow, which seriously affects the performance and stability of Spark Streaming. The existing processing methods rely on manual processing and are inefficient.

Method used

By receiving the statistical results of statistics on each batch of data in real time, we will determine whether there is data skewed based on these results, and in the case of data skewed, the data skewed is automatically processed based on the statistical results. The specific processing methods include adjusting the number of partitions and pre-processing the data of the target skew key.

Benefits of technology

The automated processing of data skew is realized, the processing efficiency is improved, manual intervention is reduced, and the adverse effects of data skew on Spark Streaming are quickly solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045603A_ABST
    Figure CN120045603A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data skew processing method and device, electronic equipment and a storage medium. The method comprises the steps of receiving a statistical result obtained by performing statistics on each batch of data by a data consumption end in real time; wherein the statistical result is used for representing index data of data skew; judging whether data skew exists or not based on a statistical result; and under the condition that the data skew is judged to exist based on the statistical result, processing the data skew according to the statistical result. According to the technical scheme, when it is judged that the data skew exists based on the statistical result counted by the data consumption end, automatic processing of the data skew is achieved according to the statistical result, and the processing process of the data skew does not need manual intervention; compared with an existing manual data skew processing method, the data skew processing method has the advantages that the processing efficiency is greatly improved, and the adverse effect of data skew on Spark Streaming is solved as soon as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and in particular, to a method, apparatus, electronic device, and storage medium for data skew processing. Background Art

[0002] Spark Streaming, as a component of the Spark framework, is specifically used for processing streaming data and is widely used because of its powerful stream processing capabilities, tight integration with the Spark ecosystem, and high throughput and fault tolerance.

[0003] Data skew is a common and tricky problem in Spark Streaming. It can cause the processing time of some nodes to be too long, and even memory overflow and other exceptions may occur, seriously affecting the performance and stability of Spark Streaming. To solve data skew, currently, it mainly relies on operation and maintenance personnel to handle it manually. Since a large amount of mechanical work needs to be done manually, the existing processing method is inefficient. Summary of the Invention

[0004] In view of this, to effectively alleviate the technical problem of low efficiency caused by manual data skew processing, embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for data skew processing.

[0005] In a first aspect, embodiments of the present invention provide a method for data skew processing. The method is applied to a data production end and includes:

[0006] Receiving in real time the statistical results of each batch of data counted by a data consumption end; wherein the statistical results are used to represent the index data of data skew;

[0007] Determining whether there is data skew based on the statistical results;

[0008] In the case where it is determined that there is data skew based on the statistical results, processing the data skew according to the statistical results.

[0009] In a possible implementation manner, the statistical results include the total count of each skewed key and the execution duration of at least one skewed task that performs the statistics;

[0010] Determining whether there is data skew based on the statistical results includes:

[0011] Determining whether there is data skew based on the total count of each skewed key and the execution duration of at least one skewed task.

[0012] In a possible implementation, determining whether there is data skew based on the total count of each skewed key and the execution duration of at least one skewed task includes:

[0013] Detecting whether the execution duration of at least one skewed task exceeds a preset duration threshold;

[0014] When it is detected that the execution duration of at least one skewed task exceeds the preset duration threshold, determining that there is data skew;

[0015] When it is detected that the execution duration of each skewed task in at least one skewed task does not exceed the preset duration threshold, calculating the statistical difference of each skewed key based on the total count of each skewed key;

[0016] Detecting whether the statistical difference of each skewed key among multiple skewed keys exceeds a preset difference threshold;

[0017] When it is detected that the statistical difference of at least one skewed key among multiple skewed keys exceeds the preset difference threshold, determining that there is data skew.

[0018] In a possible implementation, calculating the statistical difference of each skewed key based on the total count of each skewed key includes:

[0019] Calculating the mean and variance of each skewed key based on the total count of each skewed key;

[0020] Detecting whether the statistical difference of each skewed key among multiple skewed keys exceeds a preset difference threshold includes:

[0021] Detecting whether the mean of each skewed key among multiple skewed keys exceeds a preset mean threshold and whether the variance exceeds a preset variance threshold.

[0022] In a possible implementation, processing data skew according to the statistical results includes:

[0023] Processing data skew based on the total count of each skewed key and the execution duration of at least one skewed task.

[0024] In a possible implementation, processing data skew based on the total count of each skewed key and the execution duration of at least one skewed task includes:

[0025] Determining whether the number of target skewed tasks whose execution duration exceeds the preset duration threshold is greater than or equal to a preset number;

[0026] When it is determined that the number of target skew tasks is greater than or equal to a preset number, determine the number of partition increments based on the execution duration of the target skew tasks;

[0027] Process data skew by increasing the number of partition increments;

[0028] When it is determined that the number of target skew tasks is less than the preset number, obtain the target skew keys whose statistical differences exceed the preset difference threshold;

[0029] Process data skew by performing data preprocessing on the data of the target skew keys.

[0030] In a possible implementation, determining the number of partition increments based on the execution duration of the target skew tasks includes:

[0031] When the number of target skew tasks is greater than the preset number, obtain the maximum execution duration and the minimum execution duration from the execution durations of the multiple target skew tasks;

[0032] Determine the execution duration difference based on the maximum execution duration and the minimum execution duration;

[0033] Query the number of partition increments corresponding to the execution duration difference from the partition increment query table; wherein, the partition increment query table pre-stores the corresponding relationship between the execution duration difference and the number of partition increments;

[0034] When the number of target skew tasks is equal to the preset number, determine the preset number of partitions as the number of partition increments.

[0035] In a possible implementation, processing data skew according to the statistical results includes:

[0036] Input the statistical total of each skew key and the execution duration of at least one skew task into the data skew processing strategy model, and the data skew processing strategy model outputs a data skew processing strategy; wherein, the data skew processing strategy is an increase partition number strategy or a data preprocessing strategy, the increase partition number strategy includes the number of partition increments, the data preprocessing strategy includes the target skew key, and the data skew processing strategy model is obtained by training a classification model using the statistical total of each skew key and the execution duration of the skew task;

[0037] Process data skew based on the data skew processing strategy.

[0038] In a second aspect, an embodiment of the present invention provides a device for processing data skew, which is applied to a data production end, and the device includes:

[0039] A receiving module, configured to receive in real time the statistical results of the data consumer end for each batch of data; wherein, the statistical results are used to represent the index data of data skew.

[0040] A determination module, configured to determine whether there is data skew based on the statistical results.

[0041] A data skew processing module, configured to, when it is determined based on the statistical results that there is data skew, perform processing on the data skew according to the statistical results.

[0042] In a third aspect, an embodiment of the present invention provides an electronic device, which includes: a processor and a memory, and the processor is configured to execute a program for data skew processing stored in the memory to implement the data skew processing method described above.

[0043] In a fourth aspect, an embodiment of the present invention provides a storage medium, where the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the data skew processing method described above.

[0044] The data skew processing method, device, electronic device and storage medium provided by the embodiments of the present invention, the method includes: receiving in real time the statistical results of the data consumer end for each batch of data; wherein, the statistical results are used to represent the index data of data skew; determining whether there is data skew based on the statistical results; when it is determined based on the statistical results that there is data skew, performing processing on the data skew according to the statistical results. In the above technical solution, when it can be determined that there is data skew based on the statistical results statistically obtained by the data consumer end, the data skew is automatically processed according to the statistical results. The processing process of this data skew does not require manual intervention. This data skew processing method greatly improves the processing efficiency compared with the existing manual data skew processing method, and quickly solves the adverse effects brought by data skew to Spark Streaming. Description of the Drawings

[0045] Figure 1 It is a flowchart of an embodiment of a data skew processing method provided by an embodiment of the present invention.

[0046] Figure 2 It is a flowchart of an embodiment of another data skew processing method provided by an embodiment of the present invention.

[0047] Figure 3 It is a block diagram of an embodiment of a data skew processing device provided by an embodiment of the present invention.

[0048] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0049] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] For the convenience of understanding the embodiments of the present invention, the following will further explain with specific examples in conjunction with the accompanying drawings. The examples do not constitute a limitation to the embodiments of the present invention.

[0051] The embodiments of the present invention provide a method for data skew processing. This method is applied to the data production end. Refer to Figure 1 , Figure 1 which is a flowchart of an embodiment of a method for data skew processing provided by the embodiments of the present invention. Figure 1 The shown process may include the following steps:

[0052] Step 101, receive in real time the statistical results of the data consumer for each batch of data;

[0053] Generally, before the method for data skew processing, it is necessary to install the corresponding SDK (Software Development Kit) at both the data consumer end and the data production end. So that after the data consumer that imports the SDK starts the corresponding data consumption project, it begins to perform data statistics. Since Spark Streaming actually processes data in small batches one by one, therefore, in this embodiment, statistics can be performed on each batch of data to obtain statistical results. Since it is necessary to analyze whether there is data skew based on the statistical results later, the statistical results are index data that can be used to characterize data skew, so as to determine whether there is data skew through the statistical results.

[0054] Step 102, determine whether there is data skew based on the statistical results;

[0055] Since the above statistical results are index data that can standardize data skew, therefore, it can be determined whether there is data skew based on the statistical results. If it is determined that there is data skew based on the statistical results, then step 103 needs to be executed to process the data skew, while if it is determined that there is no data skew based on the statistical results, then step 103 does not need to be executed, that is, there is no need to perform data skew processing.

[0056] Step 103, in the case where it is determined that there is data skew based on the statistical results, implement the processing of the data skew according to the statistical results.

[0057] Data skew is caused by uneven data distribution, and statistical results can identify and analyze this uneven distribution. Thus, targeted processing methods can be adopted according to the statistical results to effectively process data skew, so as to improve the efficiency and stability of data processing.

[0058] The method for processing data skew provided by the embodiments of the present invention includes: receiving in real time the statistical results of each batch of data by the data consumer; wherein, the statistical results are used to represent the index data of data skew; determining whether there is data skew based on the statistical results; and when it is determined based on the statistical results that there is data skew, processing the data skew according to the statistical results. In the above technical solution, when it is determined that there is data skew based on the statistical results statistically obtained by the data consumer, the data skew can be automatically processed according to the statistical results. The process of processing this data skew does not require manual intervention. This method for processing data skew greatly improves the processing efficiency compared with the existing method of manually processing data skew, and quickly solves the adverse effects brought by data skew to Spark Streaming.

[0059] In one implementation, the statistical results include the total count of each skewed key and the execution duration of at least one skewed task for performing the statistics. Among them, the total count of each skewed key can be understood as the total number of times the data corresponding to each skewed key appears in the batch data. During the statistical process, the batch data needs to be divided into at least one skewed task to facilitate the execution of statistics. Also, the execution duration of the skewed task is also a criterion for measuring whether there is data skew. Then, in addition to the total count of each skewed key, the above statistical results also include the execution duration of at least one skewed task. Based on the above description, Figure 1 on the basis of Figure 2 Figure 2 is a flowchart of an embodiment of another method for processing data skew provided by the embodiments of the present invention. Figure 2 The process shown may include the following steps:

[0060] Step 201, receiving in real time the total count of each skewed key and the execution duration of at least one skewed task for performing the statistics of each batch of data by the data consumer;

[0061] Step 202, determining whether there is data skew based on the total count of each skewed key and the execution duration of at least one skewed task;

[0062] The above step 202 can be implemented through steps A1 to A5:

[0063] ​Step A1, detect whether the execution duration of at least one skewed task exceeds a preset duration threshold;

[0064] That is, detect whether the execution duration of each skewed task in at least one skewed task exceeds the preset duration threshold. The preset duration threshold is the duration indicating data skew. In Spark, skewed tasks are executed in parallel, and each skewed task processes a part of the data. When the data is evenly distributed, the amount of data processed by each skewed task is roughly equal, so their execution durations are also similar. However, when the data is unevenly distributed, data skew occurs, that is, the amount of data processed by some skewed tasks is much larger than that of other tasks, and the required execution duration will become longer. Therefore, by comparing the execution duration with the preset duration threshold, it can be accurately determined whether there is data skew.

[0065] Among them, the preset duration threshold can be set according to actual needs and is not limited here.

[0066] Step A2, when it is detected that the execution duration of at least one skewed task exceeds the preset duration threshold, determine that there is data skew;

[0067] In this embodiment, as long as there is at least one skewed task whose execution duration exceeds the preset duration threshold, that is, one or more skewed tasks' execution durations exceed the preset duration threshold, it can be determined that there is data skew.

[0068] Step A3, when it is detected that the execution duration of each skewed task in at least one skewed task does not exceed the preset duration threshold, calculate the statistical difference of each skewed key based on the total statistical count of each skewed key;

[0069] Since the execution duration of a skewed task is only a possible manifestation of data skew, and when it is impossible to determine whether there is data skew through the execution duration of the skewed task, it is also necessary to further determine based on the statistical difference of each skewed key calculated from the total statistical count of each skewed key. This is because the statistical difference of the skewed key can directly reflect the imbalance of data distribution, so as to more accurately judge whether there is data skew.

[0070] Step A4, detect whether the statistical difference of each skewed key in multiple skewed keys exceeds a preset difference threshold;

[0071] In one embodiment, the statistical difference includes the mean and variance. Therefore, it is necessary to use the mean and variance of each skewed key calculated from the total statistical count of each skewed key to determine whether there is data skew, that is, to detect whether the mean of each skewed key among multiple skewed keys exceeds a preset mean threshold and whether the variance exceeds a preset variance threshold.

[0072] Step A5, when it is detected that the statistical difference of at least one skewed key among multiple skewed keys exceeds a preset difference threshold, it is determined that there is data skew.

[0073] In this embodiment, when the statistical difference of one or more skewed keys among multiple skewed keys exceeds the preset difference threshold, that is, when the mean of one or more skewed keys exceeds the preset mean threshold and the variance exceeds the preset variance threshold, it is determined that there is data skew.

[0074] Step 203, when it is determined that there is data skew, based on the total statistical count of each skewed key and the execution duration of at least one skewed task, data skew is processed.

[0075] The process of specifically processing data skew based on the total statistical count of each skewed key and the execution duration of at least one skewed task can be implemented through steps B1 to B5:

[0076] Step B1, determine whether the number of target skewed tasks whose execution duration exceeds a preset duration threshold is greater than or equal to a preset number;

[0077] In this embodiment, the skewed task whose execution duration exceeds the preset duration threshold is used as the target skewed task, and the specific method for processing data skew is determined by determining whether the number of target skewed tasks is greater than or equal to the preset number. Among them, the preset number can be set to 1, that is, when it is determined that the number of target skewed tasks is greater than or equal to 1, steps B2 - B3 are executed to process data skew by increasing the number of partitions; while when it is determined that the number of target skewed tasks is less than the preset number, steps B4 - B5 are executed to process data skew by data aggregation.

[0078] Step B2, determine the number of partition increments based on the execution duration of the target skewed task;

[0079] The number of partition increments can be understood as the number of partitions that need to be increased. In this embodiment, the process of determining the above number of partition increments can be implemented through steps C1 to C4:

[0080] Step C1, when the number of target skew tasks is greater than a preset number, obtain the maximum execution duration and the minimum execution duration from the execution durations of multiple target skew tasks;

[0081] When the number of target skew tasks is greater than the preset number, it indicates that there are multiple target skew tasks, so as to obtain the maximum execution duration and the minimum execution duration from the execution durations of multiple target skew tasks. For example, if the number of target skew tasks is 3, and the corresponding execution durations are 18 seconds, 20 seconds, and 22 seconds respectively, then the obtained maximum execution duration is 22 seconds, and the minimum execution duration is 18 seconds.

[0082] Step C2, determine the execution duration difference based on the maximum execution duration and the minimum execution duration;

[0083] Use the difference between the maximum execution duration and the minimum execution duration as the execution duration difference. Continuing the previous example, the execution duration difference = 22 seconds - 18 seconds = 4 seconds.

[0084] Step C3, query the partition increment number corresponding to the execution duration difference from the partition increment query table;

[0085] Since the partition increment query table pre-stores the corresponding relationship between the execution duration difference and the partition increment number, therefore, by querying the partition increment query table, the partition increment number corresponding to the execution duration difference can be accurately determined. For the sake of understanding, as shown in Table 1:

[0086] Table 1

[0087]

[0088]

[0089] It should be noted that Table 1 only shows an example of the corresponding relationship between the execution duration difference and the partition increment number. The specific corresponding relationship between the execution duration difference and the partition increment number can be set according to actual needs and is not limited here.

[0090] Continuing the previous example, the execution duration difference obtained in Step C2 is 4 seconds. By querying the partition increment query table, it can be known that the partition increment number corresponding to the execution duration difference of 4 seconds is 4.

[0091] Step C4, when the number of target skew tasks is equal to the preset number, determine the preset partition number as the partition increment number.

[0092] When the number of target skew tasks is equal to the preset number, it indicates that there is only 1 target skew task. Then the preset partition number can be determined as the partition increment number. For example, if the preset partition number is 2, then the partition increment number determined by the number of target skew tasks is 2.

[0093] Step B3: Handle data skew by increasing the number of partitions incrementally.

[0094] In parallel computing, data skew occurs when the data volume on a certain node far exceeds that of other nodes, resulting in reduced computing efficiency or memory overflow. By increasing the number of partitions incrementally, the number of records in each partition can be made as equal as possible, thereby reducing the computing pressure on a single node and improving computing efficiency. In addition, increasing the number of partitions incrementally can also help better utilize cluster resources and avoid situations where some nodes are overloaded while others are idle, thus effectively alleviating data skew.

[0095] Step B4: Obtain the target skew keys whose statistical differences exceed the preset difference threshold.

[0096] In this embodiment, the skew keys whose statistical differences exceed the preset difference threshold are determined as the target skew keys. It can be understood that the data of the target skew keys is the cause data of data skew.

[0097] Step B5: Handle data skew by performing data preprocessing on the data of the target skew keys.

[0098] In this embodiment, by performing data preprocessing such as repartitioning and data scattering on the data of the target skew keys, the data can be distributed more evenly, thereby avoiding the phenomenon that individual skew tasks execute too slowly or memory overflow, and thus alleviating or eliminating the problem of data skew.

[0099] Through the above method, the data of the original target skew keys has been effectively processed and dispersed, enabling each node to process data relatively evenly, thereby improving the overall computing efficiency and avoiding the problems brought by data skew. Therefore, performing data preprocessing on the data of the target skew keys is an effective and commonly used method to solve the problem of data skew.

[0100] In this embodiment, in addition to the above-mentioned method for data skew handling, a machine learning method can also be used for data skew handling. Specifically, the total statistics of each skew key and the execution duration of at least one skew task are input into the data skew handling strategy model, and the data skew handling strategy model outputs the data skew handling strategy; based on the data skew handling strategy, data skew is handled.

[0101] Among them, the data skew processing strategy is the strategy of increasing the number of partitions or the data preprocessing strategy. The strategy of increasing the number of partitions includes the number of partition increments, and the data preprocessing includes the target skew key. It can be understood that through the data skew processing strategy model, the specific processing method for data skew, that is, the data skew processing strategy, can be output, so as to adopt the corresponding data skew processing strategy to process data skew. Specifically, the data skew processing can be realized by increasing the number of partition increments, or by preprocessing the data of the target skew key to realize data skew processing. For specific details, please refer to the above description and will not be elaborated here.

[0102] The above data skew processing strategy model is obtained by training a classification model using the statistical total of each skew key and the execution duration of the skew task. The specific training process is as follows:

[0103] First, obtain the training data. Among them, the training data includes the statistical total of each skew key, the execution duration of the skew task, and the corresponding label information, and this label information is the data skew processing strategy.

[0104] Then, input the above training data into the classification model to obtain the predicted data skew processing strategy output by the classification model.

[0105] After that, determine whether the classification model converges through the data skew processing strategy labeled by the training data and the predicted data skew processing strategy. If the classification model converges, determine the trained classification model as the data skew processing strategy model. If the classification model does not converge, adjust the model parameters of the classification model and continue to perform model training.

[0106] Among them, the above classification model can be a decision tree, a random forest, a logistic regression, a support vector machine, a neural network model, etc., and is not limited here.

[0107] Through the method provided in this embodiment, by adjusting the number of partitions and the data preprocessing strategy, it is ensured that the data is evenly distributed in the system, the data skew is reduced, the system performance and processing efficiency are improved, the data skew situation is detected and processed in real time, single-point overload is avoided, and the stable and reliable operation of the system is guaranteed. Especially in the case of high load, the automatic detection and adjustment mechanism reduces the need for manual intervention and reduces the complexity and cost of system operation and maintenance.

[0108] An embodiment of the present invention provides a device for data skew processing. This device is applied to the data production end. Refer to Figure 3 , which is the block diagram of an embodiment of a device for data skew processing provided by an embodiment of the present invention. As Figure 3 shown, this device includes:

[0109] A receiving module 301, configured to receive in real time the statistical results of data consumers for each batch of data; wherein the statistical results are used to represent the index data of data skew.

[0110] A determination module 302, configured to determine whether there is data skew based on the statistical results.

[0111] A data skew processing module 303, configured to, when it is determined based on the statistical results that there is data skew, process the data skew according to the statistical results.

[0112] The apparatus for processing data skew provided by an embodiment of the present invention includes: receiving in real time the statistical results of data consumers for each batch of data; wherein the statistical results are used to represent the index data of data skew; determining whether there is data skew based on the statistical results; and when it is determined based on the statistical results that there is data skew, processing the data skew according to the statistical results. In the above technical solution, when it is determined that there is data skew based on the statistical results counted by the data consumers, the data skew can be automatically processed according to the statistical results. The processing process of the data skew does not require manual intervention. This method for processing data skew greatly improves the processing efficiency compared with the existing method for manually processing data skew, and quickly solves the adverse effects brought by data skew to Spark Streaming.

[0113] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Figure 4 The illustrated electronic device 500 includes: at least one processor 501, a memory 502, at least one network interface 504, and other user interfaces 503. Each component in the electronic device 500 is coupled together through a bus system 505. It can be understood that the bus system 505 is used to realize the connection and communication between these components. The bus system 505 includes, in addition to a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 4 all kinds of buses are labeled as the bus system 505.

[0114] Among them, the user interface 503 may include a display, a keyboard, or a pointing device (such as a mouse, a trackball, a touchpad, or a touch screen, etc.).

[0115] It can be understood that the memory 502 in the embodiments of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The memory 502 described herein is intended to include but not be limited to these and any other suitable types of memory.

[0116] In some embodiments, the memory 502 stores the following elements, executable units or data structures, or subsets thereof, or extended sets thereof: the operating system 5021 and the application program 5022.

[0117] Among them, the operating system 5021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., and is used to implement various basic services and process hardware-based tasks. The application program 5022 includes various application programs, such as a media player and a browser, etc., and is used to implement various application services. The program for implementing the method of the embodiments of the present invention can be included in the application program 5022.

[0118] In the embodiments of the present invention, by calling the programs or instructions stored in the memory 502, specifically, the programs or instructions stored in the application program 5022, the processor 501 is used to execute the method steps provided by the various method embodiments, for example, including:

[0119] Receiving in real time the statistical results of the data consumption end for each batch of data; wherein, the statistical results are used to represent the index data of data skew;

[0120] Determining whether there is data skew based on the statistical results;

[0121] In the case where it is determined based on the statistical results that there is data skew, processing the data skew according to the statistical results.

[0122] In a possible implementation manner, the statistical results include the total statistical count of each skew key and the execution duration of at least one skew task for which the statistics are performed;

[0123] Determining whether there is data skew based on the statistical results includes:

[0124] Determining whether there is data skew based on the total statistical count of each skew key and the execution duration of at least one skew task.

[0125] In a possible implementation manner, determining whether there is data skew based on the total statistical count of each skew key and the execution duration of at least one skew task includes:

[0126] Detecting whether the execution duration of at least one skew task exceeds a preset duration threshold;

[0127] In the case where it is detected that the execution duration of at least one skew task exceeds the preset duration threshold, determining that there is data skew;

[0128] In the case where it is detected that the execution duration of each skew task in at least one skew task does not exceed the preset duration threshold, calculating the statistical difference of each skew key based on the total statistical count of each skew key;

[0129] Detecting whether the statistical difference of each skew key among multiple skew keys exceeds a preset difference threshold;

[0130] In the case where it is detected that the statistical difference of at least one skew key among multiple skew keys exceeds the preset difference threshold, determining that there is data skew.

[0131] In a possible implementation manner, calculating the statistical difference of each skew key based on the total statistical count of each skew key includes:

[0132] Calculating the mean and variance of each skew key based on the total statistical count of each skew key;

[0133] Detecting whether the statistical difference of each skew key among multiple skew keys exceeds the preset difference threshold includes:

[0134] Detect whether the mean value of each of multiple skewed keys exceeds a preset mean threshold and whether the variance exceeds a preset variance threshold.

[0135] In a possible implementation, processing data skew is achieved according to statistical results, including:

[0136] Processing data skew is achieved based on the total statistics of each skewed key and the execution duration of at least one skewed task.

[0137] In a possible implementation, processing data skew is achieved based on the total statistics of each skewed key and the execution duration of at least one skewed task, including:

[0138] Determine whether the number of target skewed tasks whose execution duration exceeds a preset duration threshold is greater than or equal to a preset number;

[0139] In the case where it is determined that the number of target skewed tasks is greater than or equal to the preset number, determine the number of partition increments based on the execution duration of the target skewed tasks;

[0140] Achieve processing of data skew by increasing the number of partition increments;

[0141] In the case where it is determined that the number of target skewed tasks is less than the preset number, obtain the target skewed keys whose statistical difference exceeds a preset difference threshold;

[0142] Achieve processing of data skew by aggregating the data of the target skewed keys.

[0143] In a possible implementation, determining the number of partition increments based on the execution duration of the target skewed tasks includes:

[0144] In the case where the number of target skewed tasks is greater than the preset number, obtain the maximum execution duration and the minimum execution duration from the execution durations of multiple target skewed tasks;

[0145] Determine the execution duration difference based on the maximum execution duration and the minimum execution duration;

[0146] Query the number of partition increments corresponding to the execution duration difference from the partition increment query table; wherein, the partition increment query table pre-stores the corresponding relationship between the execution duration difference and the number of partition increments;

[0147] In the case where the number of target skewed tasks is equal to the preset number, determine the preset number of partitions as the number of partition increments.

[0148] In a possible implementation, processing data skew is achieved according to statistical results, including:

[0149] Input the total count of each skewed key and the execution duration of at least one skewed task into the data skew handling policy model, and the data skew handling policy model outputs a data skew handling policy; wherein, the data skew handling policy is a policy of increasing the number of partitions or a preprocessing aggregation policy. The policy of increasing the number of partitions includes the number of increased partitions, and the preprocessing aggregation policy includes the target skewed key. The data skew handling policy model is obtained by training a classification model using the total count of each skewed key and the execution duration of the skewed task.

[0150] Implement data skew handling based on the data skew handling policy.

[0151] The method disclosed in the above embodiments of the present invention can be applied to or implemented by the processor 501. The processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 501 or the instructions in software form. The above processor 501 may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software unit may be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory 502, and the processor 501 reads the information in the memory 502 and combines its hardware to complete the steps of the above method.

[0152] It will be appreciated that the embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or any combination thereof. For hardware implementation, the processing unit may be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or any combination thereof.

[0153] For software implementation, the technologies described herein may be implemented by units that execute the functions described herein. The software code may be stored in a memory and executed by a processor. The memory may be implemented within the processor or externally to the processor.

[0154] The electronic device provided in this embodiment may be an electronic device as shown in Figure 4 and may execute all steps of the method for data skew processing as shown in Figure 1-2 so as to achieve the technical effects of the method for data skew processing as shown in Figure 1-2 For specific details, please refer to the relevant description in Figure 1-2 For the sake of brevity, it will not be elaborated herein.

[0155] The embodiment of the present invention also provides a storage medium (computer-readable storage medium). The storage medium stores one or more programs. Among them, the storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; the memory may also include a combination of the above types of memory.

[0156] When one or more programs in the storage medium can be executed by one or more processors to implement the above method for data skew processing.

[0157] The processor is used to execute the program for data skew processing stored in the memory to implement the following steps of the method for data skew processing:

[0158] Real-time receive the statistical results of the data consumer end for each batch of data; wherein, the statistical results are used to represent the index data of data skew;

[0159] Determine whether there is data skew based on the statistical results;

[0160] In the case where data skew is determined based on statistical results, handle the data skew according to the statistical results.

[0161] In a possible implementation, the statistical results include the total count of each skewed key and the execution duration of at least one skewed task for which the statistics are performed;

[0162] Determining whether there is data skew based on the statistical results includes:

[0163] Determining whether there is data skew based on the total count of each skewed key and the execution duration of at least one skewed task.

[0164] In a possible implementation, determining whether there is data skew based on the total count of each skewed key and the execution duration of at least one skewed task includes:

[0165] Detect whether the execution duration of at least one skewed task exceeds a preset duration threshold;

[0166] In the case where it is detected that the execution duration of at least one skewed task exceeds the preset duration threshold, determine that there is data skew;

[0167] In the case where it is detected that the execution duration of each skewed task in at least one skewed task does not exceed the preset duration threshold, calculate the statistical difference of each skewed key based on the total count of each skewed key;

[0168] Detect whether the statistical difference of each skewed key among multiple skewed keys exceeds a preset difference threshold;

[0169] In the case where it is detected that the statistical difference of at least one skewed key among multiple skewed keys exceeds the preset difference threshold, determine that there is data skew.

[0170] In a possible implementation, calculating the statistical difference of each skewed key based on the total count of each skewed key includes:

[0171] Calculate the mean and variance of each skewed key based on the total count of each skewed key;

[0172] Detecting whether the statistical difference of each skewed key among multiple skewed keys exceeds a preset difference threshold includes:

[0173] Detect whether the mean of each skewed key among multiple skewed keys exceeds a preset mean threshold and whether the variance exceeds a preset variance threshold.

[0174] In a possible implementation, handling the data skew according to the statistical results includes:

[0175] Based on the total count of each skewed key and the execution duration of at least one skewed task, data skew is processed.

[0176] In a possible implementation, processing data skew based on the total count of each skewed key and the execution duration of at least one skewed task includes:

[0177] Determine whether the number of target skewed tasks whose execution duration exceeds a preset duration threshold is greater than or equal to a preset number;

[0178] When it is determined that the number of target skewed tasks is greater than or equal to the preset number, determine the number of additional partitions based on the execution duration of the target skewed tasks;

[0179] Process data skew by increasing the number of additional partitions;

[0180] When it is determined that the number of target skewed tasks is less than the preset number, obtain the target skewed keys whose statistical difference exceeds a preset difference threshold;

[0181] Process data skew by aggregating the data of the target skewed keys.

[0182] In a possible implementation, determining the number of additional partitions based on the execution duration of the target skewed tasks includes:

[0183] When the number of target skewed tasks is greater than the preset number, obtain the maximum execution duration and the minimum execution duration from the execution durations of multiple target skewed tasks;

[0184] Determine the execution duration difference based on the maximum execution duration and the minimum execution duration;

[0185] Query the number of additional partitions corresponding to the execution duration difference from the additional partition number query table; wherein, the additional partition number query table pre-stores the corresponding relationship between the execution duration difference and the number of additional partitions;

[0186] When the number of target skewed tasks is equal to the preset number, determine the preset number of partitions as the number of additional partitions.

[0187] In a possible implementation, processing data skew according to the statistical results includes:

[0188] Input the total count of each skewed key and the execution duration of at least one skewed task into the data skew processing strategy model, and the data skew processing strategy model outputs a data skew processing strategy; wherein, the data skew processing strategy is a strategy of increasing the number of partitions or a preprocessing aggregation strategy, the strategy of increasing the number of partitions includes the number of partition increments, and the preprocessing aggregation strategy includes the target skewed key. The data skew processing strategy model is obtained by training a classification model using the total count of each skewed key and the execution duration of the skewed task.

[0189] Implement the processing of data skew based on the data skew processing strategy.

[0190] Those skilled in the art should also be able to further realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, they can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0191] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0192] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for processing data skew, characterized in that: The method is applied to a data production end, and the method includes: Receive in real time the statistical results of each batch of data collected by the data consumer; wherein the statistical results are used to characterize the indicator data of data skew; Determining whether there is data skew based on the statistical results; When it is determined based on the statistical result that data skew exists, the data skew is processed according to the statistical result.

2. The method according to claim 1, characterized in that The statistical result includes the total number of each tilted key and the execution time of at least one tilted task for which the statistics are executed; The determining whether there is data skew based on the statistical result includes: Whether data skew exists is determined based on the total number of statistics of each of the skewed keys and the execution time of at least one of the skewed tasks.

3. The method according to claim 2, characterized in that The determining whether there is data skew based on the total number of statistics of each of the skewed keys and the execution time of at least one of the skewed tasks includes: Detecting whether the execution time of at least one of the tilt tasks exceeds a preset time threshold; When it is detected that the execution time of at least one of the tilted tasks exceeds the preset time threshold, determining that data tilt exists; When it is detected that the execution time of each of the at least one tilted task does not exceed the preset time threshold, calculating the statistical difference of each tilted key based on the statistical total of each tilted key; Detecting whether a statistical difference of at least one of the plurality of tilt keys exceeds a preset difference threshold; When it is detected that the statistical difference of each of the tilted keys in the plurality of tilted keys exceeds the preset difference threshold, it is determined that data skew exists.

4. The method according to claim 3, characterized in that The calculating the statistical difference of each tilted key based on the statistical total of each tilted key includes: Calculate the mean and variance of each tilt key based on the statistical total of each tilt key; The detecting whether the statistical difference of each of the tilted keys in the plurality of tilted keys exceeds a preset difference threshold comprises: It is detected whether the mean of each of the tilted keys exceeds a preset mean threshold, and whether the variance exceeds a preset variance threshold.

5. The method according to claim 3, characterized in that: The processing of data skew according to the statistical results includes: Data skew is processed based on the total number of statistics of each of the skewed keys and the execution time of at least one of the skewed tasks.

6. The method according to claim 5, characterized in that The processing of data skew based on the total number of statistics of each of the skewed keys and the execution time of at least one of the skewed tasks includes: Determine whether the number of the target tilt tasks whose execution duration exceeds the preset duration threshold is greater than or equal to a preset number; In the case where it is determined that the number of the target tilted tasks is greater than or equal to the preset number, determining the partition increment number based on the execution time of the target tilted tasks; Processing data skew by increasing the partition increment number; When it is determined that the number of the target tilt tasks is less than the preset number, obtaining a target tilt key whose statistical difference exceeds the preset difference threshold; The data skew is processed by preprocessing the data of the target skew key.

7. The method according to claim 6, characterized in that The determining of the partition increment number based on the execution time of the target tilted task includes: When the number of the target tilt tasks is greater than the preset number, obtaining a maximum execution time and a minimum execution time from the execution times of the plurality of target tilt tasks; Determine an execution time difference based on the maximum execution time and the minimum execution time; Querying the partition increment corresponding to the execution time difference from the partition increment query table; wherein the partition increment query table pre-stores the corresponding relationship between the execution time difference and the partition increment; When the number of the target tilt tasks is equal to the preset number, the preset partition number is determined as the partition increment number.

8. The method according to claim 6, characterized in that The processing of data skew according to the statistical results includes: The total number of statistics of each tilted key and the execution time of at least one tilted task are input into a data tilt processing strategy model, and the data tilt processing strategy model outputs a data tilt processing strategy; wherein the data tilt processing strategy is a partition number increase strategy or a data preprocessing strategy, the partition number increase strategy includes the partition increase number, the data preprocessing strategy includes the target tilted key, and the data tilt processing strategy model is obtained by training a classification model using the total number of statistics of each tilted key and the execution time of the tilted task; The data skew is processed based on the data skew processing strategy.

9. A device for processing data skew, characterized in that: The device is applied to a data production end, and the device comprises: A receiving module, used to receive in real time the statistical results of each batch of data collected by the data consumer; wherein the statistical results are used to characterize the indicator data of data skew; A determination module, used for determining whether there is data skew based on the statistical results; The data skew processing module is used to process the data skew according to the statistical results when it is determined that there is data skew based on the statistical results.

10. An electronic device, characterized in that: include: A processor and a memory, wherein the processor is used to execute a data tilt processing program stored in the memory to implement the data tilt processing method according to any one of claims 1 to 8.

11. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the data tilt processing method according to any one of claims 1 to 8.