Data processing task running performance optimization method and device, equipment and medium
By obtaining operational performance indicator data of multiple categories and using heuristic algorithms for analysis and weighted summing, the problems of narrow evaluation dimensions and unreasonable resource allocation in the existing technology are solved, and comprehensive evaluation and optimization of data processing tasks are achieved.
Patent Information
- Application Number
- CN202510255646.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-11
AI Technical Summary
When evaluating and optimizing the system performance of data processing tasks, the existing technology relies on a single indicator or subjective experience, resulting in narrow evaluation dimensions, fuzzy problem positioning, difficulty in adapting to heterogeneous performance data, and unreasonable resource allocation, which cannot meet the business needs of high concurrency and low latency.
By obtaining operational performance index data of multiple categories, using different heuristic algorithms for analysis, calculating the severity scores of each category of indicators, and performing weighted summations, determining the comprehensive severity scores and optimization levels, and combining with the pre-constructed optimization strategy table for performance optimization.
It realizes a comprehensive evaluation and targeted analysis of data processing tasks, accurately identify performance problems, quantify overall performance risks, clarify optimization priorities, and reasonably allocate resources, which improves optimization efficiency.
Smart Images

Figure CN120295873A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method for optimizing the running performance of data processing tasks, a device for optimizing the running performance of data processing tasks, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the rapid growth of the scale and complexity of data processing tasks, how to efficiently evaluate and optimize system performance has become a key challenge. In traditional methods, performance management mostly relies on single indicators or empirical judgments, resulting in narrow evaluation dimensions, vague problem positioning, and difficulty in quantifying the comprehensive impact of multiple types of performance problems. Especially in scenarios such as distributed computing and real-time stream processing, one-sided evaluation is likely to cover up potential performance bottlenecks, lacking a systematic analysis framework, leading to scattered optimization resources and low decision-making efficiency, and it is difficult to meet the business requirements of high concurrency and low latency.
[0003] Currently, multi-dimensional indicators are collected through various monitoring tools (such as performance counters, log analysis systems), and combined with general algorithms (such as statistical analysis, rule engines) for problem detection. For example, some solutions use the weighted scoring method to aggregate performance indicators, or use machine learning models to predict anomalies. In addition, some research introduces a priority queue mechanism, based on artificial experience or simple thresholds to divide the optimization task levels to guide resource allocation.
[0004] However, the existing methods for diagnosing the running performance of tasks have the following problems: (1) relying on single indicators or general algorithms, it is difficult to adapt to the characteristics of heterogeneous performance data, resulting in missed detection or misjudgment of hidden problems; (2) the quantification method lacks dynamic adaptation to the weight differences of multi-dimensional problems and cannot accurately reflect the overall performance risk; (3) the priority division is mostly based on subjective experience or static rules, lacking an objective and unified quantification basis, which is likely to cause resource waste. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a method, device, equipment, and medium for optimizing the running performance of data processing tasks to solve the above problems.
[0006] To achieve the above purpose, the embodiments of the present invention provide a method for optimizing the running performance of data processing tasks, including:
[0007] Obtain the running performance index data of the data processing task, and determine the category of each running performance index data;
[0008] Use different heuristic algorithms to respectively perform running performance analysis on different categories of running performance index data of the data processing task, and obtain the severity scores of different running performance problems of the data processing task; wherein, different categories of running performance index data correspond to different heuristic algorithms one by one;
[0009] Weightedly sum the severity scores of different running performance problems of the data processing task to obtain the severity score of the comprehensive running performance problem of the data processing task;
[0010] Determine the level to be optimized for the data processing task by classifying the severity score of the comprehensive running performance problem of the data processing task;
[0011] If the level to be optimized for the data processing task is greater than the preset level to be optimized, then match the different running performance problems of the data processing task with a pre-constructed optimization strategy table to obtain the optimization strategies for the different running performance problems of the data processing task, so that the user can optimize the performance of the data processing task according to the optimization strategies for the different running performance problems of the data processing task; wherein, the optimization strategy table is used to represent the mapping relationship between different running performance problems of the data processing task and the corresponding optimization strategies.
[0012] Optionally, the heuristic algorithms include: data skew heuristic algorithm and Java GC heuristic algorithm;
[0013] The categories of running performance metric data include: data skew metric data and Java application performance metric data;
[0014] The data skew metric data includes: the number of subtasks and the memory usage of each subtask;
[0015] The Java application performance metric data includes: average CPU usage time, average running time, and average garbage collection consumption time.
[0016] Optionally, use different heuristic algorithms to perform running performance analysis on different categories of running performance metric data of the data processing task to obtain the severity scores of different running performance problems of the data processing task, including:
[0017] Based on the data skew heuristic algorithm, calculate the number of subtasks and the memory usage of each subtask of the data processing task to obtain the severity score of the data skew problem of the data processing task;
[0018] Based on the Java GC heuristic algorithm, calculate the average CPU usage time, average running time, and average garbage collection consumption time of the data processing task to obtain the severity score of the Java GC problem of the data processing task.
[0019] Optionally, based on the data skew heuristic algorithm, calculate the number of subtasks and the memory usage of each subtask of the data processing task to obtain the severity score of the data skew problem of the data processing task, including:
[0020] Divide multiple subtasks in a data processing task into two equal parts to obtain two subtask clusters;
[0021] Calculate the average value of the memory usage of the two subtask clusters respectively;
[0022] If there is a subtask cluster with the largest average value of memory usage and the smallest number of subtasks, then use this subtask cluster as the first subtask cluster, and determine the task cluster other than the first subtask cluster as the second subtask cluster;
[0023] Calculate the difference between the average value of the memory usage of the second subtask cluster and the average value of the memory usage of the first subtask cluster, and calculate the ratio of the difference to the average value of the memory usage of the first subtask cluster to obtain the error value of the memory usage of the data processing task;
[0024] Calculate the ratio of the average value of the memory usage of the second subtask cluster to the preset memory usage to obtain the ratio of the memory usage of the second subtask cluster;
[0025] Match the error value of the memory usage of the data processing task, the ratio of the memory usage of the second subtask cluster, and the number of subtasks of the second subtask cluster with a pre-constructed data skew severity score value table respectively to obtain the error severity score, memory usage severity score, and subtask number severity score of the data processing task; wherein, the data skew severity score value table is used to represent the mapping relationship between the error value of the memory usage of the data processing task, the ratio of the memory usage of the second subtask cluster, and the number of subtasks of the second subtask cluster and the corresponding severity scores;
[0026] Determine the minimum value among the error severity score, memory usage severity score, and subtask number severity score of the data processing task as the data skew problem severity score of the data processing task.
[0027] Optionally, based on the Java GC heuristic algorithm, calculate the average CPU usage time, average running time, and average garbage collection consumption time of the data processing task to obtain the Java GC problem severity score of the data processing task, including:
[0028] Calculate the ratio of the average garbage collection consumption time and the average CPU usage time of the data processing task to obtain the GC-CPU ratio of the data processing task;
[0029] Match the GC-CPU ratio and average running time of the data processing task with the pre-constructed Java application performance severity score value table respectively to obtain the GC-CPU ratio severity score and average running time severity score of the data processing task; wherein, the Java application performance severity score value table is used to represent the mapping relationship between the GC-CPU ratio of the data processing task and the average running time of the data processing task and the corresponding severity scores.
[0030] Determine that the minimum value of the GC-CPU ratio severity score and the average running time severity score of the data processing task is the Java GC problem severity score of the data processing task.
[0031] Optionally, perform weighted summation on the severity scores of different running performance problems of the data processing task to obtain the severity score of the comprehensive running performance problem of the data processing task, including:
[0032] Use the following formula to calculate the data skew problem severity score of the data processing task, the preset weight value of the data skew problem severity score of the data processing task, the Java GC problem severity score of the data processing task, and the preset weight value of the Java GC problem severity score of the data processing task to obtain the severity score of the comprehensive running performance problem of the data processing task;
[0033] S = W1 * R1 + W2 * R2; where S represents the severity score of the comprehensive running performance problem of the data processing task, W1 represents the preset weight value of the data skew problem severity score of the data processing task, R1 represents the data skew problem severity score of the data processing task, W2 represents the preset weight value of the Java GC problem severity score of the data processing task, and R2 represents the Java GC problem severity score of the data processing task.
[0034] Optionally, delimit the level of the severity score of the comprehensive running performance problem of the data processing task to determine the level to be optimized for the data processing task, including:
[0035] If the severity score of the comprehensive running performance problem of the data processing task is greater than the first preset severity score and less than the product of the second preset severity score and the first preset percentage, determine that the level to be optimized for the data processing task is the first level to be optimized;
[0036] If the severity score of the comprehensive running performance problem of the data processing task is greater than the product of the second preset severity score and the first preset percentage and less than the product of the second preset severity score and the second preset percentage, determine that the level to be optimized for the data processing task is the second level to be optimized;
[0037] If the severity score of the comprehensive operation performance problem of the data processing task is greater than the product of the second preset severity score and the second preset percentage and less than the product of the second preset severity score and the third preset percentage, determine that the optimization level to be optimized for the data processing task is the third optimization level;
[0038] If the severity score of the comprehensive operation performance problem of the data processing task is greater than the product of the second preset severity score and the third preset percentage, determine that the optimization level to be optimized for the data processing task is the fourth optimization level;
[0039] Among them, the first preset severity score < the second preset severity score;
[0040] The first preset percentage < the second preset percentage < the third preset percentage.
[0041] In the second aspect of the embodiment of the present invention, a data processing task operation performance optimization device is provided, including:
[0042] A data acquisition module, configured to acquire operation performance index data of a data processing task and determine the category of each operation performance index data;
[0043] A performance analysis module, configured to use different heuristic algorithms to perform operation performance analysis on different categories of operation performance index data of the data processing task respectively, and obtain the severity scores of different operation performance problems of the data processing task; among them, different categories of operation performance index data correspond to different heuristic algorithms one by one;
[0044] A weighted average module, configured to perform weighted summation on the severity scores of different operation performance problems of the data processing task to obtain the severity score of the comprehensive operation performance problem of the data processing task;
[0045] A level determination module, configured to determine the optimization level to be optimized for the data processing task by determining the level of the severity score of the comprehensive operation performance problem of the data processing task;
[0046] A strategy execution module, configured to, if the optimization level to be optimized for the data processing task is greater than the preset optimization level, match different operation performance problems of the data processing task with a pre-constructed optimization strategy table to obtain optimization strategies for different operation performance problems of the data processing task, so that the user can perform performance optimization on the data processing task according to the optimization strategies for different operation performance problems of the data processing task; among them, the optimization strategy table is used to represent the mapping relationship between different operation performance problems of the data processing task and the corresponding optimization strategies.
[0047] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and when the machine-readable instructions are executed by the processor, the above-mentioned method for optimizing the running performance of data processing tasks is executed.
[0048] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, storing computer instructions, and when the computer instructions run on a computer, the computer is enabled to execute the above-mentioned method for optimizing the running performance of data processing tasks.
[0049] Advantages of the present invention:
[0050] (1) Comprehensive performance evaluation: By obtaining operation performance index data of multiple categories, it is possible to comprehensively understand the operation status of data processing tasks from multiple dimensions, avoid the one-sidedness of single-index evaluation, and thus more accurately grasp the performance performance of tasks in different aspects.
[0051] (2) Targeted problem analysis: Using different heuristic algorithm models to process operation performance index data of different categories respectively, each algorithm model is optimized for specific types of data, can deeply dig out the operation performance problems hidden behind various types of data, and give the severity scores of the corresponding problems, improving the accuracy and pertinence of problem analysis.
[0052] (3) Quantitative comprehensive evaluation: Weighted summation of the severity scores of different operation performance problems is performed to obtain the severity score of the comprehensive operation performance problem, quantifying the complex performance problem into a specific value, facilitating the intuitive measurement of the severity of the overall performance problem of the task, and at the same time the weighted method can also reflect the differences in the impact of different types of problems on the overall performance.
[0053] (4) Clear optimization priority: The comprehensive severity score is graded to determine the optimization level to be determined, enabling users to clearly understand the performance optimization priority of data processing tasks. For tasks with a higher optimization level to be determined, they are processed first, and optimization resources are reasonably allocated.
[0054] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification, and are used to explain the embodiments of the present invention together with the following specific implementation manners, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0056] Figure 1 is a schematic flowchart of the method for optimizing the running performance of data processing tasks provided by the embodiments of the present invention;
[0057] Figure 2 It is a schematic structural diagram of a device for optimizing the running performance of data processing tasks provided by an embodiment of the present invention. Detailed implementation manners
[0058] The following will describe in detail the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only for explaining and illustrating the embodiments of the present invention, and are not used to limit the embodiments of the present invention.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments, and are not intended to limit this application.
[0060] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, "a plurality" means more than two, unless otherwise specifically defined.
[0061] Embodiment 1
[0062] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of a method for optimizing the running performance of data processing tasks provided by an embodiment of the present invention. The method includes the following steps:
[0063] S100, obtain the running performance index data of the data processing task, and determine the categories of each running performance index data;
[0064] It should be noted that a data processing task refers to a process of performing a series of operations (data collection, cleaning, transformation, analysis, and visualization) on raw data to convert it into valuable information, aiming to extract useful information from the data to support decision-making, discover patterns, or solve specific problems. However, during the process of data processing work, problems will be faced due to the data distribution characteristics and the memory management mechanism of a specific programming environment. Therefore, it is necessary to analyze the running performance of the data processing task, and the essence of analyzing the running performance of the data processing task is to analyze and calculate the corresponding running performance index data of the data processing task using different heuristic algorithms (that is: using the data skew heuristic algorithm to analyze and calculate the data skew index data of the data processing task, or using the Java GC heuristic algorithm to analyze and calculate the average CPU usage time, average running time, and average garbage collection consumption time of the data processing task).
[0065] In addition, the same heuristic algorithm is used for different types of data processing tasks to analyze and calculate the operation performance index data of the corresponding categories.
[0066] In the embodiments of the present invention, the types of data processing tasks, the types of heuristic algorithms, and the operation performance index data are not specifically limited, and can be reasonably selected according to actual needs.
[0067] In one embodiment, the types of data processing tasks include, but are not limited to: Spark tasks, Flink tasks, ETL tasks, data quality detection tasks, and script tasks.
[0068] A Spark task refers to a computing task that uses the open-source big data computing framework Apache Spark to process large-scale data.
[0069] A Flink task refers to a task that uses the open-source stream processing framework Apache Flink to process data.
[0070] An ETL task is an abbreviation for Extract, Transform, and Load. It refers to the process of extracting data from a source system, performing a series of transformation processes, and then loading it into a target data warehouse or data lake.
[0071] A data quality detection task refers to the process of evaluating and inspecting the quality of data. These tasks include checking dimensions such as the accuracy, integrity, consistency, timeliness, and uniqueness of the data.
[0072] A script task refers to an automated task written using a scripting language (such as Python, Shell, Perl, etc.).
[0073] In one embodiment, the categories of operation performance index data include, but are not limited to: data skew index data and Java application performance index data. Specifically, the data skew index data includes: the number of subtasks and the memory usage of each subtask. Specifically, the Java application performance index data includes: the average CPU usage time, the average running time, and the average garbage collection consumption time.
[0074] The average CPU usage time refers to the average value of the time spent by the CPU in processing a specific task or process within a specific time period. It reflects the busy degree and resource occupancy of the CPU when processing these tasks.
[0075] The average running time refers to the average value of the time taken for a program, task, or process to start execution and end normally. This index measures the time cost of task execution and reflects its execution efficiency.
[0076] The average garbage collection time refers to the average time taken by the garbage collector to automatically reclaim the memory space that is no longer in use in programming languages that adopt garbage collection mechanisms (such as Java and C#). Garbage collection is an important mechanism to ensure the efficient operation of a program and avoid memory leaks, and the average garbage collection time reflects the time cost of this mechanism.
[0077] S200, using different heuristic algorithms, respectively perform runtime performance analysis on different categories of runtime performance metric data of the data processing task, and obtain the severity scores of different runtime performance problems of the data processing task; among them, different categories of runtime performance metric data correspond one-to-one with different heuristic algorithms;
[0078] The heuristic algorithm for task performance analysis is a problem-solving method based on experience and heuristic rules, and is often used to handle NP-hard problems that are difficult to solve by traditional algorithms within polynomial time. Different from traditional deterministic algorithms, although it cannot give a completely accurate solution, it can provide a good approximate solution within a reasonable time. This algorithm utilizes the specific properties of the problem and heuristic rules, and gradually searches the solution space through a series of operations to approach the optimal solution. Its main features include: applying heuristic rules based on domain-specific knowledge or experience of the problem to guide the search and decision-making processes, helping to select appropriate paths and avoid invalid solutions, and accelerating the solution; not guaranteeing to find the optimal solution, but being able to obtain an approximate solution sufficient for practical problems, especially when computing resources are limited, within a reasonable time; having good scalability and being able to adapt to problems of different scales and complexities.
[0079] The heuristic algorithm model includes, but is not limited to: data skew heuristic algorithm and Java GC heuristic algorithm.
[0080] The data skew heuristic algorithm is an algorithm based on experience and insights into data characteristics, and is used to identify, mitigate or solve data skew problems. Such algorithms do not pursue the theoretically optimal solution, but rely on practical experience and understanding of the data and computing environment to find a feasible solution within a reasonable time.
[0081] The Java GC heuristic algorithm is an algorithm in which the garbage collector of the Java Virtual Machine (JVM) dynamically adjusts the garbage collection strategy based on various status information during system runtime, such as memory usage, object creation and destruction frequencies, etc. The purpose is to optimize the memory recycling efficiency as much as possible under different application scenarios and reduce the impact of garbage collection on the performance of the application program.
[0082] The data skew heuristic algorithm corresponds to the number of subtasks of the data processing task and the memory usage of each subtask.
[0083] The average CPU usage time, average running time, and average garbage collection consumption time of the Java GC heuristic algorithm corresponding to the data processing task.
[0084] In one embodiment, step S200 specifically includes:
[0085] S210, based on the data skew heuristic algorithm, calculate the number of subtasks of the data processing task and the memory usage of each subtask to obtain the severity score of the data skew problem of the data processing task;
[0086] Specifically, step S210 includes:
[0087] S211, divide the multiple subtasks in the data processing task into two equal parts to obtain two subtask clusters;
[0088] S212, calculate the average value of the memory usage of the two subtask clusters respectively;
[0089] S213, if there is a subtask cluster with the largest average value of memory usage and the smallest number of subtasks, then use this subtask cluster as the first subtask cluster, and determine the task cluster other than the first subtask cluster as the second subtask cluster;
[0090] S214, calculate the difference between the average value of the memory usage of the second subtask cluster and the average value of the memory usage of the first subtask cluster, and calculate the ratio of the difference to the average value of the memory usage of the first subtask cluster to obtain the error value of the memory usage of the data processing task;
[0091] Specifically, step S214 includes:
[0092] Use the following formula to calculate the average value of the memory usage of the first subtask cluster and the average value of the memory usage of the second subtask cluster to obtain the error value of the memory usage of the data processing task;
[0093] Among them, deviation represents the error value of the memory usage of the data processing task, avg(group_1) represents the average value of the memory usage of the first subtask cluster, and avg(group_2) represents the average value of the memory usage of the second subtask cluster.
[0094] S215, calculate the ratio of the average value of the memory usage of the second subtask cluster to the preset memory usage to obtain the ratio of the memory usage of the second subtask cluster;
[0095] S216. Respectively match the error value of the memory usage of the data processing task, the ratio of the memory usage of the second subtask cluster, and the number of subtasks in the second subtask cluster with the pre-constructed data skew severity score value table to obtain the error severity score, the memory usage severity score, and the subtask number severity score of the data processing task; wherein, the data skew severity score value table is used to represent the mapping relationship between the error value of the memory usage of the data processing task, the ratio of the memory usage of the second subtask cluster, the number of subtasks in the second subtask cluster, and the corresponding severity scores.
[0096] For ease of understanding, the pre-constructed data skew severity score value table is given exemplarily as follows, as shown in Table 1 below:
[0097] Table 1 Data Skew Severity Score Value Table
[0098] Error value Severity score Ratio of memory usage Severity score Number of tasks Severity score 1 1 1 / 8 1 10 1 2 2 1 / 4 2 20 2 4 3 1 / 2 3 50 3 8 4 1 4 100 4
[0099] For example: If the error value of the memory usage of the data processing task is 1, the ratio of the memory usage of the second subtask cluster is 0.5, and the number of subtasks in the second subtask cluster is 50, then according to Table 1, the severity score of the error value of the memory usage of the data processing task is 1, the severity score of the ratio of the memory usage of the second subtask cluster is 3, and the severity score of the number of subtasks in the second subtask cluster is 3.
[0100] S217. Determine the minimum value among the error severity score, the memory usage severity score, and the subtask number severity score of the data processing task as the data skew problem severity score of the data processing task.
[0101] For ease of understanding, the following takes actual data as an example to illustrate steps S211 - S217:
[0102] For example: Suppose there are a total of 60 subtasks in the data processing task, and they are divided into two parts. The first subtask cluster has 10 subtasks, and the memory usage of each subtask is 10MB. The second subtask cluster has 50 subtasks, and the memory usage of each subtask is 5MB.
[0103] The first step (calculate the average value):
[0104] The average value of the memory usage of the first subtask cluster = 10 * 10 / 10 = 10MB. The average value of the memory usage of the second subtask cluster = 5 * 50 / 50 = 5MB.
[0105] The second step (subtask cluster division):
[0106] The average memory usage of the first sub - task cluster (10) > the average memory usage of the second sub - task cluster (5), and the number of sub - tasks in the first sub - task cluster (10) < the number of sub - tasks in the second sub - task cluster (50). Then, the first sub - task cluster is regarded as the first sub - task cluster, and the second sub - task cluster is regarded as the second sub - task cluster.
[0107] The third step (calculate the error value):
[0108] The average memory usage of the first sub - task cluster is 10, and the average memory usage of the second sub - task cluster is 5. Then, the error value of the memory usage of the data - processing task = (10 - 5) / 5 = 1.
[0109] The fourth step (calculate the ratio):
[0110] The average memory usage of the second sub - task cluster is 5, and the preset memory usage is 10. Then, the ratio of the memory usage of the second sub - task cluster = 5 / 10 = 0.5.
[0111] The fifth step (determine the severity score):
[0112] The error value of the memory usage of the data - processing task is 1, the ratio of the memory usage of the second sub - task cluster is 0.5, and the number of sub - tasks in the second sub - task cluster is 50. Then, according to Table 1, the severity score of the error value of the memory usage of the data - processing task is 1, the severity score of the ratio of the memory usage of the second sub - task cluster is 3, and the severity score of the number of sub - tasks in the second sub - task cluster is 3.
[0113] The sixth step (determine the final severity score):
[0114] The severity score of the error value of the memory usage of the data - processing task is 1, the severity score of the ratio of the memory usage of the second sub - task cluster is 3, and the severity score of the number of sub - tasks in the second sub - task cluster is 3. The minimum value among these three severity scores is the severity score of the error value of the memory usage of the data - processing task. So, the severity score of the data skew problem of the data - processing task is 1.
[0115] S220. Based on the Java GC heuristic algorithm, calculate the average CPU usage time, average running time, and average garbage - collection consumption time of the data - processing task to obtain the Java GC problem severity score of the data - processing task.
[0116] Specifically, step S220 includes:
[0117] S221. Calculate the ratio of the average garbage collection consumption time to the average CPU usage time of the data processing task to obtain the GC-CPU ratio of the data processing task.
[0118] S222. Respectively match the GC-CPU ratio and the average running time of the data processing task with a pre-constructed Java application performance severity score value table to obtain the GC-CPU ratio severity score and the average running time severity score of the data processing task. Among them, the Java application performance severity score value table is used to represent the mapping relationship between the GC-CPU ratio of the data processing task and the average running time of the data processing task and the corresponding severity scores.
[0119] For ease of understanding, the following exemplarily gives the Java application performance severity score value table as shown in Table 2 below:
[0120] Table 2 Java application performance severity score value table
[0121] GC-CPU ratio Severity score Average running time Severity score 0.2 1 100 1 0.4 2 140 2 0.6 3 160 3 0.8 4 200 4
[0122] For example: The GC-CPU ratio of the data processing task is 0.1106, and the average running time of the data processing task is 160. According to Table 2, it can be known that 0.1106 is less than 0.2, so the GC-CPU ratio severity score of the data processing task is 1, and the average running time severity score of the data processing task is 3.
[0123] S223. Determine the minimum value of the GC-CPU ratio severity score and the average running time severity score of the data processing task as the Java GC problem severity score of the data processing task.
[0124] For ease of understanding, the following takes actual data as an example to illustrate steps S221 - S223:
[0125] Suppose we have conducted 5 tests on a data processing task, and the garbage collection consumption time (unit: second), CPU usage time (unit: second), and running time (unit: second) recorded each time are as follows:
[0126] First time: Garbage collection consumption time (2s), CPU usage time (20s), running time (150s).
[0127] Second time: Garbage collection consumption time (3s), CPU usage time (25s), running time (180s).
[0128] Third time: Garbage collection consumption time (2.5s), CPU usage time (22s), running time (160s).
[0129] Fourth time: Garbage collection consumed 3.5 s, CPU usage time was 28 s, and running time was 170 s.
[0130] Fifth time: Garbage collection consumed 1.5 s, CPU usage time was 18 s, and running time was 140 s.
[0131] Average garbage collection consumption time = (2 + 3 + 2.5 + 3.5 + 1.5) / 5 = 2.5 s.
[0132] Average CPU usage time = (20 + 25 + 22 + 28 + 18) / 5 = 22.6 s.
[0133] Average running time = (150 + 180 + 160 + 170 + 140) / 5 = 160 s.
[0134] First step (calculating the GC - CPU ratio):
[0135] The GC - CPU ratio of the data processing task = 2.5 / 22.6 = 0.1106.
[0136] Second step (determining the severity score):
[0137] The GC - CPU ratio of the data processing task is 0.1106, and the average running time of the data processing task is 160. According to Table 2, 0.1106 is less than 0.2. So, the severity score of the GC - CPU ratio of the data processing task is 1, and the severity score of the average running time of the data processing task is 3.
[0138] Third step (determining the final severity score):
[0139] The severity score of the GC - CPU ratio of the data processing task is 1, and the severity score of the average running time of the data processing task is 3. The minimum of these two severity scores is the severity score of the GC - CPU ratio of the data processing task. So, the severity score of the Java GC problem of the data processing task is 1.
[0140] S300. Weighted sum the severity scores of different running performance problems of the data processing task to obtain the severity score of the comprehensive running performance problem of the data processing task;
[0141] In one embodiment, step S300 specifically includes:
[0142] Using the following formula, calculate the severity score of the data skew problem of the data processing task, the preset weight value of the severity score of the data skew problem of the data processing task, the severity score of the Java GC problem of the data processing task, and the preset weight value of the severity score of the Java GC problem of the data processing task, to obtain the severity score of the comprehensive operation performance problem of the data processing task;
[0143] S = W1 * R1 + W2 * R2; where S represents the severity score of the comprehensive operation performance problem of the data processing task, W1 represents the preset weight value of the severity score of the data skew problem of the data processing task, R1 represents the severity score of the data skew problem of the data processing task, W2 represents the preset weight value of the severity score of the Java GC problem of the data processing task, and R2 represents the severity score of the Java GC problem of the data processing task.
[0144] For example: The preset weight value of the severity score of the data skew problem of the data processing task is 0.5, the severity score of the data skew problem of the data processing task is 1, the preset weight value of the severity score of the Java GC problem of the data processing task is 0.5, and the severity score of the Java GC problem of the data processing task is 1. Then the severity score of the comprehensive operation performance problem of the data processing task = 0.5 * 1 + 0.5 * 1 = 1.
[0145] It should be noted that the sum of the preset weight value of the severity score of the data skew problem of the data processing task and the preset weight value of the severity score of the Java GC problem of the data processing task is 1.
[0146] It should be noted that no matter how many heuristic algorithms there are, that is, it is necessary to analyze and calculate the operation performance index data of how many types of data processing tasks to obtain the severity score of the data processing task corresponding to the heuristic algorithm. Correspondingly, there will be as many weight values in step S300 here.
[0147] S400, delimit the level of the severity score of the comprehensive operation performance problem of the data processing task to determine the level to be optimized for the data processing task;
[0148] In one embodiment, step S400 specifically includes:
[0149] S410, if the severity score of the comprehensive operation performance problem of the data processing task is greater than the first preset severity score and less than the product of the second preset severity score and the first preset percentage, then determine that the level to be optimized for the data processing task is the first level to be optimized;
[0150] S420, if the severity score of the comprehensive operation performance problem of the data processing task is greater than the product of the second preset severity score and the first preset percentage and less than the product of the second preset severity score and the second preset percentage, determine that the optimization level to be determined for the data processing task is the second optimization level;
[0151] S430, if the severity score of the comprehensive operation performance problem of the data processing task is greater than the product of the second preset severity score and the second preset percentage and less than the product of the second preset severity score and the third preset percentage, determine that the optimization level to be determined for the data processing task is the third optimization level;
[0152] S440, if the severity score of the comprehensive operation performance problem of the data processing task is greater than the product of the second preset severity score and the third preset percentage, determine that the optimization level to be determined for the data processing task is the fourth optimization level;
[0153] Among them, the first preset severity score < the second preset severity score;
[0154] The first preset percentage < the second preset percentage < the third preset percentage.
[0155] For the convenience of understanding, the following is an example with actual data:
[0156] Suppose the severity score of the comprehensive operation performance problem of the data processing task is 1, the first preset severity score is 0, the second preset severity score is 2, the first preset percentage is 20%, the second preset percentage is 60%, and the third preset percentage is 80%.
[0157] Since 2 * 60% (1.2) < 1 < 2 * 80% (1.6), that is, the optimization level to be determined for the data processing task is the third optimization level.
[0158] S500, if the optimization level to be determined for the data processing task is greater than the preset optimization level, match the different operation performance problems of the data processing task with the pre-constructed optimization strategy table to obtain the optimization strategies for the different operation performance problems of the data processing task, so that the user can optimize the performance of the data processing task according to the optimization strategies for the different operation performance problems of the data processing task; among them, the optimization strategy table is used to represent the mapping relationship between the different operation performance problems of the data processing task and the corresponding optimization strategies.
[0159] It can be understood that the different operation performance problems of the data processing task correspond one-to-one with the heuristic algorithm optimization model. For example: the data skew problem of the data processing task requires the data skew heuristic algorithm.
[0160] Specifically, assume that the preset level to be optimized is the second level to be optimized, while the level to be optimized for the data processing task is the third level to be optimized. Then, the optimization strategy matching mechanism is triggered, that is, the different running performance problems of the data processing task are matched with the pre-constructed optimization strategy table to obtain the optimization strategies for the different running performance problems of the data processing task.
[0161] For the sake of easy understanding, the pre-constructed optimization strategy table is exemplarily given below, as shown in Table 3 below:
[0162] Table 3 Optimization Strategy Table
[0163]
[0164] For example: The different running performance problems of the data processing task include: data skew problem and Java GC problem. According to Table 3, it can be known that the optimization strategy corresponding to the data skew problem is: use the repartition or coalesce function to reallocate data, and the optimization strategy corresponding to the Java GC problem is: increase the Executor memory or adjust the GC strategy to G1GC.
[0165] Advantages of the present invention:
[0166] (1) Comprehensive performance evaluation: By obtaining operation performance index data of multiple categories, the operation status of the data processing task can be comprehensively understood from multiple dimensions, avoiding the one-sidedness of single-index evaluation, so as to more accurately grasp the performance performance of the task in different aspects.
[0167] (2) Targeted problem analysis: Using different heuristic algorithm models to process operation performance index data of different categories respectively, each algorithm model is optimized for specific types of data, which can deeply dig out the operation performance problems hidden behind various types of data and give the severity scores of the corresponding problems, improving the accuracy and pertinence of problem analysis.
[0168] (3) Quantitative comprehensive evaluation: Weighted summation of the severity scores of different operation performance problems is performed to obtain the severity score of the comprehensive operation performance problem, quantifying the complex performance problem into a specific value, which is convenient for intuitively measuring the severity of the overall performance problem of the task. At the same time, the weighted method can also reflect the differences in the impact of different types of problems on the overall performance.
[0169] (4) Clear optimization priority: The comprehensive severity score is graded to determine the level to be optimized, enabling users to clearly understand the performance optimization priority of the data processing task. For tasks with a higher level to be optimized, they are processed first, and optimization resources are reasonably allocated.
[0170] Embodiment 2
[0171] Based on the same inventive concept, Figure 2 As shown, the embodiment of the present invention further provides a data processing task running performance optimization device 200, comprising:
[0172] The data acquisition module 210 is used to acquire the operation performance indicator data of the data processing task and determine the category of each operation performance indicator data;
[0173] The performance analysis module 220 is used to use different heuristic algorithms to perform operation performance analysis on different types of operation performance indicator data of the data processing task, and obtain severity scores of different operation performance problems of the data processing task; wherein different types of operation performance indicator data correspond to different heuristic algorithms one by one;
[0174] A weighted average module 230 is used to perform a weighted summation of severity scores of different operational performance issues of the data processing task to obtain a severity score of the comprehensive operational performance issue of the data processing task;
[0175] A level classification module 240 is used to classify the severity scores of the comprehensive operation performance problems of the data processing tasks and determine the level of the data processing tasks to be optimized;
[0176] The strategy execution module 250 is used to match different operating performance issues of the data processing task with a pre-constructed optimization strategy table if the level of the data processing task to be optimized is greater than the preset level to be optimized, and obtain the optimization strategy for the different operating performance issues of the data processing task, so that the user can optimize the performance of the data processing task according to the optimization strategy for the different operating performance issues of the data processing task; wherein the optimization strategy table is used to characterize the mapping relationship between different operating performance issues of the data processing task and the corresponding optimization strategy.
[0177] It should be understood that the device corresponds to the above-mentioned data processing task performance optimization method embodiment, and can execute the various steps involved in the above-mentioned method embodiment. The specific functions of the device can be found in the above description. To avoid repetition, the detailed description is appropriately omitted here. The device includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the operating system (OS) of the device.
[0178] Embodiment 3
[0179] Based on the same inventive concept, an embodiment of the present invention also provides an electronic device, including: a processor and a memory, the memory storing machine-readable instructions executable by the processor, and the machine-readable instructions executing the above-mentioned data processing task performance optimization method when executed by the processor.
[0180] In a typical configuration, an electronic device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0181] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0182] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0183] Embodiment 4
[0184] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium storing computer instructions that, when run on a computer, cause the computer to execute the above-mentioned method for optimizing the running performance of data processing tasks.
[0185] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0186] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0187] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0188] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0189] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the embodiments of the present invention do not separately describe various possible combination methods.
[0190] In addition, in each embodiment of this application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0191] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0192] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for optimizing the running performance of data processing tasks, characterized in that, Including: Obtain the running performance metric data of the data processing task and determine the categories of each running performance metric data; Use different heuristic algorithms to perform running performance analysis on different categories of running performance metric data of the data processing task respectively, and obtain the severity scores of different running performance problems of the data processing task; wherein, different categories of running performance metric data correspond one-to-one with different heuristic algorithms; Perform weighted summation on the severity scores of different running performance problems of the data processing task to obtain the severity score of the comprehensive running performance problem of the data processing task; Define the level of the severity score of the comprehensive running performance problem of the data processing task to determine the level to be optimized for the data processing task; If the level to be optimized for the data processing task is greater than the preset level to be optimized, then match the different running performance problems of the data processing task with the pre-constructed optimization strategy table to obtain the optimization strategies for different running performance problems of the data processing task, so that the user can perform performance optimization on the data processing task according to the optimization strategies for different running performance problems of the data processing task; wherein, the optimization strategy table is used to represent the mapping relationship between different running performance problems of the data processing task and the corresponding optimization strategies.
2. The method for optimizing the running performance of a data processing task according to claim 1, wherein The heuristic algorithms include: data skew heuristic algorithm and Java GC heuristic algorithm; The categories of running performance metric data include: data skew metric data and Java application performance metric data; The data skew metric data includes: the number of subtasks and the memory usage of each subtask; The Java application performance metric data includes: average CPU usage time, average running time, and average garbage collection consumption time.
3. The data processing task running performance optimization method according to claim 2, wherein Using different heuristic algorithms to perform running performance analysis on different categories of running performance metric data of the data processing task respectively, and obtaining the severity scores of different running performance problems of the data processing task, including: Based on the data skew heuristic algorithm, calculate the number of subtasks and the memory usage of each subtask of the data processing task to obtain the severity score of the data skew problem of the data processing task; Based on the Java GC heuristic algorithm, calculate the average CPU usage time, average running time, and average garbage collection consumption time of the data processing task to obtain the severity score of the Java GC problem of the data processing task.
4. The data processing task running performance optimization method according to claim 3, characterized in that Based on the data skew heuristic algorithm, calculate the number of subtasks and the memory usage of each subtask of the data processing task to obtain the severity score of the data skew problem of the data processing task, including: Divide multiple subtasks in the data processing task into two equal parts to obtain two subtask clusters; Calculate the average value of the memory usage of the two subtask clusters respectively; If there is a subtask cluster with the largest average memory usage and the smallest number of subtasks, then use this subtask cluster as the first subtask cluster, and determine the task cluster other than the first subtask cluster as the second subtask cluster; Calculate the difference between the average memory usage of the second subtask cluster and the average memory usage of the first subtask cluster, and calculate the ratio of the difference to the average memory usage of the first subtask cluster to obtain the error value of the memory usage of the data processing task; Calculate the ratio of the average memory usage of the second subtask cluster to the preset memory usage to obtain the ratio of the memory usage of the second subtask cluster; Match the error value of the memory usage of the data processing task, the ratio of the memory usage of the second subtask cluster, and the number of subtasks in the second subtask cluster with the pre-constructed data skew severity score value table respectively to obtain the error severity score, memory usage severity score, and subtask number severity score of the data processing task; wherein, the data skew severity score value table is used to represent the mapping relationship between the error value of the memory usage of the data processing task, the ratio of the memory usage of the second subtask cluster, the number of subtasks in the second subtask cluster and the corresponding severity scores; Determine the minimum value among the error severity score, memory usage severity score, and subtask number severity score of the data processing task as the data skew problem severity score of the data processing task.
5. The data processing task running performance optimization method according to claim 3, wherein Based on the Java GC heuristic algorithm, calculate the average CPU usage time, average running time, and average garbage collection consumption time of the data processing task to obtain the Java GC problem severity score of the data processing task, including: Calculate the ratio of the average garbage collection consumption time and the average CPU usage time of the data processing task to obtain the GC-CPU ratio of the data processing task; Match the GC-CPU ratio and the average running time of the data processing task with the pre-constructed Java application performance severity score value table respectively to obtain the GC-CPU ratio severity score and the average running time severity score of the data processing task; wherein, the Java application performance severity score value table is used to represent the mapping relationship between the GC-CPU ratio of the data processing task, the average running time of the data processing task and the corresponding severity scores; Determine the minimum value between the GC-CPU ratio severity score and the average running time severity score of the data processing task as the Java GC problem severity score of the data processing task.
6. The data processing task running performance optimization method according to claim 3, wherein Perform a weighted sum of the severity scores of different running performance problems of the data processing task to obtain the severity score of the comprehensive running performance problem of the data processing task, including: Use the following formula to calculate the data skew problem severity score of the data processing task, the preset weight value of the data skew problem severity score of the data processing task, the Java GC problem severity score of the data processing task, and the preset weight value of the Java GC problem severity score of the data processing task to obtain the severity score of the comprehensive running performance problem of the data processing task; S = W1 * R1 + W2 * R2; where S represents the severity score of the comprehensive operation performance problem of the data processing task, W1 represents the preset weight value of the severity score of the data skew problem of the data processing task, R1 represents the severity score of the data skew problem of the data processing task, W2 represents the preset weight value of the severity score of the Java GC problem of the data processing task, and R2 represents the severity score of the Java GC problem of the data processing task.
7. The method for optimizing the running performance of a data processing task according to claim 1, wherein Grade the severity score of the comprehensive operation performance problem of the data processing task to determine the optimization level to be optimized for the data processing task, including: If the severity score of the comprehensive operation performance problem of the data processing task is greater than the first preset severity score and less than the product of the second preset severity score and the first preset percentage, determine that the optimization level to be optimized for the data processing task is the first optimization level to be optimized; If the severity score of the comprehensive operation performance problem of the data processing task is greater than the product of the second preset severity score and the first preset percentage and less than the product of the second preset severity score and the second preset percentage, determine that the optimization level to be optimized for the data processing task is the second optimization level to be optimized; If the severity score of the comprehensive operation performance problem of the data processing task is greater than the product of the second preset severity score and the second preset percentage and less than the product of the second preset severity score and the third preset percentage, determine that the optimization level to be optimized for the data processing task is the third optimization level to be optimized; If the severity score of the comprehensive operation performance problem of the data processing task is greater than the product of the second preset severity score and the third preset percentage, determine that the optimization level to be optimized for the data processing task is the fourth optimization level to be optimized; Among them, the first preset severity score < the second preset severity score; The first preset percentage < the second preset percentage < the third preset percentage.
8. An apparatus for optimizing the running performance of a data processing task, characterized in that, Including: A data acquisition module for acquiring the operation performance index data of the data processing task and determining the category of each operation performance index data; A performance analysis module for using different heuristic algorithms to perform operation performance analysis on different categories of operation performance index data of the data processing task to obtain the severity scores of different operation performance problems of the data processing task; among them, different categories of operation performance index data correspond one-to-one with different heuristic algorithms; A weighted average module for performing weighted summation on the severity scores of different operation performance problems of the data processing task to obtain the severity score of the comprehensive operation performance problem of the data processing task; A grade determination module for grading the severity score of the comprehensive operation performance problem of the data processing task to determine the optimization level to be optimized for the data processing task; A policy execution module, which is configured to, if the to-be-optimized level of a data processing task is greater than a preset to-be-optimized level, match different running performance problems of the data processing task with a pre-constructed optimization policy table to obtain optimization policies for different running performance problems of the data processing task, so that a user can perform performance optimization on the data processing task according to the optimization policies for different running performance problems of the data processing task; wherein, the optimization policy table is used to represent the mapping relationship between different running performance problems of the data processing task and corresponding optimization policies.
9. An electronic device, characterized in that, It includes: A processor and a memory, where the memory stores machine-readable instructions executable by the processor, and when the machine-readable instructions are executed by the processor, the data processing task running performance optimization method according to any one of claims 1-7 is executed.
10. A computer-readable storage medium stores computer instructions, characterized in that, When the computer instructions run on a computer, the computer is caused to execute the data processing task running performance optimization method according to any one of claims 1-7.