A data aggregation optimization method and system based on partition data reordering, a terminal and a storage medium

By partitioning and reordering intermediate files during the mapping phase, the redundant access problem of Shuffle operations in Spark distributed computing is solved, communication efficiency and throughput are improved, and the overall performance of Spark is enhanced.

CN120849443BActive Publication Date: 2025-11-21GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511359848.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-11-21
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

Existing technologies suffer from redundant access issues during Shuffle operations in Spark distributed computing, resulting in high communication overhead and low efficiency.

Method used

By partitioning and merging intermediate files during the mapping phase, calculating the size ratio of partitioned data, reordering the data, and reducing the number of accesses to intermediate files during the data aggregation phase, a data aggregation optimization method based on partitioned data reordering is adopted.

Benefits of technology

It reduces the communication overhead of Shuffle operations and improves the efficiency of Spark distributed computing, especially when dealing with massive amounts of small tasks and small data blocks, significantly improving job execution time and throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849443B_ABST
    Figure CN120849443B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data aggregation optimization, and discloses a data aggregation optimization method and system based on partitioned data reordering, a terminal and a storage medium.The method comprises the following steps: when different nodes execute a mapping task, all intermediate files of the nodes are acquired, partition processing is performed to obtain multiple partitioned data, all the partitioned data are subjected to merging processing to obtain multiple target intermediate files, the target partition data size of each target intermediate file is calculated to obtain a target partition calculation result, reordering processing is performed on all the target intermediate files to obtain a final intermediate file, the final intermediate file of each node is acquired, and all the final intermediate files are subjected to aggregation processing to obtain a data aggregation optimization result.The intermediate files generated in advance are merged, and the partitioned data are reordered according to the proportion of each partitioned data, so that the access frequency of the intermediate files by each node is reduced in the data aggregation stage, and the efficiency of Spark distributed calculation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data aggregation and optimization technology, and in particular to a data aggregation and optimization method, system, terminal, and computer-readable storage medium based on partitioned data reordering. Background Technology

[0002] The shuffle operation (used to aggregate data across partitions, grouping data with the same key into the same partition) is a critical stage in Spark distributed computing for cross-node data redistribution, and its performance directly impacts job efficiency. Due to the large amount of disk I / O, network transmission, and load balancing involved, the shuffle operation becomes a performance bottleneck, especially when processing massive amounts of small tasks and small data blocks.

[0003] Currently, methods for reducing the number of Shuffle files using file pre-merging include Riffle (also known as Optimized Shuffle Service for Large-Scale Data Analysis) and OPS (Optimized Shuffle Management System for Apache Spark). Riffle primarily addresses the disk I / O bottleneck caused by the massive number of fragmented intermediate files during the Spark Shuffle phase. It improves data reading efficiency through a dynamic file merging mechanism, and its core function is to reduce the number of files accessed by Reduce tasks (data aggregation tasks). OPS aims to optimize Shuffle performance through decoupling architecture and reduce disk I / O latency by utilizing memory pools and network acceleration. While Riffle significantly reduces the number of files and improves read throughput, it does not optimize the order of data blocks within the merged file, requiring Reduce tasks to still scan the entire file to extract target partition data, resulting in redundant access. OPS, while reducing transmission latency, relies entirely on memory storage for intermediate data, which can lead to memory overflow. Furthermore, it does not sort data blocks in the memory pool by partition, requiring Reduce tasks to still perform a full scan to locate the target data, thus failing to solve the redundant access problem. Furthermore, when a job is divided into multiple smaller tasks and the partition data blocks are small, the number of data packets generated during network transmission increases and their size decreases, increasing bandwidth usage and data transmission latency, thereby affecting the overall execution efficiency of the job.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The main objective of this invention is to provide a data aggregation optimization method, system, terminal, and storage medium based on partitioned data reordering. This invention aims to solve the problem of redundant access in existing technologies when performing data aggregation tasks, which results in large communication overhead from Shuffle operations and thus low efficiency of Spark distributed computing.

[0006] To achieve the above objectives, the present invention provides a data aggregation optimization method based on partitioned data reordering, the method comprising the following steps:

[0007] When different nodes execute mapping tasks, all intermediate files of the node are obtained according to the data aggregation optimization instructions. Each intermediate file is partitioned to obtain multiple partition data. All partition data are then merged to obtain multiple target intermediate files.

[0008] The target partition data of each target intermediate file is calculated to obtain the target partition calculation result. Based on the target partition calculation result, all target intermediate files are reordered to obtain the final intermediate file of the node.

[0009] Obtain the final intermediate file for each node, aggregate all the final intermediate files of the nodes, and obtain the data aggregation optimization result.

[0010] Optionally, the data aggregation optimization method based on partitioned data reordering, wherein when different nodes are executing mapping tasks, all intermediate files of the node are obtained according to the data aggregation optimization instruction, and each intermediate file is partitioned to obtain multiple partitioned data, specifically includes:

[0011] When different nodes execute mapping tasks, the intermediate files and index files corresponding to all mapping tasks of the node are obtained according to the data aggregation optimization instructions;

[0012] The corresponding partition ID is obtained from the index file, and all the intermediate files are partitioned according to the partition ID to obtain multiple partition data.

[0013] Optionally, the data aggregation optimization method based on partitioned data reordering, wherein merging all the partitioned data to obtain multiple target intermediate files specifically includes:

[0014] Obtain the partition number of all the partition data, and classify all the partition data according to the partition number to obtain the classification result;

[0015] Based on the classification results, partition data with the same partition number are merged to obtain multiple target intermediate files.

[0016] Optionally, the data aggregation optimization method based on partitioned data reordering, wherein calculating the data size of the target partition data for each target intermediate file to obtain the target partition calculation result specifically includes:

[0017] The data size of the target partition data of each target intermediate file is calculated to obtain the data value of each target intermediate file, and the total data value of all target partition data is calculated.

[0018] The proportion of each data value is calculated based on the total data value to obtain the partition data proportion value of each target intermediate file, and the target partition calculation result is obtained based on the proportion values ​​of all partition data.

[0019] Optionally, the data aggregation optimization method based on partitioned data reordering, wherein the step of reordering all the target intermediate files according to the target partition calculation results to obtain the final intermediate file of the node specifically includes:

[0020] Set a percentage threshold, compare the percentage values ​​of all partition data in the target partition calculation result with the percentage threshold, and obtain the comparison result;

[0021] Sort all the data percentage values ​​of the first partition in the comparison results by size to obtain the partition data sorting result, and set all the data percentage values ​​of the second partition in the comparison results to a preset position in the partition data sorting result to obtain the target partition data sorting result;

[0022] Based on the sorting results of the target partition data, all the target intermediate files are merged to obtain the final intermediate file of the node;

[0023] Wherein, the first partition data percentage is the partition data percentage greater than or equal to the percentage threshold, and the second partition data percentage is the partition data percentage less than the percentage threshold.

[0024] Optionally, the data aggregation optimization method based on partitioned data reordering, wherein setting the proportion of all second partition data in the comparison result at a preset position in the partitioned data sorting result specifically includes:

[0025] If there is a partition data percentage value in the second partition that is less than the percentage threshold, then the second partition data percentage value is set in the preset position of the partition data sorting result;

[0026] If there are multiple partition data percentage values ​​less than the percentage threshold for the second partition data percentage value, then obtain the partition number order of all target intermediate files corresponding to the second partition data percentage value, and set all the second partition data percentage values ​​in the preset position of the partition data sorting result according to the partition number order.

[0027] Optionally, the data aggregation optimization method based on partitioned data reordering, wherein obtaining the final intermediate file of each node and aggregating all the final intermediate files of the nodes to obtain the data aggregation optimization result specifically includes:

[0028] Obtain the final intermediate file for each node, and extract the final partition data for each partition of the final intermediate files for all nodes;

[0029] Obtain the final partition number of all final partition data, and aggregate the final partition data with the same final partition number to obtain the data aggregation optimization result. The aggregation process includes summation, averaging and maximum value processing.

[0030] Optionally, the data aggregation optimization method based on partitioned data reordering, wherein the data aggregation optimization system based on partitioned data reordering includes:

[0031] The data merging module is used to obtain all intermediate files of the nodes according to the data aggregation optimization instructions when different nodes are performing mapping tasks, perform partitioning processing on each intermediate file to obtain multiple partition data, and merge all the partition data to obtain multiple target intermediate files.

[0032] The data sorting module is used to calculate the data size of the target partition data of each target intermediate file, obtain the target partition calculation result, and re-sort all the target intermediate files according to the target partition calculation result to obtain the final intermediate file of the node.

[0033] The data aggregation module is used to obtain the final intermediate file of each node, aggregate the final intermediate files of all nodes, and obtain the data aggregation optimization result.

[0034] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a data aggregation optimization program based on partitioned data reordering stored in the memory and executable on the processor, wherein when the data aggregation optimization program based on partitioned data reordering is executed by the processor, it implements the steps of the data aggregation optimization method based on partitioned data reordering as described above.

[0035] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a data aggregation optimization program based on partitioned data reordering, and when the data aggregation optimization program based on partitioned data reordering is executed by a processor, it implements the steps of the data aggregation optimization method based on partitioned data reordering as described above.

[0036] In this invention, when different nodes are executing mapping tasks, all intermediate files of the node are obtained according to the data aggregation optimization instructions. Each intermediate file is partitioned to obtain multiple partition data, and all partition data are merged to obtain multiple target intermediate files. The data size of the target partition data in each target intermediate file is calculated to obtain the target partition calculation result. Based on the target partition calculation result, all target intermediate files are reordered to obtain the final intermediate file of the node. The final intermediate file of each node is obtained, and all final intermediate files of the nodes are aggregated to obtain the data aggregation optimization result. This invention merges the generated intermediate files in advance during the mapping stage and sorts them according to the proportion of each partition data, thereby reducing the number of times each node accesses the intermediate files during the data aggregation stage, thus reducing the communication overhead generated by data aggregation and improving the efficiency of Spark distributed computing. Attached Figure Description

[0037] Figure 1 This is a flowchart of a preferred embodiment of the data aggregation optimization method based on partitioned data reordering of the present invention;

[0038] Figure 2 This is a schematic diagram of the overall architecture of the data aggregation optimization method based on partitioned data reordering of the present invention;

[0039] Figure 3 This is a schematic diagram of the pre-merging of intermediate files in an embodiment of the present invention;

[0040] Figure 4 This is a schematic diagram of partitioned data sorting in an embodiment of the present invention;

[0041] Figure 5 This is a schematic diagram of the job execution time of Spark, Riffle, OPS and SRO under the Large dataset in this embodiment of the invention;

[0042] Figure 6 This is a schematic diagram of the job execution time of Spark, Riffle, OPS and SRO under the Huge dataset in this embodiment of the invention;

[0043] Figure 7 This is a schematic diagram of the average throughput of OPS and SRO jobs under the Large dataset in this embodiment of the invention;

[0044] Figure 8 This is a schematic diagram of the average throughput of OPS and SRO jobs under the Huge dataset in this embodiment of the invention;

[0045] Figure 9 This is a structural diagram of a preferred embodiment of the data aggregation optimization system based on partitioned data reordering of the present invention;

[0046] Figure 10 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0048] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.

[0049] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0050] The preferred embodiment of the data aggregation optimization method based on partitioned data reordering described in this invention, such as... Figure 1 As shown, the data aggregation optimization method based on partitioned data reordering includes the following steps:

[0051] Step S10: When different nodes are performing mapping tasks, all intermediate files of the node are obtained according to the data aggregation optimization instruction. Each intermediate file is partitioned to obtain multiple partition data. All partition data are then merged to obtain multiple target intermediate files.

[0052] Specifically, in this embodiment of the invention, to address the problem of redundant access in existing technologies during data aggregation tasks, which leads to high communication overhead from Shuffle operations and consequently low efficiency in Spark distributed computing, a data aggregation optimization method based on partitioned data reordering is proposed. The terminal executing this method is an SRO (Shuffle Reduce optimized), and the corresponding architecture diagram is shown below. Figure 2 As shown, each node has multiple executors that perform mapping tasks, dividing the intermediate file data into multiple partitions according to partitioning rules and serializing the data in each partition. Once all executors in a node have completed processing a partition's data, the Shuffle optimizer initiates a merge operation, merging the data from these partitions to generate a new intermediate file (i.e., a template intermediate file). When all partition data in a node has been processed, the Shuffle optimizer calculates the percentage of each partition's data relative to the total data, using this as a metric for partition data sorting. It then performs a sorting operation, sorting and merging the data blocks from different partitions to generate a unique intermediate file for the node (i.e., the final intermediate file). Finally, aggregation tasks are executed. Each aggregation task reads data from the final intermediate files of each node, performs aggregation operations on all values ​​with the same key, and generates the final output (i.e., the data aggregation optimization result).

[0053] First, the generated intermediate files undergo pre-merging processing, specifically as follows: Figure 3As shown, a mergeFiles (data structure) is predefined within each node to store the data generated by each partition, and a partitionSizes (partition data size list) is defined to record the size of each partition's data. Then, when different nodes execute mapping tasks, they obtain the shuffleFile (intermediate files) and indexFile (index files) corresponding to all mapping tasks of the node according to the data aggregation optimization instructions. The corresponding ID partitionId (partition ID) is obtained from the index file, where partitionId is part of the intermediate data "address" indicating which logical partition the required data is allocated to. All intermediate files are partitioned according to the partition ID to obtain multiple partitionData (partition data), for example, partition I data and partition II data. Finally, the partition data generated by multiple executors needs to be merged to generate a mergeFile (target intermediate file), and the p of the partition is calculated. The `artitionSize` (data size) is calculated and saved to `partitionSizes`. Specifically, the partition numbers of all partitioned data are obtained, and all partitioned data are categorized according to the partition numbers to obtain categorization results. Based on the categorization results, partitioned data with the same partition number are merged to obtain multiple target intermediate files. Each target intermediate file contains all data for that node and partition, and `mergeFile` is saved to `mergeFiles`. For example, if partition I and partition II data are processed sequentially, then, according to the order of processing completion, the partition I data completed by different mapping tasks are merged first to generate a merged file containing only partition I data. Then, partition II data is merged in the same way to generate a new merged file. At this time, the mapping tasks and file pre-merging are performed simultaneously, generating merged files of the same partition first, providing complete partition data blocks for the subsequent generation of the final intermediate file, and preparing for subsequent partition data reordering.

[0054] Step S20: Calculate the data size of the target partition data for each target intermediate file to obtain the target partition calculation result. Reorder all the target intermediate files according to the target partition calculation result to obtain the final intermediate file of the node.

[0055] Specifically, in this embodiment of the invention, after obtaining the target intermediate file of a node, the data size of the target partition data of each target intermediate file is calculated to obtain the data value of each target intermediate file, such as... Figure 4In section (a), there are seven rectangles with different numbers (60, 40, 15, 25, 80, 130, and 50), representing seven different target partitions. The number on each rectangle represents the size of each target partition in the intermediate file. Next, the partition data needs to be reordered. First, a newPartitionSizes structure is defined to record the percentage of each partition's data in the total data size (i.e., the partition data percentage). The total data value of all target partitions is calculated. Then, based on the total data value, the percentage of each data value is calculated to obtain the partition data percentage of each target intermediate file, and this percentage is saved to newPartitionSizes. Figure 4 In section (b), there are 7 rectangles with different percentages (32.5%, 20%, 15%, 12.5%, 10%, 3.75%, and 6.25%), representing 7 different target partition data. The percentage on each rectangle represents the proportion of each target partition data in the target intermediate file. The percentage is used as the sorting index, and the partition data block with the larger percentage is placed at the beginning of the final intermediate file.

[0056] Subsequently, a percentage threshold (e.g., 10%) is set in this invention. If the percentage of partition data is lower than this percentage threshold, the target intermediate file corresponding to the percentage of partition data is placed at the end of the final intermediate file. Then, a new partitionList (partition data list) is defined to record the partitionId (partition number) where the percentage of partition data is less than the percentage threshold. The purpose is to place the target intermediate file corresponding to the partition number where the percentage of partition data is less than the percentage threshold directly after the final intermediate file, without participating in the subsequent partition data sorting. During this process, the target partition data corresponding to the target intermediate file whose partition data percentage is less than the percentage threshold will be removed from newPartitonSizes and added to partitionList. Specifically, the percentage values ​​of all partition data in the target partition calculation result are compared with the percentage threshold to obtain a comparison result. The percentage values ​​of all first partition data in the comparison result (the first partition data percentage refers to the percentage value of partition data that is greater than or equal to the percentage threshold) are sorted by size to obtain a partition data sorting result. The percentage values ​​of all second partition data in the comparison result (the second partition data percentage refers to the percentage value of partition data that is less than the percentage threshold) are set in the preset position of the partition data sorting result to obtain the target partition data sorting result.

[0057] If there are multiple partition data percentage values ​​less than the stated percentage threshold for the second partition data percentage value, then the partition number order of all target intermediate files corresponding to the second partition data percentage value is obtained, without changing the relative order among all target intermediate files. For example, Figure 4 The target partition data with sizes of 15 and 25 were reordered, but the relative order between the two target partition data was not changed. Instead, they were placed at the end of the final intermediate file.

[0058] Finally, all the target intermediate files are merged according to the sorting results of the target partition data to obtain the finalFile (final intermediate file) of the node. The generation of finalFile is the end point of the mapping stage, and also the starting point of the cross-partition data aggregation stage and the data source of the data aggregation stage.

[0059] Step S30: Obtain the final intermediate file of each node, and aggregate all the final intermediate files of the nodes to obtain the data aggregation optimization result.

[0060] Specifically, in this embodiment of the invention, the final intermediate file of each node is obtained, and the final partition data of each partition of the final intermediate file of all nodes is extracted; the final partition number of all final partition data is obtained, and the final partition data with the same final partition number are aggregated to obtain the data aggregation optimization result, wherein the aggregation process includes summation, averaging and maximum value processing.

[0061] Furthermore, based on the simplified model analysis framework, the effectiveness of the sorting strategy of this invention in optimizing overall communication complexity is analyzed, as well as the computational and space overhead generated by the sorting strategy when executing mapping tasks. A Shuffle job has on a certain node... The partition, the first The amount of data generated by each partition is Total data volume for:

[0062] ;

[0063] When each intermediate file is written to disk, the underlying storage divides the data into several contiguous blocks, assuming each block is of size [size missing]. When a data aggregation task reads data from the same partition, if this data is scattered across... On a discontinuous block, it is necessary to perform Random reads. The fixed overhead number of blocks, such as metadata or file headers, is [number]. When there is no sorting, let the original block divergence be... That is, each partition is evenly distributed in On each block, the original block divergence is:

[0064] ;

[0065] Set the threshold to Divide all partitions into two categories. A large partition is a set of partitions whose data percentage is greater than or equal to a certain threshold. A small partition is a set of partitions whose data share is less than a percentage threshold. This refers to the size of the large partition. If the size of the smaller partition is given, then:

[0066] ;

[0067] ;

[0068] ;

[0069] For large partitions of data exceeding the threshold, the fixed number of overhead blocks such as metadata or file headers must be satisfied. Reorder and write the data in descending order based on data volume. Reduced to:

[0070] ;

[0071] For small partitions of data less than or equal to the threshold constant.

[0072] The random I / O overhead for writing or reading a block is The complexity of raw Shuffle communication without sorting The communication complexity after sorting Communication overhead revenue Raw Shuffle Time Cost and the time cost after sorting for:

[0073] ;

[0074] ;

[0075] = ;

[0076] ;

[0077] ;

[0078] The threshold strategy introduces additional overhead at the mapping end. First, it involves calculating the size of all partitions and then dividing them. and The cost is Sort the large partitions as follows Merging intermediate files from a large partition has a worst-case read / write complexity of [missing value]. Additional time overhead at the mapping end and space overhead for:

[0079] ;

[0080] ;

[0081] Net Time Benefits and optimal threshold for:

[0082] ;

[0083] ;

[0084] Only when At this time, the sorting strategy brings overall speedup. At that time, more partitions participate in the sorting. Maximum, but higher mapping cost; when Almost no partitions are involved in the sorting, resulting in the lowest cost, but For scenarios with severe data skew, a smaller threshold should be selected to minimize the shuffle overhead of large partitions; while for scenarios with uniform data distribution, a higher threshold should be selected or no sorting strategy should be used to avoid unnecessary additional overhead.

[0085] To visually demonstrate the advantages of Spark's SRO (Spark Resource Management) invention, its performance is evaluated using job execution time and throughput. Job execution time directly reflects the overall performance of the optimized system in actual task execution, while throughput illustrates Spark's transmission capacity when processing large-scale data, reflecting the amount of data the system can process and transmit per unit time. This directly relates to the system's resource utilization and data flow efficiency. Especially when processing massive amounts of data, an increase in throughput often leads to a significant performance improvement.

[0086] Regarding job execution time performance: Using two datasets of different sizes, namely the Large dataset and the Huge dataset, five jobs—Repartition, Aggregate, Join, Sort, and WordCount—were used to verify and compare the job execution times of Spark, Riffle, OPS, and SRO. Repartition repartitions the data, redistributing it from one partition distribution to another, involving large-scale data transfer and reorganization. Aggregate groups data with the same key value together for subsequent aggregation calculations. Join matches and combines data from two or more datasets based on certain conditions. Sort sorts the dataset according to a key. WordCount counts the frequency of each word in a text file, including word extraction, grouping, and counting the occurrences of each word. The Large and Huge datasets corresponding to each job are shown in Tables 1 and 2, respectively.

[0087] Table 1: Datasets for Repartition, Sort, and WordCount

[0088]

[0089] Table 2: Datasets for Aggregate and Join

[0090]

[0091] When using the Large dataset, the execution time of these 5 jobs is as follows: Figure 5 As shown, the average execution time of SRO is reduced by 18.6% compared to the original Spark, 8.9% compared to Riffle, and 3.2% compared to OPS; while when using the Huge dataset, the execution times of these five jobs are as follows: Figure 6 As shown, SRO's average execution time is reduced by 23.7% compared to the original Spark, 10.6% compared to Riffle, and 10.1% compared to OPS.

[0092] Regarding throughput performance: Using the Large and Huge datasets from the job execution time performance analysis, we compared and validated the throughput of Spark and SRO using four jobs: Aggregate, Join, Sort, and WordCount. When using the Large dataset, the average throughput of these four jobs was as follows: Figure 7As shown, when using the Huge dataset, the average throughput of these four jobs is as follows: Figure 8 As shown, by Figure 7 and Figure 8 The performance improvements after optimization are clearly visible; specifically, compared to Spark, SRO shows an average 34% increase in aggregation throughput and an average 32% increase in connection throughput. The throughput improvement is even more pronounced with larger datasets, primarily due to intermediate file pre-merging and partition data reordering. These optimizations reduce intermediate data storage and access, thereby improving data processing efficiency. In contrast, sorting and word frequency statistics show an average throughput improvement of 20%.

[0093] Furthermore, such as Figure 9 As shown, based on the above-described data aggregation optimization method based on partitioned data reordering, the present invention also provides a data aggregation optimization system based on partitioned data reordering, wherein the data aggregation optimization system based on partitioned data reordering includes:

[0094] The data merging module 51 is used to obtain all intermediate files of the node according to the data aggregation optimization instruction when different nodes are performing mapping tasks, perform partitioning processing on each intermediate file to obtain multiple partition data, and merge all the partition data to obtain multiple target intermediate files.

[0095] The data sorting module 52 is used to calculate the data size of the target partition data of each target intermediate file, obtain the target partition calculation result, and re-sort all the target intermediate files according to the target partition calculation result to obtain the final intermediate file of the node.

[0096] The data aggregation module 53 is used to obtain the final intermediate file of each node, aggregate the final intermediate files of all nodes, and obtain the data aggregation optimization result.

[0097] Furthermore, such as Figure 10 As shown, based on the above-mentioned data aggregation optimization method based on partitioned data reordering, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 10 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0098] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a data aggregation optimization program 40 based on partitioned data reordering, which can be executed by the processor 10 to implement the data aggregation optimization method based on partitioned data reordering in this application.

[0099] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the data aggregation optimization method based on partitioned data reordering.

[0100] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.

[0101] In one embodiment, when the processor 10 executes the data aggregation optimization program 40 based on partition data reordering in the memory 20, the following steps are performed:

[0102] When different nodes execute mapping tasks, all intermediate files of the node are obtained according to the data aggregation optimization instructions. Each intermediate file is partitioned to obtain multiple partition data. All partition data are then merged to obtain multiple target intermediate files.

[0103] The target partition data of each target intermediate file is calculated to obtain the target partition calculation result. Based on the target partition calculation result, all target intermediate files are reordered to obtain the final intermediate file of the node.

[0104] Obtain the final intermediate file for each node, aggregate all the final intermediate files of the nodes, and obtain the data aggregation optimization result.

[0105] Specifically, when different nodes are executing mapping tasks, all intermediate files of the node are obtained according to the data aggregation optimization instruction, and each intermediate file is partitioned to obtain multiple partition data, which includes:

[0106] When different nodes execute mapping tasks, the intermediate files and index files corresponding to all mapping tasks of the node are obtained according to the data aggregation optimization instructions;

[0107] The corresponding partition ID is obtained from the index file, and all the intermediate files are partitioned according to the partition ID to obtain multiple partition data.

[0108] The step of merging all the partition data to obtain multiple target intermediate files specifically includes:

[0109] Obtain the partition number of all the partition data, and classify all the partition data according to the partition number to obtain the classification result;

[0110] Based on the classification results, partition data with the same partition number are merged to obtain multiple target intermediate files.

[0111] Specifically, the step of calculating the data size of the target partition data for each target intermediate file to obtain the target partition calculation result includes:

[0112] The data size of the target partition data of each target intermediate file is calculated to obtain the data value of each target intermediate file, and the total data value of all target partition data is calculated.

[0113] The proportion of each data value is calculated based on the total data value to obtain the partition data proportion value of each target intermediate file, and the target partition calculation result is obtained based on the proportion values ​​of all partition data.

[0114] The step of reordering all the target intermediate files based on the target partition calculation results to obtain the final intermediate file of the node specifically includes:

[0115] Set a percentage threshold, and compare the percentage values ​​of all partition data in the target partition calculation result with the percentage threshold to obtain the comparison result;

[0116] Sort all the data percentage values ​​of the first partition in the comparison results by size to obtain the partition data sorting result, and set all the data percentage values ​​of the second partition in the comparison results to a preset position in the partition data sorting result to obtain the target partition data sorting result;

[0117] Based on the sorting results of the target partition data, all the target intermediate files are merged to obtain the final intermediate file of the node;

[0118] Wherein, the first partition data percentage is the partition data percentage greater than or equal to the percentage threshold, and the second partition data percentage is the partition data percentage less than the percentage threshold.

[0119] Specifically, setting the percentage values ​​of all second partition data in the comparison result to a preset position in the partition data sorting result includes:

[0120] If there is a partition data percentage value in the second partition that is less than the percentage threshold, then the second partition data percentage value is set in the preset position of the partition data sorting result;

[0121] If there are multiple partition data percentage values ​​less than the percentage threshold for the second partition data percentage value, then obtain the partition number order of all target intermediate files corresponding to the second partition data percentage value, and set all the second partition data percentage values ​​in the preset position of the partition data sorting result according to the partition number order.

[0122] The step of obtaining the final intermediate file for each node and aggregating all the final intermediate files for all nodes to obtain the data aggregation optimization result specifically includes:

[0123] Obtain the final intermediate file for each node, and extract the final partition data for each partition of the final intermediate files for all nodes;

[0124] Obtain the final partition number of all final partition data, and aggregate the final partition data with the same final partition number to obtain the data aggregation optimization result. The aggregation process includes summation, averaging and maximum value processing.

[0125] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a data aggregation optimization program based on partitioned data reordering, and the data aggregation optimization program based on partitioned data reordering, when executed by a processor, implements the steps of the data aggregation optimization method based on partitioned data reordering as described above.

[0126] In summary, this invention provides a data aggregation optimization method, system, terminal, and storage medium based on partitioned data reordering. The method includes: when different nodes are executing mapping tasks, obtaining all intermediate files of the node according to data aggregation optimization instructions; partitioning each intermediate file to obtain multiple partition data; merging all partition data to obtain multiple target intermediate files; calculating the data size of the target partition data in each target intermediate file to obtain the target partition calculation result; reordering all target intermediate files according to the target partition calculation result to obtain the final intermediate file of the node; obtaining the final intermediate file of each node; and aggregating all the final intermediate files of the nodes to obtain the data aggregation optimization result. This invention merges the generated intermediate files in advance during the mapping stage and sorts them according to the proportion of each partition data, thereby reducing the number of times each node accesses the intermediate files during the data aggregation stage, thus reducing the communication overhead generated by data aggregation and improving the efficiency of Spark distributed computing.

[0127] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0128] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0129] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A data aggregation optimization method based on partitioned data reordering, characterized in that, The data aggregation optimization method based on partitioned data reordering includes: When different nodes execute mapping tasks, all intermediate files of the node are obtained according to the data aggregation optimization instructions. Each intermediate file is partitioned to obtain multiple partition data. All partition data are then merged to obtain multiple target intermediate files. The data size of the target partition data of each target intermediate file is calculated to obtain the target partition calculation result. Based on the target partition calculation result, all the target intermediate files are reordered to obtain the final intermediate file of the node. The step of calculating the data size of the target partition data for each of the target intermediate files to obtain the target partition calculation result specifically includes: The data size of the target partition data of each target intermediate file is calculated to obtain the data value of each target intermediate file, and the total data value of all target partition data is calculated. The proportion of each data value is calculated based on the total data value to obtain the partition data proportion value of each target intermediate file, and the target partition calculation result is obtained based on the proportion values ​​of all partition data. The step of reordering all the target intermediate files based on the target partition calculation results to obtain the final intermediate file of the node specifically includes: Set a percentage threshold, compare the percentage values ​​of all partition data in the target partition calculation result with the percentage threshold, and obtain the comparison result; Sort all the data percentage values ​​of the first partition in the comparison results by size to obtain the partition data sorting result, and set all the data percentage values ​​of the second partition in the comparison results to a preset position in the partition data sorting result to obtain the target partition data sorting result; Based on the sorting results of the target partition data, all the target intermediate files are merged to obtain the final intermediate file of the node; Wherein, the first partition data percentage is a partition data percentage greater than or equal to the percentage threshold, and the second partition data percentage is a partition data percentage less than the percentage threshold; Obtain the final intermediate file for each node, and aggregate all the final intermediate files of the nodes to obtain the data aggregation and optimization result.

2. The data aggregation optimization method based on partitioned data reordering according to claim 1, characterized in that, When different nodes execute mapping tasks, all intermediate files of the nodes are obtained according to data aggregation optimization instructions. Each intermediate file is then partitioned to obtain multiple partition data, specifically including: When different nodes execute mapping tasks, the intermediate files and index files corresponding to all mapping tasks of the node are obtained according to the data aggregation optimization instructions; The corresponding partition ID is obtained from the index file, and all the intermediate files are partitioned according to the partition ID to obtain multiple partition data.

3. The data aggregation optimization method based on partitioned data reordering according to claim 1, characterized in that, The step of merging all the partitioned data to obtain multiple target intermediate files specifically includes: Obtain the partition number of all the partition data, and classify all the partition data according to the partition number to obtain the classification result; Based on the classification results, partition data with the same partition number are merged to obtain multiple target intermediate files.

4. The data aggregation optimization method based on partitioned data reordering according to claim 1, characterized in that, Setting the percentage values ​​of all second partition data in the comparison result to a preset position in the partition data sorting result specifically includes: If there is a partition data percentage value in the second partition that is less than the percentage threshold, then the second partition data percentage value is set in the preset position of the partition data sorting result; If there are multiple partition data percentage values ​​less than the percentage threshold for the second partition data percentage value, then obtain the partition number order of all target intermediate files corresponding to the second partition data percentage value, and set all the second partition data percentage values ​​in the preset position of the partition data sorting result according to the partition number order.

5. The data aggregation optimization method based on partitioned data reordering according to claim 1, characterized in that, The process of obtaining the final intermediate file for each node and aggregating all the final intermediate files for all nodes to obtain the data aggregation and optimization result specifically includes: Obtain the final intermediate file for each node, and extract the final partition data for each partition of the final intermediate files for all nodes; Obtain the final partition number of all final partition data, and aggregate the final partition data with the same final partition number to obtain the data aggregation optimization result. The aggregation process includes summation, averaging and maximum value processing.

6. A data aggregation optimization system based on partitioned data reordering, characterized in that, The data aggregation optimization system based on partitioned data reordering is applied to the data aggregation optimization method based on partitioned data reordering according to any one of claims 1-5, wherein the data aggregation optimization system based on partitioned data reordering includes: The data merging module is used to obtain all intermediate files of the nodes according to the data aggregation optimization instructions when different nodes are performing mapping tasks, perform partitioning processing on each intermediate file to obtain multiple partition data, and merge all the partition data to obtain multiple target intermediate files. The data sorting module is used to calculate the data size of the target partition data of each target intermediate file, obtain the target partition calculation result, and re-sort all the target intermediate files according to the target partition calculation result to obtain the final intermediate file of the node. The data aggregation module is used to obtain the final intermediate file of each node, aggregate the final intermediate files of all nodes, and obtain the data aggregation optimization result.

7. A terminal, characterized in that, The terminal includes: a memory, a processor, and a data aggregation optimization program based on partitioned data reordering stored in the memory and executable on the processor. When the data aggregation optimization program based on partitioned data reordering is executed by the processor, it implements the steps of the data aggregation optimization method based on partitioned data reordering as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a data aggregation optimization program based on partitioned data reordering, which, when executed by a processor, implements the steps of the data aggregation optimization method based on partitioned data reordering as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN114237510A

  • MapReduce optimization for partitioned intermediate output

    US10574508B1