Multi-stage MapReduce operation performance joint tuning method and system

By adopting the multi-stage MapReduce job performance joint tuning method on the GPU cluster, using cached redundant data and distributed computing and coding technology, the problem of insufficient processing capabilities of multi-stage MapReduce jobs on the GPU cluster is solved, and efficient massive data processing is achieved.

CN120144488APending Publication Date: 2025-06-13STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510212550.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to implement efficient multi-stage MapReduce job processing on a GPU cluster composed of a large number of nodes, especially in the case of massive data.

Method used

The multi-stage MapReduce job performance joint tuning method is adopted to cache redundant data, perform distributed computing and encoding data, and send encoded data in broadcast mode, reducing the communication overhead during the data reshuffle process.

Benefits of technology

It greatly shortens the completion time of multi-stage MapReduce jobs, optimizes the operation performance of MapReduce, and improves the system's fault tolerance and data transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144488A_ABST
    Figure CN120144488A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-stage MapReduce job performance joint adjusting and optimizing method and system, and belongs to the technical field of big data processing. Input data are read in from a distributed file system, Map function operation is executed, generated data are collected, overflow merging is carried out after cache overflow, and shuffle file fragments are obtained; performing intermediate data transmission on the shuffle file fragments; and merging the intermediate data, executing Reduce function operation, and finally writing a result obtained by the Reduce function operation into the distributed file system. The communication overhead in the data shuffling process can be greatly reduced by caching redundant data, carrying out distributed calculation coding on the data and sending the coded data in a broadcast mode. Therefore, the distributed computing coding and the redundant computing Reduce function operation are combined, so that the completion time of the multi-stage MapReduce operation is greatly shortened, and the MapReduce operation performance is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] MapReduce has become a popular distributed computing model. MapReduce is a programming model proposed by Google to address data-intensive problems and is used for parallel computing of large-scale data sets. When dealing with big data, MapReduce only requires users to implement the Map function and the Reduce function, and leaves distributed file management, job scheduling, fault tolerance, and inter-machine communication to the MapReduce system, which makes MapReduce increasingly become the mainstream programming model and distributed computing framework of cloud computing platforms.

[0003] Hadoop divides the input data set of a MapReduce job into several independent data blocks, and the Map tasks of the MapReduce job process these data blocks in parallel. After the MapReduce framework performs a "shuffle" operation (including partitioning, sorting, etc.) on the output of Map, the Reduce tasks of the MapReduce job pull and merge the outputs of the Map tasks from each data node. The input and output files of a MapReduce job are generally stored in a distributed file system (Hadoop Distributed FileSystem, HDFS). The MapReduce framework is responsible for task scheduling and monitoring, as well as re-executing failed tasks.

[0004] Hadoop is an open-source implementation version of MapReduce. Another important part of Hadoop is the distributed file system HDFS, which stores data sets distributively. The Hadoop platform is multi-user oriented, and each legitimate user can submit jobs to the Hadoop platform. How to quickly respond to user requests, how to improve the utilization rate of the cluster, and how to optimize the execution efficiency of MapReduce jobs have become current research hotspots. Among them, the optimization of the global completion time for a set of MapReduce jobs is one of the research hotspots. For a set of MapReduce jobs, by sharing resources and scheduling execution among jobs, the global completion time of the job set can be optimized, thereby making the user's response time shorter. Moreover, the optimized scheduling and resource sharing of multi-stage MapReduce jobs can make full use of cluster resources and improve the utilization rate of the cluster.

[0005] With the development of heterogeneous computing technology, more and more applications have begun to attempt to accelerate data processing by leveraging heterogeneous processors. Heterogeneous computing refers to enhancing computing performance by utilizing new types of hardware coprocessors on traditional computing platforms based on the Central Processing Unit (CPU). Some dedicated processors, such as the Graphics Processing Unit (GPU), Field-Programmable Gate Array (FPGA), and Cell processor, etc., feature high computing density, low power consumption, and high performance. Combining them with the multifunctional CPU can share the pressure on the CPU during high-intensity computing and make up for the deficiencies in the CPU's computing power. The huge advantages of heterogeneous architectures in High Performance Computing (HPC) have already emerged. In 2010, China's "Tianhe-1" high-performance computer ranked first on the global supercomputer Top500 list, and its unique CPU and GPU collaborative heterogeneous computing structure contributed significantly. The computing performance of Titian at the Oak Ridge National Laboratory in the United States has reached 17.59 pflops, and the vast majority of the computing power comes from the 18,688 GPU accelerators deployed on the cluster. The successful application of heterogeneous computing in the field of scientific computing has prompted the consideration of whether the parallel computing characteristics of GPUs can be applied to the commercial big data computing field.

[0006] The existing optimization of data center network transmission for MapReduce tasks mainly focuses on two aspects: traffic scheduling at the network level and task placement and scheduling at the application level. However, with the explosive growth of network scale, users' requirements for service quality are also increasing. This rapid development trend highlights the important position of data centers as information service infrastructures on the one hand, and exposes many difficulties in optimizing data center network resources under new applications and new computing models on the other hand.

[0007] In addition, when dealing with complex business processes, it is generally difficult to complete them with a single MapReduce job. Usually, multiple MapReduce jobs are required to complete the task. Since the output of the previous job is the input of the next job, the input / output (I / O) consumption, disk read / write, and data network transmission consumption between jobs are very large. Although the multi-stage MapReduce framework can integrate multiple logical functions into a single MapReduce job, if these logical operations are only executed sequentially without considering their characteristics, the performance of the job cannot reach the optimal level. Therefore, a reasonable execution plan can reduce the number of jobs, reduce the amount of data transferred from upper-level logical operations to lower-level logical operations, and avoid the output of unnecessary intermediate results, thereby improving the overall performance of the process.

[0008] Although many studies on MapReduce models on GPUs have emerged before, most of them are based on single-GPU implementations, or are extended on different types of GPUs, or are optimized using new GPU features. They do not truly utilize the advantages of large-scale clusters to improve the computing performance of MapReduce, nor do they consider the processing ability in the case of massive data. Therefore, it is necessary to implement a MapReduce model on a GPU cluster composed of a large number of nodes that can provide real-time processing of large-scale data. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to provide a method and system for jointly optimizing the performance of multi-stage MapReduce jobs to solve the technical problems of data transmission time affecting job completion time and data center network performance in view of the above-mentioned deficiencies in the prior art.

[0010] The present invention adopts the following technical solutions:

[0011] A method for jointly optimizing the performance of multi-stage MapReduce jobs includes the following steps:

[0012] S1. Read input data from a distributed file system, execute Map function operations, collect the data generated by the Map function operations, and after cache spilling and merging of the generated data, obtain shuffle file segments;

[0013] S2. Perform intermediate data transmission on the shuffle file segments;

[0014] S3. Merge the received intermediate data files, execute Reducer function operations, and write the obtained results into a distributed file system.

[0015] Preferably, in step S1, the cache overwriting specifically is as follows:

[0016] Before performing cache overwriting, sort and pre-join the intermediate output data, as well as perform data compression. When the cache occupancy reaches the threshold, replace the data in the cache onto the hard disk to achieve cache overwriting. During multiple cache overwriting processes, multiple overwritten files appear on the disk.

[0017] Preferably, after the cache overwriting is completed, merge all the overwritten files into one Map function output data file.

[0018] Preferably, in step S1, the operation of the Map function specifically is as follows:

[0019] The operation of the Map function executes the Map function, stores the generated data in the memory buffer. Before storing it in the memory buffer, first divide the data until the threshold of the buffer capacity is reached.

[0020] Preferably, after obtaining the Map function output data file, each node in the distributed file system obtains Q Map function output data files after the Map function operation on the locally existing files.

[0021] Preferably, in step S3, the operation of the Reducer function grabs the processed intermediate data from each terminal and stores the data in the memory buffer after decompression.

[0022] Preferably, step S3 specifically is as follows:

[0023] When the memory buffer reaches the set threshold, overflow the intermediate data onto the hard disk, and perform a preliminary merge on the shuffle file segments in the memory and on the hard disk to merge them into an independent shuffle file as the input data for the Reduce function operation;

[0024] Process the input data according to the written reduce function and store the obtained result in the memory;

[0025] Then write out the result obtained from the operation of the Reduce function to the distributed file system.

[0026] Preferably, processing the input data according to the written reduce function specifically is as follows:

[0027] There are K distributed computing nodes in the distributed file system. Node k executes the operation of the Reduce function, k ∈ {1, …, K}, and obtains the output result u q .

[0028] Preferably, the output result u q Specifically is as follows:

[0029] First, obtain all N intermediate values v of the Map functions required through Shuffle data exchange q,1 ,…,v q,N , and then calculate the output result u q = h q (v q,1 ,…,v q,N ), where h q is the Reduce function

[0030] In a second aspect, an embodiment of the present invention provides a multi-stage MapReduce job performance joint tuning system, including:

[0031] A reading module that reads input data from a distributed file system, performs Map function operations, collects the data generated by the Map function operations, and after cache spilling and merging the generated data, obtains shuffle file segments;

[0032] A transmission module that performs intermediate data transmission on the shuffle file segments;

[0033] A tuning module that merges the received intermediate data files, performs Reducer function operations, and writes the obtained results into the distributed file system.

[0034] In a third aspect, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above multi-stage MapReduce job performance joint tuning method.

[0035] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium including a computer program. When the computer program is executed by a processor, it implements the steps of the above multi-stage MapReduce job performance joint tuning method.

[0036] Compared with the prior art, the present invention has at least the following beneficial effects:

[0037] A method for jointly optimizing the performance of multi-stage MapReduce jobs. The calculation process consists of multiple calculation stages and network transmission stages. The output of the previous round of Reduce results will be used as the input for the next round of Map calculations. Calculating each Reduce function on multiple nodes provides data redundancy for the subsequent Map function calculations, which helps improve the fault tolerance of the system and reduce the communication load of the next round of data shuffling. In addition, by caching redundant data, performing distributed computing encoding on the data, and sending the encoded data in a broadcast manner, the communication overhead during the data shuffling process can be greatly reduced. Therefore, this project combines distributed computing encoding with redundant computing Reduce functions, thereby greatly shortening the completion time of multi-stage MapReduce jobs and optimizing the running performance of MapReduce.

[0038] Furthermore, by combining Reduce redundant calculation with distributed computing encoding, the communication transmission performance in multi-stage MapReduce is optimized to achieve the goal of shortening the overall execution time of multi-stage jobs.

[0039] Furthermore, when multiple consecutive operations need to be processed, the overall performance of the job process can be improved by reducing the number of generated MapReduce jobs and thus reducing the number of disk reads and writes.

[0040] Furthermore, the function of the distributed file system is to store data (usually a single data file exceeds 10GB) dispersedly on each storage node. Such a design has many benefits: First, dispersing the data on each node can speed up the combined data reading speed; Second, since the MapReduce model adopts a strategy of "computing approaching storage and merging computing and storage nodes" in its design, using distributed storage can take advantage of data locality, reduce network I / O and latency, and improve the processing speed.

[0041] It can be understood that the beneficial effects of the second aspect above can refer to the relevant descriptions in the first aspect above and will not be elaborated here.

[0042] In summary, the present invention jointly optimizes the MapReduce performance by adjusting various parameters of the cluster through a combination of simulation operation and correction optimization, so that a multi-stage GPU-MapReduce program can exert its maximum performance on the cluster and achieve fast and efficient processing of massive data.

[0043] Next, through the accompanying drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings

[0044] Figure 1 It is a schematic flowchart of the present invention;

[0045] Figure 2 It is a schematic diagram of the Apache Hadoop architecture;

[0046] Figure 3 shows the uncoded strategy and the coded strategy in the Shuffle stage. Among them, (a) shows the communication overhead in the Shuffle stage of MapReduce in the uncoded case, and (b) shows redundant storage of files;

[0047] Figure 4 It is a comparison diagram of the communication overhead between the coded strategy and the uncoded strategy;

[0048] Figure 5 It is a flowchart of the GPU-MapReduce job running on the cluster;

[0049] Figure 6 It is the execution process of the Map function and the Reduce function;

[0050] Figure 7 It is a schematic diagram of the computer device provided by an embodiment of the present invention.

[0051] Figure 8 It is a block diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0053] In the description of the present invention, it should be understood that the terms "including" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0054] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0055] It should be further understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, in the present invention, the character " / " generally indicates an "or" relationship between the preceding and following related objects.

[0056] It should be understood that although terms such as first, second, third, etc. may be used in the embodiments of the present invention to describe preset ranges and the like, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0057] Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".

[0058] Various structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures and their relative sizes and positional relationships are only exemplary, and in practice, there may be deviations due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0059] The present invention provides a method for jointly optimizing the performance of multi-stage MapReduce jobs. By caching redundant data, encoding the data, and sending the encoded data in a broadcast manner, the communication overhead in MapReduce network transmission can be greatly reduced. In the calculation application of multi-stage MapReduce jobs, the calculation process consists of multiple calculation stages and communication stages. Applying encoding methods to optimize the calculation time or communication load for each calculation stage or communication stage respectively is a feasible solution, but this solution idea ignores the association between consecutive stages and will lead to a decline in overall performance. Therefore, the present invention constructs a multi-stage joint encoding method by utilizing the data flow association characteristics between different stages. On this basis, considering the data flow between different stages, it studies how to select calculation nodes for deploying calculation tasks so as to process data as close as possible to reduce the data transmission overhead.

[0060] MapReduce is a coarse-grained parallelism: large computing tasks are distributed to each computing node. While GPU is a fine-grained parallelism: computing tasks are distributed to numerous computing cores of the processor and rely on multi-threading to complete. Although the parallelisms of the two are one macroscopic and the other microscopic, it can be found that they have an implicit common point: both adopt a way similar to Single Instruction Multiple Data (SIMD) to run. When the Map task runs on each node, the map function uses the same processing program for each piece of data in the data set and processes one piece at a time. Therefore, it is entirely possible to parallelize this process at a finer granularity and map it to the multi-threading of the GPU to complete. In addition, there are a large number of tuning parameters in the MapReduce framework, and many of them are directly related to the running efficiency of the job. Study how to tune the parameters to make the MapReduce program perform at its best on the cluster as much as possible.

[0061] At the highest level of abstraction, the programming model of MapReduce only contains two programming interfaces: the two functions map(key1, valuel) and reduce(key2, list(value2)). Programmers can implement the business logic (data processing algorithms) they need by writing these two functions. All input and output data of MapReduce are organized in the form of "key-value pairs" (<key, value>). The user-defined map function running on the Map nodes (Mappers) processes each "key-value pair" data <K1, V1> in the input data set and generates multiple "key-value pair" form results <K2, V2>, which are called intermediate results. The intermediate results of each Mapper will be transmitted to the Reduce nodes (Reducers) through the network, and this process is vividly called "Shuffle". Similarly, the user-defined reduce function processes the results <K2, list(V2)> with the same key value and outputs the final result <K3, V3>.

[0062] Embodiment 1

[0063] Please refer to Figure 1 , a method for jointly tuning the performance of a multi-stage MapReduce job according to the present invention, includes the following steps:

[0064] S1. Read the input data from the distributed file system, execute the Map function operation, collect the data generated by the Map function operation, and perform spill merging after the generated data is spilled and written to obtain the Map function output data file;

[0065] The Map phase includes the following processes

[0066] Node k, k ∈ {1, …, K} calculates the set of input files stored on itself For the set of input files M k For each input file ω n in it, node k executes the Map function to calculate Q intermediate Map values, that is To obtain the complete output result value, each input file is stored on at least one node

[0067] S101. Data Reading (Read):

[0068] Read the input data from the distributed file system (DFS), and each "key-value pair" is used as a record of an input data record

[0069] S102. Execute Map (Map):

[0070] Execute the map function written by the user to generate intermediate output data

[0071] S103. Data Collection (Collect):

[0072] Store the data generated by map into the memory buffer. Before entering the buffer, the data will be partitioned until it reaches the threshold of the buffer capacity

[0073] S104. Cache Spill (Spill):

[0074] When the cache occupancy reaches or is about to reach the threshold, it is necessary to replace the data in the cache to the hard disk. This process is called spill. Before spill, the data will be sorted (Sort) and pre-joined (Combine) similar to reduce, as well as data compression (Compression). During multiple spill processes, multiple spill files will appear on the disk

[0075] S105. Spill Merge (Spill-Merge):

[0076] After the cache spill is completed, it is necessary to merge all the spill files into a single map output data file

[0077] Map function operation

[0078] In a MapReduce cluster consisting of K distributed computing nodes (servers), N input files need to be calculated through any Q output functions to obtain output results, where And N≥K.

[0079] Suppose there are N input files ω 1 ,ω 2 ,…,ω N , for Q output functions φ 1 ,φ 2 ,…,φ Q Then the output result is u q =φq(v 1 ,ω 2 ,…,ω N ), where q∈{1,2,…,Q}. From the execution process of MapReduce, we can see that the calculation result of the output function is decomposed into the calculation of intermediate values ​​in the Map phase and the calculation of the output result by merging intermediate values ​​in the Reduce phase, that is, φ q (ω 1 ,ω 2 ,…,ω N )=h q (g q,1 (ω 1 ),…,g q,N (ωN)). Among them, g q,1 (ω 1 ),…,g q,N (ω N ) is the Map intermediate value calculation function g q,1 ,…,g q,N For the corresponding input files ω 1 ,ω 2 ,…,ω N The calculated intermediate value; in the Reduce stage, all the N intermediate values ​​merged are calculated by the Reduce function hqq to obtain the output result. The execution process of the Map function and the Reduce function is as follows Figure 6 shown.

[0080] Map function operation Responsible for calculating the input file ω n ; It will input the file ω n Calculate Q Maps to calculate the intermediate value v 1,n =g 1,n (ω n ),…,v Q,n =g Q,n (ω n ),n∈{1,2,…,N}.

[0081] After the Map phase, each node has obtained the intermediate values of the Q output functions calculated from the local existing files. Therefore, for any symmetric assignment of a fixed file distribution and the Reduce function, although the nodes need to relabel the Reduce functions they are responsible for, the data transfer volume during the Shuffle process remains unchanged. In other words, the communication load is independent of the assignment of the Reduce function.

[0082] Next, the MapReduce framework is generalized, and a cascaded MapReduce computing framework is proposed. That is, after the Map phase, each Reduce function is calculated by s nodes, where s ∈ {1, 2, …, K}. Since many MapReduce jobs require multiple rounds of Map and Reduce calculations, where the Reduce results of the previous round are used as the input for the calculation of the Map function in the next round, and calculating each Reduce function on multiple nodes can provide data redundancy for the calculation of the subsequent Map function, which helps to improve fault tolerance and reduce the communication load for the next round of data transformation. Only considering the case where Q / K ∈ N enables the symmetric assignment of Reduce tasks to maintain load balance.

[0083] The computation-communication function of the cascaded MapReduce computing framework is defined as:

[0084]

[0085] where r is the computation overhead, L is the communication overhead, s is the node, and s ∈ {1, 2, …, K}.

[0086] S2. Perform intermediate data transfer for the shuffle file segments;

[0087] The Shuffle phase includes the following processes:

[0088] Assign the nodes k, k ∈ {1, …, K} to be responsible for the output results of the set of output functions To reduce the network load without considering the cascading process, it is required that each output function is only responsible for one node, that is and

[0089] To calculate the output result u q of each output function φ q in the output function set Wk, the node k needs all N Map intermediate values v q,1 , …, v q,N , where some of the intermediate values are known through the calculation in its own Map phase, and the remaining intermediate values need to be obtained through network transmission from other nodes.

[0090] For each node k, the locally known intermediate values can be encoded to obtain X k , and the encoding function is denoted as ψ k , that is After that, node k broadcasts X k to all other nodes.

[0091] S201. Intermediate data transfer (Copy):

[0092] The Reducer fetches the processed intermediate data from each end, decompresses it, and the data will be stored in the in-memory shuffle buffer. When the buffer reaches a certain threshold, it will also be spilled to the hard disk. Therefore, at the end of the Copy phase, there are shuffle file fragments on both the memory and the hard disk;

[0093] S202. Shuffle file merging (Pre-Merge):

[0094] Before the start of the reduce phase, the shuffle file fragments in the memory and on the hard disk are preliminarily merged;

[0095] S3. Merge the received intermediate data files, perform the Reducer function operation, and write the obtained results into the distributed file system.

[0096] The Reduce phase includes the following processes:

[0097] Node k, k ∈ {1, …, K} is responsible for the output results of the output function set where |W k | = Q / K. After the Shuffle phase, without considering packet loss, node k has received X 1 , X 2 , …, X K , and then combined with the locally known intermediate values, decode to obtain all N Map intermediate values v k in each output function φ q in the output function set W q,1 , …, v q,N ; the decoding function is denoted as ξ k q , that is Finally, node k executes the Reduce function to calculate the output result u q = h q (v q,1 , …, v q,N ), φ q ∈ W k .

[0098] The Reduce function h qResponsible for counting the output function φ q The calculation result of

[0099] First, all N Map intermediate values v required need to be obtained through Shuffle data exchange q,1 ,…,v q,N , and then calculate the output result u q = h q (v q,1 ,…,v q,N ), q ∈ {1, 2, …, Q}.

[0100] S301. Intermediate data merging (Shuffle-Merge):

[0101] Completely merge the intermediate data fetched from each source into an independent shuffle file as the input data for the reduce function;

[0102] S302. Execute Reduce (Reduce):

[0103] Process the data obtained by the Reducer according to the reduce function written by the user, and store the obtained result in memory;

[0104] S303. Data writing (Write):

[0105] Write the result obtained by the reduce process to the distributed file system.

[0106] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, method, or program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "platform" here.

[0107] Embodiment 2

[0108] In another embodiment of the present invention, a multi-stage MapReduce job performance joint tuning system is provided, which can be used to implement the above multi-stage MapReduce job performance joint tuning method. Specifically, the multi-stage MapReduce job performance joint tuning system includes a reading module, a transmission module, and a tuning module.

[0109] Among them, the reading module reads the input data from the distributed file system, executes the Map function operation, collects the data generated by the Map function operation, and performs spill merging on the generated data after cache spilling to obtain shuffle file fragments;

[0110] A transmission module for performing intermediate data transmission on shuffle file segments;

[0111] An optimization module for merging the received intermediate data files, performing Reducer function operations, and writing the obtained results into a distributed file system.

[0112] Embodiment 3

[0113] The present invention provides a terminal device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Graphics Processing Units (GPU), Tensor Processing Units (TPU), Digital Signal Processors (DSP), Application Specific Integrated Circuits (ASIC), Field-Programmable Gate Arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the multi-stage MapReduce job performance joint optimization method, including:

[0114] Reading input data from a distributed file system, performing Map function operations, collecting the data generated by the Map function operations, performing spill merging after cache spilling on the generated data to obtain shuffle file segments; performing intermediate data transmission on the shuffle file segments; merging the received intermediate data files, performing Reducer function operations, and writing the obtained results into a distributed file system.

[0115] Please refer to Figure 7, the terminal device is a computer device. The computer device 60 in this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the computer program 63 is executed by the processor 61, it implements the multi-stage MapReduce job performance joint tuning method in the embodiment. To avoid repetition, it will not be elaborated here one by one. Alternatively, when the computer program 63 is executed by the processor 61, it implements the functions of each model / unit in the multi-stage MapReduce job performance joint tuning system in the embodiment. To avoid repetition, it will not be elaborated here one by one.

[0116] The computer device 60 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art can understand that Figure 8 merely examples of the computer device 60, which do not constitute a limitation on the computer device 60, may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0117] The so-called processor 61 may be a central processing unit (CPU), or may also be other general-purpose processors, graphics processing units (GPU), tensor processing units (TPU), digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0118] The memory 62 may be an internal storage unit of the computer device 60, such as the hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk equipped on the computer device 60, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0119] Further, the memory 62 may also include both the internal storage unit of the computer device 60 and external storage devices. The memory 62 is used to store computer programs as well as other programs and data required by the computer device. The memory 62 may also be used to temporarily store data that has been output or is to be output..

[0120] Please refer to Figure 8 , the terminal device 600 is an electronic device, and the electronic device is presented in the form of a general computing device. The components of the electronic device may include but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0121] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the method section of this specification above. For example, the processing unit 610 can execute as Figure 1 the steps shown in

[0122] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.

[0123] The storage unit 620 may also include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0124] The bus 630 may represent one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any one of the multiple bus structures.

[0125] The electronic device 600 may also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 650. Moreover, the electronic device 600 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 660. The network adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0126] Embodiment 4

[0127] The present invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in the terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here may include both the built-in storage medium in the terminal device, and of course may also include the extended storage medium supported by the terminal device. It may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space, and these instructions may be one or more computer programs (including program codes). It should be noted that more specific examples of the computer-readable storage medium here include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0128] The computer-readable storage medium also includes a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, radio frequency, etc., or any suitable combination of the above.

[0129] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network or a wide area network, or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0130] One or more instructions stored in the computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the method for jointly optimizing the performance of a multi-stage MapReduce job in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor in the following steps:

[0131] Read the input data from the distributed file system, perform the Map function operation, collect the data generated by the Map function operation, perform spill merge after the generated data is spilled to disk, and obtain the shuffle file segment; perform intermediate data transmission on the shuffle file segment; merge the received intermediate data files, perform the Reducer function operation, and write the obtained result into the distributed file system.

[0132] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0133] There are various specific implementation frameworks for MapReduce. Well-known ones include the MapReduce system of Google and some open-source implementation forms such as Hadoop, Phinix, Mars, etc. Our GPU-MapReduce system is built on the open-source Hadoop project. Therefore, the Apache Hadoop framework will be introduced in detail here.

[0134] As the original creator of MapReduce, Google published its distributed file system, MapReduce computing framework, and the distributed database built on it in the form of a paper. Apache Hadoop implemented the distributed file system and MapReduce parallel computing framework based on this paper.

[0135] The two basic components of the Hadoop framework are the distributed file system (HDFS) and the parallel computing framework MapReduce. HDFS is responsible for large-scale distributed data storage and provides read and write operation interfaces. MapReduce is responsible for processing the data in HDFS.

[0136] Please refer to Figure 2 , HDFS is a distributed file system, and its prototype is GFS of Google. A distributed file system refers to connecting the storage resources on each physical node through a network, dispersing the data storage in each physical node, and forming a unified file directory through metadata management on it. The distributed file system has the same functions logically as the file system on a single-machine operating system, but the actual storage capacity is much larger than that of the file system on a single machine.

[0137] There are various implementation forms of distributed file systems, such as NFS (Network File System) of Sun Microsystems and Global File System of Red Hat. These file systems require reliability support from underlying hardware, and the fault tolerance and backup mechanisms of data are guaranteed by hardware. Therefore, they are built on expensive hardware infrastructures. However, the application scenarios of GFS and HDFS are very different:

[0138] 1. The underlying hardware is relatively inexpensive but less reliable consumer-grade storage hardware, and the probability of hardware failure is very high when running in a 7×24 environment. Therefore, it is necessary to rely on software redundancy mechanisms to ensure the reliability of data storage;

[0139] 2. The stored files are all very large files, usually exceeding 10GB or even hundreds of TB for a single file. Such large-scale files cannot be borne by a single storage device. Therefore, the files must be stored dispersedly on multiple storage devices.

[0140] The distributed file system is the basis for the "data locality" and "computation approaching data" strategies in the MapReduce computing model. In the MapReduce framework of Hadoop, there are also two types of nodes. One is called "JobTracker", and the other is called "TaskTracker". The JobTracker is the scheduler for MapReduce jobs on the entire cluster. It is responsible for monitoring the running progress and exceptions of each task, as well as task allocation. The TaskTracker node is each unit where MapReduce tasks are carried out. The TaskTracker fetches data from the distributed file system and performs calculations according to the requirements of the map or reduce function. Usually in a MapReduce cluster, the datanode node is also the tasktracker at the same time. Therefore, the tasktracker usually only performs calculations on local data and has no data communication with other nodes during the calculation process, but only reports the running status of the task to the TaskTracker. This is the data locality strategy. Since only the running status is communicated instead of a large amount of data being transferred, the goal of "computation approaching data" is achieved.

[0141] For this scenario in Hadoop, the present invention provides three solutions: iterative MapReduce, dependency-combined MapReduce, and chained MapReduce, which are specifically as follows:

[0142] The iterative MapReduce is specifically as follows:

[0143] Algorithms such as PageRank and K-means require multiple iterations and can only be completed by multiple MapReduce jobs. During the map / reduce iteration process, a while loop is used to manage the iteration, which ends until a certain judgment condition is met; the output result of the previous MapReduce during the iteration process is used as the input of the next MapReduce.

[0144] Iterative MapReduce has a major drawback: in each iteration, all jobs (tasks) are recreated, which incurs a very high cost; in each iteration, data is written to and read from local storage, resulting in relatively high I / O and network transmission costs. In response to the characteristics of iterative computing, there have already been some open-source projects that have made improvements in this area, modifying traditional MapReduce to adapt to iterative computing, such as Haloop and Twister.

[0145] Characteristics of iterative computing: The input data can be divided into two categories, namely static and variable data, and in most cases, the static data is much larger than the variable data. After the static and variable data are computed, new variable data is obtained, and then the new data is recomputed with the static data, and so on in a loop until the iteration end condition is met and the task ends.

[0146] In response to its characteristics, Haloop and Twister have optimized the MapReduce framework: modifying the original job structure to enable it to run multiple map-reduce pairs, so that there is no need to start multiple jobs (the loop body of iterative computing occurs within the job, that is, job reuse), and job reuse can greatly improve the iteration efficiency; map / reduce static data caching (resident in memory or cached to disk) allows map tasks to quickly enter the computing process.

[0147] However, the model abstractions of Haloop and Twister are not high enough, and the supported computing models are limited. Iterative MapReduce is still in development.

[0148] The dependency relationship combined MapReduce is as follows:

[0149] For example, a task requires three MapReduce jobs: jobl, job2, and job3. Among them, jobl and job2 are independent of each other, and job3 can only be executed after jobl and job2 are completed. Such multiple jobs with dependencies are called dependent combined MapReduce. Hadoop provides an execution and control mechanism for this combined relationship, configuring the dependencies between jobs during job initialization. This method is mainly for controlling at the job flow level. For example: job3.addDepending(job1) establishes the dependency that job3 can only be executed after jobl is completed.

[0150] The traditional MapReduce framework divides a job into two stages: Map and Reduce. In the Map stage, each MapTask reads a Block of data (default configuration), calls the map() function for processing, and then writes the result to the local disk; in the Reduce stage, each ReduceTask remotely reads the data required at the Reduce end from the nodes where each MapTask is located, calls the reduce() function for data processing, and writes the final result to Hdfs. In a MapReduce job, disk I / O occurs twice, which can improve reliability but reduce system performance. For this reason, the traditional MapReduce framework is not very suitable for processing multiple consecutive jobs and iterative jobs. This type of job will repeatedly perform disk I / O under the MapReduce framework, resulting in poor overall performance of the job. When processing data, Hadoop tries to ensure data locality, that is, the data is stored on a certain node, and this node will perform the data calculation for this part, which can reduce the data transmission on the network, reduce the demand for network bandwidth, and speed up the processing speed; "local calculation" is the most effective means to save network bandwidth and shorten processing time. By default, each MapTask reads one Block by default, and the data corresponding to each Block is on the same node as the MapTask; if the default configuration is modified so that each MapTask can read multiple Blocks, these Blocks can be on any node in the cluster. At this time, it is necessary to copy the data that is not on the node where the MapTask is located to this node, which will affect the performance of the job. The network transmission generated by the Reduce end copying the data it needs from other nodes is inevitable. When multiple consecutive operations need to be processed, the overall performance of the job flow can be improved by reducing the number of generated MapReduce jobs, thereby reducing the number of disk I / O operations.

[0151] Chained MapReduce is as follows:

[0152] Chained MapReduce means that, outside the core MapReduce, an auxiliary map process is added, and then this auxiliary map process and the core MapReduce process are combined into a chained MapReduce to complete the entire job. Hadoop provides specialized ChainMapper and ChainReducer to handle chained tasks. ChainMapper allows multiple sub-map tasks to be added to a map task, and ChainReducer can add multiple sub-map tasks after the Reducer execution.

[0153] The output of the previous map serves as the input of the next map, and the output of the last map serves as the input of the reduce. The final result is output to HDFS by the reduce. However, in chained MapReduce, the logical operations performed in the mapper are still executed sequentially without considering the relationships between jobs, which still has a certain impact on performance.

[0154] The calculation based on MapReduce is encoded as follows:

[0155] Taking the specific example in Figure 3 as an example. Suppose there are 6 files and 3 nodes, and each file contains three colors: red, green, and blue. Now, all red colors are given to node 1, all green colors are given to node 2, and all blue colors are given to node 3. Assume that files 1 and 2 are stored in node 1, then node 1 also needs the red colors of files 3, 4, 5, and 6; files 3 and 4 are stored in node 2, then node 2 also needs the green colors of files 1, 2, 5, and 6; files 5 and 6 are stored in node 3, then node 3 also needs the blue colors of files 1, 2, 3, and 4.

[0156] Figure 3(a) shows the communication overhead in the Shuffle phase of MapReduce in the case of no encoding. Node 1 unicasts the green and blue colors stored in its files 1 and 2 to nodes 2 and 3 respectively; Node 2 unicasts the red and blue colors stored in its files 3 and 4 to nodes 1 and 3 respectively; Node 3 unicasts the red and green colors stored in its files 5 and 6 to nodes 1 and 2 respectively; then the task ends. The communication overhead of this process is 3 × 4 = 12, that is, each node unicasts 4 times, and there are 3 nodes.

[0157] The present invention contemplates redundant storage of files, i.e., each file is stored on multiple nodes respectively. As shown in Figure 3(b), node 1 stores files 1, 2, 3, 4, and it only lacks the red parts of 5 and 6. Node 2 and node 3 unicast the red parts of files 5 and 6 to it respectively; node 2 stores files 3, 4, 5, 6, and it only lacks the green parts of 1 and 2. Node 1 and node 3 unicast the green parts of files 1 and 2 to it respectively; node 3 stores files 1, 2, 5, 6, and it only lacks the blue parts of 3 and 4. Node 1 and node 3 unicast the blue parts of files 3 and 4 to it respectively; at this time, the task ends. The communication overhead of this process is 3×2 = 6, which is half lower than that without redundant storage.

[0158] Meanwhile, in the redundant storage method, it is found that the green part of file 1 and the red part of file 3 unicast by node 1 are respectively needed by one of node 2 and node 3, and the other is already available; the same is true for the observation of other nodes. Therefore, it is considered to multicast the data that needs to be unicast in each node through the encoding method of exclusive-or bits, and other nodes can decode the required data by combining the local data after receiving the encoded data, thus reducing the communication overhead. Through such an encoding method, it is found that the communication overhead is 3×1 = 3, which is half lower than that of the redundant storage method.

[0159] There are only two main steps in this whole process:

[0160] The first step is the distribution of map tasks, that is, the input file set is evenly distributed on different worker nodes in a specific way, and the total number of nodes where each input file is stored is the redundancy number;

[0161] The second step is how to perform data exchange in the Shuffle stage, and the algorithm specifies that mutual multicast transmission needs to be carried out within each set of worker nodes with a length of redundancy number + 1.

[0162] Figure 4 It is a comparison graph of the communication overhead between the encoding strategy and the non-encoding strategy. As Figure 4 can be seen, the computational load and the communication load are inversely proportional, that is, when the computational overhead increases by r times, the communication overhead in the encoding case is reduced by r times compared with that in the non-encoding case. Among them, the multiple by which the computational overhead increases is the redundancy number; when the redundancy number is r, the same multiple of computational overhead is correspondingly increased.

[0163] The GPU-MapReduce cluster consists of a control node and worker nodes. In a multi-machine cluster with a minimum configuration, it includes a master node (Master), and all the remaining nodes are computational nodes (Slave).

[0164] The Master nodes are not configured with GPU acceleration cards because they do not run computationally intensive tasks. However, the Master nodes maintain the metadata of the entire cluster file system and rely heavily on memory. Therefore, the Master nodes are configured with a large amount of memory.

[0165] The worker nodes are both storage nodes and computing nodes. Each worker node is equipped with a multi-core CPU and a GPU acceleration card. Each acceleration card has its own video memory. The acceleration card is connected to the host through the PCI-e bus, and the nodes are connected to each other through a network.

[0166] Data in the distributed file system is replicated multiple times (usually 3 times) and scattered across multiple different physical nodes because the underlying environment of the distributed file system consists of a large number of low-end storage hardware with poor reliability and a relatively high probability of node failure. Therefore, a redundancy mechanism must be adopted to address data loss caused by node failure.

[0167] On the distributed file system, a large file is split into multiple file blocks and placed on different nodes respectively. The default size of the data block is 64MB (this value can be set). For example, if the total size of a file is 10GB, it will be split into 160 blocks and distributed to 160 Map tasks for processing. The Map tasks read the data according to data splits and read one split each time.

[0168] The map function processes data according to "records". In the most commonly used data format, a record is a line of data in a split. By default, the map function processes one record at a time, continuously reads and executes until all the data on the node is processed.

[0169] MapReduce job execution process based on software and hardware cooperation

[0170] Please refer to Figure 5 , for a new GPU-MapReduce job to run in a production environment, it needs to go through 5 steps:

[0171] 1. Initialization;

[0172] 2. Simulation run, parameter collection;

[0173] 3. Parameter tuning;

[0174] 4. Cluster job deployment;

[0175] 5. Cluster startup, job execution.

[0176] Initialization section, submit a MapReduce job. On a GPU-MapReduce cluster, a well-written MapReduce job cannot be directly scheduled to run on the physical cluster. Instead, it will first be submitted to a Slave node for simulation. The purpose of doing this is, first of all, for security reasons: some MapReduce jobs are risky to the cluster. Usually, a physical cluster runs multiple MapReduce jobs simultaneously. The addition of certain jobs may cause the cluster load to soar instantly, thus affecting the normal operation of other jobs. The Master node will judge whether a job will have a significant impact on the cluster based on the results of the simulated run, so as to avoid a large jitter in the overall performance of the cluster.

[0177] The submitted MapReduce job will be simulated in a pseudo-distributed manner, and the open-source Ganglia monitoring system is used to monitor the performance of the cluster. It will continuously monitor the usage of CPU, memory, hard disk, etc. during the running of the job. In addition, the NVIDIA Visual Profiler tool is used to collect parameters for the GPU acceleration card of a single node. Generally speaking, since the amount of data for the simulated run is very small, just a very small subset of the total data, the time required for a single simulated run is very short, usually completing the entire simulated run and performance parameter collection work within 5 minutes.

[0178] After obtaining the performance parameters, estimate and set the optimization parameters, and distribute them to each node of the cluster in the form of a configuration file through the network and take effect in the cluster. After the job configuration is completed, start the job on the entire cluster to complete the computing task.

[0179] Since there are certain errors in the parameters during the initial optimization process, the obtained configuration may not be optimal. To correct the errors in the configuration, parameters will continue to be collected when the job is actually running on the cluster. After the job runs several times, the configuration can be corrected again, that is, corrective optimization.

[0180] In summary, the multi-stage MapReduce job performance joint tuning method and system of the present invention can greatly reduce the communication overhead during the data shuffling process by caching redundant data, performing distributed computing encoding on the data, and sending the encoded data in a broadcast manner. Therefore, this project combines distributed computing encoding with the redundant computing Reduce function, thus greatly shortening the completion time of multi-stage MapReduce jobs and optimizing the running performance of MapReduce.

[0181] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0182] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0183] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0184] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are only illustrative. For example, the division of the above-mentioned module or unit is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0185] The unit described as a separated component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0186] In addition, in each embodiment of the present invention, each functional unit may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0187] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0188] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0189] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the process in Figure 1one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.

[0190] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 the steps of the functions specified in one block or multiple blocks.

[0191] The above content is only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.

Claims

1. A multi-stage MapReduce job performance joint tuning method, characterized in that: The following steps are involved: S1, read input data from the distributed file system, execute the Map function operation, collect the data generated by the Map function operation, and merge the generated data after the cache overflow to obtain the shuffle file fragment; S2, intermediate data transmission of shuffle file segments; S3. Merge the received intermediate data files, execute the Reducer function operation, and write the obtained results into the distributed file system.

2. The multi-stage MapReduce job performance joint tuning method according to claim 1 is characterized in that: In step S1, the cache overflow is specifically as follows: Before performing cache overflow, the intermediate output data is sorted and pre-joined, and the data is compressed. When the cache occupancy reaches a closed value, the data in the cache is replaced on the hard disk to implement cache overflow. During multiple cache overflow processes, multiple overflow files appear on the disk.

3. The multi-stage MapReduce job performance joint tuning method according to claim 2 is characterized in that: After the cache overflow is completed, all overflow files are merged into one Map function output data file.

4. The multi-stage MapReduce job performance joint tuning method according to claim 1, characterized in that: In step S1, the Map function operation is performed, specifically: Map function operation executes the Map function and stores the generated data in the memory buffer. Before storing in the memory buffer, the data is first divided until the threshold of the buffer capacity is reached.

5. The multi-stage MapReduce job performance joint tuning method according to claim 4 is characterized in that: After obtaining the Map function output data files, each node in the distributed file system obtains Q Map function output data files after the Map function operation on the local existing files.

6. The multi-stage MapReduce job performance joint tuning method according to claim 1, characterized in that: In step S3, the Reducer function operation grabs the processed intermediate data from each terminal, and stores the data in the memory buffer after decompression.

7. The multi-stage MapReduce job performance joint tuning method according to claim 1, characterized in that: Step S3 is specifically as follows: When the memory buffer reaches the set threshold, the intermediate data is overflowed to the hard disk, and the shuffle file fragments in the memory and hard disk are preliminarily merged into an independent shuffle file as the input data for the Reduce function operation; Process the input data according to the written reduce function and store the results in memory; The results of the Reduce function operation are then written to the distributed file system.

8. The multi-stage MapReduce job performance joint tuning method according to claim 7 is characterized in that: The input data is processed according to the written reduce function, specifically: There are K distributed computing nodes in the distributed file system. Node k executes the Reduce function operation, k∈{1,…,K}, and obtains the output result u q .

9. The multi-stage MapReduce job performance joint tuning method according to claim 8, characterized in that: Output result q Specifically: First, obtain all the required N Map function intermediate values ​​v through Shuffle data exchange q,1 ,…,v q,N , then calculate the output result u q =h q (v q,1 ,…,v q,N ), h q is the Reduce function.

10. A multi-stage MapReduce job performance joint tuning system, characterized in that: include: The reading module reads input data from the distributed file system, executes the Map function operation, collects the data generated by the Map function operation, and merges the generated data after the cache overflow to obtain the shuffle file fragment; The transmission module performs intermediate data transmission of shuffle file fragments; The tuning module merges the received intermediate data files, executes the Reducer function operation, and writes the obtained results to the distributed file system.

Citation Information

Patent Citations

  • MapReduce calculation process optimization method based on B-tree data structure

    CN110377601A

  • MapReduce overflow improvement method based on MPI-IO

    CN114116293A

  • Coding MapReduce-oriented Shuffle performance optimization method and system under Rack architecture

    CN114844781A

  • Transmission error correction method and system oriented to coding MapReduce framework

    CN116633485A

  • Systems and / or methods for leveraging in-memory storage in connection with the shuffle phase of mapreduce

    US20160034205A1