MapReduce data computing acceleration method and system
By using an in-memory database in MapReduce to save Map processing results and simplify the intermediate processing process, the problem of long processing time of MapReduce is solved, and data processing efficiency and cluster resource utilization are improved.
Patent Information
- Application Number
- CN201910156488.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-03-01
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2039-03-01
AI Technical Summary
The existing MapReduce process time is long, which affects the data processing efficiency and cluster resource utilization of the big data platform.
The processing results are saved in memory on the Map side, reducing the read and write to the disk, and simplifying the intermediate processing flow between Map and Reduce, and using in-memory databases such as levelDB for data exchange.
It significantly improves the computing performance of MapReduce, shortens processing time, improves the data processing capability of the cluster, reduces computing costs, and provides users with calculation results faster.
Smart Images

Figure CN111638924B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of big data, and in particular relates to a MapReduce data computing acceleration method and system. Background Art
[0002] With the development of business, there are more and more users. With the development of technology, all kinds of data can be obtained. We have entered the era of big data. After years of development, the big data department has become an essential basic department for every company. The big data platform has to process hundreds of GB, hundreds of PB or even more data every day, and run millions of data processing tasks every day, which requires huge cluster resources. The existing offline processing tasks are mainly completed using MapReduce (a programming model for parallel computing of large-scale data sets).
[0003] The MapReduce model mainly consists of three stages: Map, Shuffle, and Reduce. Map is responsible for filtering and dividing data, converting raw data into key-value pairs; Reduce is merging, processing values with the same key value and then outputting new key-value pairs as the final result. In order for Reduce to process the results of Map in parallel, the output of Map must be sorted and split, and then handed over to the corresponding Reduce. The process of further sorting the Map output and handing it over to Reduce is Shuffle.
[0004] The Shuffle process on the Map side is to partition (Partition), sort (Sort) the Map result, then merge the outputs belonging to the same partition (partition) and write them on the disk (Spill), and finally merge the files distributed on different disks to get a partitioned and ordered file (Merge). The process is roughly: Map input -> Partition -> Sort -> Spill -> Merge -> Map output.
[0005] The Shuffle process on the Reduce side is mainly divided into two stages: copying Map output (Copy) and sorting and merging (Merge). The process is roughly as follows: Copy -> Merge.
[0006] Although the existing MapReduce can process massive amounts of data, the processing time has always been relatively long. How to improve the processing speed of MapReduce, enhance the data processing capabilities of the cluster, enable users to obtain calculation results faster, and provide data calculation services to more users has become a difficult problem that needs to be solved urgently. Summary of the invention
[0007] The technical problem to be solved by the embodiments of the present invention is to overcome the defect of long processing time of existing MapReduce and provide a MapReduce data calculation acceleration method and system.
[0008] The embodiment of the present invention solves the above technical problems through the following technical solutions:
[0009] A MapReduce data computing acceleration method, the method comprising performing the following steps at a Map end:
[0010] Start the map task;
[0011] After performing map processing on the input data, the key-value pair records are output;
[0012] Call the partition function to determine the reduce task to which the output key-value pair record is assigned;
[0013] Write the output key-value pair records into the memory space corresponding to the assigned reduce task;
[0014] The method further includes executing the following steps at the Reduce end:
[0015] Start the reduce task;
[0016] After the map task is completed, read the key-value pair records in the memory space corresponding to the reduce task;
[0017] Perform reduce processing on the read key-value pair records and output the processing results.
[0018] Preferably, the output key-value pair records are written into the memory space corresponding to the assigned reduce task through the memory database;
[0019] The key-value pair records in the memory space corresponding to the reduce task are read through the memory database.
[0020] Preferably, the in-memory database includes levelDB;
[0021] The method further comprises:
[0022] Initialize a levelDB instance, wherein the levelDB instance corresponds to a reduce task;
[0023] Storing connection information of the levelDB instance;
[0024] The output key-value pair records are written into the memory space corresponding to the assigned reduce task through the in-memory database, including:
[0025] Get the connection information of the levelDB instance;
[0026] Connect the map task to the levelDB instance;
[0027] Determine whether the levelDB instance has the key of the output key-value pair record. If so, call the append write API of the levelDB instance to append the write. If not, call the create key-value pair API of the levelDB instance to create a new key and insert the output key-value pair record.
[0028] Close the connection between the map task and the levelDB instance;
[0029] Reading the key-value pair records in the memory space corresponding to the reduce task through the memory database includes:
[0030] Read all key-value pair records in the levelDB instance through the traversal API of the levelDB instance;
[0031] The method also includes destroying the levelDB instance after performing reduce processing on the read key-value pair records and outputting the processing results.
[0032] Preferably, the connection information of the levelDB instance is stored in a Redis database.
[0033] Preferably, the memory space corresponding to the reduce task is the local memory space or the remote memory space of the Reduce end.
[0034] Preferably, the key-value pair records output by a map task are written into the memory space corresponding to one or more reduce tasks;
[0035] The memory space corresponding to a reduce task is written with key-value pair records output by one or more map tasks.
[0036] A MapReduce data computing acceleration system, the system comprising: a Map end and a Reduce end;
[0037] The Map side includes:
[0038] The map task module is used to start the map task, map the input data and output key-value pair records;
[0039] The partition module is used to call the partition function to determine the reduce task to which the output key-value pair record is assigned;
[0040] The writing module is used to write the output key-value pair records into the memory space corresponding to the assigned reduce task;
[0041] The Reduce end includes:
[0042] Reduce task module, used to start reduce tasks;
[0043] The read module is used to read the key-value pair records in the memory space corresponding to the reduce task after the map task is completed;
[0044] The Reduce task module is also used to perform reduce processing on the read key-value pair records and then output the processing results.
[0045] Preferably, the writing module writes the output key-value pair records into the memory space corresponding to the assigned reduce task through the memory database;
[0046] The reading module reads the key-value pair records in the memory space corresponding to the reduce task through the memory database.
[0047] Preferably, the in-memory database includes levelDB;
[0048] The system further comprises:
[0049] An initialization module, used to initialize a levelDB instance and store connection information of the levelDB instance. The levelDB instance corresponds to a reduce task;
[0050] The writing module is used for:
[0051] Get the connection information of the levelDB instance;
[0052] Connect the map task to the levelDB instance;
[0053] Determine whether the levelDB instance has the key of the output key-value pair record. If so, call the append write API of the levelDB instance to append the write. If not, call the create key-value pair API of the levelDB instance to create a new key and insert the output key-value pair record.
[0054] Close the connection between the map task and the levelDB instance;
[0055] The reading module is used for:
[0056] Read all key-value pair records in the levelDB instance through the traversal API of the levelDB instance;
[0057] The system further comprises:
[0058] The destruction module is used to destroy the levelDB instance after performing reduce processing on the read key-value pair records and outputting the processing results.
[0059] Preferably, the connection information of the levelDB instance is stored in a Redis database.
[0060] Preferably, the memory space corresponding to the reduce task is the local memory space or the remote memory space of the Reduce end.
[0061] Preferably, the key-value pair records output by a map task are written into the memory space corresponding to one or more reduce tasks;
[0062] The memory space corresponding to a reduce task is written with key-value pair records output by one or more map tasks.
[0063] On the basis of being in accordance with the common sense in the art, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present invention.
[0064] The positive improvement effects of the embodiments of the present invention are:
[0065] The inventors have conducted an in-depth analysis of the existing MapReduce processing process and found that there are two main reasons why it takes a long time to process massive data:
[0066] (1) The results of Map processing must be written to disk, which reduces the overall performance of MapReduce.
[0067] (2) The intermediate processing flow is too long: After Map input, it goes through six stages, namely Partition, Sort, Spill, Merge, Copy, and Merge, before reaching Reduce processing.
[0068] Based on the above analysis, the embodiment of the present invention stores the Map processing results in the memory, reduces the reading and writing of the disk, and significantly improves the system performance; the embodiment of the present invention also optimizes the intermediate process of MapReduce, simplifies the intermediate processing flow from Map input to Reduce processing, reduces the calculation process, saves time, and thus improves the data processing capability of the cluster, allowing users to obtain calculation results faster, providing data calculation services to more users, and meeting existing needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 A flowchart of the steps executed on the Map side by a MapReduce data computing acceleration method according to Embodiment 1 of the present invention;
[0070] Figure 2 A flowchart of the steps executed on the Reduce side of a MapReduce data computing acceleration method according to Embodiment 1 of the present invention;
[0071] Figure 3 A flowchart of the steps executed on the Map side by a MapReduce data computing acceleration method according to Embodiment 2 of the present invention;
[0072] Figure 4 A flowchart of the steps executed on the Reduce side of a MapReduce data computing acceleration method according to Embodiment 2 of the present invention;
[0073] Figure 5 A schematic diagram of a preferred implementation flow of a MapReduce data computing acceleration method in practical application according to Embodiment 2 of the present invention;
[0074] Figure 6 This is a schematic block diagram of a MapReduce data computing acceleration system according to Embodiment 3 of the present invention;
[0075] Figure 7 This is a schematic block diagram of a MapReduce data computing acceleration system according to Embodiment 4 of the present invention. DETAILED DESCRIPTION
[0076] The present invention is further described below by way of examples, but the present invention is not limited to the scope of the examples.
[0077] Example 1
[0078] This embodiment provides a MapReduce data computing acceleration method. Figure 1 As shown, the method includes performing the following steps on the Map side (the server or cluster where the map task is deployed):
[0079] Step 11: Start the map task;
[0080] Step 12: Output a key-value pair record after performing map processing on the input data, wherein the input data may be allocated by a MapReduce framework, and the map processing includes but is not limited to converting the input data into the key-value pair record;
[0081] Step 13: Call the partition function to determine the reduce task to which the output key-value pair record is assigned, wherein determining the reduce task to which the output key-value pair record is assigned may include but is not limited to determining to which reduce task the key-value pair record is to be assigned according to the key of the output key-value pair record. For example, each key-value pair output by the map has a partition value, and the partition value is obtained by calculating the hash value of the key and taking the modulus of the number of Reduce tasks by default. If the partition value of a key-value pair is 1, it means that the key-value pair will be handed over to the first reducer task for processing;
[0082] Step 14: Write the output key-value pair records into the memory space corresponding to the assigned reduce task.
[0083] like Figure 2 As shown, the method further includes executing the following steps on the Reduce side (the server or cluster where the reduce task is deployed):
[0084] Step 21: Start the reduce task;
[0085] Step 22: After the map task is completed, read the key-value pair records in the memory space corresponding to the reduce task, wherein the reduce task can specifically continuously obtain the heartbeat of the map task through the heartbeat to obtain the progress of the map task, and then determine whether the map task has been completed (the Reduce end has a map task allocation list, which includes all the key-value pair records assigned to the reduce task for processing. Each map task will report the written key-value pair records regularly, and the progress of the map task and whether it is completed can be determined by comparing the list);
[0086] Step 23: Perform reduce processing on the read key-value pair records and output the processing results, wherein the reduce processing may include but is not limited to converting all the read key-value pair records into a key, corresponding to a traverser, such as a list in Java, and passing it to the reduce function, and the reduce function receives the converted data for merging operations.
[0087] This embodiment saves the map output in the memory, reduces the reading and writing of the disk, and significantly improves the computing performance of the MapReduce framework; at the same time, it also simplifies the intermediate processing flow from Map input to Reduce processing from the original 6 (Partition, Sort, Spill, Merge, Copy, Merge) to 3 (steps 13, 14, 22), greatly shortening the calculation process, reducing system complexity, and saving time. In the big data platform, the computing performance of MapReduce has been greatly improved under the condition that the cluster resources remain unchanged, thereby ensuring the timeliness of the big data platform and data warehouse, while improving the computing efficiency and reducing the computing cost, so that the big data platform can run more computing tasks and serve more users.
[0088] Example 2
[0089] This embodiment is a further improvement on Embodiment 1. In this embodiment, step 14 can write the output key-value pair record into the memory space corresponding to the assigned reduce task through the memory database; step 22 can read the key-value pair record in the memory space corresponding to the reduce task through the memory database. In this embodiment, each reduce task corresponds to a memory space, that is, the reduce task and the memory space can be regarded as a one-to-one relationship. The memory space corresponding to the reduce task can be the local memory space or the remote memory space of the Reduce end, and the remote memory space of the Reduce end can even include the local memory space of the Map end. Of course, from the perspective of network overhead, setting the memory space corresponding to the reduce task as the local memory space of the Reduce end can reduce the network access of the reduce process in the next stage, thereby improving computing efficiency and system stability. In step 14, the Map side remotely writes (remoteWrite) the data into the local memory space of the Reduce side, replacing the Sort, Spill and Merge in the prior art; the Reduce side locally reads (localRead) the data to be processed by the reduce side in step 22, replacing the Copy and Merge in the prior art, further simplifying the calculation process and shortening the calculation time.
[0090] Specifically, the memory database preferably includes levelDB. The method further includes:
[0091] Initialize a levelDB instance, wherein the levelDB instance corresponds to a reduce task;
[0092] Stores connection information for the levelDB instance.
[0093] Since Redis has the advantages of fast access speed and high concurrency, the connection information of the levelDB instance can be stored in the Redis database. When storing, the key is jobid_reduceid and the value is the corresponding connection information. Of course, this embodiment is not limited to this. The connection information can also be stored using memcached (a high-performance distributed memory object caching system), zookeeper (a distributed, open source distributed application coordination service), etc.
[0094] Accordingly, if Figure 3 As shown, step 14 writes the output key-value pair records into the memory space corresponding to the assigned reduce task through the memory database, specifically including:
[0095] Step 141: Obtain connection information of the levelDB instance;
[0096] Step 142: Connect the map task to the levelDB instance;
[0097] Step 143: Determine whether the levelDB instance has the key of the output key-value pair record, if yes, execute step 144, if no, execute step 145;
[0098] Step 144: call the append write API of the levelDB instance to append write, and then execute step 146;
[0099] Step 145: Call the new key-value pair API of the levelDB instance to create a new key and insert the output key-value pair record;
[0100] Step 146: Close the connection between the map task and the levelDB instance.
[0101] Accordingly, if Figure 4 As shown, step 22 of reading the key-value pair records in the memory space corresponding to the reduce task through the memory database specifically includes:
[0102] Step 221: read all key-value pair records in the levelDB instance through the traversal API of the levelDB instance;
[0103] The method further comprises executing, after step 23:
[0104] Step 24: Destroy the levelDB instance.
[0105] In this embodiment, the key-value pair records output by a map task can be written into the memory space (level DB instance) corresponding to one or more reduce tasks, that is, a map task can write into the memory space corresponding to one or more reduce tasks. The memory space (level DB instance) corresponding to a reduce task can also be written into the key-value pair records output by one or more map tasks, that is, the memory space corresponding to a reduce task can be written by one or more map tasks. That is, the memory space corresponding to the map task and the reduce task (or reduce task) can be regarded as a many-to-many relationship.
[0106] This embodiment writes levelDB from the Map side and reads the local levelDB from the Reduce side. No performance-consuming operation such as sorting occurs, and levelDB has good random write and batch read performance, which greatly improves the execution efficiency of MapReduce.
[0107] Referring to the above description, taking the example of starting n map tasks (hereinafter referred to as Map1-Map n) and m reduce tasks (hereinafter referred to as Reduce1-Reduce n) at the same time and deploying a levelDB instance locally on each Reduce end (hereinafter referred to as levelDB1-levelDB m) as an example, a preferred implementation process of this embodiment in practical application is provided, as shown in FIG. Figure 5 As shown:
[0108] Phase 1: First, Map1 inputs data, then Map1 performs map processing on the input data, calls the partition function, and determines the assigned reduce task. The same is true for Map2-Map n (the process sequence of each Map in the figure is represented by a different line type), which will not be repeated here.
[0109] Phase 2: Assuming that Map1-Map n in the previous phase are all assigned to Reduce1, corresponding to level DB1, then obtain the connection information of level DB1 from Redis;
[0110] Phase 3: First, Map1 connects to level DB1, and then determines whether the key of the data in levelDB1 exists. If yes, it calls the API to append the data. If not, it calls the API to create a new key and insert the data. The same is true for Map2-Map n, which will not be repeated here.
[0111] Phase 4: After the data of Map1-Map n are written into levelDB1, Reduce1 first reads the data in the local level DB1, then performs reduce processing on the data, and finally outputs the processing results.
[0112] Example 3
[0113] This embodiment provides a MapReduce data computing acceleration system. Figure 6 As shown, the system includes: a Map end 21 (a server or cluster for deploying map tasks) and a Reduce end 22 (a server or cluster for deploying reduce tasks).
[0114] The Map terminal 21 includes:
[0115] A map task module 211 is used to start a map task, perform map processing on input data and then output a key-value pair record, wherein the input data may be allocated by a MapReduce framework, and the map processing includes but is not limited to converting the input data into the key-value pair record;
[0116] The partition module 212 is used to call the partition function to determine the reduce task to which the output key-value pair record is assigned, wherein determining the reduce task to which the output key-value pair record is assigned may include but is not limited to determining to which reduce task the key-value pair record is to be assigned according to the key of the output key-value pair record. For example, each key-value pair output by the map has a partition value, and the partition value is obtained by calculating the hash value of the key and taking the modulus of the number of Reducetasks by default. If the partition value of a key-value pair is 1, it means that the key-value pair will be handed over to the first reducer task for processing;
[0117] A writing module 213, used to write the output key-value pair record into the memory space corresponding to the assigned reduce task;
[0118] The Reduce end 22 includes:
[0119] Reduce task module 221, used to start the reduce task;
[0120] The reading module 222 is used to read the key-value pair records in the memory space corresponding to the reduce task after the map task is completed, wherein the reduce task can specifically continuously obtain the heartbeat of the map task through the heartbeat to obtain the progress of the map task, and then determine whether the map task has been completed (the Reduce end 22 has a map task allocation list, which includes all the key-value pair records assigned to the reduce task for processing. Each map task will report the written key-value pair records regularly, and the progress of the map task and whether it is completed can be determined by comparing the list);
[0121] The Reduce task module 221 is also used to perform reduce processing on the read key-value pair records and output the processing results, wherein the reduce processing may include but is not limited to converting all the read key-value pair records into a key, corresponding to a traverser, such as a list in Java, and passing it to the reduce function, and the reduce function receives the converted data for merging operations.
[0122] This embodiment saves the map output in the memory, reduces the reading and writing of the disk, and significantly improves the computing performance of the MapReduce framework; at the same time, it also simplifies the intermediate processing flow from Map input to Reduce processing from the original 6 (Partition, Sort, Spill, Merge, Copy, Merge) to 3 (partition module 212, write module 213 and read module 222), greatly shortening the calculation process, reducing system complexity, and saving time. In the big data platform, the computing performance of MapReduce has been greatly improved under the condition that the cluster resources remain unchanged, thereby ensuring the timeliness of the big data platform and data warehouse, while improving the computing efficiency and reducing the computing cost, so that the big data platform can run more computing tasks and serve more users.
[0123] Example 4
[0124] This embodiment is a further improvement on Embodiment 3. In this embodiment, the write module 213 writes the output key-value pair records into the memory space corresponding to the assigned reduce task through the memory database; the read module 222 reads the key-value pair records in the memory space corresponding to the reduce task through the memory database. In this embodiment, each reduce task corresponds to a memory space, that is, the reduce task and the memory space can be regarded as a one-to-one relationship. The memory space corresponding to the reduce task can be the local memory space or the remote memory space of the Reduce end 22. The remote memory space of the Reduce end 22 can even include the local memory space of the Map end 21. Of course, from the perspective of network overhead, setting the memory space corresponding to the reduce task as the local memory space of the Reduce end can reduce the network access of the reduce process in the next stage, thereby improving computing efficiency and system stability. Figure 7 As shown, the Map end 21 remotely writes (remoteWrite) the data into the local memory space of the Reduce end 22 in the write module 213, replacing the Sort, Spill and Merge in the prior art; the Reduce end 22 locally reads (localRead) the data that needs to be reduced in the read module 222, replacing the Copy and Merge in the prior art, further simplifying the calculation process and shortening the calculation time.
[0125] Specifically, the memory database preferably includes levelDB. The system also includes: an initialization module, which is used to initialize the levelDB instance and store the connection information of the levelDB instance. The levelDB instance corresponds to the reduce task. Since Redis has the advantages of fast access speed and high concurrency, the connection information of the levelDB instance can be stored in the Redis database. When storing, the key is jobid_reduceid and the value is the corresponding connection information. Of course, this embodiment is not limited to this. The connection information can also be stored using memcached (a high-performance distributed memory object caching system), zookeeper (a distributed, open source distributed application coordination service), etc.
[0126] Accordingly, the writing module 213 is used for:
[0127] Get the connection information of the levelDB instance;
[0128] Connect the map task to the levelDB instance;
[0129] Determine whether the levelDB instance has the key of the output key-value pair record. If so, call the append write API of the levelDB instance to append the write. If not, call the create key-value pair API of the levelDB instance to create a new key and insert the output key-value pair record.
[0130] Close the connection between the map task and the levelDB instance;
[0131] Accordingly, the reading module 222 is used for:
[0132] Read all key-value pair records in the levelDB instance through the traversal API of the levelDB instance;
[0133] The system further includes: a destruction module, which is used to destroy the levelDB instance after performing reduce processing on the read key-value pair records and outputting the processing results.
[0134] In this embodiment, the key-value pair records output by a map task can be written into the memory space (level DB instance) corresponding to one or more reduce tasks, that is, a map task can write into the memory space corresponding to one or more reduce tasks. The memory space (level DB instance) corresponding to a reduce task can also be written into the key-value pair records output by one or more map tasks, that is, the memory space corresponding to a reduce task can be written by one or more map tasks. That is, the memory space corresponding to the map task and the reduce task (or reduce task) can be regarded as a many-to-many relationship.
[0135] In this embodiment, the Map end 21 writes to the levelDB, and the Reduce end 22 reads the local levelDB. No performance-consuming operation such as sorting occurs, and the levelDB has good random writing and batch reading performance, which greatly improves the execution efficiency of MapReduce.
[0136] The preferred implementation process of the system of this embodiment in practical applications can refer to the relevant description and Figure 5 .
[0137] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that these are only examples, and the protection scope of the present invention is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, but these changes and modifications all fall within the protection scope of the present invention.
Claims
1. A MapReduce data computing acceleration method, characterized in that: The method comprises executing the following steps on the Map side: Start the map task; After performing map processing on the input data, the key-value pair records are output; Call the partition function to determine the reduce task to which the output key-value pair record is assigned; Write the output key-value pair records into the memory space corresponding to the assigned reduce task; The method further includes executing the following steps at the Reduce end: Start the reduce task; After the map task is completed, read the key-value pair records in the memory space corresponding to the reduce task; Perform reduce processing on the read key-value pair records and output the processing results; Write the output key-value pair records into the memory space corresponding to the assigned reduce task through the in-memory database; Reading the key-value pair records in the memory space corresponding to the reduce task through the memory database; The memory space corresponding to the reduce task is the local memory space or remote memory space of the Reduce end.
2. The MapReduce data computing acceleration method according to claim 1, characterized in that: The memory database includes levelDB; The method further comprises: Initialize a levelDB instance, wherein the levelDB instance corresponds to a reduce task; Storing connection information of the levelDB instance; The output key-value pair records are written into the memory space corresponding to the assigned reduce task through the in-memory database, including: Get the connection information of the levelDB instance; Connect the map task to the levelDB instance; Determine whether the levelDB instance has the key of the output key-value pair record. If so, call the append write API of the levelDB instance to append the write. If not, call the create key-value pair API of the levelDB instance to create a new key and insert the output key-value pair record. Close the connection between the map task and the levelDB instance; Reading the key-value pair records in the memory space corresponding to the reduce task through the memory database includes: Read all key-value pair records in the levelDB instance through the traversal API of the levelDB instance; The method also includes destroying the levelDB instance after performing reduce processing on the read key-value pair records and outputting the processing results.
3. The MapReduce data computing acceleration method according to claim 2, characterized in that: The connection information of the levelDB instance is stored in the Redis database.
4. The MapReduce data computing acceleration method according to claim 1, characterized in that: The key-value pair records output by a map task are written into the memory space corresponding to one or more reduce tasks; The memory space corresponding to a reduce task is written with key-value pair records output by one or more map tasks.
5. A MapReduce data computing acceleration system, characterized in that: The system includes: a Map side and a Reduce side; The Map side includes: The map task module is used to start the map task, map the input data and output key-value pair records; The partition module is used to call the partition function to determine the reduce task to which the output key-value pair record is assigned; The writing module is used to write the output key-value pair records into the memory space corresponding to the assigned reduce task; The Reduce end includes: Reduce task module, used to start reduce tasks; The read module is used to read the key-value pair records in the memory space corresponding to the reduce task after the map task is completed; The Reduce task module is also used to perform reduce processing on the read key-value pair records and then output the processing results; The writing module writes the output key-value pair records into the memory space corresponding to the assigned reduce task through the memory database; The reading module reads the key-value pair records in the memory space corresponding to the reduce task through the memory database; The memory space corresponding to the reduce task is the local memory space or remote memory space of the Reduce end.
6. The MapReduce data computing acceleration system according to claim 5, characterized in that: The memory database includes levelDB; The system further comprises: An initialization module, used to initialize a levelDB instance and store connection information of the levelDB instance. The levelDB instance corresponds to a reduce task; The writing module is used for: Get the connection information of the levelDB instance; Connect the map task to the levelDB instance; Determine whether the levelDB instance has the key of the output key-value pair record. If so, call the append write API of the levelDB instance to append the write. If not, call the create key-value pair API of the levelDB instance to create a new key and insert the output key-value pair record. Close the connection between the map task and the levelDB instance; The reading module is used for: Read all key-value pair records in the levelDB instance through the traversal API of the levelDB instance; The system further comprises: The destruction module is used to destroy the levelDB instance after performing reduce processing on the read key-value pair records and outputting the processing results.
7. The MapReduce data computing acceleration system according to claim 6, characterized in that: The connection information of the levelDB instance is stored in the Redis database.
8. The MapReduce data computing acceleration system according to claim 5, characterized in that: The key-value pair records output by a map task are written into the memory space corresponding to one or more reduce tasks; The memory space corresponding to a reduce task is written with key-value pair records output by one or more map tasks.
Citation Information
Patent Citations
A data processing method and apparatus
CN109101188A