Spark platform Shuffle process compression algorithm decision method
A technology of compression algorithms and decision-making methods, applied in computing, special data processing applications, instruments, etc., can solve problems such as instability, high threshold, increase system complexity, etc., to achieve the effect of improving efficiency and accuracy, and reducing threshold
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Publication Date
- 2018-01-19
Smart Images

Figure 1 
Figure 2 
Figure 3
Abstract
Description
technical field
[0001] The invention relates to the field of performance optimization of the Shuffle process of a big data processing platform, in particular to a method for decision-making the optimal compression algorithm configuration of the Shuffle process of the Spark platform. Background technique
[0002] With the advent of the big data era, the corresponding big data processing new technologies continue to develop, and a variety of big data processing platforms have also emerged, the most eye-catching of which is Apache Spark.
[0003] Spark is a distributed big data parallel processing platform based on memory computing. It integrates batch processing, real-time stream processing, interactive query and graph computing, avoiding the need to deploy resources brought by different clusters in various computing scenarios. waste.
[0004] Spark's memory-based computing properties make it inherently advantageous for iterative computing, and it is especially suitable for i...
Examples
Embodiment Construction
[0075] Below in conjunction with accompanying drawing and specific implementation case, further illustrate the present invention, it should be understood that these implementation cases are only used to illustrate the present invention and are not intended to limit the scope of the present invention, after reading the present invention, those skilled in the art will understand various aspects of the present invention Modifications in equivalent forms all fall within the scope defined by the appended claims of this application.
[0076] Such as figure 1 As shown, the present invention first collects cluster basic data, and collects performance data of the Spark platform for user applications, including cluster runtime environment, hardware configuration, memory occupancy, network uplink and downlink speed and other measurement indicators. The specific detailed performance indicators are as follows: Table 4 shows.
[0077] Table 4 is the performance index table
[0078]
[0...