A method for performance modeling and simulation of big data systems
By analyzing the operation logs of big data jobs and building corresponding models, the problem of insufficient performance prediction accuracy of big data systems in the existing technology is solved, and accurate prediction and stability improvement of the performance of new big data jobs is achieved.
Patent Information
- Application Number
- CN202211267167.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-10-17
AI Technical Summary
The existing big data system performance prediction model declines when facing cluster changes and job changes, and it is difficult to provide response indicators for a single resource request.
By analyzing the operation logs of multiple big data jobs on different clusters, establishing a big data job portrait, and building a load prediction model and resource response model. Using these models, predict the performance of new big data jobs and verify the load to resource response through simulation.
It realizes accurate prediction of the performance of new big data jobs, improves the accuracy and stability of big data job performance prediction, and can adapt to cluster and job changes.
Smart Images

Figure CN115629857B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data operations, and specifically relates to a big data system performance modeling and simulation method. Background Art
[0002] In the digital age, data has become an important production factor in production activities in all walks of life. With the increasingly perfect digital construction, it is more convenient to obtain large-scale data. The rules implied in big data often have high reference value and can serve as an important basis for production decision-making. Under the trend of big data-driven development, big data systems have emerged. In order to deploy and optimize big data systems with higher cost-effectiveness, enterprises and organizations often need to evaluate the performance of big data operations in big data systems.
[0003] The performance modeling research of computer systems, especially the computer system performance model based on queuing theory, has been applied in the performance modeling of big data systems. The computer system performance model based on queuing theory abstracts computing resources into service desks and the resource requests that use computing resources into customers. Through appropriate transformation, the computer performance model based on queuing theory can obtain appropriate queuing model indicators such as average arrival rate, demand arrival interval distribution and service time distribution, and then calculate indicators such as the time expectation and throughput rate of demand response through the corresponding queuing theory model. However, most of the computer performance models based on queuing theory queuing models give the response expectation of computer system processing requests from a statistical level, but rarely give the response indicators of a single resource request.
[0004] In recent years, the rise of artificial intelligence technologies such as machine learning has provided new ideas for big data systems from the perspective of data-driven models. Machine learning models such as neural networks, support vector regression, linear regression, random forests, and XGBoost have been widely used in performance regression tasks for big data jobs and have achieved good results. However, data-driven big data performance models are highly bound to the clusters that generate training data and big data jobs, and their prediction effects decline when faced with cluster changes and job changes.
[0005] In general, the mechanisms used by most current big data job performance prediction models, such as queuing theory and machine learning, are still at the statistical analysis level or are deeply bound to existing job data and clusters running big data jobs, and their generalization performance is insufficient. Summary of the invention
[0006] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and to provide a method, system, device and storage medium for modeling and simulating the performance of a big data system, by analyzing the operation logs of multiple existing big data jobs on multiple different clusters to establish a big data job portrait, and then further establish a load prediction model for big data tasks and a resource response model for various computing resources in the big data cluster. When predicting the performance of a new big data job, the method predicts the load application of each major data task based on the behavioral models such as the task graph and scheduling method provided by the user and the cluster configuration, and implements the call of the load to the various computing resource response models through simulation, thereby realizing the performance prediction of the new big data job.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] The present invention discloses a big data system performance modeling and simulation method, comprising the following steps:
[0009] S1. Big data job log collection and analysis: Run multiple known big data jobs on multiple big data clusters, collect and analyze big data job operation logs, and form a big data job profile including job operation configuration and performance indicators;
[0010] S2. Construction of simulation model library: Based on the portraits of major data jobs, extract influential configuration parameters and the big data system status during the operation of big data jobs, build multiple big data task load prediction models and hardware resource response models of different computing resources, and form a simulation model library;
[0011] S3. Big data job behavior analysis: The user analyzes the big data job whose performance is to be predicted, obtains the software behavior such as the task relationship and scheduling method of the big data job and the big data cluster configuration running the big data job, and inputs the software behavior and cluster configuration into the system;
[0012] S4. Generation and execution of big data job simulation files: The big data job simulation files are constructed by combining the software lines and load application conditions of the big data job to be predicted, as well as the simulation template and resource response model. The big data job simulation files are compiled and run to obtain the performance prediction results of the big data job to be predicted.
[0013] As a preferred technical solution, the step S1 is specifically as follows:
[0014] S1-1. Collect multiple different types of test benchmarks and general big data operations;
[0015] S1-2. Deployment of computing resource usage collection scripts or collection tools on multiple heterogeneous big data clusters with different node topologies and nodes with different hardware configurations;
[0016] S1-3. Start the computing resource usage collection script or collection tool deployed in step S1-2 on the multiple big data clusters described in step S1-2, and then start all the big data jobs collected in step S1-1 in sequence, and only when the previous big data job is completed can the next big data job be started;
[0017] S1-4. Collect the logs generated by the big data job run in step S1-3 in the big data system and the computing resource usage during the job execution obtained by the computing resource usage collection script or collector started in step S1-3 to form a complete big data job log;
[0018] S1-5. Analyze the big data job log and extract the configuration parameter group D of the big data cluster;
[0019] S1-6. Analyze the big data job log to obtain the task set T that constitutes the big data job; for a big data job m, if it consists of n tasks, the task set of the big data job is represented by T m = {R m,1 ,R m,2 ,…,R m,n}, where R m,i is the i-th task of big data job m;
[0020] S1-7. Analyze the big data job log and combine the big data cluster configuration parameter group D to obtain the node physical configuration parameter set c on which each big data task runs i,j ; For each big data task R i,j , the node physical configuration parameter set c that needs to be extracted i,j Including the parameters of the computing unit, memory, disk and network of the node where the task is running;
[0021] S1-8. Analyze the big data job log and extract the big data system configuration parameter set s when each big data task is running i,j ; For each big data task R i,j , the big data system configuration parameter set s that needs to be extracted i,j Including the distributed file system configuration parameters of the node where the corresponding task runs, the distributed computing framework configuration parameters, the distributed scheduling framework configuration parameters, and the big data system operating environment configuration parameters;
[0022] S1-9. Analyze the big data job log and extract the performance indicator set p for each big data task i,j ; For each big data task R i,j , the big data task performance indicator set p that needs to be extracted i,jIncluding the type of task, the running time of the task, the number and size of reads and writes to the file system during the task running process, the computing unit usage track and memory usage track during the task running process;
[0023] S1-10. For each big data task, combine its node physical configuration parameter set, big data system configuration parameter set and big data task performance indicator set to form a big data task portrait; for big data task R i,j , whose portrait r i,j =(c i,j ,s i,j ,p i,j );
[0024] S1-11. Analyze the big data job log and extract the overall image of the big data job; for big data job T i , its overall image z i Including the type of the job, the business configuration parameter set of the job, the task types included in the job and the quantity distribution of different types of tasks, the scheduling method of the job, the waiting time for the job to start, and the overall running time of the job;
[0025] S1-12. Combine the overall profile of the big data job and the profile of the big data tasks included in the big data job to form a big data job profile; for the big data job T i , whose portrait t i =(z i ,{r i,1 ,r i,2 ,…,r i,n}), the generated big data job portrait will be stored in the big data job portrait library.
[0026] As a preferred technical solution, in step S1-5, the configuration parameter group D of the big data cluster includes the topological structure of the cluster, the status of each link in the cluster, and the cluster physical configuration information of each node in the cluster;
[0027] The cluster physical configuration information of each node in the cluster includes configuration information of computing units, memory, disks and networks;
[0028] The computing unit includes a CPU, a GPU and an FPGA.
[0029] As a preferred technical solution, in step S2, the load prediction model construction process specifically includes the following steps:
[0030] S2-1-1. Extract all resource requirements of different types of tasks under different job types from the big data job portraits stored in the big data job portrait library; the resource requirements are any one of the four requirements of computing unit usage time, memory usage size, disk read and write size, network transmission start and end points, and transmission data volume;
[0031] S2-1-2. For each resource requirement a, extract the big data system configuration parameter set s of the big data task to which the resource requirement belongs from the big data job portrait stored in the big data job portrait library i,j and node physical configuration parameter set c i,j And the business configuration parameter set of the big data job to which it belongs is converted into the business configuration parameter set v of the big data task i,j , forming the load demand; the load demand after extraction is expressed as a four-tuple (a, s a ,c a ,v a ), the load requirements of all load types γ of the big data tasks of type θ under the big data job of type τ are combined into a set to form a load data set A of the corresponding type τ,θ,γ ;
[0032] S2-1-3. For each load data set A τ,θ,γ Implement dimensionality reduction;
[0033] S2-1-4. Use load data set A τ,θ,γ Construct a load forecasting model based on A τ,θ,γ In each record (a,u), u is the input and a is the output. The machine learning model is used to establish the load prediction model Γ τ,θ,γ , and stored in the system load prediction model library.
[0034] As a preferred technical solution, in step S2-1-3, the dimension reduction process takes the data volume in the resource demand a in the load demand as the output, and the big data system configuration parameter set s in the load demand a All parameters in the node physical configuration parameters c a and business configuration parameter set v a All parameters in are used as input, and the impact factors of different configuration parameters on resource requirements are calculated using the dimensionality reduction algorithm, and the configuration parameters with higher impact factors are retained; after dimensionality reduction, the load data set A τ,θ,γ Each record in is represented as a tuple (a,u), where u is a configuration parameter that has a significant influence on this type of load.
[0035] As a preferred technical solution, in step S2, the hardware resource response model construction process specifically includes the following steps:
[0036] S2-2-1. Construct a queue system simulation model library; the queue system simulation model library includes multiple basic queue system simulation models; the basic queue system simulation model is a stateful queue with a single service station, and its components include queue status, service time distribution of the queue service station, and service time of the currently serving customers; the basic queue simulation model includes a resource demand expected response time calculation method, which calculates the expected waiting time of new customers based on the queue status and service time distribution and the service time of the currently serving customers combined with the relevant calculation methods of queuing theory;
[0037] S2-2-2. Extract the resource requirements of each task, the hardware type and processing time for processing the corresponding requirements from the big data job portrait generated in step 2-8 to form a demand response. The extracted demand response is represented by a triple (resource requirement, demand processor, processing time);
[0038] S2-2-3. Establish a cascade queuing simulation model for different types of demand processors, including computing units, memory, disk, and network.
[0039] As a preferred technical solution, in the step S2-2-3, the cascaded queuing system simulation model is composed of multiple basic queuing system simulation models cascaded in a certain structure; the cascade structure of the cascaded queuing simulation model needs to be designed in accordance with the commonly used demand processor; the cascaded queuing simulation model after the structural design is completed needs to be verified using the demand response data obtained in step S2-2-2, and the cascade structure is adjusted based on the verification structure until the error between the simulated response time of the cascaded queuing simulation model to the resource demand and the actual demand processing time in the request response is no longer significantly reduced. After the cascade structure adjustment is completed, the cascaded queuing simulation model representing the demand processor is stored in the system hardware resource response model library.
[0040] As a preferred technical solution, step S3 is specifically as follows:
[0041] S3-1. The user determines the big data cluster configuration that will be run for the big data job to be run and the big data job to be predicted;
[0042] S3-2. The user determines the big data system configuration parameters of each node on the cluster, including the distributed file system configuration parameters of each node, the distributed computing framework configuration parameters, the distributed scheduling framework configuration parameters, and the big data system operating environment configuration parameters, and structures all the configuration parameters in this step and merges them into the cluster configuration file formed in step S3-1, and then inputs the cluster configuration file into the system;
[0043] S3-3. The user determines the big data task type, task quantity, task dependency, scheduling method, and some big data task-level business configuration parameters that the big data job to be predicted should include.i,j ,After forming a directed acyclic task graph with nodes and configuration parameters, a big data job description file is formed in a structured form in combination with the scheduling method and input into the system;
[0044] S3-4. For each big data task in the big data job to be predicted, the system searches the cluster configuration document of step S3-2 according to the task running node label determined in step S3-3, obtains the node physical configuration information of the corresponding task running and the big data system configuration information on the node, and respectively forms the node physical configuration parameter set c of the task. i,j And the big data system configuration parameter set s for this task i,j ;
[0045] S3-5. The system combines the configuration parameters of the big data task in the big data job to be predicted in step S3-3 and step S3-4 to obtain the parameter set k of the big data task i,j =(s i,j ,c i,j ,v i,j );
[0046] S3-6. For each big data task in the big data job to be predicted, the system selects a load prediction model Γ for predicting different load types from the system load prediction model library based on the type of each big data task and the type of big data job to which each big data task belongs. τ,θ,γ and the big data task parameter set k i,j The model input parameters in predict resource requirements; the resource requirements to be predicted include computing unit usage requirements, memory usage requirements, disk read and write requirements, and network transmission requirements;
[0047] S3-7. The system compares the resource requirements obtained in step S3-6 with the physical configuration parameters c of the corresponding node in the big data task parameter set i,j Some relevant parameters in the middle, confirm that the node where the big data task runs meets the resource requirements. If the node cannot meet the resource requirements, it is necessary to decompose the current big data task according to the node configuration, update the task graph and recalculate the business configuration parameters v′ of the new big data task i,j and running nodes, and find the physical configuration parameter set c′ corresponding to the new big data task running node in the cluster configuration file of step S3-3 i,j and the big data system configuration parameter set s′ i,j , forming the big data task parameter set k of the new big data task i,j =(s′ i,j ,c′ i,j ,v′ i,j ), and then repeat steps S3-6 and S3-7 for each new big data task obtained by decomposition.
[0048] As a preferred technical solution, the big data cluster configuration includes the topological structure of the big data cluster, the link status between the nodes of the big data cluster and the physical configuration of each node in the big data cluster. The physical configuration of each node includes the computing unit configuration, memory configuration, disk configuration and network configuration, forming a structured cluster configuration file.
[0049] As a preferred technical solution, step S4 is specifically as follows:
[0050] S4-1. The user selects a programming language for executing big data job simulation. The system retrieves the simulation template of the corresponding programming language according to the programming language selected by the user. The simulation template provides a simulation framework and outputs the performance indicators of the entire simulation process after the simulation framework is executed;
[0051] S4-2. The system analyzes the task graph adjusted in step S3-7 and the big data task scheduling method specified by the user in step S3-3, and arranges the operation relationship of the big data tasks included in the big data job to be predicted in the task graph, and specifies the big data task scheduling method in the simulation template;
[0052] S4-3. For each big data task in the big data job to be predicted, the system searches for the physical configuration of the node running the big data task in the cluster configuration file input by the user in step S3-2, and selects the cascade queuing simulation model of the processor corresponding to the node physical configuration from the system hardware resource response model library, including the resource response model of the computing unit, the resource response model of the memory, the resource response model of the disk, and the resource response model of the network, and fills its calling interface into the interface specified by the simulation template obtained in step 4-1;
[0053] S4-4. For each big data task included in the big data job to be predicted, the system fills the resource requirements of the big data task obtained in step S3-7 and the type of processor required to process the corresponding resource requirements of the node running the corresponding big data task into the corresponding position of the simulation module of the corresponding big data task in the simulation template according to the load mapping interface specified in the simulation template obtained in step S4-1;
[0054] S4-5 system according to step S3-2 user input cluster configuration file fill step S4-1 to obtain the simulation template other information to be filled;
[0055] S4-6 system completes the simulation template after filling the integration analysis and compilation work to generate an executable simulation file;
[0056] S4-7. The system runs the executable big data job simulation file and outputs the performance indicators of the big data job to be predicted after the execution of the file is completed.
[0057] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0058] (1) The present invention uses the cascade queue simulation model, which has a relatively low computational complexity, as a simulation model of computer computing resources, which can speed up the simulation speed in the big data job simulation process.
[0059] (2) The present invention constructs a hardware resource response model library for various types of computing resources. Users can select the corresponding resource response model according to the physical computing resource configuration of each node on the big data cluster where the big data job to be predicted needs to run, thereby realizing the performance prediction of big data jobs in heterogeneous big data clusters.
[0060] (3) The present invention uses a machine learning model as a big data task resource demand prediction model, which realizes the intelligent and precise prediction of big data job task load, thereby further improving the accuracy of big data job performance prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0062] Figure 1 This is a process diagram of a big data system performance modeling and simulation method of the present invention.
[0063] Figure 2 The structure diagram of an example of the cascade queue simulation model referred to in the present invention.
[0064] Figure 3 This is a schematic diagram of a task graph formed by a user analyzing a Hadoop WordCount big data job for performance to be predicted in Example 1.
[0065] Figure 4 This is a schematic diagram of a task graph formed by a user analyzing a Spark Sort big data job with performance to be predicted in Example 2. DETAILED DESCRIPTION
[0066] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0067] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0068] Example 1
[0069] This embodiment 1 is based on Hadoop MapReduce WordCount big data job performance modeling, combined with the most widely used big data system Hadoop, to explain the proposed big data job performance prediction method. Figure 1 As shown, the present invention proposes a big data system performance modeling and simulation method, including the collection and analysis of big data job logs, simulation model construction, big data job behavior analysis, and the generation and execution of big data job simulation files. When performing performance prediction for the WordCount big data job of Hadoop, the specific implementation steps are as follows:
[0070] S1. Collection and analysis of big data job logs: Run multiple Hadoop WordCount big data jobs on multiple Hadoop big data clusters, collect and analyze Hadoop WordCount big data job operation logs, and form a Hadoop WordCount big data job portrait including job operation configuration and performance indicators. The specific method steps are:
[0071] S1-1. Collect multiple Hadoop WordCount big data jobs;
[0072] S1-2. Deploy computing resource usage collection scripts or collection tools on multiple heterogeneous Hadoop big data clusters with different node topologies and nodes with different hardware configurations, such as the Linux system performance analysis tool sar that can be used to collect general hardware operating status, slabtop that can be used to collect memory status information, and blktrace that can be used to collect disk information. At the same time, ensure that the historical log server JobHistoryServer in the Hadoop big data system of each Hadoop big data cluster remains open;
[0073] S1-3. Start the computing resource usage collection script or collection tool deployed in step S1-2 on the multiple Hadoop big data clusters described in step S1-2, and then start all the Hadoop WordCount big data jobs collected in step S1-1 in sequence, and only when the previous Hadoop WordCount big data job is completed can the next Hadoop WordCount big data job be started;
[0074] S1-4. Collect the logs generated by the Hadoop WordCount big data job in the Hadoop big data system in step S1-3 and the CPU, memory, disk, network and other computing resource usage during the job execution obtained by the computing resource usage collection script or collector started in step S1-3, and form a complete Hadoop WordCount big data job log;
[0075] S1-5. Manually analyze all Hadoop WordCount big data job logs and extract the configuration parameter group D of the Hadoop big data cluster. The configuration parameter group of the big data cluster includes the topology of the cluster, the status of each link in the cluster, and the cluster physical configuration information of each node in the cluster, including the configuration information of basic computer components such as CPU, memory, disk and network;
[0076] S1-6. Analyze all Hadoop WordCount big data job logs using Apache Rumen big data job log analysis tool and manual analysis methods to obtain the task set T that constitutes the Hadoop WordCount big data job. For a Hadoop WordCount big data job m, if it consists of n tasks, the task set of the Hadoop WordCount big data job can be expressed as T m = {R m,1 ,R m,2 ,…,Rm,n}, where R m,i The i-th task of Hadoop WordCount big data job m;
[0077] S1-7. Analyze all Hadoop WordCount big data job logs using Apache Rumen big data job log analysis tool and manual analysis methods, and combine Hadoop big data cluster configuration parameter group D to obtain the node physical configuration parameter set c on which each Hadoop big data task runs i,j For each Hadoop big data task R i,j , the node physical configuration parameter set c that needs to be extracted i,j Including the CPU parameters, memory parameters, disk parameters and network parameters of the node where the task runs;
[0078] S1-8. Analyze all Hadoop WordCount big data job logs using Apache Rumen big data job log analysis tool and manual analysis methods, and extract the Hadoop big data system configuration parameter set when each Hadoop big data task is running. i,j For each big data task R i,j , the big data system configuration parameter set s that needs to be extracted i,j Including the configuration parameters of the distributed file system HDFS of the node where the corresponding task is running, the configuration parameters of the distributed computing framework MapReduce, the configuration parameters of the distributed scheduling framework YARN, and the configuration parameters of the Hadoop big data system operating environment including but not limited to JVM configuration and operating system related configuration;
[0079] S1-9. Analyze all Hadoop WordCount big data job logs using Apache Rumen big data job log analysis tool and manual analysis methods, and extract the performance indicator set of each Hadoop big data task. i,j For each Hadoop big data task R i,j , the performance indicator set p of Hadoop big data tasks to be extracted i,j It includes the task type, task running time, the number and size of reads and writes to the distributed file system HDFS and the operating system file system during the task running process, the CPU usage track and memory usage track during the task running process, and calculates the CPU usage median and the peak and median memory usage;
[0080] S1-10. For each Hadoop big data task, combine its node physical configuration parameter set, Hadoop big data system configuration parameter set and Hadoop big data task work performance indicator set to form a Hadoop big data task profile. i,j , whose portrait r i,j =(c i,j ,s i,j ,p i,j );
[0081] S1-11. Use Apache Rumen big data job log analysis tool and manual analysis method to analyze all Hadoop WordCount big data job logs and extract the overall image of Hadoop WordCount big data job. i , its overall image z i Including the type of the job (WordCount), the business configuration parameter set of the job, the task types (Map and Reduce) included in the job and the quantity distribution of different types of tasks, the scheduling method of the job, the waiting time for the job to start, and the overall running time of the job;
[0082] S1-12. Combine the overall image of the Hadoop WordCount big data job and the image of the Hadoop big data task included in the Hadoop WordCount big data job to form a Hadoop WordCount big data job image. i , whose portrait t i =(z i ,{r i,1 ,r i,2 ,…,r i,n}). The generated Hadoop WordCount big data job portrait will be stored in the Hadoop big data job portrait library.
[0083] S2. Construction of simulation model library: Based on the profiles of each Hadoop WordCount big data job, extract the influential configuration parameters and the Hadoop big data system status during the running of the Hadoop WordCount big data job, build multiple Hadoop big data task load prediction models for Hadoop WordCount big data jobs and hardware resource response models of different computing resources, and add them to the Hadoop load prediction model library and hardware resource response model library.
[0084] Furthermore, the specific construction steps of the Hadoop WordCount load prediction model in the Hadoop load prediction model library include:
[0085] S2-1-1. Extract all resource requirements of Map tasks and Reduce tasks from the Hadoop WordCount big data job portrait stored in the Hadoop big data job portrait library. Resource requirements include four types of resource requirements: CPU usage time, memory usage size, disk read and write size, network transmission start and end locations, and transmission data volume;
[0086] S2-1-2. For each resource requirement a, extract the Hadoop big data system configuration parameter set s of the big data task to which the resource requirement belongs from the HadoopWordCount big data job portrait stored in the Hadoop big data job portrait library i,j and node physical configuration parameter set c i,j And the business configuration parameter set of the Hadoop WordCount big data job is converted into the business configuration parameter set v of the Hadoop big data task i,j , forming the load demand. The load demand after extraction is expressed as a four-tuple (a, s a ,c a ,v a ). All load requirements of type γ of Hadoop big data tasks of type θ under the Hadoop WordCount big data job are grouped together to form a load data set A of the corresponding type. WordCount,θ,γ . Wherein, θ is Map or Reduce, and γ is any one of the CPU usage requirement, memory usage requirement, disk read / write requirement, and network transmission requirement in step S2-1-1;
[0087] S2-1-3. For each load data set A WordCountθ,γ Implement dimensionality reduction. The dimensionality reduction process takes the data volume in the resource requirement a in the load requirement as output, and the big data system configuration parameter set s in the load requirement a All parameters in the node physical configuration parameters c a and business configuration parameter set v a All parameters in are used as input, and the impact factors of different configuration parameters on resource requirements are calculated using principal component analysis, random forest and XGBoost. The median is taken as the impact factor of the corresponding parameter, and the configuration parameters with the impact factor in the top 50% are retained. WordCountθ,γ Each record in is represented as a tuple (a,u), where u is a configuration parameter that has a significant influence on this type of load;
[0088] S2-1-4. Use load data set A WordCountθ,γ Build a load forecasting model. WordCountθ,γ In each record (a,u), u is the input and a is the output. The Hadoop WordCount load prediction model Γ is established using the Stacking ensemble learning model of multiple support vector regression and linear regression. WordCountθ,γ , and stored in the system Hadoop load prediction model library.
[0089] Furthermore, the specific steps of constructing the hardware resource response model library include:
[0090] S2-2-1. Construct a queue system simulation model library. The queue system simulation model library includes multiple basic queue system simulation models. The basic queue system simulation model is a stateful queue with a single service station. Its components include the queue state (the number of customers in the queue), the service time distribution of the queue service station, and the service time of the current customer being served. The basic queue simulation model includes a method for calculating the expected response time of resource demand. This method can calculate the expected waiting time of new customers based on the queue state, service time distribution, and the service time of the current customer being served in combination with the simulated queue system state diagram in the queue theory;
[0091] S2-2-2. Extract the resource requirements of each task, the hardware type and processing time for processing the corresponding requirements from the Hadoop WordCount big data job portrait generated in step 2-8 to form a demand response. The extracted demand response is represented by a triple (resource requirement, demand processor, processing time);
[0092] S2-2-3. Build a cascade queue simulation model for different types of demand processors, including computing units, memory, disks, and networks. The cascade queue system simulation model consists of multiple basic queue system simulation models cascaded in a certain structure. The cascade structure of the cascade queue simulation system needs to be designed according to the commonly used demand processors. For example, the cascade queue simulation model structure for a dual-core CPU can be designed as Figure 2The structure shown, the specific cascade structure can refer to the principle of multi-core CPU processing request; the concurrent paths of the cascade queuing simulation model of the memory can refer to the number of channels of the memory and other indicators, and the specific cascade structure can refer to the memory access principle; the parallel paths of the cascade queuing simulation model of the disk can refer to the number of concurrent seeks of the disk arm and other indicators, and the specific cascade structure can refer to the disk access principle; the cascade queuing simulation model of the network can refer to the network topology and other conditions, and the specific cascade structure can refer to the relevant principles of network data transmission. The cascade queuing simulation model after the structural design is completed needs to be verified using the demand response data obtained in step S2-2-2, and the cascade structure is adjusted based on the verification structure until the error between the response simulation time of the cascade queuing simulation model to the resource demand and the actual demand processing time in the request response is no longer significantly reduced. After the cascade structure adjustment is completed, the cascade queuing simulation model representing the demand processor is stored in the system resource response model library.
[0093] S3. Big data job behavior analysis: The user analyzes the task relationship and scheduling method of the Hadoop WordCount big data job to be predicted, and configures the Hadoop big data cluster that runs the Hadoop WordCount big data job to be predicted and inputs it into the system. The specific method steps are:
[0094] S3-1. The user determines the Hadoop big data cluster configuration on which the Hadoop WordCount big data job to be run and whose performance needs to be predicted will be run, including the topological structure of the Hadoop big data cluster, the link status between the nodes of the Hadoop big data cluster, and the physical configuration of each node in the Hadoop big data cluster, including the CPU configuration, memory configuration, disk configuration, and network configuration, etc., to form a structured cluster configuration file;
[0095] S3-2. The user determines the Hadoop big data system configuration parameters of each node in the cluster, including the distributed file system HDFS configuration parameters of each node, the distributed computing framework MapReduce configuration parameters, the distributed scheduling framework YARN configuration parameters, including but not limited to the JVM configuration and the operating system configuration parameters of the Hadoop big data system operating environment, and structures all the configuration parameters in this step and merges them into the cluster configuration file formed in step S3-1, and then inputs the cluster configuration file into the system;
[0096] S3-3. The user determines that the Hadoop WordCount big data job to be predicted must include the number of Map tasks and Reduce tasks, task dependencies, scheduling methods, and some big data task-level business configuration parameters v i,j, a directed acyclic Hadoop WordCount big data task graph with nodes and configuration parameters is formed, and then a HadoopWordCount big data job description file is formed in XML format and input into the system. Users can obtain the above parameters by analyzing the big data files that need to be processed by the Hadoop WordCount big data job. In general, when processing the Hadoop WordCount job, for each file block in the big data file, Hadoop forms a Map task, and the output of all Map tasks is input into a Reduce task. For each Map task, its business configuration parameter v i,j It at least contains the number of words in the data block and the label of the Hadoop big data cluster node where the corresponding file block of the Map task is located; for the Reduce task, its business configuration parameter v i,j At least it contains the total number of words processed and a label of the Hadoop cluster node that is estimated to run the Reduce task. Since there is no obvious dependency between the Map tasks of the Hadoop WordCount big data task in terms of input and output, they can run in parallel. Based on the above analysis, users can form Figure 3 The Hadoop WordCount big data task graph shown in the figure. After the processing is completed, the Hadoop WordCount big data task graph and scheduling method are formed into an XML document with XML tags and rules pre-specified by the system and then input into the system;
[0097] S3-4. For each Hadoop big data task in the Hadoop WordCount big data job to be predicted, the system searches for the cluster configuration document of step S3-2 according to the task running node label determined in step S3-3, obtains the node physical configuration information of the corresponding task running and the Hadoop big data system configuration information on the node, and respectively forms the node physical configuration parameter set c of the task i,j And the Hadoop big data system configuration parameter set s for this task i,j ;
[0098] S3-5. The system combines the configuration parameters of the Hadoop big data task in the Hadoop WordCount big data job to be predicted obtained in step S3-3 and step S3-4 to obtain the parameter set k of the Hadoop big data task i,j =(s i,j ,c i,j ,v i,j );
[0099] S3-6. For each Hadoop big data task in the Hadoop WordCount big data job to be predicted, the system selects the Hadoop WordCount load prediction model Γ for predicting different load types from the system Hadoop load prediction model library according to the type of each Hadoop big data task (Map or Reduce) WordCountθ,γ and Hadoop big data task parameter set k i,j The model input parameters in predict resource requirements. The resource requirements to be predicted include CPU usage requirements, memory usage requirements, disk read and write requirements, and network transmission requirements.
[0100] S3-7. The system compares the resource requirements obtained in step S3-6 with the physical configuration parameters c of the corresponding node in the Hadoop big data task parameter set i,j Some relevant parameters in the Hadoop big data task are used to confirm that the node where the Hadoop big data task runs can meet the resource requirements. If the node cannot meet the resource requirements, the current Hadoop big data task needs to be decomposed and the Hadoop WordCount big data task graph needs to be updated according to the node configuration, and the business configuration parameters v′ of the new Hadoop big data task need to be recalculated. i,j and running nodes, and find the physical configuration parameter set c′ corresponding to the new Hadoop big data task running node in the cluster configuration file of step S3-3 i,j and Hadoop big data system configuration parameter set s′ i,j , forming the Hadoop big data task parameter set k of the new Hadoop big data task i,j =(s′ i,j , c′ i,j ,v′ i,j ), and then repeat steps S3-6 and S3-7 for each new Hadoop big data task obtained by decomposition.
[0101] S4. Generation and execution of big data job simulation file: Combine the software line and load application of the Hadoop WordCount big data job to be predicted, as well as the simulation template and resource response model to form a Hadoop WordCount big data job simulation file. Compile and run the Hadoop big data job simulation file to obtain the performance prediction result of the Hadoop WordCount big data job to be predicted. The specific method steps are:
[0102] S4-1. The user selects the programming language to execute the Hadoop WordCount big data job simulation. The system calls the Hadoop big data job simulation template of the corresponding programming language according to the programming language selected by the user. The simulation template provides the Hadoop big data job simulation framework and can output the performance indicators of the entire simulation process after the simulation framework is executed;
[0103] S4-2. The system analyzes the task graph adjusted in step S3-7 and the Hadoop big data task scheduling method specified by the user in step S3-3, and arranges the Hadoop big data task behavior simulation module in the simulation template according to the operation relationship of the Hadoop big data task included in the Hadoop WordCount big data job to be predicted in the task graph, and specifies the Hadoop big data task scheduling method in the simulation template, including but not limited to first-come-first-served, capacity scheduling, and fair scheduling;
[0104] S4-3. For each Hadoop big data task in the Hadoop WordCount big data job whose performance is to be predicted, the system searches for the physical configuration of the node running the Hadoop big data task in the cluster configuration file input by the user in step S3-2, and selects the cascade queuing simulation model of the processor corresponding to the node physical configuration from the system resource response model library, including the CPU resource response model, the memory resource response model, the disk resource response model and the network resource response model, and fills its calling interface into the interface specified by the simulation template obtained in step 4-1;
[0105] S4-4. For each Hadoop big data task included in the Hadoop WordCount big data job to be predicted, the system fills the resource requirements of the Hadoop big data task obtained in step S3-7 and the required processor type of the node running the corresponding Hadoop big data task to process the corresponding resource requirements into the corresponding position of the simulation module of the Hadoop big data task in the simulation template according to the load mapping interface specified in the simulation template obtained in step S4-1;
[0106] S4-5 system according to step S3-2 user input cluster configuration file fill step S4-1 to obtain the simulation template other information to be filled;
[0107] S4-6. The system performs integration analysis and compilation on the completed simulation template to generate executable Hadoop big data job simulation files;
[0108] S4-7. The system runs the executable Hadoop WordCount big data job simulation file, and outputs the performance indicators of the Hadoop WordCount big data job to be predicted after the execution of the file is completed, and forms a report.
[0109] Example 2
[0110] This embodiment 2 is based on Spark's Sort big data job performance modeling, combined with the currently widely used big data system Spark, to explain the proposed big data job performance prediction method. Figure 1 As shown, the present invention proposes a big data system performance modeling and simulation method, including the collection and analysis of big data job logs, simulation model construction, big data job behavior analysis, and the generation and execution of big data job simulation files. When predicting the performance of Spark's Sort big data job, the specific implementation steps are as follows:
[0111] S1. Collection and analysis of big data job logs: Run multiple known SparkSort big data jobs on multiple Spark big data clusters, collect and analyze Spark Sort big data job operation logs, and form a Spark Sort big data job profile including job operation configuration and performance indicators. The specific method steps are:
[0112] S1-1. Collect multiple Spark Sort big data jobs;
[0113] S1-2. Deploy computing resource usage collection scripts or collection tools on multiple heterogeneous Spark big data clusters with different node topologies and different hardware configuration nodes, such as sar, a Linux system performance analysis tool that can be used to collect general hardware operating status, slabtop, which can be used to collect memory status information, and blktrace, which can be used to collect disk information. At the same time, ensure that the historical log server SparkHistoryServer in the Spark big data system in each Spark big data cluster remains open;
[0114] S1-3. Start the computing resource usage collection script or collection tool deployed in step S1-2 on the multiple Spark big data clusters described in step S1-2, and then start all the Spark Sort big data jobs collected in step S1-1 in sequence, and only when the previous Spark Sort big data job is completed can the next Spark Sort big data job be started;
[0115] S1-4. Collect the logs generated by the Spark Sort big data job run in the Spark big data system in step S1-3 contained in the Spark big data system history log server SparkHistoryServer and the CPU, memory, disk, network and other computing resource usage during the job execution obtained by the computing resource usage collection script or collector started in step S1-3 to form a complete Spark Sort big data job log;
[0116] S1-5. Manually analyze all Spark Sort big data job logs and extract the configuration parameter group D of the Spark big data cluster. The configuration parameter group of the big data cluster includes the topology of the cluster, the status of each link in the cluster, and the cluster physical configuration information of each node in the cluster, including the configuration information of basic computer components such as CPU, memory, disk and network;
[0117] S1-6. Use Databricks big data job log analysis tool and manual analysis method to analyze all Spark Sort big data job logs and obtain the task set T that constitutes the Spark Sort big data job. For a SparkSort big data job m, if it consists of n tasks, the task set of the Spark Sort big data job can be expressed as T m = {R m,1 ,R m,2 ,…,R m,n}, where R m,i The i-th task of Spark Sort big data job m;
[0118] S1-7. Combine the Databricks big data job log analysis tool and manual analysis methods to analyze all Spark Sort big data job logs, and combine the Spark big data cluster configuration parameter group D to obtain the node physical configuration parameter set c on which each Spark big data task runs i,j For each Spark big data task R i,j , the node physical configuration parameter set c that needs to be extracted i,j Including the CPU parameters, memory parameters, disk parameters and network parameters of the node where the task runs;
[0119] S1-8. Use Databricks big data job log analysis tools and manual analysis methods to analyze all Spark Sort big data job logs and extract the Spark big data system configuration parameter set when each Spark big data task is running. i,j For each big data task R i,j, the big data system configuration parameter set s that needs to be extracted i,j Including the distributed file system HDFS configuration parameters of the node where the corresponding task is running, the distributed computing framework Spark configuration parameters, the distributed scheduling framework YARN configuration parameters, and the Spark big data system operating environment configuration parameters including but not limited to JVM configuration and operating system related configuration;
[0120] S1-9. Combined use of Databricks big data job log analysis tools and manual analysis methods to analyze all Spark Sort big data job logs and extract the performance indicator set p for each Spark big data task i,j For each Spark big data task R i,j , the Spark big data task performance indicator set p that needs to be extracted i,j It includes the task type, task running time, the number and size of reads and writes to the distributed file system HDFS and the operating system file system during the task running process, the CPU usage track and memory usage track during the task running process, and calculates the CPU usage median and memory usage peak;
[0121] S1-10. For each Spark big data task, combine its node physical configuration parameter set, Spark big data system configuration parameter set and Spark big data task work performance indicator set to form a Spark big data task profile. i,j , whose portrait r i,j =(c i,j ,s i,j ,p i,j );
[0122] S1-11. Use Databricks big data job log analysis tool and manual analysis method to analyze all Spark Sort big data job logs and extract the overall image of Spark Sort big data job. i , its overall image z i Including the type of the job (Sort), the business configuration parameter set of the job, the task types included in the job (Sample, Shuffle Write and Shuffle Read) and the quantity distribution of different types of tasks, the scheduling method of the job, the waiting time for the job to start and the overall running time of the job;
[0123] S1-12. Combine the overall image of the Spark Sort big data job and the image of the Spark big data task included in the Spark Sort big data job to form a Spark Sort big data job image. i , whose portrait t i =(z i ,{r i,1 ,r i,2 ,…,r i,n}). The generated Spark Sort big data job portrait will be stored in the Spark big data job portrait library.
[0124] S2. Construction of simulation model library: Based on the profiles of each Spark Sort big data job, extract the influential configuration parameters and the Spark big data system status during the running of the Spark Sort big data job, build multiple Spark big data task load prediction models for Spark Sort big data jobs and hardware resource response models for different computing resources, and add them to the Spark load prediction model library and hardware resource response model library.
[0125] The specific steps for building the Spark Sort load prediction model in the Spark load prediction model library include:
[0126] S2-1-1. Extract all resource requirements of Sample task, Shuffle Write task and Shuffle Read task from the Spark Sort big data job portrait stored in the Spark big data job portrait library. Resource requirements include any one of the four resource requirements: CPU usage time, memory usage size, disk read and write size, network transmission start and end points, and transmission data volume;
[0127] S2-1-2. For each resource requirement a, extract the Spark Sort big data system configuration parameter set s of the big data task to which the resource requirement belongs from the Spark Sort big data job portrait stored in the Spark big data job portrait library i,j and node physical configuration parameter set c i,j And the business configuration parameter set of the Spark Sort big data job to which it belongs is converted into the business configuration parameter set v of the Spark big data task i,j , forming the load demand. The load demand after extraction is expressed as a four-tuple (a, s a ,c a ,v a). All load requirements of type γ of the Spark big data tasks of type θ under the Spark Sort big data job are grouped together to form a load data set A of the corresponding type. Sortθ,γ . Wherein, θ is Sample, Shuffle Write or Shuffle Read, and γ is any one of the CPU usage requirement, memory usage requirement, disk read / write requirement and network transmission requirement in step S2-1-1;
[0128] S2-1-3. For each load data set A Sortθ,γ Implement dimensionality reduction. The dimensionality reduction process takes the data volume in the resource requirement a in the load requirement as output, and the big data system configuration parameter set s in the load requirement a All parameters in the node physical configuration parameters c a and business configuration parameter set v a All parameters in are used as input, and the impact factors of different configuration parameters on resource requirements are calculated using KL transformation, LightGBM and CatBoost. The median is taken as the impact factor of the corresponding parameter, and the top 50% of the configuration parameters with the highest impact factor are retained. sortθ,γ Each record in is represented as a tuple (a,u), where u is a configuration parameter that has a significant influence on this type of load;
[0129] S2-1-4. Use load data set A Sortθ,γ Build a load forecasting model. Sortθ,γ In each record (a,u), u is the input and a is the output. The Spark Sort load prediction model Γ is established using multiple support vector regression Boosting ensemble learning models. Sortθ,γ , and stored in the system Spark load prediction model library.
[0130] Furthermore, the specific steps of constructing the hardware resource response model library include:
[0131] S2-2-1. Construct a queue system simulation model library. The queue system simulation model library includes multiple basic queue system simulation models. The basic queue system simulation model is a stateful queue with a single service station. Its components include the queue state (the number of customers in the queue), the service time distribution of the queue service station, and the service time of the current customer being served. The basic queue simulation model includes a method for calculating the expected response time of resource demand. This method can calculate the expected waiting time of new customers based on the queue state, service time distribution, and the service time of the current customer being served in combination with the simulated queue system state diagram in the queue theory;
[0132] S2-2-2. Extract the resource requirements of each task, the hardware type and processing time for processing the corresponding requirements from the Spark Sort big data job portrait generated in step 2-8 to form a demand response. The extracted demand response is represented by a triple (resource requirement, demand processor, processing time);
[0133] S2-2-3. Build a cascade queue simulation model for different types of demand processors, including computing units, memory, disks, and networks. The cascade queue system simulation model consists of multiple basic queue system simulation models cascaded in a certain structure. The cascade structure of the cascade queue simulation system needs to be designed according to the commonly used demand processors. For example, the cascade queue simulation model structure for a dual-core CPU can be designed as Figure 2 The structure shown, the specific cascade structure can refer to the principle of multi-core CPU processing request; the concurrent paths of the cascade queuing simulation model of the memory can refer to the number of channels of the memory and other indicators, and the specific cascade structure can refer to the memory access principle; the parallel paths of the cascade queuing simulation model of the disk can refer to the number of concurrent seeks of the disk arm and other indicators, and the specific cascade structure can refer to the disk access principle; the cascade queuing simulation model of the network can refer to the network topology and other conditions, and the specific cascade structure can refer to the relevant principles of network data transmission. The cascade queuing simulation model after the structural design is completed needs to be verified using the demand response data obtained in step S2-2-2, and the cascade structure is adjusted based on the verification structure until the error between the response simulation time of the cascade queuing simulation model to the resource demand and the actual demand processing time in the request response is no longer significantly reduced. After the cascade structure adjustment is completed, the cascade queuing simulation model representing the demand processor is stored in the system resource response model library.
[0134] S3. Big data job behavior analysis: The user analyzes the software behaviors such as the task relationship and scheduling method of the Spark Sort big data job to be predicted and the Spark Sort big data cluster configuration that runs the Spark big data job to be predicted and inputs it into the system. The specific method steps are:
[0135] S3-1. The user determines the Spark big data cluster configuration where the Spark Sort big data job to be run and the performance to be predicted will be run, including the topology of the Spark big data cluster, the link status between the nodes of the Spark big data cluster, and the physical configuration of each node in the Spark big data cluster, including the CPU configuration, memory configuration, disk configuration, and network configuration, etc., to form a structured cluster configuration file;
[0136] S3-2. The user determines the Spark big data system configuration parameters of each node on the cluster, including the distributed file system HDFS configuration parameters of each node, the distributed computing framework Spark configuration parameters, the distributed scheduling framework YARN configuration parameters, including but not limited to the Spark big data system operating environment configuration parameters of the JVM configuration and the operating system configuration, and structures all the configuration parameters in this step and merges them into the cluster configuration file formed in step S3-1, and then inputs the cluster configuration file into the system;
[0137] S3-3. The user determines the number of Sample tasks, Shuffle Write tasks, and Shuffle Read tasks, task dependencies, scheduling methods, and some big data task-level business configuration parameters that the Spark Sort big data job whose performance is to be predicted must include. i,j , after forming a directed acyclic Spark Sort big data task graph with nodes and configuration parameters, a Spark Sort big data job description file is formed in XML format and input into the system. Users can obtain the above parameters by analyzing the big data files that the Spark Sort big data job needs to process. Generally speaking, when processing a Spark Sort job, for each file block in a big data file, Spark forms a Sample task. The output of each Sample task will be input into a different Shuffle Write task, and finally the corresponding Shuffle Read task will receive the output of the Shuffle Write task. For each Sample task, its business configuration parameter v i,j It at least includes the number of items to be sorted in the data block, the number of samples, and the label of the Spark big data cluster node where the file block corresponding to the Sample task is located; for the Shuffle Write task, its business configuration parameter v i,j At least it contains the number of items to be sorted that are exchanged during the Shuffle process and the label of the Spark cluster node that is estimated to run the Shuffle Write task; for the Shuffle Read task, its business configuration parameter v i,j At least it includes the number of items to be sorted after Shuffle and the labels of the Spark cluster nodes that are expected to run the Shuffle Read task. Since there is no obvious dependency between the Sample tasks, Shuffle Write tasks, and Shuffle Write tasks of the Spark Sort big data task in terms of input and output, they can run in parallel. Based on the above analysis, users can form Figure 4The Spark Sort big data task graph is shown. The Spark Sort big data task graph and scheduling method after processing are formed into an XML document with XML tags and rules pre-specified by the system and then input into the system;
[0138] S3-4. For each Spark big data task in the Spark Sort big data job to be predicted, the system searches for the cluster configuration document of step S3-2 according to the task running node label determined in step S3-3, obtains the node physical configuration information of the corresponding task running and the Spark big data system configuration information on the node, and respectively forms the node physical configuration parameter set c of the task. i,j And the Spark big data system configuration parameter set s for this task i,j ;
[0139] S3-5. The system combines the configuration parameters of the Spark big data task in the Spark Sort big data job to be predicted obtained in step S3-3 and step S3-4 to obtain the parameter set k of the Spark big data task i,j =(s i,j ,c i,j ,v i,j );
[0140] S3-6. For each Spark big data task in the Spark Sort big data job to be predicted, the system selects the Spark Sort load prediction model Γ for predicting different load types from the system Spark load prediction model library according to the type of each Spark big data task (Sample, Shuffle Write or Shuffle Read) sortθ,γ and Spark big data task parameter set k i,j The model input parameters in predict resource requirements. The resource requirements to be predicted include CPU usage requirements, memory usage requirements, disk read and write requirements, and network transmission requirements.
[0141] S3-7. The system compares the resource requirements obtained in step S3-6 with the physical configuration parameters c of the corresponding node in the Spark big data task parameter set i,j Some relevant parameters in the Spark big data task are checked to confirm that the node where the Spark big data task runs can meet the resource requirements. If the node cannot meet the resource requirements, it is necessary to decompose the current Spark big data task according to the node configuration and update the Spark Sort big data task graph and recalculate the business configuration parameters v′ of the new Spark big data task. i,j and running nodes, and find the physical configuration parameter set c′ corresponding to the new Spark big data task running node in the cluster configuration file of step S3-3i,j and Spark big data system configuration parameter set s′ i,j , forming the Spark big data task parameter set k of the new Spark big data task i,j =(s′ i,j ,c′ i,j ,v′ i,j ), and then repeat steps S3-6 and S3-7 for each new Spark big data task obtained by decomposition.
[0142] S4. Generation and execution of big data job simulation file: Combine the software line and load application of the Spark Sort big data job to be predicted, as well as the simulation template and resource response model to form a Spark Sort big data job simulation file. Compile and run the Spark big data job simulation file to obtain the performance prediction result of the Spark Sort big data job to be predicted. The specific method steps are:
[0143] S4-1. The user selects the programming language for executing the Spark Sort big data job simulation. The system retrieves the Spark big data job simulation template of the corresponding programming language according to the programming language selected by the user. The simulation template provides the Spark big data job simulation framework and can output the performance indicators of the entire simulation process after the simulation framework is executed;
[0144] S4-2. The system analyzes the task graph adjusted in step S3-7 and the Spark big data task scheduling method specified by the user in step S3-3, and arranges the Spark big data task behavior simulation module in the simulation template according to the running relationship of the Spark big data tasks contained in the Spark Sort big data job to be predicted in the task graph, and specifies the Spark big data task scheduling method in the simulation template, including but not limited to first-come-first-served, capacity scheduling, and fair scheduling;
[0145] S4-3. For each Spark big data task in the Spark Sort big data job whose performance is to be predicted, the system searches for the physical configuration of the node running the Spark big data task in the cluster configuration file input by the user in step S3-2, and selects the cascade queuing simulation model of the processor corresponding to the node physical configuration from the system hardware resource response model library, including the CPU resource response model, the memory resource response model, the disk resource response model and the network resource response model, and fills its calling interface into the interface specified by the simulation template obtained in step 4-1;
[0146] S4-4. For each Spark big data task included in the Spark Sort big data job whose performance is to be predicted, the system fills the resource requirements of the Spark big data task obtained in step S3-7 and the demand processor type of the node running the corresponding Spark big data task to process the corresponding resource requirements into the corresponding position of the simulation module of the corresponding Spark big data task in the simulation template according to the load mapping interface specified in the simulation template obtained in step S4-1;
[0147] S4-5 system according to step S3-2 user input cluster configuration file fill step S4-1 to obtain the simulation template other information to be filled;
[0148] S4-6. The system performs integration analysis and compilation on the completed simulation template to generate an executable Spark big data job simulation file;
[0149] S4-7. The system runs the executable Spark Sort big data job simulation file, and outputs the performance indicators of the Spark Sort big data job to be predicted after the execution of the file is completed, and forms a report.
[0150] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0151] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0152] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A method for modeling and simulating performance of big data systems. It is characterized in that The following steps are involved: S1. Big data job log collection and analysis: Run multiple known big data jobs on multiple big data clusters, collect and analyze big data job operation logs, and form a big data job profile including job operation configuration and performance indicators; S2. Construction of simulation model library: Based on the portraits of major data jobs, extract influential configuration parameters and the big data system status during the operation of big data jobs, build multiple big data task load prediction models and hardware resource response models of different computing resources, and form a simulation model library; the hardware resource response model construction process specifically includes the following steps: S2-2-1. Construct a queue system simulation model library; the queue system simulation model library includes multiple basic queue system simulation models; the basic queue system simulation model is a single-server stateful queue, and its components include queue status, service time distribution of the queue server, and service time of the currently serving customers; the basic queue simulation model includes a resource demand expected response time calculation method, which calculates the expected waiting time of new customers based on the queue status, service time distribution, and service time of the currently serving customers; S2-2-2. Extract the resource requirements of each task, the hardware type and processing time for processing the corresponding requirements from the generated big data job portrait to form a demand response. The extracted demand response is represented by a triple (resource requirement, demand processor, processing time); S2-2-3. Establish a cascade queuing simulation model for computing units, memory, disks, and networks; S3. Big data job behavior analysis: The user analyzes the big data job whose performance is to be predicted, obtains the software behavior of the big data job and the big data cluster configuration running the big data job, and inputs the software behavior and cluster configuration into the system, including the software behavior task relationship and scheduling method; S4. Generation and execution of big data job simulation files: The big data job simulation files are constructed by combining the software behavior and load application of the big data job to be predicted, as well as the simulation template and resource response model. The big data job simulation files are compiled and run to obtain the performance prediction results of the big data job to be predicted.
2. A big data system performance modeling and simulation method according to claim 1, It is characterized in that The step S1 is specifically as follows: S1-1. Collect multiple different types of test benchmarks and big data operations; S1-2. Deployment of computing resource usage collection scripts or collection tools on multiple heterogeneous big data clusters with different node topologies and nodes with different hardware configurations; S1-3. Start the computing resource usage collection script or collection tool deployed in step S1-2 on the multiple big data clusters described in step S1-2, and then start all the big data jobs collected in step S1-1 in sequence, and only when the previous big data job is completed can the next big data job be started; S1-4. Collect the logs generated by the big data job run in step S1-3 in the big data system and the computing resource usage during the job execution obtained by the computing resource usage collection script or collector started in step S1-3 to form a complete big data job log; S1-5. Analyze the big data job log and extract the configuration parameter group D of the big data cluster; S1-6. Analyze the big data job log to obtain the task set T that constitutes the big data job; for a big data job m, if it consists of n tasks, the task set of the big data job is represented by T m = {R m,1 ,R m,2 ,…,R m,n }, where R m,i is the i-th task of big data job m; S1-7. Analyze the big data job log and combine the big data cluster configuration parameter group D to obtain the node physical configuration parameter set c on which each big data task runs i,j ; For each big data task R i,j , the node physical configuration parameter set c that needs to be extracted i,j Including the parameters of the computing unit, memory, disk and network of the node where the task is running; S1-8. Analyze the big data job log and extract the big data system configuration parameter set s when each big data task is running i,j ; For each big data task R i,j , the big data system configuration parameter set s that needs to be extracted i,j Including the distributed file system configuration parameters of the node where the corresponding task runs, the distributed computing framework configuration parameters, the distributed scheduling framework configuration parameters, and the big data system operating environment configuration parameters; S1-9. Analyze the big data job log and extract the performance indicator set p for each big data task i,j ; For each big data task R i,j , the big data task performance indicator set p that needs to be extracted i,j Including the type of task, the running time of the task, the number and size of reads and writes to the file system during the task running process, the computing unit usage track and memory usage track during the task running process; S1-10. For each big data task, combine its node physical configuration parameter set, big data system configuration parameter set and big data task performance indicator set to form a big data task portrait; for big data task R i,j , whose portrait r i,j =(c i,j ,s i,j ,p i,j ); S1-11. Analyze the big data job log and extract the overall profile of the big data job; for big data job T i , its overall image z i Including the type of the job, the business configuration parameter set of the job, the task types included in the job and the quantity distribution of different types of tasks, the scheduling method of the job, the waiting time for the job to start, and the overall running time of the job; S1-12. Combine the overall profile of the big data job and the profile of the big data tasks included in the big data job to form a big data job profile; for the big data job T i , whose portrait t i =(z i ,{r i,1 ,r i,2 ,…,r i,n }), the generated big data job portrait will be stored in the big data job portrait library.
3. A big data system performance modeling and simulation method according to claim 2, It is characterized in that In step S1-5, the configuration parameter group D of the big data cluster includes the topological structure of the cluster, the status of each link in the cluster, and the cluster physical configuration information of each node in the cluster; The cluster physical configuration information of each node in the cluster includes configuration information of computing units, memory, disks and networks; The computing unit includes a CPU, a GPU and an FPGA.
4. A big data system performance modeling and simulation method according to claim 2, It is characterized in that In step S2, the load prediction model construction process specifically includes the following steps: S2-1-1. Extract all resource requirements of different types of tasks under different job types from the big data job portraits stored in the big data job portrait library; the resource requirements are any one of the four requirements of computing unit usage time, memory usage size, disk read and write size, network transmission start and end points, and transmission data volume; S2-1-2. For each resource requirement a, extract the big data system configuration parameter set s of the big data task to which the resource requirement belongs from the big data job portrait stored in the big data job portrait library i,j and node physical configuration parameter set c i,j And the business configuration parameter set of the big data job to which it belongs is converted into the business configuration parameter set v of the big data task i,j , forming the load demand; the load demand after extraction is expressed as a four-tuple (a, s a ,c a ,v a ), the load requirements of all load types γ of the big data tasks of type θ under the big data job of type τ are combined into a set to form a load data set θ of the corresponding type τ,θ,γ ; S2-1-3. For each load data set A τ,θ,γ Implement dimensionality reduction; the dimensionality reduction process takes the data volume in the resource demand a in the load demand as the output, and the big data system configuration parameter set s in the load demand a All parameters in the node physical configuration parameters c a and business configuration parameter set v a All parameters in are used as input, and the impact factors of different configuration parameters on resource requirements are calculated using the dimensionality reduction algorithm, and the configuration parameters with higher impact factors are retained; after dimensionality reduction, the load data set A τ,θ,γ Each record in is represented as a tuple (a,u), where u is a configuration parameter that has a significant influence on this type of load; S2-1-4. Use load data set A τ,θ,γ Build a load forecasting model based on A τ,θ,γ In each record (a,u), u is used as input and a is used as output. The load prediction model Γ is established using the machine learning model. τ,θ,γ , and stored in the system load prediction model library.
5. A big data system performance modeling and simulation method according to claim 1, It is characterized in that In step S2-2-3, the cascade queue simulation model is composed of multiple basic queue system simulation models cascaded in a certain structure; the cascade structure of the cascade queue simulation model needs to be designed according to the commonly used demand processor; After the structural design is completed, the cascade queue simulation model needs to be verified using the demand response data obtained in step S2-2-2, and the cascade structure is adjusted based on the verification structure until the error between the response simulation time of the cascade queue simulation model to resource demand and the actual demand processing time in the request response is no longer significantly reduced. After the cascade structure adjustment is completed, the cascade queue simulation model representing the demand processor will be stored in the system hardware resource response model library.
6. A big data system performance modeling and simulation method according to claim 4, It is characterized in that The step S3 is specifically as follows: S3-1. The user determines the big data cluster configuration that will be run for the big data job to be run and the big data job to be predicted; S3-2. The user determines the big data system configuration parameters of each node on the cluster, including the distributed file system configuration parameters of each node, the distributed computing framework configuration parameters, the distributed scheduling framework configuration parameters, and the big data system operating environment configuration parameters, and structures all the configuration parameters in this step and merges them into the cluster configuration file formed in step S3-1, and then inputs the cluster configuration file into the system; S3-3. The user determines the big data task type, task quantity, task dependency, scheduling method, and some big data task-level business configuration parameters that the big data job to be predicted should include. i,j ,After forming a directed acyclic task graph with nodes and configuration parameters, a big data job description file is formed in a structured form in combination with the scheduling method and input into the system; S3-4. For each big data task in the big data job to be predicted, the system searches the cluster configuration document of step S3-2 according to the task running node label determined in step S3-3, obtains the node physical configuration information of the corresponding task running and the big data system configuration information on the node, and respectively forms the node physical configuration parameter set c of the task. i,j And the big data system configuration parameter set s for this task i,j ; S3-5. The system combines the configuration parameters of the big data task in the big data job to be predicted in step S3-3 and step S3-4 to obtain the parameter set k of the big data task i,j =(s i,j ,c i,j ,v i,j ); S3-6. For each big data task in the big data job to be predicted, the system selects a load prediction model Γ for predicting different load types from the system load prediction model library based on the type of each big data task and the type of big data job to which each big data task belongs. τ,θ,γ and the big data task parameter set k i,j The model input parameters in predict resource requirements; the resource requirements to be predicted include computing unit usage requirements, memory usage requirements, disk read and write requirements, and network transmission requirements; S3-7. The system compares the resource requirements obtained in step S3-6 with the physical configuration parameters c of the corresponding node in the big data task parameter set i,j Some relevant parameters in the middle, confirm that the node where the big data task runs meets the resource requirements. If the node cannot meet the resource requirements, it is necessary to decompose the current big data task according to the node configuration, update the task graph and recalculate the business configuration parameters v′ of the new big data task i,j and running nodes, and find the physical configuration parameter set c′ corresponding to the new big data task running node in the cluster configuration file of step S3-3 i,j and the big data system configuration parameter set s′ i,j , forming the big data task parameter set k of the new big data task i,j =(s′ i,j ,c′ i,j ,v′ i,j ), and then repeat steps S3-6 and S3-7 for each new big data task obtained by decomposition.
7. A big data system performance modeling and simulation method according to claim 6, It is characterized in that The big data cluster configuration includes the topological structure of the big data cluster, the link status between the nodes of the big data cluster and the physical configuration of each node in the big data cluster. The physical configuration of each node includes computing unit configuration, memory configuration, disk configuration and network configuration, forming a structured cluster configuration file.
8. A big data system performance modeling and simulation method according to claim 6, It is characterized in that The step S4 is specifically as follows: S4-1. The user selects a programming language for executing big data job simulation. The system retrieves the simulation template of the corresponding programming language according to the programming language selected by the user. The simulation template provides a simulation framework and outputs the performance indicators of the entire simulation process after the simulation framework is executed; S4-2. The system analyzes the task graph adjusted in step S3-7 and the big data task scheduling method specified by the user in step S3-3, and arranges the operation relationship of the big data tasks included in the big data job to be predicted in the task graph, and specifies the big data task scheduling method in the simulation template; S4-3. For each big data task in the big data job to be predicted, the system searches for the physical configuration of the node running the big data task in the cluster configuration file input by the user in step S3-2, and selects the cascade queuing simulation model of the processor corresponding to the node physical configuration from the system hardware resource response model library, including the resource response model of the computing unit, the resource response model of the memory, the resource response model of the disk, and the resource response model of the network, and fills its calling interface into the interface specified by the simulation template obtained in step 4-1; S4-4. For each big data task included in the big data job to be predicted, the system fills the resource requirements of the big data task obtained in step S3-7 and the type of processor required to process the corresponding resource requirements of the node running the corresponding big data task into the corresponding position of the simulation module of the corresponding big data task in the simulation template according to the load mapping interface specified in the simulation template obtained in step S4-1; S4-5 system according to step S3-2 user input cluster configuration file fill step S4-1 to obtain the simulation template other information to be filled; S4-6 system completes the simulation template after filling the integration analysis and compilation work to generate an executable simulation file; S4-7. The system runs the executable big data job simulation file and outputs the performance indicators of the big data job to be predicted after the execution of the file is completed.
Citation Information
Patent Citations
Method and system for predicting relationships between business load and performance and between resource configuration and performance
CN107301466A
Systems and methods for generating performance prediction model and estimating execution time for applications
US20170169336A1