Distributed System Resource Optimization Allocation Method Based on LSTM and Genetic Algorithm

By combining the LSTM time prediction model and genetic algorithm, the allocation of user job resources in the distributed system is optimized, which solves the problem of job execution time that cannot be shortened due to improper resource allocation in the prior art, and achieves the optimization of resource allocation and the shortening of execution time.

CN114528094BActive Publication Date: 2025-05-30ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210041802.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-05-30
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

In a distributed computing environment, it is difficult for the prior art to effectively allocate the resources of user jobs, resulting in the inability to minimize job execution time, and allocating too many resources will increase communication overhead.

Method used

Using a method combining LSTM-based time prediction model and genetic algorithm, we analyze and predict user job information and optimize resource allocation to ensure that each job gets the most appropriate amount of resources.

Benefits of technology

The resource allocation optimization of user jobs is achieved, which significantly shortens the execution time of jobs and avoids the increase in communication overhead caused by excessive resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114528094B_ABST
    Figure CN114528094B_ABST
Patent Text Reader

Abstract

A resource allocation method based on an LSTM time prediction model and a genetic algorithm, comprising: 1) training a job execution time prediction model based on an LSTM network; 2) using a genetic algorithm to allocate reasonable amounts of resources to each job in a batch job; changing the fitness function of the genetic algorithm to an LSTM-based time prediction model, and through selection, crossover, and mutation of the genetic algorithm, iterating to obtain the amount of resources suitable for each job; 3) using a resource allocation algorithm based on the genetic algorithm to give different amounts of resources to different jobs; when the Spark distributed computing framework receives a job, it will calculate according to the amount of cluster resources that different jobs can use to obtain the shortest processing time of the job. After submitting the information of the batch jobs to be processed, the present invention can give an optimized resource allocation scheme for each job, thereby achieving the optimization goal of the shortest running time of the batch jobs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of neural networks, task scheduling, and optimization algorithms. A time prediction model is designed through a neural network. By combining the time prediction model with an optimization algorithm and applying it to resource allocation, the user jobs can be allocated an appropriate amount of resources to achieve the optimization goal of the shortest execution time of the user jobs. Background Art

[0002] In a distributed computing environment, the amount of resources allocated to a job determines the execution speed of the job. Research shows that allocating too many resources to a job not only fails to shorten the running time of the job, but also increases the communication overhead between job running nodes, resulting in a longer execution time of the job. Therefore, it is very necessary to study the resource allocation method of the distributed system to allocate the most suitable amount of resources for each job. Summary of the Invention

[0003] The present invention aims to overcome the above-mentioned disadvantages of the prior art and provides a method for optimizing resource allocation in a distributed system based on LSTM and genetic algorithms.

[0004] The present invention uses an LSTM time prediction model and combines it with a genetic algorithm to allocate resources for user jobs, thus solving the disadvantage of not allocating the amount of resources for user jobs in the previous default scheduling environment.

[0005] The method for optimizing resource allocation in a distributed system based on LSTM and genetic algorithms of the present invention includes the following steps:

[0006] 1) Train a job execution time prediction model based on the LSTM network. The input of the LSTM network is the information of the job, and the output information is the running time of the job.

[0007] 2) Use the genetic algorithm to allocate a reasonable amount of resources for each job in a batch job. Modify the fitness function of the genetic algorithm to be based on the LSTM time prediction model, and through the selection, crossover, and mutation of the genetic algorithm, iterate out the appropriate amount of resources for each job.

[0008] 3) Use the resource allocation algorithm based on the genetic algorithm to give different amounts of resources for different jobs. When the Spark distributed computing framework receives a job, it will calculate according to the amount of cluster resources that different jobs can use to obtain the shortest processing time of the job.

[0009] Further, step 1) specifically includes:

[0010] 1.1) During the operation of the cluster, analyze the influencing factors of the running time of user jobs, and finally determine five influencing factors for the execution time of user jobs: job type, job data volume, number of CPU cores used by the job, memory size used by the job, and number of nodes used by the job;

[0011] 1.2) Run different jobs in a real distributed cluster (the parameters describing the jobs are job type, data volume, number of CPU cores used by the job, memory size, and number of nodes), and collect the job running time as the training and test data for the time prediction model;

[0012] 1.3) The inputs of the time prediction model based on LSTM are job type, job data volume, number of CPU cores used by the job, memory size, and number of nodes, and the model output is the running time of the job;

[0013] The loss function adopted by the model is the mean square error (MSE), and the calculation method is as follows:

[0014]

[0015] where y i represents the actual running time of the job, represents the predicted execution time of the job, and m is the number of job samples;

[0016] 1.4) Select the hyperparameters of the model; for the value of the learning rate, use the step-by-step experiment method; first, conduct experiments on the classic learning rate values, and use the corresponding loss values in the iterative process to determine the order of magnitude of the optimal learning rate; then, adjust the values of the learning rate within this order of magnitude and conduct further experiments to finally obtain the optimal learning rate; for the number of iterations, conduct experiments with different numbers of iterations, and take the data with the smallest corresponding loss value as the optimal number of iterations; select different numbers of network layers to run the model, and take the data with the smallest corresponding loss value as the optimal number of network layers; select different Dropout rates to run the model, and take the data with the smallest corresponding loss value as the optimal Dropout rate;

[0017] The number of hidden layer nodes is determined using the following empirical formula and experiments;

[0018]

[0019] where n h , n i , n o represent the number of hidden layer nodes, input layer nodes, and output layer nodes of the neural network respectively. The optimization search algorithm for determining the number of hidden layer nodes includes the following steps:

[0020] (a) Determine the initial value range of the number of hidden layer nodes;

[0021] (b) Narrow the value range;

[0022] (c) Expand the value range;

[0023] (d) Determine the optimal number of hidden layer nodes.

[0024] Further, step 2) specifically includes:

[0025] 2.1) Conduct chromosome coding design; the chromosome is used to describe the user job information that the cluster needs to process, and binary coding is adopted; in the chromosome, each job occupies the same number of bits, which respectively represent the job type, the data volume of the job, the number of CPU cores used by the job, the memory size, and the number of nodes;

[0026] 2.2) Generate an initial population according to the batch jobs to be processed; generate individuals according to the generation rules of the chromosome; since the type and data volume of each job are fixed, the corresponding coding values in the chromosome are determined, and the other bits of the coding are randomly generated 0 or 1; if this individual does not conform to the application background, such as the memory is 0, or the number of CPU cores is 0, then it is discarded;

[0027] 2.3) Use the time prediction model based on LSTM as the fitness function of the genetic algorithm;

[0028] 2.4) The selection operation selects excellent individuals according to the fitness function to enter the next iteration. In the present invention, the roulette wheel selection strategy is selected, which is one of the most basic selection strategies, and the probability of an individual in the population being selected is proportional to the value of the corresponding fitness function of the individual; accumulate the fitness values of all individuals in the population and then normalize them, and select the individual corresponding to the area where the random number falls. In the present invention, that is, find out the resource allocation scheme that can make the execution time of the batch job shorter;

[0029] 2.5) The crossover operation randomly selects and replaces part of the structures of two parent individuals with a certain probability to generate new individuals, which is an important method to obtain excellent new individuals. The crossover operation of the present invention randomly exchanges the rest of the chromosome on the premise of keeping the data and type of each job unchanged. Crossover is the main method to generate new individuals in the genetic algorithm, so the crossover probability generally takes a relatively large value; however, once the value is too large, it will also destroy the good state in the population and have an adverse impact on evolution; if the value is too small, the speed of generating new individuals will be too slow. In the present invention, a suitable crossover probability is selected and used;

[0030] 2.6) Mutation randomly changes the values of certain genes of individuals in the population with a very small mutation probability. According to the characteristics of resource allocation, the mutation operation in the present invention stipulates that random mutation is performed on the coding other than the data and type parts of the operations in the chromosome. For the mutation probability: if the mutation probability takes a relatively large value, although more new individuals can be generated, it may also lead to the destruction of some good individuals, making the performance of the genetic algorithm close to that of the random search algorithm; if the mutation probability takes too small a value, it will cause mutation;

[0031] 2.7) Iterate steps 2.4), 2.5) and 2.6). After iterating a certain number of times, the optimal resource allocation scheme for each operation is obtained.

[0032] Further, step 3) specifically includes:

[0033] 3.1) Run different job types in a real distributed cluster. For different job data volumes, corresponding to different numbers of nodes, memory sizes, and CPU cores, obtain the running time of the jobs. After obtaining a certain amount of data, construct an LSTM-based time prediction model;

[0034] 3.2) Use the genetic algorithm in step 2) to find a suitable resource allocation scheme for each job in the batch jobs;

[0035] 3.3) Allocate the specified resource allocation scheme for each job in the Spark cluster and execute the jobs.

[0036] Further, the input of the LSTM network described in step 1) is the information of the jobs, including job type, data volume, required memory, CPU cores, and number of nodes.

[0037] The present invention mainly includes two parts, namely the LSTM-based time prediction method and the genetic algorithm-based resource allocation method. For the time prediction model, first train the LSTM time prediction model based on the historical running data of the jobs to find the most suitable hyperparameters. The time prediction method can predict the job running time according to the characteristics of the jobs and the amount of resources they use. In the genetic algorithm-based resource allocation method, the job running time is used as the fitness function of the genetic algorithm. The chromosome represents the information of the jobs, that is, job type, data volume, required memory, CPU cores, and number of nodes, and binary coding is used for chromosome coding. After submitting the information of the batch jobs to be processed, the present invention can give the optimized resource allocation scheme for each job, so as to achieve the optimization goal of the shortest running time of the batch jobs.

[0038] The advantages of the present invention are: considering the consumption of data transmission between nodes, using as few nodes as possible to complete the jobs; at the same time, according to the characteristics of different jobs, finding the most suitable resource allocation scheme for the jobs to obtain the optimal execution time. Description of the Drawings

[0039] Figure 1 is the flowchart of the present invention.

[0040] Figure 2 is the time prediction model of the present invention.

[0041] Figure 3 The mean squared error when the learning rate of the time prediction model is 0.1.

[0042] Figure 4 The mean squared error when the learning rate of the time prediction model is 0.01.

[0043] Figure 5 The mean squared error of the time prediction model for different numbers of iterations.

[0044] Figure 6 is the genetic algorithm of the present invention.

[0045] Figure 7 is the encoding method of the genetic algorithm of the present invention. Detailed Description of the Invention

[0046] The present invention will be further described below with reference to the drawings.

[0047] This embodiment proposes a distributed system resource optimization allocation method based on LSTM and genetic algorithm for a 1G WordCount job, including the following steps:

[0048] 1) Use a real cluster to operate on tasks of different types and different data volumes to obtain the execution time of the corresponding tasks. Establish a time prediction model based on the LSTM network, put historical data into the model for training, and optimize the LSTM network parameters at the same time to obtain a time prediction model suitable for this historical data;

[0049] 2) Design a genetic algorithm based on cluster resource allocation. By encoding the jobs and initializing the population, replace the fitness function with the time prediction model, and then through selection, crossover, and mutation, perform reasonable resource allocation through the iterative optimization of the genetic algorithm to obtain the shortest processing time of the batch jobs.

[0050] 3) Modify the default scheduling method of Spark. After the user submits a job, use the genetic algorithm to allocate the resources used by the task job, so that each job can obtain a suitable cluster resource allocation, shortening the job execution time.

[0051] Step 1) proposes a user job time prediction model based on the LSTM recurrent neural network, specifically including:

[0052] 1.1) During the operation of the cluster, analyze the influencing factors of the running time of user jobs, and finally determine five influencing factors for the execution time of user jobs: job type, job data volume, number of CPU cores used by the job, memory size used by the job, and number of nodes used by the job.

[0053] 1.2) Run different jobs in a real distributed cluster (the parameters describing the jobs are job type, data volume, number of CPU cores used by the job, memory size, and number of nodes), and collect the job running time as the training and test data for the time prediction model.

[0054] 1.3) As Figure 2 The input x of the model 1 , x 2 , x 3 , x 4 , x 5 are the job type, job data volume, number of CPU cores used by the job, memory size used by the job, and number of nodes used by the job respectively. y is the output of the model.

[0055] The loss function adopted by this model is the mean square error (MSE), and the calculation method is as follows:

[0056]

[0057] Among them, y i represents the actual running time of the job, represents the predicted execution time of the job, and m is the number of job samples.

[0058] 1.4) Conduct model hyperparameter selection. During the process of model establishment, the selection of hyperparameters has a very crucial impact on the prediction results of the model. Therefore, when determining the final training model, it is necessary to conduct experimental comparisons to determine the model. The present invention also proposes a method for the selection of hyperparameters. For the learning rate and the number of iterations, the present invention adopts a step-by-step experimental method. First, conduct experiments on the classic learning rate values, and use the corresponding loss values during the iteration process to determine the order of magnitude of the optimal learning rate. Subsequently, adjust the values of the learning rate within this order of magnitude and conduct further experiments, such as Figure 3 , 4, and finally obtain the optimal learning rate of 0.02. Through Figure 5It is determined that the number of network iterations is 300 times. Regarding the number of network layers, increasing the number of network layers will improve the test accuracy of the model. However, for LSTM, blindly increasing the number of network layers will make the model too complex. Therefore, through testing, the most appropriate number of model layers is selected as 2 layers. Regarding the number of hidden layers and the Dropout rate, the hidden layers in a neural network can help the network model learn the hidden associations between data. If the number of hidden layer nodes is too small, the model will not be able to fully explore the implicit relationships between various parameters, resulting in poor prediction performance. If the number of hidden layer nodes is too large, overfitting is likely to occur, and at the same time, the network becomes too complex, increasing the training time of the network. An empirical formula for determining the number of hidden layer nodes is given:

[0059]

[0060] where n h , n i , n o represent the number of hidden layer nodes, the number of input layer nodes, and the number of output layer nodes of the neural network respectively. The optimization search algorithm for determining the number of hidden layer nodes includes the following steps:

[0061] (a) Determine the initial value range of the number of hidden layer nodes. From formula (2), it can be seen that the initial value range is [a, b]. In the present invention, n i = 5, n o = 1. Therefore, it is calculated that a = 3 and b = 16. Therefore, the initial value range of the number of hidden layer nodes is [3, 16].

[0062] (b) Narrow the value range. Calculate the first test point x 1 = 0.618×(b - a) + a = 0.618×13 + 3 = 11 and the second test point x 2 = 0.382×(b - a) + a = 0.382×13 + 3 = 8 through the golden section ratio formula. Through experiments, it is obtained that the network loss error corresponding to 11 hidden layer nodes is less than the corresponding value of 8 hidden layers. Therefore, the interval is narrowed to [8, 16].

[0063] (c) Expand the value range. Use the golden section method to calculate the expansion value c such that 16 = 0.618×(c - a) + a, then c = 24. Therefore, the expansion interval is [16, 24].

[0064] (d) Determine the optimal number of hidden layer nodes. Combining the results of (b) and (c), determine the value range of the number of hidden layer nodes as [8, 24]. Conduct experiments in this value range to obtain the MSE, MAE, and MAPE corresponding to each number of nodes, as shown in Table 1. It can be seen that when the number of hidden layer nodes is 24, the network performance is the best. Therefore, the number of hidden layer nodes is taken as 24.

[0065] The Dropout rate is set to effectively reduce the probability of overfitting and play a role in regularization. The present invention selects an appropriate Dropout rate. Through comparative experiments, the DropOut value is determined to be 0.1.

[0066] Step 2) uses a genetic algorithm to allocate task resource amounts for different user jobs, specifically including:

[0067] 2.1) Conduct chromosome coding design. As Figure 7 Chromosomes are used to describe the user job information that the cluster needs to process, and binary coding is adopted. In a chromosome, each job occupies the same number of bits, which respectively represent the job type, the data volume of the job, the number of CPU cores used by the job, the memory size, and the number of nodes.

[0068] 2.2) Generate an initial population according to the batch jobs to be processed. Generate individuals according to the generation rules of chromosomes. Since the type and data volume of each job are fixed, the corresponding coding values in the chromosome are determined, and the other bits of the coding are randomly generated 0s or 1s. If this individual does not conform to the application background, such as the memory being 0 or the number of CPU cores being 0, it is discarded.

[0069] 2.3) Use a time prediction model based on LSTM as the fitness function of the genetic algorithm.

[0070] 2.4) The selection operation selects excellent individuals according to the fitness function to enter the next iteration. In the present invention, the roulette wheel selection strategy is selected, which is one of the most basic selection strategies. The probability of an individual in the population being selected is proportional to the value of the corresponding fitness function of the individual. Accumulate the fitness values of all individuals in the population and then normalize them, and select the individual corresponding to the area where the random number falls. In the present invention, that is, find a resource allocation plan that can make the execution time of the batch job shorter;

[0071] 2.5) The crossover operation randomly selects and replaces part of the structures of two parent individuals with a certain probability to generate new individuals, which is an important method to obtain excellent new individuals. The crossover operation of the present invention randomly exchanges the remaining parts of the chromosomes on the premise of keeping the data and type of each job unchanged. Crossover is the main method to generate new individuals in the genetic algorithm, so the crossover probability generally takes a relatively large value. However, once the value is too large, it will also destroy the good state in the population and have an adverse impact on evolution; if the value is too small, the speed of generating new individuals will be too slow. In the present invention, an appropriate crossover probability is selected for use;

[0072] 2.6) Mutation randomly changes the values of some genes of individuals in the population with a very small mutation probability. According to the characteristics of resource allocation, the mutation operation in the present invention stipulates that random mutation is performed on the coding other than the data and type parts of the operations in the chromosome. For the mutation probability: if the mutation probability takes a relatively large value, although more new individuals can be generated, it may also lead to the destruction of some good individuals, making the performance of the genetic algorithm close to that of the random search algorithm; if the mutation probability takes too small a value, it will cause mutation;

[0073] 2.7) Iterate steps 2.4), 2.5) and 2.6). After iterating a certain number of times, the optimal resource allocation scheme for each operation is obtained.

[0074] Step 3) Verify the effectiveness of the algorithm system through the Spark big data distributed framework, specifically including:

[0075] 3.1) Build 5 nodes in a real cluster, namely, the Master has 2 CPU cores, 5G of memory, and 80G of disk; Slave1 has 2 CPU cores, 5G of memory, and 40G of disk, has 2 CPU cores, 5G of memory, and 80G of disk; Slave3 has 1 CPU core, 5G of memory, and 40G of disk; Slave4 has 1 CPU core, 5G of memory, and 40G of disk; Run WordCount and Sort with different data volumes of jobs generated based on BigDataBench in the real distributed cluster, and then conduct experiments to obtain the running time of the jobs for different numbers of nodes, memory sizes, and CPU cores. After obtaining a certain amount of data, build the corresponding time prediction model based on step 1).

[0076] 3.2) Use the time prediction model as the fitness function in the genetic algorithm of step 2), and obtain the appropriate job resource amount for each job through the iteration of the genetic algorithm.

[0077] 3.3) When the cluster receives a job submission, first submit the job type and job data volume to the genetic algorithm for running, combine the cluster resource amount to obtain the optimal resource allocation strategy, and then submit the corresponding job resource amount to the Spark cluster. For example, for a 1G WordCount job, through the method of the present invention, it can be obtained that 1 CPU core and 3G of memory of the Master node, 1 CPU core and 3G of memory of the Slave1 node, and 1 CPU core and 3G of memory of the Slave2 node are allocated, and the running time is reduced by 9.89%.

[0078] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments, and the protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept of the present invention.

Claims

1. A method for optimizing resource allocation in a distributed system based on LSTM and genetic algorithm, characterized in that, it includes the following steps: 1) Train a job execution time prediction model based on the LSTM network; the input of the LSTM network is the information of the job, where the information of the job includes job type, data volume, required memory, number of CPU cores, and number of nodes, and the output information is the running time of the job; 2) Use the genetic algorithm to allocate reasonable resource amounts for each job in the batch job; change the fitness function of the genetic algorithm to the time prediction model based on LSTM, and through the selection, crossover, and mutation of the genetic algorithm, iterate out the resource amount suitable for each job; 3) Use the resource allocation algorithm based on the genetic algorithm to give different resource amounts for different jobs; when the Spark distributed computing framework receives a job, it will calculate according to the cluster resource amounts that different jobs can use to obtain the shortest processing time of the job; Step 1) specifically includes: 1.1) During the operation of the cluster, analyze the influencing factors of the running time of user jobs, and finally determine five influencing factors for the execution time of user jobs: job type, data volume of the job, number of CPU cores used by the job, memory size used by the job, and number of nodes used by the job; 1.2) Run different jobs in a real distributed cluster, describe the parameters of the job as job type, data volume, number of CPU cores used by the job, memory size, and number of nodes, and collect the running time of the job as the training and test data of the time prediction model; 1.3) The inputs of the time prediction model based on LSTM are job type, job data volume, number of CPU cores used by the job, memory size, and number of nodes respectively, and the model output is the running time of the job; The loss function adopted by the model is the mean square error (MSE), and the calculation method is as follows: Among them, y i represents the actual running time of the job, represents the predicted execution time of the job, and m is the number of job samples; 1.4) Select the hyperparameters of the model; for the value of the learning rate, use the method of step-by-step experiments; first, conduct experiments on the classic learning rate values, and use the corresponding loss values in the iterative process to determine the order of magnitude of the optimal learning rate; then, adjust the value of the learning rate in this order of magnitude and conduct further experiments to finally obtain the optimal learning rate; for the number of iterations, conduct experiments with different numbers of iterations, and take the data with the smallest corresponding loss value as the optimal number of iterations; select different numbers of network layers to run the model, and take the data with the smallest corresponding loss value as the optimal number of network layers; select different Dropout rates to run the model, and take the data with the smallest corresponding loss value as the optimal Dropout rate; The number of hidden layer nodes is determined by the following empirical formula and experiments; where n h , n i , n o respectively represent the number of hidden layer nodes, the number of input layer nodes, and the number of output layer nodes of the neural network; the optimization search algorithm for determining the number of hidden layer nodes includes the following steps: (a) Determine the initial value range of the number of hidden layer nodes; (b) Narrow the value range; (c) Expand the value range; (d) Determine the optimal number of hidden layer nodes.

2. According to the method for optimizing resource allocation in a distributed system based on LSTM and genetic algorithm described in claim 1, characterized in that : Step 2) specifically includes: 2.1) Conduct chromosome coding design; the chromosome is used to describe the user job information that the cluster needs to process, and binary coding is adopted; in the chromosome, each job occupies the same number of bits, which respectively represent the job type, the data volume of the job, the number of CPU cores used by the job, the memory size, and the number of nodes; 2.2) Generate an initial population according to the batch jobs to be processed; generate individuals according to the generation rules of chromosomes; since the type and data volume of each job are fixed, the corresponding coding values in the chromosome are determined, and the other bits of the coding are randomly generated 0s or 1s; if this individual does not conform to the application background, that is, when the memory is 0 or the number of CPU cores is 0, it is discarded; 2.3) Use the time prediction model based on LSTM as the fitness function of the genetic algorithm; 2.4) The selection operation selects individuals with excellent performance according to the fitness function to enter the next iteration; the roulette wheel selection strategy is selected, which is one of the most basic selection strategies. The probability of an individual in the population being selected is proportional to the value of the corresponding fitness function of the individual; the fitness values of all individuals in the population are accumulated and then normalized, and the individual corresponding to the area where the random number falls is selected, that is, a resource allocation scheme that can make the batch job execution time shorter is found; 2.5) The crossover operation randomly selects and replaces part of the structures of two parent individuals with a certain probability to generate new individuals. The crossover operation randomly exchanges the rest of the chromosome on the premise of keeping the data and type of each job unchanged; 2.6) According to the characteristics of resource allocation, the mutation operation stipulates that the coding other than the data and type parts of the job in the chromosome is randomly mutated; 2.7) Iterate steps 2.4), 2.5), and 2.6). After iterating a certain number of times, the optimal resource allocation scheme for each job is obtained.

3. The distributed system resource optimization allocation method based on LSTM and genetic algorithm according to claim 1, characterized in that : Step 3) specifically includes: 3.1) Run different job types, job data volumes, corresponding to different numbers of nodes, memory sizes, and CPU cores in a real distributed cluster to obtain the running time of the jobs. After obtaining a certain amount of data, construct a time prediction model based on LSTM; 3.2) Use the genetic algorithm in step 2) to find a suitable resource allocation scheme for each job in the batch job; 3.3) Allocate the specified resource allocation scheme for each job in the Spark cluster and execute the job.