A scheduling task configuration method and apparatus
Through the deep reinforcement learning algorithm DDPG and Encoder-Decoder model, the problems of low accuracy and poor efficiency of Spark configuration parameters are solved, and the efficient performance improvement of Spark scheduling tasks is achieved.
Patent Information
- Application Number
- CN202110009185.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-05
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-01-05
AI Technical Summary
In the prior art, the Spark configuration parameters have low accuracy and poor efficiency, resulting in the impact of the scheduling task processing performance.
Deep reinforcement learning algorithms DDPG and Encoder-Decoder models are used to collect task information from the data mart historical execution log, and the best Spark configuration parameters are obtained through Actor network and Critic network training.
Accurate and fast Spark parameter configuration is realized, greatly improving the performance of scheduling tasks.
Smart Images

Figure CN113760497B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a scheduling task configuration method and apparatus. Background Art
[0002] Spark is a general parallel framework similar to Hadoop MapReduce open-sourced by UC Berkeley AMP lab. Due to the characteristic of storing intermediate results in memory, Spark runs iterative and interactive programs 10 times faster than the traditional disk computing framework Hadoop. Optimization of Spark configuration parameters has always been one of the research hotspots in big data systems. Since there are many configuration parameters (more than 100), the performance is greatly affected by the configuration parameters, and application programs have different characteristics. Therefore, using the default configuration far from achieves the best performance.
[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0004] The prior art performs manual configuration parameters and automatic configuration parameters. The disadvantage of the manual configuration parameter method is that it is too time-consuming and requires the user to have a deep understanding of the Spark operation mechanism, the meaning, function, and value range of the parameters. The user needs to manually increase or decrease the Spark parameter values, then configure Spark and run the application program to find the parameter value that makes the execution time the shortest. Since the optimal configuration parameters are different for different cluster environments, different application programs, and different input data sets, the manual configuration parameter method is a time-consuming and boring task. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a scheduling task configuration method and apparatus, which can solve the problem that the processing performance of scheduling tasks is affected due to low accuracy and poor efficiency of setting Spark configuration parameters.
[0006] To achieve the above object, according to one aspect of the embodiments of the present invention, a scheduling task configuration method is provided, including collecting relevant information of a task from the historical execution log of a data mart according to the id of a spark scheduling task; inputting the information into the Actor network in a preset DDPG model to obtain an action space value corresponding to the information; wherein, the Actor network in the DDPG model adopts an Encoder-Decoder model; inputting the action space value and the information into the Critic network in the preset DDPG model for training to obtain a corresponding reward value; and then calculating a preset network loss function according to the reward value to obtain a DDPG model corresponding to the maximum reward value as a configuration model;
[0007] Finally, obtain the target information of the current Spark scheduling task, obtain the corresponding Spark configuration parameters through the configured model, and then execute the current Spark scheduling task based on the Spark configuration parameters.
[0008] Optionally, collect relevant information about the task from the data mart historical execution log, including:
[0009] Spark configuration parameters, relevant metrics feedback by the data mart after the Spark scheduling task is completed, and the operators and task information included in each stage of the directed acyclic graph generated when starting the Spark scheduling task.
[0010] Optionally, input the information into the Actor network in the preset DDPG model to obtain the action space value corresponding to the information, including:
[0011] Use the encoder in the Encoder-Decoder model to splice the operators and task information included in each stage as the input layer for embedding to obtain an embedding vector matrix of n*m dimensions corresponding to each stage; then process the matrix through a convolutional neural network to output the processed matrix.
[0012] Optionally, before processing the matrix through a convolutional neural network, include:
[0013] Generate a two-dimensional matrix from the embedding vector matrices of n*m dimensions corresponding to all stages through a preset splicing model respectively.
[0014] Optionally, output the processed matrix, including:
[0015] Input the processed matrix into a preset recurrent neural network to output the matrix processed by the recurrent neural network.
[0016] Optionally, after inputting into a preset recurrent neural network, include:
[0017] Input the matrix output by the recurrent neural network into the attention layer to identify the importance label of each state space and the influence degree label of the operators in each state space on the relevant metrics.
[0018] Optionally, input the information into the Actor network in the preset DDPG model to obtain the action space value corresponding to the information, including:
[0019] Using the decoder in the Encoder-Decoder model, the encoded information is decoded by a convolutional neural network to obtain the configuration parameters to be processed, and the final configuration parameters are randomly extracted from the configuration parameters to be processed as the action space value corresponding to the information.
[0020] In addition, the present invention also provides a scheduling task configuration device, including an acquisition module for collecting relevant information of a task from the historical execution log of a data mart according to the id of a spark scheduling task; inputting the information into the Actor network in a preset DDPG model to obtain the action space value corresponding to the information; wherein, the Actor network in the DDPG model adopts an Encoder-Decoder model; inputting the action space value and the information into the Critic network in the preset DDPG model for training to obtain the corresponding reward value; and then calculating a preset network loss function according to the reward value to obtain the DDPG model corresponding to the maximum reward value as the configuration model; finally; a processing module for obtaining the target information of the current spark scheduling task, obtaining the corresponding spark configuration parameters through the configuration model, and then executing the current spark scheduling task based on the spark configuration parameters.
[0021] One embodiment of the above invention has the following advantages or beneficial effects: The present invention calculates the optimal parameters configured in the execution of a spark task through the reinforcement learning algorithm DDPG, the Encoder-Decoder model, and the evaluation indicators of the execution effect of historical spark tasks (execution time, cumulative memory resource usage, task read / write data volume, disk IO size, etc.). Thus, the present invention realizes accurate and fast spark parameter configuration, and greatly improves the performance of spark scheduling tasks.
[0022] The further effects of the above non-conventional optional methods will be described in combination with specific embodiments below. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention. Among them:
[0024] Figure 1 is a schematic diagram of the main process of the scheduling task configuration method according to the first embodiment of the present invention;
[0025] Figure 2 is a schematic diagram of a stage in a directed acyclic graph according to an embodiment of the present invention;
[0026] Figure 3 is a schematic diagram of the scheduling task configuration method according to the second embodiment of the present invention;
[0027] Figure 4 is a schematic diagram of a two-dimensional matrix according to an embodiment of the present invention;
[0028] Figure 5 is a schematic diagram of the main modules of a scheduling task configuration device according to an embodiment of the present invention;
[0029] Figure 6 is an exemplary system architecture diagram to which an embodiment of the present invention can be applied;
[0030] Figure 7 is a schematic structural diagram of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. Detailed implementation manners
[0031] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0032] Figure 1 is a schematic diagram of the main process of a scheduling task configuration method according to the first embodiment of the present invention, as Figure 1 shown, the scheduling task configuration method includes:
[0033] Step S101, collect relevant information of the task from the data mart historical execution log according to the id of the spark scheduling task.
[0034] In an embodiment of the present invention, during the execution of a Spark task, a directed acyclic graph (DAG) corresponding to the task is generated before each task starts. During subsequent execution, the task is executed according to the directed acyclic graph. Finally, execution time of the task, cumulative memory resource usage, task read / write data volume, disk IO size and other metrics can be obtained. The data mart environment can be regarded as Environment, then the parameter configuration problem of the spark task can be regarded as an MDP decision problem.
[0035] Among them, each time a task is executed, the intelligent agent Agent packages and submits the application to the data mart according to the configuration parameters of the task. The data mart executes the task based on the application and obtains a feedback signal. The feedback information may include execution time, cumulative used memory resource amount, task read / write data volume, disk IO size, etc. The intelligent agent Agent modifies the configuration parameters according to the feedback information of the data mart, repackages and submits again. The data mart executes again and obtains feedback information until the feedback information obtained by the data mart can meet the preset index requirements.
[0036] In addition, a directed acyclic graph (DAG) is used to model the relationship of the resilient distributed dataset (RDD) in Spark, describing the dependency relationship of the RDD. This relationship is also called lineage. The dependency relationship of the RDD is maintained by Dependency. The corresponding implementation of the DAG in Spark is DAGScheduler. The resilient distributed dataset (RDD) is a highly restricted shared memory model. A directed acyclic graph (DAG) refers to a directed graph without loops.
[0037] A data mart, also known as a data market, is a repository that collects data from operational data and other data sources serving a particular group of professionals.
[0038] In some embodiments, collecting relevant information about tasks from the historical execution logs of the data mart may include spark configuration parameters, relevant metrics feedback by the data mart after the completion of the spark scheduling task execution, and the operator and task information (such as Figure 2 shown) included in each stage of the directed acyclic graph generated when starting the spark scheduling task.
[0039] Among them, the spark configuration parameters may include the maximum number of CPUs (threads) used by the driver, the size of the heap memory applied for by a single executor, the maximum number of concurrent tasks (tasks) of a single executor, the number of executors, etc. The driver is a piece of application program submitted by the user during the spark development process, used to create the spark context, divide the dataset and generate a directed acyclic graph, and coordinate with other groups in the spark, coordinate resources, etc. The executor is a process for executing the application program in the spark. A single execution process includes multiple execution tasks.
[0040] It should be noted that the present invention is a configuration model trained based on the DDPG model (the full name of DDPG is Deep Deterministic Policy Gradient, that is, Deep Deterministic Policy Gradient), and the DDPG model is an Actor-Critic framework, that is, the DDPG model includes an Actor network and a Critic network, where the Actor network is the Policy Gradient algorithm and the Critic network is the Q-learning algorithm. The present invention creatively uses the Encoder-Decoder model to process the relevant information of the tasks collected from the historical execution logs of the data mart in the Actor network.
[0041] The relevant metrics fed back by the data mart after the spark scheduling task is completed are used as the reward value reward. For example: after the task is completed: the comprehensive weighted value of metrics such as the completion execution time, the cumulative memory resource usage, the task read / write data volume, and the disk IO size is used as the reward reward.
[0042] In a preferred embodiment, the operators included in each stage of the DAG generated each time the spark task is started and the task information included in each state (for example: the number of Tasks, Duration, GC Time, ShuffleRead data volume, ShuffleWrite data volume, etc.) are used as the state space state, that is, the state space state can be input into the Actor network to obtain the corresponding action space value.
[0043] Step S102, input the information into the Actor network in the preset DDPG model to obtain the action space value corresponding to the information; input the action space value and the information into the Critic network in the preset DDPG model for training to obtain the corresponding reward value; and then calculate the preset network loss function according to the reward value to obtain the DDPG model corresponding to the maximum reward value as the configuration model.
[0044] Get the final.
[0045] In some embodiments, encoding the information through an encoder Encoder includes:
[0046] Concatenate the operators and task information included in each stage as the input layer for embedding to obtain an embedding vector matrix of n*m dimensions corresponding to each stage. Then, use the CNN layer to process the matrix, and then output the processed matrix through the state output layer.
[0047] Among them, a stage is that a Spark task forms a directed acyclic graph according to the dependencies between RDDs (Spark resilient datasets) and divides it into multiple interdependent execution stages, which are composed of a group of parallel tasks.
[0048] Preferably, the state space parameters (including the stage operator: sti and the task information of each stage: ei) are converted into one-hot format data through the InputLayer, thus realizing the splicing of the state space parameters.
[0049] It should be noted that the state space parameters spliced by the InputLayer are embedded through the EmbbedingLayer (embedding layer) to obtain an n*m-dimensional matrix of embedding vectors. Among them, the size of the embedding vector can be preset according to the data volume and effect. The EmbbedingLayer can embed the output of the previous InputLayer, reduce the dimension of the previous input layer, and also reduce the number of learned parameters.
[0050] In addition, the CNN layer network can fully discover and learn the relationships between each stage, so as to better guide the action space action to obtain the maximum reward value reward. The final state output layer is used for encoder encoding to obtain the real-time state.
[0051] In a further embodiment, before processing the matrix by the CNN layer, it includes:
[0052] The matrices of n*m-dimensional embedding vectors corresponding to all stages are input into a preset stage layer to generate a two-dimensional matrix based on a preset splicing model. That is to say, in order to fully learn the connections between each stage in the state space, a stage layer (StageLayer) needs to be added after the EmbbedingLayer to convert the EmbbedingLayer into two-dimensional matrix data.
[0053] As another further embodiment, before outputting the processed matrix through the state output layer, the matrix processed by the CNN layer can be input into a preset recurrent neural network. In order to better capture the characteristics of the current parameter tuning based on the task parameter tuning of each round, a recurrent neural network is connected after the CNN layer in the present invention. Preferably, a GRU Layer is used to adjust the information flow based on the internal mechanism of the "gate" to solve the short-term memory problem.
[0054] Preferably, after being input into a preset recurrent neural network, the matrix output by the recurrent neural network can be further input into an attention layer (i.e., the attention layer) to identify the importance label of each state space and the influence degree label of the operator of each state space on the relevant index. That is to say, on the premise of the GRU Layer, an additional attention layer is added, so as to identify the importance in each state and the influence degree of the operator in each state on the final result index.
[0055] As some other embodiments, obtaining the corresponding action space value from the encoded information through a decoder Decoder may include: decoding the encoded information using a CNN to obtain the to-be-processed configuration parameters, and then randomly extracting from the to-be-processed configuration parameters to obtain the configuration parameters of Spark. Specific implementation:
[0056] The two-dimensional matrix in the Encoder process is decoded using a CNN to obtain the to-be-processed configuration parameter ap (as shown in Figure 2 ), that is, the predicted action space action for the current training is obtained, and then some valid parameters are selected from the action space action (randomly selected in this embodiment to cut off the correlation between parameters, prevent overfitting, and thus better learn) and stored in aval (as shown in Figure 2 ), and updated in each round. Finally, the aval result is added to the Spark task for running, and the weighted sum of the feedback indicators of the data mart is obtained as the reward value reward.
[0057] In some other embodiments, when inputting the action space value and the state space value into a Critic model to fit and obtain a Q value (i.e., the reward value reward), the output action value and the state are jointly input into the Critic model, and the Q value (i.e., the reward value reward) is obtained after being trained by a multi-layer network. Among them, the Q value is a function value used to guide the learning direction of the model in the reinforcement learning process, that is, the reward value reward.
[0058] It should be noted that when obtaining the configuration model by maximizing the final reward value based on a preset network loss function, the network loss function is:
[0059]
[0060] Among them, s represents the state of an agent at a certain moment. a represents the action executed at a certain moment. Q(s,a) represents the discounted future reward when the agent takes a certain action in a certain state and then takes the optimal action. λ is a hyperparameter to be adjusted and is the weight of the supervised learning model. θ is the network parameter of the Actor. μ is the action value of the corresponding actor in the state at a certain moment. N is the number of model training iterations.
[0061] The first half of the formula is the gradient obtained from the network calculation difference between the Q value updated each time and the next state next_state, and the second half is the parameter calculated in the Actor network, which is used to guide the Actor network parameters to maximize the final reward value.
[0062] Step S103: Obtain the target information of the current spark scheduling task, obtain the corresponding spark configuration parameters through the configured model, and then execute the current spark scheduling task based on the spark configuration parameters.
[0063] In summary, the present invention realizes a solution for configuring the best spark parameters by using the deep reinforcement learning DDPG algorithm and combining the Encoder-Decoder model to construct a directed acyclic graph in the parsed spark task.
[0064] Figure 2 It is a schematic diagram of the scheduling task configuration method according to the second embodiment of the present invention. The scheduling task configuration method may include:
[0065] First, the present invention uses the deep reinforcement learning DDPG algorithm. In data preparation, all information of the task collected from the historical execution log of the data mart is obtained through the spark scheduling task id. For example: the task uses the order as the main table and associates with the commodity table to obtain other dimension information such as department information, commodity name, purchasing and sales staff account information, etc. Preferably, spark configuration parameters such as the maximum number of cpus (threads) used by the driver, the size of the in-memory heap applied for by a single executor, the maximum number of concurrent tasks of a single executor, and the number of executors are used as the action space action. The relevant metrics feedback by the data mart after the spark scheduling task is executed are used as the reward reward. The operators included in each stage in the DAG generated each time the spark task is started and the task information included in each state (for example: the number of Tasks, Duration, GC Time, ShuffleRead data volume, ShuffleWrite data volume, etc. information) are used as the state space state.
[0066] Then, through the established Actor network, the Encoder process is performed first and then the Decoder process. The Encoder process specifically includes:
[0067] Through InputLayer, the state space parameters (including each stage information: sti, and the task information contained in each stage when the commodity trading task script is executed), including numerical type variables (such as: (4,2,6,45,4...)) and discrete scalars are processed by one-hot (for example (1,0,0,0,0,0,0,0,0), (0,1,0,0,0,0,0,0), (0,0,1,0,0,0,0,0))) and then spliced. The previous layer of InputLayer is embedded in the EmbbedingLayer to obtain an embedding vector n*m dimensional matrix (the size of the embedding vector depends on the amount of data and the effect). In order to fully learn the mutual connection between each stage in the state space in the same layer of the network, add a layer of StageLayer to splice the Embbeding layer into a two-dimensional matrix data. (The embedding matrix of each state is obtained in the Embbeding layer, and the StageLayer splices the embedding matrix into a large two-dimensional matrix (such as Figure 3 As shown)). Then, in the next layer of StageLayer, CNN layer is used to learn two-dimensional matrix data. The CNN layer network can fully discover and learn the relationship between each stage, so as to better guide the action to maximize the reward. In order to better capture the characteristics of the current parameter adjustment based on each round of task parameter adjustment, another layer of recurrent neural network is connected to the CNN layer. GRU Layer can be used (that is, based on the internal mechanism of "gate", to adjust the information flow and solve the short-term memory problem.) Adding another layer of attention layer on the premise of GRU Layer can identify the importance of each state and the influence of the operator in each state on the final result indicator. Finally, a state output layer is connected, and the state output layer is added for encoder encoding to obtain the real-time state. At this point, the encoder encoding process is completed.
[0068] Among them, the Decoder process includes: after the encoding is completed, CNN is used to decode to obtain ap, and then some parameters are randomly extracted from it (in order to cut off the correlation between parameters, prevent overfitting, and better learn), and aval is obtained. These effective actions are to obtain the final output value: the configuration parameters of spark.
[0069] In addition, the Critic model is a component of the DDPG reinforcement learning algorithm. It fits the parameters output by the Actor network, that is, the output action value and state are jointly input into the Critic model, and the Q value is obtained after training through multiple layers of the network. Then, based on the network loss function:
[0070]
[0071] After a certain number of iterations, the trained network is saved as a target network Netqt, and the network trained in a single iteration is denoted as Netqeval. Each time aval is used as input, two Q values, Qqt and Qeval, are calculated, which correspond to the first half above. The second half is the parameter calculated in the Actor, which is used to guide the Actor network parameters to maximize the final reward value.
[0072] Repeat the above process until the entire parameter tuning process ends, that is, the end of this round of training. Then, add other id task data for training to obtain the final output of the trained configuration model. Thus, the target information of the current spark scheduling task is obtained, and the corresponding spark configuration parameters are obtained through the trained configuration model to execute the current spark scheduling task.
[0073] Figure 5 It is a schematic diagram of the main module of the scheduling task configuration device according to an embodiment of the present invention, as Figure 5 shown. The scheduling task configuration device 500 includes an acquisition module 501 and a processing module 502. Among them, the acquisition module 501 collects relevant information of the task from the historical execution logs of the data mart according to the id of the spark scheduling task; inputs the information into the Actor network in the preset DDPG model to obtain the action space value corresponding to the information; wherein, the Actor network in the DDPG model adopts an Encoder-Decoder model; inputs the action space value and the information into the Critic network in the preset DDPG model for training to obtain the corresponding reward value; further calculates the preset network loss function according to the reward value to obtain the DDPG model corresponding to the maximum reward value as the configuration model; finally; the processing module 502 obtains the target information of the current spark scheduling task, obtains the corresponding spark configuration parameters through the configuration model, and then executes the current spark scheduling task based on the spark configuration parameters.
[0074] In some embodiments, the acquisition module 501 collects relevant information of the task from the historical execution logs of the data mart, including:
[0075] Spark configuration parameters, relevant metrics fed back by the data mart after the Spark scheduling task is completed, and the operators and task information included in each stage of the directed acyclic graph generated when the Spark scheduling task is started.
[0076] In some embodiments, the obtaining module 501 inputs the information into the Actor network in a preset DDPG model to obtain the action space value corresponding to the information, including:
[0077] Using the encoder in the Encoder-Decoder model, splicing the operators and task information included in each stage as the input layer for embedding to obtain an embedding vector corresponding to each stage, an n*m-dimensional matrix; and then processing the matrix through a convolutional neural network to output the processed matrix.
[0078] In some embodiments, before the obtaining module 501 processes the matrix through a convolutional neural network, it includes:
[0079] Respectively passing the embedding vectors corresponding to all stages, the n*m-dimensional matrices, through a preset splicing model to generate a two-dimensional matrix.
[0080] In some embodiments, the obtaining module 501 outputs the processed matrix, including:
[0081] Inputting the processed matrix into a preset recurrent neural network to output the matrix processed by the recurrent neural network.
[0082] In some embodiments, after the obtaining module 501 is input into a preset recurrent neural network, it includes:
[0083] Inputting the matrix output by the recurrent neural network into an attention layer to identify the importance label of each state space and the influence degree label of the operator of each state space on the relevant metrics.
[0084] In some embodiments, the obtaining module 501 inputs the information into the Actor network in a preset DDPG model to obtain the action space value corresponding to the information, including:
[0085] Using the decoder in the Encoder-Decoder model, decoding the encoded information using a convolutional neural network to obtain the to-be-processed configuration parameters, and then randomly extracting the final configuration parameters from the to-be-processed configuration parameters as the action space value corresponding to the information.
[0086] It should be noted that there is a corresponding relationship between the scheduling task configuration method and the scheduling task configuration device of the present invention in specific implementation contents, so the repeated contents will not be described again.
[0087] Figure 6 FIG. 600 shows an exemplary system architecture to which the scheduling task configuration method or the scheduling task configuration device according to an embodiment of the present invention can be applied.
[0088] As Figure 6 shown, the system architecture 600 may include terminal devices 601, 602, 603, a network 604, and a server 605. The network 604 is used to provide a medium for communication links between the terminal devices 601, 602, 603 and the server 605. The network 604 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0089] Users can use the terminal devices 601, 602, 603 to interact with the server 605 through the network 604 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 601, 602, 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0090] The terminal devices 601, 602, 603 may be various electronic devices having a scheduling task configuration screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0091] The server 605 may be a server that provides various services, such as a background management server that supports shopping websites browsed by users using the terminal devices 601, 602, 603 (only as an example). The background management server may analyze and process data such as product information query requests received, and feedback the processing results (such as target push information, product information - only as examples) to the terminal devices.
[0092] It should be noted that the scheduling task configuration method provided by the embodiments of the present invention is generally executed by the server 605. Correspondingly, the computing device is generally disposed in the server 605.
[0093] It should be understood that Figure 6 the numbers of the terminal devices, the network, and the server in
[0094] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. Figure 7 shown, which shows a schematic structural diagram of a computer system 700 of a terminal device suitable for implementing the embodiments of the present invention. Figure 7 The shown terminal device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0095] AsFigure 7 As shown in Figure 7 , the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the computer system 700 are also stored. The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0096] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.
[0097] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709 and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above-described functions defined in the system of the present invention are executed.
[0098] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0100] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes an acquisition module and a processing module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases.
[0101] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes collecting relevant information of the task from the data mart historical execution log according to the id of the spark scheduling task; inputting the information into the Actor network in the preset DDPG model to obtain the action space value corresponding to the information; wherein, the Actor network in the DDPG model adopts an Encoder-Decoder model; inputting the action space value and the information into the Critic network in the preset DDPG model for training to obtain the corresponding reward value; then calculating the preset network loss function according to the reward value to obtain the DDPG model corresponding to the maximum reward value as the configuration model; obtaining finally; acquiring the target information of the current spark scheduling task, obtaining the corresponding spark configuration parameters through the configuration model, and then executing the current spark scheduling task based on the spark configuration parameters.
[0102] According to the technical solution of the embodiments of the present invention, the problem that the processing performance of the scheduling task is affected due to the low precision and poor efficiency of setting the Spark configuration parameters in the prior art can be solved.
[0103] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A scheduling task configuration method, characterized in that, Including: Collect relevant information of a task from the data mart historical execution log according to the id of the spark scheduling task; Input the information into the Actor network in a preset DDPG model to obtain the action space value corresponding to the information; wherein, the Actor network in the DDPG model adopts an Encoder-Decoder model; this process includes: using the encoder in the Encoder-Decoder model to splice the operators and task information included in each stage in the information as the input layer for embedding to obtain an embedding vector matrix of n*m dimensions corresponding to each stage; then processing the matrix through a convolutional neural network to output the processed matrix; a stage is a stage in the directed acyclic graph generated when starting a spark scheduling task; Input the action space value and the information into the Critic network in a preset DDPG model for training to obtain the corresponding reward value; then calculate a preset network loss function according to the reward value to obtain the DDPG model corresponding to the maximum reward value as the configuration model; wherein, the reward value is a function value used to guide the learning direction of the model in the reinforcement learning process; the preset network loss function includes a first half and a second half. The first half is the gradient obtained based on the network calculation difference between the reward value updated each time and the next state, and the second half is the parameter calculated in the Actor network, which is used to guide the Actor network parameters to make the final reward value reach the maximum; Obtain the target information of the current spark scheduling task finally, obtain the corresponding spark configuration parameters through the configuration model, and then execute the current spark scheduling task based on the spark configuration parameters.
2. The method according to claim 1, wherein Collect relevant information of a task from the data mart historical execution log, including: Spark configuration parameters, relevant metrics feedback by the data mart after the spark scheduling task is completed, and the operators and task information included in each stage in the directed acyclic graph generated when starting the spark scheduling task.
3. The method according to claim 1, wherein Before processing the matrix through a convolutional neural network, including: Respectively generate a two-dimensional matrix from the embedding vector matrices of n*m dimensions corresponding to all stages through a preset splicing model.
4. The method according to claim 1, wherein Output the processed matrix, including: Input the processed matrix into a preset recurrent neural network to output the matrix processed by the recurrent neural network.
5. The method according to claim 4, characterized in that, After inputting into the preset recurrent neural network, including: Input the matrix output by the recurrent neural network into an attention layer to identify the importance label of each state space and the influence degree label of the operator of each state space on the relevant metrics; wherein, the relevant metrics are the relevant metrics feedback by the data mart after the spark scheduling task is completed.
6. The method according to any one of claims 1-5, characterized in that Input the information into the Actor network in a preset DDPG model to obtain the action space value corresponding to the information, including: Using the decoder in the Encoder-Decoder model, the encoded information is decoded by a convolutional neural network to obtain the configuration parameters to be processed, and the final configuration parameters are randomly selected from the configuration parameters to be processed as the action space value corresponding to the information.
7. A scheduling task configuration device, characterized in that, It includes: An acquisition module, configured to collect relevant information of a task from the historical execution log of the data mart according to the id of the spark scheduling task; Input the information into the Actor network in a preset DDPG model to obtain the action space value corresponding to the information; wherein, the Actor network in the DDPG model adopts an Encoder-Decoder model; this process includes: using the encoder in the Encoder-Decoder model, splicing the operators and task information included in each stage in the information as the input layer for embedding to obtain an embedding vector matrix of n*m dimensions corresponding to each stage; and then processing the matrix through a convolutional neural network to output the processed matrix; a stage is a stage in the directed acyclic graph generated when starting a spark scheduling task; Input the action space value and the information into the Critic network in the preset DDPG model for training to obtain the corresponding reward value; and then calculate the preset network loss function according to the reward value to obtain the DDPG model corresponding to the maximum reward value as the configuration model; finally; wherein, the reward value is a function value used to guide the learning direction of the model in the reinforcement learning process; the preset network loss function includes a first half and a second half. The first half is the gradient obtained based on the network calculation difference between the reward value updated each time and the next state, and the second half is the parameter calculated in the Actor network, which is used to guide the parameters of the Actor network to make the final reward value reach the maximum; A processing module, configured to obtain the target information of the current spark scheduling task, obtain the corresponding spark configuration parameters through the configuration model, and then execute the current spark scheduling task based on the spark configuration parameters.
8. An electronic device, characterized in that, It includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
A task scheduling method and device
CN109725988A
Spark operation time prediction method and device based on graph convolution network
CN111126668A