Multi-cloud job scheduling method for cumulative data processing application
By abstracting the job scheduling problem in a multi-cloud environment into a Markov decision process, constructing a CDP-SS scheduling method, and using the PPO algorithm to optimize scheduling decisions, this method solves the scheduling problem of cumulative data processing applications in a multi-cloud environment and achieves a scheduling effect with low SLA violation rate and low resource cost.
Patent Information
- Application Number
- CN202510050918.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-10-10
AI Technical Summary
Existing cloud job scheduling methods are difficult to effectively schedule cumulative data processing applications in multi-cloud environments, cannot optimize job scheduling globally, lack adaptability and dynamic adjustment capabilities, and cannot meet the dual goals of low application SLA violation rate and low resource cost.
The job scheduling problem in a multi-cloud environment is abstracted into a Markov decision process. Resource-level, job-level, and application-level information are integrated to construct a CDP-SS scheduling method. The PPO algorithm is used to design the Actor and Critic network. Reinforcement learning is used to optimize scheduling decisions. Taking into account job resource allocation quality, data distribution, and resource cost factors, a complete scheduling solution is constructed.
It achieves the dual goals of low application SLA violation rate and low resource cost in a multi-cloud environment, and improves the scheduling efficiency and cost control of cumulative data processing applications.
Smart Images

Figure CN120762827A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-cloud job scheduling, and in particular relates to a job scheduling method for data accumulation-type applications in a multi-cloud environment. Background Art
[0002] In recent years, the widespread adoption of cloud computing technology has driven the emergence of numerous cloud service providers. These providers offer computing resources and services over the Internet, enabling on-demand access and providing services to users with flexible payment models. With the increasing popularity of cloud computing, a single cloud environment is unable to meet the increasingly diverse application needs, prompting users to seek more flexible and diverse resources to accommodate the diverse characteristics, scale, and demand scenarios of various applications. At the same time, single cloud environments are also limited in terms of cost-effectiveness. Users are eager to choose the most cost-effective cloud service combination based on their actual needs to reduce costs, which has driven the shift from a single cloud to a multi-cloud environment. The emergence of multi-cloud environments aims to provide users with more flexible and lower-cost cloud services by integrating the resources of multiple cloud service providers.
[0003] Cumulative data processing (CDP) applications process data that is generated cumulatively over a long period of time. Data operations consist of two phases: preprocessing and aggregate analysis. The accumulated data first undergoes independent preprocessing and deformation, and then all preprocessed data is aggregated for complex data analysis operations, which must be completed within a certain deadline. Cumulative data processing is a common type of application in cloud environments. This type of application has a long lifecycle, which makes it possible for applications to fully utilize low-cost resources in different clouds in stages. Therefore, cumulative data processing applications are naturally suitable for scheduling and running in multi-cloud environments.
[0004] In terms of job scheduling methods, traditional cloud job scheduling is mostly based on rule design, and these algorithms need to be adjusted according to the scheduling scene. Existing cloud job scheduling methods mainly include two categories: heuristic / meta-heuristic algorithm-based and machine learning-based. The heuristic / meta-heuristic algorithm-based scheduling algorithm has been widely studied. In each round of job scheduling, this kind of method tries to search for a near-optimal scheduling scheme according to the current job to be scheduled. However, due to the characteristics of long-time and batch data generation of cumulative data processing applications, the jobs contained in the applications are distributed in multiple scheduling rounds, and these methods are difficult to coordinate the scheduling of jobs in multiple rounds, difficult to optimize in the global range, and also lack of adaptability, unable to dynamically adjust the scheduling method according to the change of the environment. Most of the cloud job scheduling methods based on machine learning use machine learning or deep learning models to predict the resource demand of the job or the available idle resources of the data center, and make job scheduling based on the prediction results. Although this kind of job scheduling decision considers the future resource demand of the application, it cannot learn from historical scheduling decisions to optimize future scheduling decisions. Reinforcement learning-based job scheduling is the latest development of cloud job scheduling. Reinforcement learning optimizes subsequent scheduling decisions by learning the quality of historical scheduling decisions through feedback. However, existing work has not designed a reinforcement learning model for the characteristics of long-time and cumulative processing of cumulative data processing applications. The existing reinforcement learning scheduling method for scientific workflows cannot adapt to the characteristics of job generation with uncertainty and dynamic changes in application execution progress of cumulative data processing applications. SUMMARY
[0005] To solve the above problems, the present application proposes a job scheduling method CDP-SS for CDP applications in a multi-cloud environment. The CDP-SS scheduling method abstracts the CDP job scheduling problem in a multi-cloud environment as a Markov decision process. In the definition of the state space, CDP-SS integrates resource-level, job-level and application-level information into the definition of the state space, so as to completely express the resource supply and demand state of the multi-cloud data center and the dynamic change of the application execution progress caused by the uncertainty of job generation. In the definition of the reward space, CDP-SS considers the quality of resource allocation of the job, the distribution of intermediate result data of the application among the multi-cloud data centers, the risk of application timeout and the resource cost factors, and accurately describes the quality of job scheduling decision in a multi-cloud environment. This method constructs a relatively complete scheduling scheme for CDP applications in a multi-cloud environment to meet the dual goals of low application SLA (Service Level Agreement) violation rate and low resource cost.
[0006] The proposed method for scheduling CDP applications consists of five steps: initialization, Markov decision process modeling, CDP-SS network structure modeling, CDP-SS model training, and cumulative data processing application job scheduling. This method is based on the PPO scheduling method and has six basic parameters: the discount factor γ for calculating future rewards, the actor network learning rate α1, the critic network learning rate α2, the generalized advantage estimate λ, the clip parameter for policy gradient updates, and the batch size for training. γ is set to 0.95, α1 is set to 0.0001-0.001, α2 is set to 0.0001-0.001, λ is set to 0.95, the clip parameter is set to 0.2, and the batch size is set to 256-1024.
[0007] (1) Initialization. Data is initialized using the information from the cumulative data processing application in the log. The data preprocessing phase includes multiple data preprocessing jobs; the data aggregation analysis phase includes only one aggregation analysis job, which is started and submitted for scheduling after all preprocessing jobs are completed.
[0008] 1.1) The cumulative data processing application information submitted by each user is defined as: DCP i =(PreOps i ,AggOps i ,DataSource i ,DStartTime i ,DEndTime i ,Deadline i ). Among them, CDP i Indicates the i-th CDP application; PreOps i and AggOps i Respectively represent the data operations performed in the data preprocessing stage and the aggregation analysis stage; DataSource i Indicates the data source processed by the application; DStartTime i and DEndTime i Respectively represent the start and end time of the data processed by the application; Deadline i It indicates the time when the user expects the application to end.
[0009] 1.2) Use DCS = {DC1, DC2, .....DC n} indicates that the multi-cloud environment consists of multiple cloud data centers. Each data center DC i Contains a group of physical servers, with machine i,jRepresents the jth server in the i-th data center. Each cloud data center DC i Provide tn i A virtual machine image, defined as Therefore, any virtual machine image Defined as in, Indicates the computing power of the virtual machine CPU. Indicates the memory size of the virtual machine. Indicates the rental price per unit time of the virtual machine.
[0010] 1.3) A CDP application consists of multiple preprocessing jobs and one aggregation analysis job. The i-th CDP application job set is represented as in AggJob represents the jth preprocessing job of the i-th CDP application. i Indicates the aggregation job of the i-th CDP application.
[0011] 1.4) Let CDPjob represent CDP_JOBS i For any job in the set, the processing data size of CDPJob is datasize, the computation amount is length, the computing power of the allocated virtual machine instances is vflops, the number is vnum, and the job turnaround time is MK(CDPJob). Then the job turnaround time of CDPJob is expressed as:
[0012] MK(CDPJob)=t_Wait(CDPJob)+t_DTrans(CDPJob)+t_Exe(CDPJob)#(1)
[0013] t_Wait(CDPJob) represents the time a job waits for scheduling, and is calculated using the following formula:
[0014] t_Wait(CDPJob)=ATime(CDPJob)-STime(CDPJob)#(2)
[0015] ATime(CDPJob) represents the arrival time of the job, and STime(CDPJob) represents the scheduling time of the job. t_DTrands(CDPJob) represents the transmission time of the job processing data, which is calculated using Formula (3) based on the network bandwidth between the application agent that submits the job and the cloud data center.
[0016]
[0017] Among them, bd(DC p ,DC q ) represents the data center DC pand DC q The network bandwidth between the two nodes is represented by the value of sizeof(data), and the sizeof(data) represents the size of the transmitted data. t_Exe(COPJob) represents the execution time of the job, which is calculated as follows:
[0018]
[0019] 1.5) For any job, its virtual machine resource cost j_vcost is expressed as follows:
[0020] j_vcost=vnum×consumetime×vprice#(5)
[0021] Among them, vnum and vprice represent the number of virtual machine instances allocated to the job and the price per unit time respectively, and consumetime represents the resource consumption time of the job, which is the sum of its data transmission time and execution time.
[0022] Apply CDP i Data transmission cost data_cost i It is the cumulative cost of transmitting intermediate results between the data center where all preprocessing jobs are located and the data center where the aggregation analysis job is located, and is expressed as:
[0023]
[0024] Among them, ∑(·) represents the summation operation, and the subscript of ∑(·) is the range of the summation operation, Represents preprocessing jobs The size of the intermediate result data generated, Represents preprocessing jobs Aggregation Analysis Job AggJob i The price of transferring unit-scale data between data centers. If both the data preprocessing and aggregation analysis phases are performed in the same data center, the data transfer cost is zero.
[0025] In summary, the application of COD i The total cost of preprocessing operations p_cost i , the cost of the aggregate analysis operation a_cost i and the total application cost i Respectively expressed as:
[0026]
[0027] a_cost i =j_vcost(AggJob i )+date_cost i #(8)
[0028] cost i =p_cost i +a_cost i #(9)
[0029] (2) Markov decision process modeling
[0030] 2.1) State space. The state space includes the cloud resource state, the state of pending jobs, and the state of the historical job queue.
[0031] 2.1.1) Cloud Resource State. The defined cloud resource state includes the virtual machine state State_VM and the physical server state State_PM, expressed as: State_MC = (State_VM, State_PM). State_VM records the resource allocation and usage status of virtual machine instances deployed in multiple cloud data centers. For each virtual machine instance, its state is described by a five-tuple vector as follows:
[0032] vm_state=(vCPUcapacity,vMemcapacity,rentperiod,vCPUutil,vMemutil)
[0033] vCPUcapacity and vMemcapacity represent the rated CPU computing capacity and memory resource size of the VM, respectively, and are obtained based on the VM instance type (ctype). rentperiod represents the lease duration of the VM instance. vCPUutil represents the ratio of the CPU resource requirement (vCPUReq) of the job using the VM instance to the rated CPU computing capacity (vCPUcapacity) of the VM instance. Similarly, vMemutil is calculated. vCPUutil and vMemutil represent the actual utilization efficiency of the VM's rated resources and reflect the rationality of job resource allocation.
[0034] Based on the state of virtual machine instances, the state of the virtual machine in a data center, State_VM, is expressed as a matrix formed by sequentially concatenating the states of all virtual machine instances in multiple data centers. In the State_VM matrix, the states of virtual machine instances are arranged in the order of data center-physical server. For each physical server, the matrix rows are allocated according to the maximum number of deployable virtual machine instances. When the number of actually deployed virtual machine instances is less than the maximum value, the row is filled with zero values. Each row of the matrix represents the state of three dimensions of a physical server in the data center, including the virtual machine share allocation rate VMutil, the CPU resource allocation rate CPUuil, and the memory resource allocation rate Memutil. VMutil is expressed as the ratio of the number of deployed virtual machine instances on the physical server to the total rated number of virtual machine instances.
[0035] 2.1.2) Status of jobs to be scheduled. Suppose there are N jobs in the queue of jobs to be scheduled, then the status of the jobs to be scheduled Statee_job is represented by an N×9 matrix. State_job=(jobtype,length,,datasize,vnum,vCPUReq,vMemReq,density,app_urgency,app_perlength). Among them, one row of the matrix represents the status of a job. The status of each job contains seven dimensions. jobtype represents the type of job, which is divided into preprocessing jobs and aggregation analysis jobs; length represents the computing scale of the job; datasize represents the data scale processed by the job; vnum represents the number of virtual machines required by the job; vCPUReq and vMemReq represent the job's requirements for CPU and memory respectively; density represents the amount of computing that the job must complete per unit time by the application deadline, calculated as density=length / remaintime, and residual represents the time from the current moment to the expected deadline of the application. app_urgency and app_perjob represent the status of the application to which the job belongs. app_progress represents the relative remaining time of the application, and app_perjob represents the expected job execution progress of the application, which are expressed as follows:
[0036]
[0037] Deadline, Starttime, and Curtime represent the expected deadline, start time, and current time of the application. num_remainingjob represents the number of jobs belonging to the application in the queue of jobs to be scheduled.
[0038] 2.1.3) Historical job queue state. Denoted as State_Req, State_Req includes five features of the generated jobs in each unit time period, which are total data size, total computation size, virtual machine demand quantity, virtual machine CPU demand sum and virtual machine memory demand sum, denoted as
[0039]
[0040] 2.2) Action space. The scheduling method proposed in this method traverses the job queue to be scheduled in the order of arrival, and determines the scheduling decision of each job to be scheduled in turn by using the reinforcement learning agent. Therefore, the action space of reinforcement learning is represented as:
[0041]
[0042] where none represents that the job is not scheduled, represents that the job is scheduled to the data center DC i , and the virtual machine type allocated to it is
[0043] 2.3) Reward. Define the reward components: virtual machine resource allocation quality R vm , data aggregation degree R data , application execution risk R risk , and new job cost R cost . The reward for each scheduling decision is represented as the cumulative sum of the above four reward items.
[0044] 2.3.1) Virtual machine resource allocation quality. The virtual machine resource allocation quality of the job is defined as follows:
[0045]
[0046] where vCPUReq and vMemReq are obtained from the resource request of the job, and vCPUAlloc and vMemAlloc are the resource sizes of the actually allocated virtual machines. α and β are weight coefficients, and the sum is 1. vCPUMax and vMemMax represent the highest configuration of CPU and memory in all types of virtual machines, and vCPUMin and vMemMin represent the minimum configuration.
[0047] 2.3.2) Data aggregation degree. When the scheduled job is a preprocessing job , the application data aggregation degree reward is defined as follows:
[0048]
[0049] Wherein, ∑(·) represents a summation operation, and the superscript m and subscript j=1 of ∑(·) represent that the value of j is an integer from 1 to m. Indicates the application of CDP i Completed preprocessing jobs, m represents the total number of completed preprocessing jobs in the application, and dc(·) represents the data center where the job is located. nor_bd(·) represents the normalized bandwidth between the cloud data centers where two jobs are located, that is, the ratio of the actual network bandwidth to the maximum network bandwidth. For two jobs in the same data center, nor_bd(·) is set to 1.
[0050] 2.3.3) Application execution risk. The reward function is set to the average risk of all generated applications, expressed as:
[0051]
[0052] in, Is CDP applied in the job queue i The expected turnaround time of unscheduled jobs, N is the number of generated applications, and t represents the current time point.
[0053] 2.3.4) New operating costs. The bonus value is expressed as:
[0054]
[0055] Here, nor_vc(·) represents the normalized cost of the currently scheduled job, which is the ratio of the resource cost of the currently scheduled job to the average cost of all scheduled jobs.
[0056] (3) CDP-SS model network structure construction. The CDP-SS model is designed based on the PPO algorithm. The actor network and the critic network have similar network structures. The state input of each network consists of four parts: State_VM, State_PM, State_job, and State_Req.
[0057] 3.1) VM-autoencoder Module. The State_VM matrix is partitioned into sub-matrices, each corresponding to the VM state on a physical server. The VMautoencoder encoder is used to generate a five-dimensional latent feature vector for each sub-matrix. Ultimately, the VM latent feature vectors for all physical servers are sequentially concatenated to form the LatentVMencoding matrix for the overall VM latent feature of the multi-cloud data center.
[0058] The encoder design of VMautoencoder uses a combination of two convolutional layers and one fully connected layer. Specifically, both convolutional layers are configured with a convolution kernel of size 1×3, a stride of 1, a fixed number of output channels of 16, and ReLU as the activation function. After these two layers of convolution operations, the feature matrix is compressed into a one-dimensional vector by the flatten operation, and then input into the fully connected layer to extract a feature vector containing five elements. The decoder of VMautoencoder takes the five-element feature vector output by the encoder as input. The vector is first preliminarily processed by a fully connected layer containing 160 neurons, and then enters two deconvolution layers to reconstruct the feature matrix. The parameters of these two deconvolution layers are consistent with the corresponding convolution layer parameters in the encoder to ensure accurate transmission and reconstruction of information, and the activation function is also ReLU.
[0059] 3.2) Concat Module. The Concat module is responsible for concatenating LatentVMencoding, State_PM, and State_job to form a complete state matrix for the multi-cloud data center, which serves as the input to DC_Encoder. First, the data center virtual machine state latent feature matrix, LatentVMencoding, is horizontally concatenated with the physical machine state feature matrix, State_PM. Second, the concatenated feature matrix is vertically concatenated with the state feature matrix, State_job, of the job to be scheduled.
[0060] 3.3) DC_Encoder module. DC_Encoder is designed as a neural network consisting of two convolutional layers and one pooling layer. Specifically, the first convolution layer uses a convolution kernel of size 2×2, a stride of 1, and 16 output channels. The second convolution layer uses a convolution kernel of size 3×3, a stride of 1, and 16 output channels. The ReLU activation function is applied after each convolution layer. This method uses three adaptive-sized pooling kernels to perform multiple pooling on the feature matrix, with bin sizes of 1x1, 2x2, and 3x3, respectively. The pooling results are concatenated to obtain a fixed-size output vector, namely LatentDCencoding.
[0061] 3.4) Req_Predict Model. The job status prediction model, Req_Predict, takes State_Req as input and constructs a hidden layer using a two-layer GRU structure and two fully connected layers to extract time series features from the data. Specifically, the GRU's time window size is set to 10. Each layer of the two-layer GRU contains 32 neurons, and the two layers of the neural network are processed progressively. After State_Req passes through the two-layer GRU, its output dimension is 10×64, which is expanded into a one-dimensional vector with 640 features.
[0062] The model then adopts two fully connected layers for feature extraction and transformation. The number of neurons in these two fully connected layers is 64 and 3 respectively, and the output of the last fully connected layer is the predicted job state value of the next time unit, denoted as PredictedReq. These predicted values specifically include the total number of virtual machine requirements, the total amount of computation, and the total amount of data, denoted as
[0063] 3.5) Act-Dec module. According to the obtained latent vector of LatentDCencoding and the predicted PredictedReqvector, these two vectors are spliced, and the spliced vector is then mapped to a k-dimensional vector through a fully connected network. Where k is the number of actions in the action space. Then, through the softmax function activation, the selection probability of each possible action is finally output.
[0064] (4) CDP-SS model training. In the network structure proposed in this method, the VMautopencoder and Req_Predictor modules adopt independent offline training, and the trained predictor and encoder modules are integrated into the Actor and Critic networks. For DC_Encoder and Act_Dec, an end-to-end training method is adopted.
[0065] 4.1) Initialize the model. Initialize the Actor and Critic in the model, and set the discount factor γ, the learning rate of the Actor network α1, the learning rate of the Critic network α2, the advantage estimate λ, the Clip parameter, and the batch size hyperparameter.
[0066] 4.2) Collect data. Load the CDP application into the multi-cloud environment for scheduling. At each scheduling time, according to the current state s and policy model π θ output the action a and its logarithmic probability logp, as well as the state value v based on the value model V φ After executing the action, the new state s', reward r, and termination flag done are observed. Each time the CDP application job is scheduled, a sample experience data Experience is recorded, Experience is represented by the tuple (s, a, logp, v, s', r, done), and these sample experience data are stored in the Rollout pool. At the same time, the current state is updated to s'.
[0067] 4.3) Calculate the advantage estimate. Randomly select 256 to 1024 experience data from the Rollout pool as training samples. Use the output of the value network as the state value V(s t of each state s t). To evaluate how good an action is relative to the average, the advantage function The formula is: where G t is the discounted cumulative reward from time step t, γ is the discount factor, r t+1+k is the reward obtained at time step t+1+k.
[0068] 4.4) Optimization objective function: calculate the policy ratio where π θ (a t |s t ) is the current policy, is the old policy. The PPO algorithm uses a specially designed objective function, which is in the form of:
[0069]
[0070] where E denotes the expected value, is the estimate of the advantage function, and ε is a small positive number (0<ε<0.4), and the clip function limits the range of changes in the policy ratio r t (θ). When r t (θ) exceeds the range of 1-ε to 1+ε, clip will be cut to this interval, and the excess value will be set to the maximum boundary value to prevent the policy from updating too large and causing instability.
[0071] 4.5) Update the policy. Use the gradient ascent method to update the policy parameters, i.e. where α is the learning rate, and θ is the parameter value of the policy network, including all weights and bias values in the input layer, hidden layer and output layer of the neural network structure in sections 3.3) and 3.5). denotes the gradient of L(θ) with respect to the parameter θ, which is a vector, and each component of the vector is the partial derivative of L(θ) with respect to the corresponding component of θ. Calculate the loss function L(φ) of the value network E[(V(s t )-r) 2 ], update the parameters φ of the value network using the gradient descent algorithm to minimize the loss function L(φ). The value network has the same network structure as the policy network, and the parameter φ includes all weights and bias values in the neural network structure.
[0072] 4.6) Repeat the steps. Repeat the above steps 4.2)-4.4) using the new policy parameters until the number of iterations reaches 1000. Finally, empty the playback pool Rollout and save the model CDP-SS parameters after all the update cycles are completed.
[0073] (5) Cumulative data processing application job scheduling
[0074] 5.1) Let M cumulative data processing applications run in a multi-cloud environment consisting of K data centers. The scheduling problem for cumulative data processing applications is defined as follows:
[0075]
[0076] Among them, SLA (CDP i ) represents the number of applications that do not violate the SLA, Max(·) represents maximization, Min(·) represents minimization, and cost(·) represents the total cost of the application. Application scheduling has two objectives: maximizing the number of applications that meet deadlines and minimizing the resource cost of application execution.
[0077] 5.2) When the cumulative data processing application starts scheduling, the state information of the scheduling environment is initialized according to the results of the Markov decision process modeling, including the cloud resource status, the status of the jobs to be scheduled, and the status of the historical job queue.
[0078] 5.3) Load the trained CDP-SS model for scheduling decisions. During each CDP application's preprocessing job or aggregation job scheduling cycle, the model receives the cloud environment's State_VM, State_PM, State_job, and State_Req state information. This state information is input into the CDP-SS model's Actor Network, which outputs a probability distribution for scheduling actions and selects the action with the highest probability as the scheduling policy. The scheduling policy corresponds to an element in the action set defined in Section 2.2), indicating that the job should not be scheduled or should be scheduled to a specific VM type in a specific cloud data center.
[0079] 5.4) Record a sample experience data item based on the job schedule and store it in the replay pool Rollout. This sample experience data item is used to train the CDP-SS model from Sections 4.3) to 4.6) to improve its scheduling performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 This is an execution process diagram of an accumulation-type data processing application.
[0081] Figure 2 This is the architecture diagram of the CDP-SS model.
[0082] Figure 3 Flowchart of the present invention.
[0083] Figure 4 Flowchart for modeling a Markov decision process.
[0084] Figure 5Flowchart for CDP-SS model training. Figure 6 Flowchart for CDP-SS model training. DETAILED DESCRIPTION
[0085] The present invention will be described below with reference to the accompanying drawings and specific embodiments.
[0086] The cumulative data processing application proposed in this invention is scheduled in a multi-cloud environment. Figure 1 This is the execution process of a CDP application. CDP application execution is divided into multiple preprocessing jobs, which are executed using the existing cloud's job scheduling mechanism. The data preprocessing phase includes multiple data preprocessing jobs; the data aggregation and analysis phase includes only one aggregation and analysis job. This job is initiated and scheduled after all preprocessing jobs are completed. Job submission is automatically completed by the application agent based on the user's application description. In the preprocessing job description, the JobTrigger describes the trigger for periodic or irregular data preprocessing job submission, such as periodic submission or when data accumulation reaches a specified amount. When the trigger conditions are met, the application agent automatically generates a preprocessing job and submits it to the cloud scheduler. This job preprocesses the newly generated data. All intermediate results generated by the data preprocessing jobs are stored in local storage in the data center where they are running. After the data generation deadline, the application agent generates a data aggregation and analysis job and submits it to the multi-cloud environment for scheduling and execution. This job reads the intermediate results generated by all preprocessing jobs and processes them to produce the final result.
[0087] Figure 2This is the architecture diagram of the CDP-SS model of the present invention. The state input of each network consists of four parts: State_VM, State_PM, State_job, and State_Req. The VMautoencoder module extracts the latent features of the cloud virtual machine resource state (Latent VM encoding) through the autoencoder to achieve state dimensionality reduction. The Concat module splices the latent state features of the cloud virtual machine resources with State_PM and State_job to form the overall state feature expression of the multi-cloud data center. This feature expression generates the latent feature expression of the multi-cloud data center state (Latent DC Encoding) through the DC_Encoder module. The Req_Predict module uses State_Req as feature input to predict the total resource demand and data processing scale (Predicted Req) of the scheduled job in the next period. The Actor and Critic networks have the same network structure in the above four modules and are both used to extract the state features of the multi-cloud environment. In the final output module Act_Dec, the Actor network takes Latent DC Encoding and Predicted Req as input and outputs the action probability vector action, while the Critic network outputs a single-value state value.
[0088] The following combination Figure 3 SUMMARY OF THE INVENTION The overall process describes the specific implementation method of the present invention. In this implementation method, the basic parameters are set as follows: γ is set to 0.95, α1 is set to 0.0001, α2 is set to 0.0002, GAE Lambda is set to 0.95, Clip Parameter is set to 0.2, and Batch Size is set to 128.
[0089] The specific implementation method is divided into the following steps:
[0090] (1) Initialization.
[0091] 1.1) Define CDP application information:
[0092]
[0093] 1.2) DCS = {DC1, DC2, DC3} represents three cloud data centers. DC1 contains five physical servers, machine 1,1 Indicates the first server in DC1, providing three types of virtual machine images in DC2 and DC3 are similar, but have different configurations and prices.
[0094] 1.3) CDP1 is expressed as CDP2 is expressed as
[0095] 1.4) Order The processing data scale is 100GB, the computing amount is 50, the computing capacity of the allocated virtual machine instance is 50, and the number is 2. The operation turnaround time can be expressed as:
[0096]
[0097] for Preprocessing jobs, similar to calculations.
[0098] Assume that the data processing scale of AggJob1 is 50 GB, the computation workload is 25, the computing capacity of the allocated virtual machine instance is 50, and the number is 1. The turnaround time MK(AggJob1) of the aggregation analysis job is calculated as follows:
[0099]
[0100] 1.5) For Its virtual machine resource cost is expressed as:
[0101] j_vcost=2×0.51×1=1.02$#(20)
[0102] for The virtual machine resource costs for preprocessing and aggregation jobs are calculated similarly. The data transfer cost for AggJob1 is:
[0103] data_cost1=50×0.001=0.05$#(21)
[0104] In summary, the total preprocessing cost of CDP1, p_cost1, is the sum of j_vcost (3.2$), p_cost2 (4.5$), a_costi (0.85$), and cost1 (5.35$). The cost of CDP2 is calculated similarly.
[0105] (2) Markov decision process modeling
[0106] 2.1) State space.
[0107] 2.1.1) The cloud resource state is represented as: State_MC = (State_VM, State_PM). Each row in State_VM represents the state of a virtual machine. One row of states is as follows: vm_state = (24, 64, 12, 0.8, 0.6). The states of other virtual machines are defined in a similar way. In the State_VM matrix, each row is the state of a physical server, machine 1,1 The status is (0.6, 0.75, 0.9).
[0108] 2.1.2) Status of jobs to be scheduled. Suppose there are 5 jobs in the job queue to be scheduled, then the status of the jobs to be scheduled State_job is represented by a 5×9 matrix.
[0109]
[0110] 2.1.3) Historical job queue status, let the time window size be set to 4, and State_Req is expressed as:
[0111]
[0112] 2.2) Action Space. The action space of reinforcement learning is expressed as:
[0113]
[0114] 2.3) Reward. Assume that the agent will preprocess the job The action selection is
[0115] 2.3.1) Quality of virtual machine resource allocation. The resource request is: vCPUReq = 10, vMemReq = 24, while the actual allocated virtual machine resources are vCPUAlloc = 24, vMemAlloc = 64. The maximum configuration of the virtual machine is vCPUMax = 96, vMemMax = 256, and the minimum configuration is vCPUMin = 24, vMemMin = 64. The weight coefficients α = 0.5 and β = 0.5 are used.
[0116] The quality of virtual machine resource allocation for a job is defined as follows:
[0117]
[0118] 2.3.2) Data aggregation. Assumption All are done in data center DC1, and Completed in DC2. Calculation Data aggregation, where nor_bd(DC1,DC2)=0.5:
[0119] R data =0.5×(1+0.5)=0.75#(23)
[0120] 2.3.3) Application execution risk. Assume that the current time is 2023, 01, 01, 10:00:00, and the deadline of CDP1 is 2023, 01, 01, 22:30:00, where t′ exp (CDP1) = 4. Calculate application execution risk:
[0121]
[0122] 2.3.4) New job cost. Assume the average cost of all scheduled jobs is avg_cost = $1.5. Calculate the new job cost reward:
[0123]
[0124] Assume that the agent chooses to Assign to action The total reward is: R = 0.63 + 0.75 + 2.125 - 0.68 = 2.825.
[0125] (3) Construction of CDP-SS model network structure.
[0126] 3.1) VM-autoencoder module. The State_VM matrix is divided according to the physical server, and each sub-matrix represents the state of the virtual machine on a physical server. In the encoder, the sub-matrix is flattened into a one-dimensional vector after two layers of convolution operations, and then input into the fully connected layer to extract the five-dimensional feature vector. In the decoder, the five-dimensional feature vector passes through the fully connected layer and enters the deconvolution layer for feature matrix reconstruction. The virtual machine hidden state feature vectors of all physical machines are spliced in sequence to form the overall virtual machine hidden feature matrix LatentVMencoding of the multi-cloud data center. The machine in the matrix 1,1 The hidden layer feature vector is (0.74, 0.81, 1.34, 1.45, 0.94).
[0127] 3.2) Concat module: Concatenates LatentVM encoding, State_PM, and State_job into a 12×9 feature matrix.
[0128] 3.3) DC_Encoder module. After two layers of convolution and pooling operations, the complete state matrix is converted into a 224-dimensional fixed-size output vector LatentDCencoding = (3.4, 3.2, 4.5, 6.7, ...).
[0129] 3.4)Req_Predict model. After the double-layer GRU, the State_Req is unfolded into a one-dimensional vector with 640 features, and then two fully connected layers are used for feature extraction and conversion, outputting the predicted next time unit job state PredictedReq = (4, 300, 400).
[0130] 3.5) Act-Dec module. The concatenated vector is mapped to a k-dimensional vector through a fully connected network, and then activated through a softmax function to output the selection probability of each possible action (0.2, 0.3, 0.6, 0.4, …), and the action with the maximum probability is selected and executed.
[0131] (4) CDP-SS model training.
[0132] 4.1) Initialize the model. Set the parameters of the Actor and Critic in the model, the discount factor γ is 0.95, the learning rate of the Actor network α1 is 0.0001, the learning rate of the Critic network α2 is 0.0002, the advantage estimate λ is 0.95, the clip parameter is 0.2, and the batch size is 64 hyperparameters.
[0133] 4.2) Collect data. Initialize the parameters of the policy model and the value model, and initialize the Rollout replay pool to empty. At each scheduling, according to the current state s and the policy model π θ output the action a and its logarithmic probability logp. After executing the action a, observe the new state s', the reward R and the termination flag done. Store the sample (s, a, logp, R, s', done) into the Rollout replay pool, and update the current state to s'.
[0134] 4.3) Calculate the advantage estimate. Suppose at a certain time step t, the policy π θ has selected the action and according to the samples in the replay pool, the advantage of this action is calculated to be 2.
[0135] 4.4) Optimization objective function: in this embodiment, take ∈ = 0.2, use the probability ratio r t (θ) and the advantage function to construct the objective function L(θ).
[0136] 4.5) Update the policy. Use the gradient ascent method to update the policy parameters, i.e. take the learning rate α = 0.001. Use the gradient descent algorithm to update the parameters φ of the value network.
[0137] Repeat steps 4.6. Repeat steps 5.2)-5.4) using the new policy parameters until 1000 iterations have been reached. Finally, clear the replay pool Rollout and save the parameters of the CDP-SS model after all update cycles are completed.
[0138] (5) Cumulative data processing application job scheduling
[0139] 5.1) Let 20 cumulative data processing applications run in a multi-cloud environment consisting of three data centers. The scheduling problem for cumulative data processing applications is defined as follows:
[0140]
[0141] 5.2) When these 20 CDP applications start scheduling, the scheduling environment is initialized according to the status information in 2.1).
[0142] 5.3) Load the trained CDP-SS model for scheduling decision making. The scheduling action is The scheduling action is The scheduling action is The scheduling action of AggJob1 is The scheduling action is The scheduling action is Other job scheduling decisions are similar.
[0143] 5.4) Further train and optimize the CDP-SS model based on empirical data.
[0144] The performance test compares this method with the mainstream scheduling algorithms in the current cloud environment to demonstrate the performance advantages of the method proposed in this invention in CDP application scheduling. The comparison method is as follows:
[0145] (1) Random algorithm: Select jobs to be scheduled from the job queue and allocate virtual machine resources in a random manner.
[0146] (2) FCFS: Jobs are scheduled in the order they are submitted to the system. Jobs that arrive first are given priority to obtain computing resources and start execution.
[0147] (3) PSOMC: Based on particle swarm optimization and membrane computing algorithms, it finds the optimal scheduling solution based on the information of the jobs to be scheduled and the available virtual machine resources, with the job completion time and cost as the goals.
[0148] (4) DB-ACO: Based on the ant colony algorithm, it solves the optimal scheduling solution with the premise of meeting deadline and budget constraints and the goal of minimizing the execution cost of workflow jobs.
[0149] (5) HCDRL: Based on deep reinforcement learning technology, it realizes the collaborative scheduling of multiple workflows with workflow job queues and virtual machine resource usage as status and job execution time, cost and fairness as rewards.
[0150] The performance test was run on a computer with an Intel Core i7-9700 processor, 32GB of memory, a 1TB hard drive, and Windows operating system.
[0151] In model evaluation, this method measures performance using cost and SLA violation rate. The cost metric quantifies the average cost of scheduling CDPs. The SLA violation rate metric measures the timeliness of the scheduling model when processing scheduled jobs, specifically assessing whether the application can be scheduled within the specified deadline. The calculation formula is as follows:
[0152]
[0153] Where m is the generated CDP i The number of cost i Including CDP i The total cost of the scheduled preprocessing and aggregation jobs in .
[0154] The performance test used batch processing jobs from Alibaba logs. The total data size processed by the application was calculated by multiplying the memory usage and execution time of all included preprocessing jobs. In the experiment, application processing data follows a uniform distribution. When the generated data reaches a given size, the application preprocessing job submission is triggered. For each generated CDP application, the number of preprocessing jobs generated is equal to the number of tasks in the corresponding job in the log. The resource requirements of the preprocessing and aggregation analysis jobs are determined by the average CPU and memory usage of all included tasks and the total number of tasks. The computational load of the job is calculated by multiplying the floating-point computing power of the CPU resources actually used by the job by the job execution time. The application deadline is set as the sum of α times the total execution time of the corresponding application in the log and the application submission time. α is set as a random number in the range [0.1, 1] to simulate the different execution urgency requirements of CDP applications.
[0155] The experimental comparison results of the present invention are as follows:
[0156] Table 1 Cost comparison of CDP-SS model with other algorithms
[0157]
[0158] Table 2 Comparison of SLA violation rates of the CDP-SS model with other algorithms
[0159]
[0160] In terms of cost comparison, the CDP-SS model demonstrates significant advantages over other algorithms. Specifically, the cost of the CDP-SS model is 4.459, which is approximately 20.95% lower than the cost of the Random algorithm, approximately 19.67% lower than the cost of the FCFS algorithm, approximately 13.57% lower than the cost of the DB-ACO algorithm, approximately 10.73% lower than the cost of the PSOMC algorithm, and approximately 8.96% lower than the cost of the HCDRL algorithm. This demonstrates that the CDP-SS model excels in cost control.
[0161] The CDP-SS model also performed well in terms of SLA violation rates. Its SLA violation rate was 3.32%, significantly lower than the Random algorithm's 31.37% (a decrease of approximately 28.05 percentage points), the FCFS algorithm's 34.53% (a decrease of approximately 31.21 percentage points), the DB-ACO algorithm's 15.86% (a decrease of approximately 12.54 percentage points), the PSOMC algorithm's 14.26% (a decrease of approximately 10.94 percentage points), and the HCDRL algorithm's 10.76% (a decrease of approximately 7.44 percentage points). These data clearly demonstrate that the CDP-SS model has a significant advantage in reducing SLA violation rates and can effectively improve service quality and user experience.
[0162] Finally, it should be noted that the above examples are only used to illustrate the present invention and are not intended to limit the technology described in the present invention. All technical solutions and improvements that do not depart from the spirit and scope of the invention should be included in the scope of the claims of the present invention.
Claims
1. A multi-cloud job scheduling method for cumulative data processing applications, characterized by: It consists of five steps: initialization, Markov decision process modeling, CDP-SS network structure model construction, CDP-SS model training, and cumulative data processing application job scheduling; it has six basic parameters: the discount factor γ for calculating future rewards, the actor network learning rate α1, the critic network learning rate α2, the generalized advantage estimate λ, the clip parameter for policy gradient updates, and the batch size of the training process. Among them, γ is 0.95, α1 is 0.0001-0.001, α2 is 0.0001-0.001, λ is 0.95, the clip parameter is 0.2, and the batch size is 256-1024. (1) Initialization: Use the information of the accumulated data processing application in the log to initialize the data; the data preprocessing stage includes multiple data preprocessing jobs; the data aggregation analysis stage includes only one aggregation analysis job, which starts the submission schedule after all preprocessing jobs are completed; 1.1) The information of the cumulative data processing application submitted by each user is defined as: CDP i =(PreOps i ,AggOps i ,DataSource i ,DStartTime i ,DEndTime i ,Deadline i ); among them, CDP i Indicates the i-th CDP application; PreOps i and AggOps i Respectively represent the data operations performed in the data preprocessing stage and the aggregation analysis stage; DataSource i Indicates the data source processed by the application; DStartTime i and DEndTime i Respectively represent the start and end time of the data processed by the application; Deadline i It indicates the end time of the application that the user expects; 1.2) Use DCS = {DC1, DC2, .....DC n} indicates that the multi-cloud environment consists of multiple cloud data centers; each data center DC i Contains a group of physical servers, with machine i,j Represents the jth server in the i-th data center; each cloud data center DC i Provide tni A virtual machine image, defined as Therefore, any virtual machine image Defined as in, Indicates the computing power of the virtual machine CPU. Indicates the memory size of the virtual machine. Indicates the rental price of the virtual machine per unit time; 1.3) A CDP application consists of multiple preprocessing jobs and one aggregation analysis job. The i-th CDP application job set is represented as in AggJob represents the jth preprocessing job of the i-th CDP application. i represents the aggregation job of the i-th CDP application; 1.4) Let CDPJob represent CDP_JOBS i For any job in the set, the processing data size of CDPJob is datasize, the computation amount is length, the computing power of the allocated virtual machine instances is vflops, the number is vnum, and the job turnaround time is MK(CDPJob). Then the job turnaround time of CDPJob is expressed as: MK(CDPJob)=t_Wait(CDPJob)+t_DTrans(CDPJob)+t_Exe(CDPJob)#(1) t_Wait(CDPJob) represents the time a job waits for scheduling, and is calculated using the following formula: t_Wait(CDPJob)=ATime(CDPJob)-STime(CDPJob)#(2) ATime(CDPJob) represents the arrival time of the job, STime(CDPJob) represents the scheduling time of the job; t_DTrans(CDPJob) represents the transmission time of the job processing data, which is calculated using formula (3) based on the network bandwidth between the application agent submitting the job and the cloud data center; Among them, bd(DC p ,DC q ) represents the data center DC p and DC q The network bandwidth between the two nodes is represented by the sizeof(data), the sizeof(data) represents the size of the transmitted data, and t_Exe(CDPJob) represents the execution time of the job, which is calculated as follows: 1.5) For any job, its virtual machine resource cost j_vcost is expressed as follows: j_vcost=vnum×consumetime×vprice#(5) in, vnum and vprice represent the number of virtual machine instances allocated to the job and the price per unit time, respectively. consumetime represents the resource consumption time of the job, which is the sum of its data transmission time and execution time. Apply CDP i Data transmission cost data_cost i It is the cumulative cost of transmitting intermediate results between the data center where all preprocessing jobs are located and the data center where the aggregation analysis job is located, and is expressed as: Among them, ∑(·) represents the summation operation, and the subscript of ∑(·) is the range of the summation operation, Represents a preprocessing job The size of the intermediate result data generated, Represents a preprocessing job Aggregation analysis job AggJob i The price of transmitting unit-scale data between data centers. When the data preprocessing and aggregation analysis stages are performed in the same data center, the data transmission cost is zero. In summary, application of CDP i The total cost of preprocessing operations p_cost i , the cost of the aggregate analysis operation a_cost i and the total application cost i Respectively expressed as: a_cost i =j_vcost(AggJob i )+data_cost i #(8) cost i =p_cost i +a_cost i #(9) (2) Markov decision process modeling 2.1) State space: The state space includes the cloud resource state, the state of pending jobs, and the state of the historical job queue; 2.1.1) Cloud Resource State: The defined cloud resource state includes the virtual machine state State_VM and the physical server state State_PM, expressed as: State_MC = (State_VM, State_PM). State_VM records the resource allocation and usage status of virtual machine instances deployed in multiple cloud data centers. For each virtual machine instance, its state is described by a five-tuple vector as follows: vm_state=(vCPUcapacity,vMemcapacity,rentperiod,vCPUutil,vMemutil) in, vCPUcapacity and vMemcapacity represent the rated CPU computing capacity and memory resource scale of the virtual machine, respectively, and are obtained based on the type vtype of the virtual machine instance. Rentperiod represents the lease duration of the virtual machine instance. vCPUutil represents the ratio of the CPU resource demand vCPUReq of the job using the virtual machine instance to the rated CPU computing capacity vCPUcapacity of the virtual machine instance. Similarly, vMemutil is calculated. vCPUutil and vMemutil represent the actual utilization efficiency of the rated resources of the virtual machine, which can reflect the rationality of job resource allocation. Based on the state of virtual machine instances, the state of a data center's virtual machine, State_VM, is expressed as a matrix consisting of the states of all virtual machine instances in multiple data centers in sequence. In the State_VM matrix, the virtual machine instance states are arranged in the order of data center-physical server. For each physical server, the matrix rows are allocated according to the maximum number of deployable virtual machine instances. When the number of actually deployed virtual machine instances is less than the maximum value, the row is padded with zero values. Each row of the matrix represents the state of three dimensions of a physical server in the data center, including the virtual machine share allocation ratio VMutil, the CPU resource allocation ratio CPUutil, and the memory resource allocation ratio Memutil. VMutil is expressed as the ratio of the number of deployed virtual machine instances on the physical server to the total rated number of virtual machine instances. 2.1.2) Status of Jobs to be Scheduled: Suppose there are N jobs in the queue of jobs to be scheduled, then the status of the jobs to be scheduled, State_job, is represented by an N×9 matrix; State_job = (jobtype, length, datasize, vnum, vCPUReq, vMemReq, density, app_urgency, app_perlength); where each row of the matrix represents the status of a job; the status of each job consists of seven dimensions; jobtype represents the type of job, which can be divided into preprocessing jobs and aggregation analysis jobs; length represents the computational scale of the job; datasize represents the data size processed by the job; vnum represents the number of virtual machines required by the job; vCPUReq and vMemReq represent the CPU and memory requirements of the job, respectively; density represents the computational load that the job must complete per unit time by the application deadline, calculated as density = length / remaintime, where residualtime represents the duration from the current time to the expected deadline of the application; app_urgency and app_perjob represent the status of the application to which the job belongs; app_progress represents the relative remaining time of the application, and app_perjob represents the expected job execution progress of the application, as shown below: Deadline, Starttime, and Curtime represent the expected deadline, start time, and current time of the application; num_remainingjob represents the number of jobs belonging to the application in the queue of jobs to be scheduled; 2.1.3) Historical job queue status: State_Req represents the five characteristics of the jobs generated in each unit time period, namely, the total data size, the total computing size, the number of virtual machines required, the total CPU requirements of the virtual machines, and the total memory requirements of the virtual machines, which are expressed as 2.2) Action Space: Traverse the queue of jobs to be scheduled in order of arrival, and use the reinforcement learning agent to determine the scheduling decision for each job to be scheduled. Therefore, the action space of reinforcement learning is expressed as: in, None means the job is not scheduled. Indicates that the job is scheduled to the data center DC i , the virtual machine type assigned to it is 2.3) Rewards; Define reward items: VM resource allocation quality R vm , data aggregation R data , Application Execution Risk R risk , New operating cost R cost The reward for any scheduling decision is expressed as the cumulative sum of the above four reward items; 2.3.1) VM resource allocation quality: The VM resource allocation quality of a job is defined as follows: in, vCPUReq and vMemReq are obtained from the job's resource request, while vCPUAlloc and vMemAlloc are the actual resource sizes allocated to the virtual machine. α and β are weight coefficients, which sum to 1. vCPUMax and vMemMax represent the maximum CPU and memory configurations for all types of virtual machines, respectively, while vCPUMin and vMemMin represent the minimum configurations. 2.3.2) Data aggregation; when the scheduled job is a pre-processing job , then the application data aggregation reward is defined as follows: Wherein, ∑(·) represents a summation operation, the superscript m and subscript j=1 of ∑(·) represent that the value range of j is an integer from 1 to m, and PreJob_f i j Indicates the application of CDP i Completed preprocessing jobs, m represents the total number of completed preprocessing jobs in the application, dc(·) represents the data center where the job is located; nor_bd(·) represents the normalized bandwidth between the cloud data centers where two jobs are located, that is, the ratio of the actual network bandwidth to the maximum network bandwidth; for two jobs located in the same data center, nor_bd(·) is set to 1; 2.3.3) Application execution risk: The reward function is set to the average risk of all generated applications, expressed as: in, Is CDP applied in the job queue i The expected turnaround time of unscheduled jobs, N is the number of generated applications, and t represents the current time point; 2.3.4) New operating costs; the bonus value is expressed as: Here, nor_vc(·) represents the normalized cost of the currently scheduled job, which is the ratio of the resource cost of the currently scheduled job to the average cost of all scheduled jobs. (3) Construction of the CDP-SS model network structure; The CDP-SS model is designed based on the PPO algorithm; The Actor network and the Critic network have similar network structures; The state input of each network consists of four parts: State_VM, State_PM, State_job, and State_Req; 3.1) VM-autoencoder module: The State_VM matrix is partitioned into sub-matrices, each corresponding to the virtual machine state on a physical server. A five-dimensional latent feature vector is generated for each sub-matrix using the VMautoencoder. Finally, the latent feature vectors of the virtual machine states of all physical servers are sequentially concatenated to form the LatentVMencoding matrix for the overall virtual machine latent feature of the multi-cloud data center. The VMautoencoder's encoder design uses a combination of two convolutional layers and one fully connected layer. Specifically, both convolutional layers are configured with a convolution kernel of size 1×3, a stride of 1, a fixed number of output channels of 16, and use ReLU as the activation function. After these two convolution operations, the feature matrix is compressed into a one-dimensional vector by the flatten operation and then input into the fully connected layer to extract a feature vector containing five elements. The VMautoencoder's decoder takes the five-element feature vector output by the encoder as input. This vector is first processed by a fully connected layer containing 160 neurons and then enters two deconvolution layers to reconstruct the feature matrix. The parameters of these two deconvolution layers are consistent with the corresponding convolution layer parameters in the encoder to ensure accurate transmission and reconstruction of information. The activation function is also ReLU. 3.2) Concat Module: The Concat module is responsible for concatenating LatentVMencoding, State_PM, and State_job to form a complete state matrix for the multi-cloud data center, which serves as the input to DC_Encoder. First, the data center virtual machine state latent feature matrix LatentVMencoding is horizontally concatenated with the physical machine state feature matrix State_PM. Second, the concatenated feature matrix is vertically concatenated with the state feature matrix of the job to be scheduled, State_job. 3.3) DC_Encoder module; DC_Encoder is designed as a neural network consisting of two convolutional layers and one pooling layer. Specifically, the first convolution layer uses a 2×2 convolution kernel with a stride of 1 and 16 output channels. The second convolution layer uses a 3×3 convolution kernel with a stride of 1 and 16 output channels. The ReLU activation function is applied after each convolution layer. The feature matrix is pooled multiple times using three adaptive pooling kernels with bin sizes of 1x1, 2x2, and 3x3, respectively. The pooling results are concatenated to obtain a fixed-size output vector, namely LatentDCencoding. 3.4) Req_Predict Model: The job status prediction model, Req_Predict, takes State_Req as input and constructs a hidden layer using a two-layer GRU structure and two fully connected layers to extract time series features from the data. Specifically, the GRU's time window size is set to 10. Each layer of the two-layer GRU contains 32 neurons, and the two-layer neural network is used for progressive processing. After State_Req passes through the two-layer GRU, its output dimension is 10×64, which is expanded into a one-dimensional vector with 640 features. The model then uses two fully connected layers for feature extraction and conversion; the number of neurons in these two fully connected layers is 64 and 3 respectively. The output of the last fully connected layer is the predicted value of the job status in the next time unit, represented by PredictedReq; these predicted values specifically include the total number of virtual machine requirements, the total computing amount, and the total data amount, represented as 3.5) Act-Dec module: Based on the obtained LatentDCencoding hidden vector and the predicted PredictedReq vector, these two vectors are concatenated. The concatenated vector is then passed through a fully connected network layer and mapped into a k-dimensional vector. Where k is the number of actions in the action space. It is then activated by the softmax function and finally outputs the selection probability of each possible action. (4) CDP-SS model training: In the proposed network structure, the VMautoencoder and Req_Predictor modules are trained independently offline, and the trained predictor and encoder modules are integrated into the Actor and Critic networks; the DC_Encoder and Act_Dec are trained end-to-end. 4.1) Initialize the model; initialize the Actor and Critic in the model, and set the discount factor γ, the Actor network learning rate α1, the Critic network learning rate α2, the advantage estimate λ, the Clip parameter, and the batch size hyperparameter; 4.2) Collect data; load the CDP application into the multi-cloud environment for scheduling. At each scheduling, according to the current state s and the policy model π θ Output action a and its logarithmic probability logp, as well as the value model V φ After executing the action, the new state s′, reward r and termination flag done are observed. Each time a CDP application job is scheduled, a sample experience data Experience is recorded. Experience is represented by a tuple (s, a, logp, v, s′, r, done). These sample experience data are stored in the replay pool Rollout, and the current state is updated to s′. 4.3) Calculate advantage estimate; randomly extract 256 to 1024 experience data from the replay pool as training samples; use the output of the value network as the output of each state s t The state value V(s t ); In order to evaluate the quality of an action relative to the average level, it is necessary to calculate the advantage function The calculation formula is: Among them G t is the discounted cumulative return starting from time step t, γ is the discount factor, r t+1+k is the reward obtained at time step t+1+k; 4.4) Optimizing the objective function: calculating the strategy ratio where π θ (a t |s t ) is the current policy, It is the old strategy; the PPO algorithm uses a specially designed objective function, the objective function is in the form of: Where E represents the expected value, is an estimate of the advantage function, ε is a small positive number with a value of 0<ε<0.4, and the clip function limits the strategy ratio r t (θ) changes in the range, when r t When (θ) exceeds the range of 1-∈ to 1+∈, clip will cut it to this interval and set the exceeding value to the maximum boundary value; 4.5) Update the strategy; use the gradient ascent method to update the strategy parameters, that is, Where α is the learning rate, and θ is the parameter value of the policy network, including all weights and bias values in the input layer, hidden layer, and output layer of the neural network structure in Sections 3.3) and 3.5); Represents the gradient of L(θ) with respect to the parameter θ, which is a vector whose each component is the partial derivative of L(θ) with respect to the corresponding component of θ; calculate the loss function of the value network L(φ) = E[(V(s t )-r) 2 ], using the gradient descent algorithm to update the parameter φ of the value network to minimize the loss function L(φ); the network structure of the value network is the same as that of the policy network, and the parameter φ includes all weights and bias values in the neural network structure; 4.6) Repeat steps; repeat steps 4.2)-4.4) above with the new policy parameters until the number of iterations reaches 1000; finally, clear the replay pool Rollout and save the model CDP-SS parameters after all update cycles are completed; (5) Cumulative data processing application job scheduling 5.1) Let M cumulative data processing applications run in a multi-cloud environment consisting of K data centers. The scheduling problem for cumulative data processing applications is defined as follows: Among them, SLA (CDP i ) represents the number of SLA-compliant applications, Max(·) represents maximization, Min(·) represents minimization, and cost(·) represents the total cost of the application. Application scheduling has two objectives: maximizing the number of applications that meet deadline constraints and minimizing the resource cost of application execution. 5.2) When the cumulative data processing application starts scheduling, the state information of the scheduling environment is initialized based on the results of the Markov decision process modeling, including the state of cloud resources, the state of pending jobs, and the state of the historical job queue; 5.3) Load the trained CDP-SS model for scheduling decisions. During each CDP application preprocessing job or aggregation job scheduling cycle, the model receives the cloud environment's State_VM, State_PM, State_job, and State_Req state information. This state information is input into the CDP-SS model's Actor network, which outputs a probability distribution for scheduling actions and selects the action with the highest probability as the scheduling policy. The scheduling policy corresponds to an element in the action set defined in Section 2.2, indicating that the job should not be scheduled or should be scheduled to a VM type in a specific cloud data center. 5.4) Record a sample experience data Experience according to the job scheduling, and store the sample experience data in the replay pool Rollout; the sample experience data is used for CDP-SS model training in Sections 4.3) to 4.6).