Cloud edge collaborative workflow real-time scheduling method and system based on deep reinforcement learning

Through the method of deep reinforcement learning, the task-server association relationship is constructed and the server status is predicted using the LSTM neural network. Combined with the Actor-Critic network for real-time scheduling, the instability problem of task scheduling in the cloud-edge collaborative environment is solved, and more efficient task allocation and security are achieved.

CN120762855APending Publication Date: 2025-10-10HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510925345.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-05
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively cope with real-time, changing, and complex scenarios in a cloud-edge collaborative environment, resulting in task scheduling cost overruns and execution failures, especially when edge computing resources fluctuate, affecting the security of applications such as autonomous driving.

Method used

A method based on deep reinforcement learning is adopted to build sub-task association relationships and server status monitoring of workflow task sets. Z-Score normalization, Laplace matrix decomposition, K-Means++ clustering and LSTM neural network prediction are used, combined with the Actor-Critic network to make real-time task scheduling decisions and optimize the allocation of tasks on cloud-edge servers.

Benefits of technology

It achieves the goal of avoiding task scheduling cost overruns and failures in a dynamic cloud-edge computing environment, achieving shorter execution times and better task scheduling solutions, and improving the adaptability and security of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120762855A_ABST
    Figure CN120762855A_ABST
Patent Text Reader

Abstract

The invention provides a workflow real-time scheduling method and system based on deep reinforcement learning for the field of cloud edge collaboration. According to the method, firstly, based on the available resource quantity of a cloud side server and a workflow task submitted by a user, key attributes and association relationships of all subtasks of the workflow are extracted, and based on normalization and a Laplacian matrix decomposition method, normalization attributes of the task are calculated; then, on the basis of a K-Means + + algorithm and the characteristic value of the cloud side server, the sub-tasks are clustered according to similarity and then divided into different execution environments, so that resource and cost hyper-branched is avoided; then, an online deep learning model is constructed based on an LSTM neural network, and the bandwidth, CPU and memory resources of the cloud edge server are predicted in real time; and thirdly, constructing an initial state quantity of a deep reinforcement learning (DDPG) algorithm based on a server state prediction value generated in real time by a neural network, and calculating a scheduling scheme with the shortest execution time of the current workflow under the cost constraint in real time. And finally, the computer system executes workflow task scheduling based on the generated optimal scheduling scheme, and sends an execution result back to the user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of cloud-edge collaborative workflow task scheduling, and specifically designs to sample and predict the real-time resource state of cloud-edge servers based on online deep learning, and uses a deep reinforcement learning DDPG method to generate a real-time scheduling scheme for workflow tasks. BACKGROUND

[0002] With the widespread development of cloud computing, edge computing, Internet of Things and other industries, cloud-edge collaboration is widely used in real service scenarios, such as the metaverse, Internet of Vehicles, large model inference deployment, etc. Cloud computing and edge computing have the outstanding characteristics of sufficient computing resources and proximity to users, respectively. By making good use of both cloud computing and edge computing computing environments, the user's upload task or task group can be simultaneously satisfied in terms of low latency, low cost, efficient computing, high privacy, and other different dimensions. For example, in the processing of automatic driving tasks in the Internet of Vehicles environment, the automatic driving task can be decomposed into a workflow composed of positioning tasks, picture processing tasks, behavior decision tasks, etc. The positioning task usually requires low latency and high privacy protection level, and can be executed on the edge; the picture processing task usually contains more picture analysis and other data computing tasks, and can be executed on the cloud; based on the analysis results of the positioning and picture processing tasks, the system implements behavior decision on the vehicle side to provide reference for the next step of the automatic driving system.

[0003] Deep Reinforcement Learning (DRL) combines the perception ability of deep learning and the decision-making ability of reinforcement learning, allowing agents to interact with complex environments and try and error, providing solutions for the perception and decision-making of complex systems. DRL has made great progress in the field of scheduling and is widely used in resource allocation, task offloading, task scheduling, etc. With the continuous development of DRL, its modeling and decision-making capabilities for complex environments have been continuously enhanced, and new applications and breakthroughs have been made in complex scenarios such as Internet of Vehicles, metaverse large models, etc. However, although DRL can perform well in scheduling decisions, it is still difficult to make real-time decisions when faced with real-time and complex scenarios. In the face of flexible and variable real-world scenarios, especially in edge computing scenarios where bandwidth and computing resources are in real-time change, the long data transmission time and execution time caused by environmental factors can cause cost overruns and even task scheduling execution failures. For example, in the automatic driving task mentioned above, if the data transmission of the edge end where the positioning service is located is timed out or even dropped, it will directly affect the final automatic driving decision, which may cause serious traffic accidents. In order to ensure the safety of related applications, it is necessary to study the workflow task scheduling method in the cloud-edge collaborative environment to adapt to the actual scheduling environment.

[0004] Current workflow task scheduling methods for cloud-edge collaborative environments fall into two main categories: heuristic-based scheduling and reinforcement learning-based scheduling. Heuristic-based scheduling methods employ specific optimization strategies to optimize decisions within the solution space to obtain feasible solutions. These algorithms are generally suitable for simple scenarios and can quickly find solutions, but they are prone to getting stuck in local optima and require continuous rule adjustments for complex, dynamic, real-time scenarios. Reinforcement learning-based scheduling methods use continuous trial-and-error interaction with the environment to iteratively learn strategies. They can also adjust strategies and rewards promptly in the face of unexpected changes, resulting in greater adaptability and the ability to achieve globally optimal solutions. Currently, most research focuses on heuristic algorithms. While these methods can quickly generate solutions and minimize the impact of real-world resource fluctuations on decision-making, they still struggle to avoid cost overruns and execution failures caused by resource fluctuations. Therefore, reinforcement learning-based scheduling methods offer greater adaptability. Summary of the Invention

[0005] Purpose of the invention: The purpose of the present invention is to provide a real-time scheduling method and system for cloud-edge collaborative workflows based on deep reinforcement learning. In the face of cloud-edge server resource fluctuations in real scenarios, it can avoid cost overruns and task execution failures and obtain a better task scheduling solution.

[0006] Technical Solution: To achieve the above-mentioned invention objectives, the present invention provides a method and system for real-time scheduling of cloud-edge collaborative workflows based on deep reinforcement learning, including the following steps:

[0007] Step 1: Based on the workflow tasks uploaded by the user, the association relationship and DAG structure between the subtasks of the workflow task set are constructed based on the graph adjacency matrix, and the attributes of each subtask in the workflow task set are read;

[0008] Step 2: Based on the rented cloud-edge collaborative computing environment, communicate with the service environment's automated monitoring system to obtain the server's current status attributes;

[0009] Step 3: Based on the subtask attributes of the workflow task set obtained in Step 1, the numerical features are normalized using the Z-Score normalization method. The task association relationships are decomposed using the Laplace matrix. Based on the above two parts of workflow characteristics, the K-Means++ algorithm is used to perform n = 4 task clustering. The resulting clusters are divided into cloud environments or edge environments based on the minimum total execution time under the budget constraint.

[0010] Step 4: Use LSTM neural network to predict server status online. Use an input sequence window with a capacity of W, input a real-time updated LSTM neural network model, and predict the status of each server at the next time step. Each sequence is saved in an experience replay buffer R DNN , the network model is asynchronously updated using a dynamic real-time update strategy to achieve more accurate prediction;

[0011] Step 5: initialize the deep reinforcement learning base parameters, take the state quantity sampled by the sliding window as the current state quantity, initialize each training round and perform action selection for t time steps, store the four-tuple (s t ,a t ,r t ,s t+1 ) value experience replay buffer, randomly take samples from the buffer to train and update the Actor and Critic networks, repeat the above process until the maximum iteration number or termination condition is reached;

[0012] Step 6: take the output of step 5 as the final scheduling scheme, which is used for scheduling user-submitted workflow tasks on the current cloud execution environment, execute the tasks according to the scheduling scheme, and return the task execution results to the user;

[0013] Preferably, the attributes of the sub-tasks of the workflow task set in step 1 include the number of instructions of the task, the estimated execution time of the task, and the estimated data input / output quantity of the task;

[0014] Preferably, the server attributes in step 2 include the online status of the cloud edge server, the number of CPU cores of the server, the frequency of the CPU, the size of the memory, the size of the bandwidth, and the response delay;

[0015] Preferably, the process in step 3 can be further described as:

[0016] Step 31: for the numerical parameters of the number of instructions, the estimated execution time, and the estimated data input / output quantity of the sub-tasks, data normalization is performed using the following Z-Score normalization formula:

[0017]

[0018] Step 32: for the association relationship between the sub-tasks, the following formula is used to calculate the normalized Laplacian matrix and perform eigenvalue decomposition, extract the first k non-zero eigenvectors as low-dimensional embedding;

[0019] L = D - A sym

[0020] Step 33: for the normalized numerical data and the dimensionality-reduced association relationship data extracted above, K-Means++ clustering method is used to perform N=4 clustering, the algorithm is randomly selected as the center, the distance between the sample point and the cluster center is calculated according to the following formula, and the nearest cluster center is selected, the task is added to the cluster where the nearest cluster center is located:

[0021]

[0022] Step 34: According to the state parameters of the cloud edge server, the average value of the resource amount is calculated, and the clustering division scheme that can minimize the overall execution time of the workflow task is calculated under the cost constraint, and a suitable execution environment is selected for each cluster and the subtasks inside it;

[0023] As a preferred, the server real-time state prediction process based on online deep learning in step 4 can be further described as:

[0024] Step 41: Initialize the main LSTM neural network parameter θ main and the parameter θ tar get of the target LSTM neural network, initialize the network hyperparameters such as learning rate Learning_rate, buffer update parameter τ prediction , batch size Batch_size;

[0025] Step 42: Use a sliding window to sample the server state of the last W time windows with a time step length of 10 milliseconds, and use the normalized value as the neural network input;

[0026] Step 43: For the server Sever i , the real-time state s i,t = <cpu, mem, bandwidth> is sampled in real time;

[0027] Step 44: For the sampled real-time state s i,t , use the Z-Score standardization method for normalization, and store s′ i,t in the experience replay buffer R DNN , and keep the last N=10000 data;

[0028] Step 45: Extract the previous L states from R DNN , input the main LSTM neural network model, and predict the next time step state Pass the state into deep reinforcement learning as the initial state s1 of action selection;

[0029] Step 46: Sample the real-time state s i,t+Δt , and calculate the neural network loss L based on the following formula:

[0030]

[0031] Step 47: When the update time step or the loss threshold L Max is reached, randomly extract Batch_size samples from R DNN , and calculate the average loss LBatch , calculate new network parameters θ', update network parameters using the following formula and synchronize to the main neural network and target neural network to avoid short-term dramatic fluctuations in network parameters dramatic changes;

[0032] θ' = τ prediction θ + (1 - τ prediction )θ', τ ∈ (0, 1)

[0033] As preferred, the scheduling scheme decision process based on reinforcement learning described in step 5 can be further described as:

[0034] Step 51: Randomly initialize Actor network μ(s|θ μ ) and Critic network Q(s|θ Q ) parameters θ μ and θ Q , create the corresponding Actor network, Critic network and the corresponding target network;

[0035] Step 52: Randomly initialize an experience replay buffer R with capacity N, used to store action decision samples, i.e. four-tuple (s t , a t , r t , s t+1 );

[0036] Step 53: Initialize reinforcement learning hyperparameters, including discount factor γ, batch size B, learning rate;

[0037] Step 54: For each time step, get the initialized environment state s1, containing workflow task state information, and each server state information obtained by sliding window sampling;

[0038] Step 55: Actor network selects action a μ according to μ(s|θ t ) network and current time step state s t , after execution, state s t becomes s t+1 , get reward r t , store four-tuple (s t , a t , r t , s t+1 ) in experience replay buffer R;

[0039] Step 56: Randomly draw B samples from the cache R for reward evaluation and update Actor and Critic network;

[0040] Step 57: Critic network Q(s|θ Q) The network calculates the target Q value and uses the following formula to calculate the loss function, and the new parameter θ is obtained by gradient descent Q’ :

[0041]

[0042] Step 58: The Actor network optimizes the function by maximizing the Q value of the Critic network, calculates the optimization target according to the following formula, and obtains the new parameter θ by gradient ascent μ’ :

[0043]

[0044] Step 59: Update the parameters of the target network according to the following formula:

[0045] θ Q’ =τθ Q +(1-τ)θ Q’ ,τ∈(0,1)

[0046] θ μ’ =τθ μ +(1-τ)θ μ’ ,τ∈(0,1)

[0047] Based on the same inventive concept, the present invention provides a computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the computer system implements the steps of real-time scheduling of workflow tasks in a cloud-edge collaborative environment based on deep reinforcement learning.

[0048] Based on the same inventive concept, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of real-time scheduling of workflow tasks in a cloud-edge collaborative environment based on deep reinforcement learning.

[0049] Beneficial effects: The present invention provides a real-time scheduling method for workflow tasks in a cloud-edge collaborative environment based on deep reinforcement learning. The method proposes a new task-server adaptability calculation method to divide workflow tasks into appropriate execution environments to avoid cost overruns and scheduling failures in task scheduling. At the same time, the online deep learning strategy is used to predict the status of the cloud-edge server in real time, providing a basis for the task scheduling decision of the reinforcement learning algorithm. Because the present invention adopts the above-mentioned technical solution, the solution can achieve optimal scheduling of workflow tasks in a dynamically fluctuating cloud-edge computing environment, and meet shorter execution time while ensuring cost constraints. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1Method flowchart of the embodiment of the present application;

[0051] Figure 2 Task-server adaptability sensing flowchart of the embodiment of the present application;

[0052] Figure 3 Server state prediction method flowchart based on online reinforcement learning of the embodiment of the present application;

[0053] Figure 4 Reinforcement learning task scheduling flowchart of the embodiment of the present application; DETAILED DESCRIPTION

[0054] The present application will be further clarified by the following examples with reference to the accompanying drawings. It should be understood that these examples are only used to illustrate the present application and not intended to limit the scope of the present application. After reading the present application, those skilled in the art can make various modifications to the present application, and all such modifications fall within the scope of the appended claims.

[0055] As shown in Figure 1 The embodiment of the present application discloses a workflow task real-time scheduling method in a cloud-edge collaborative environment based on deep reinforcement learning, mainly including the following steps:

[0056] Step 1: Extract the workflow task attributes. According to the user uploaded workflow task, the association relationship and DAG structure between the subtasks of the workflow task set are constructed based on the graph adjacency matrix, and the attributes of each subtask in the workflow task set are read;

[0057] Specifically, the attributes of the subtasks of the workflow task set in step 1 include the number of instructions of the task, the estimated execution time of the task, and the estimated data output of the task;

[0058] Step 2: Extract the cloud-edge server state attributes. According to the rented cloud-edge collaborative computing environment, communicate with the automatic monitoring system of the service environment to obtain the current state attributes of the server;

[0059] Specifically, the server attributes in step 2 include the online state of the cloud-edge server, the number of CPU cores of the server, the frequency of the CPU, the size of the memory, the size of the bandwidth, and the response delay;

[0060] Step 3: Task-server adaptability sensing. According to the subtask attributes of the workflow task set obtained in step 1, the numerical features are normalized using the Z-Score standardization method, and the task association relationship is decomposed using the Laplacian matrix. According to the above two parts of the workflow features, the K-Means++ algorithm is used to perform n=4 task clustering, and the obtained clustering is divided into cloud environment or edge environment according to the standard of minimum total execution time under the budget constraint;

[0061] like Figure 2 As shown, the task-server fitness perception in step 3 of the present invention includes the following process:

[0062] Step 31: For the numerical parameters of each workflow subtask, such as the number of task instructions, estimated execution time, and estimated data input and output, the following Z-Score normalization formula is used to perform data normalization:

[0063]

[0064] Step 32: For the association between subtasks, use the following formula to calculate the normalized Laplacian matrix and perform eigendecomposition, extracting the first k non-zero eigenvectors as low-dimensional embedding;

[0065] L=DA sym

[0066] Step 33: For the normalized numerical data and the dimensionality-reduced correlation data extracted above, use the K-Means++ clustering method to perform clustering with N=4;

[0067] Step 34: Randomly select the algorithm to gather the center;

[0068] Step 35: Calculate the distance between the sample point and the cluster center according to the following formula, select the nearest cluster center, and add the task to the cluster where the nearest cluster center is located:

[0069]

[0070] Step 36: Based on the state parameters of the cloud edge servers, calculate the average resource amount, calculate the clustering partitioning scheme that can minimize the overall execution time of the workflow tasks under the cost constraint, and select the appropriate execution environment for each cluster and its internal subtasks;

[0071] Step 4: Use LSTM neural network to predict server status online. Use an input sequence window with a capacity of W, input a real-time updated LSTM neural network model, and predict the status of each server at the next time step. Each sequence is stored in an experience replay buffer R DNN , using a dynamic real-time update strategy to asynchronously update the network model to achieve more accurate predictions;

[0072] like Figure 3 As shown, the LSTM-based online deep learning prediction process in step 4 of the present invention can be further described as follows:

[0073] Step 41: Initialize the main LSTM neural network parameters θ mainand the parameters θ of the target LSTM neural network target , initialize network hyperparameters such as learning rate Learning_rate, slow update parameter τ prediction , Batch size Batch_size, each time step length is 10 milliseconds, a sliding window is used to sample the server status of the latest W time windows, and the normalized value is used as the neural network input;

[0074] Step 42: For Server i , real-time sampling status s i,t =<cpu,mem,bandwidth> ;

[0075] Step 43: For the real-time status of the sample i,t , normalized using the Z-Score standardization method, s′ i,t Store experience replay cache R DNN , retain the latest N = 10000 data;

[0076] Step 44: From R DNN Extract the first L states, input them into the main LSTM neural network model, and predict the state of the next time step Pass the state into deep reinforcement learning as the initial state s1 for action selection;

[0077] Step 45: Sample real-time status i,t+Δt , the neural network loss L is calculated based on the following formula:

[0078]

[0079] Step 46: When the update time step or loss threshold L is reached Max When, from R DNN Randomly extract Batch_size samples from the target LSTM neural network and calculate the average loss L Batch , calculate the new network parameters θ', use the following formula to slowly update the network parameters and synchronize them to the main neural network and target neural network to avoid drastic changes in network parameters caused by short-term drastic fluctuations;

[0080] θ'=τ prediction θ+(1-τ prediction )θ',τ∈(0,1)

[0081] Step 5: Scheduling decision of reinforcement learning. Initialize the basic parameters of deep reinforcement learning, use the state obtained by sliding window sampling as the current state, perform action selection for t time steps after each training round initialization, and convert the four-tuple (s t ,a t ,r t ,st+1 ) Store the value experience replay cache, randomly take samples from the cache to train and update the Actor and Critic networks, and repeat the above process until the maximum number of iterations or the termination condition is reached;

[0082] like Figure 4 As shown, the scheduling solution decision process based on reinforcement learning in step 5 of the present invention can be further described as follows:

[0083] Step 51: Randomly initialize the Actor network μ(s|θ μ ) and Critic network Q(s|θ Q ) parameter θ μ and θ Q , create the corresponding Actor network, Critic network and the corresponding target network;

[0084] Step 52: Randomly initialize an experience replay buffer R with a capacity of N to store action decision samples, i.e., a four-tuple (s t ,a t ,r t ,s t+1 );

[0085] Step 53: Initialize reinforcement learning hyperparameters, including discount factor γ, batch size B, and learning rate r;

[0086] Step 54: For each time step, get the initialized environment state s t , including workflow task status information and each server status information obtained by sliding window sampling;

[0087] Step 55: Actor network according to μ(s|θ μ ) network and the current time step state s t Select action a t , after execution state s t becomes s t+1 , get reward r t , the quaternion (s t ,a t ,r t ,s t+1 ) is stored in the experience replay buffer R;

[0088] Step 56: Randomly extract B samples from the cache R for reward evaluation and updating the Actor and Critic networks;

[0089] Step 57: Critic network according to Q(s|θ Q ) The network calculates the target Q value and uses the following formula to calculate the loss function, and the new parameter θ is obtained by gradient descent Q’ :

[0090]

[0091] Step 58: The Actor network optimizes the function by maximizing the Q value of the Critic network, calculates the optimization target according to the following formula, and obtains the new parameter θ by gradient ascent μ’ :

[0092]

[0093] Step 59: Update the parameters of the target network according to the following formula:

[0094] θ Q’ =τθ Q +(1-τ)θ Q’ ,τ∈(0,1)

[0095] θ μ’ =τθ μ +(1-τ)θ μ’ ,τ∈(0,1)

[0096] Step 510: Repeat the above process until the cutoff condition is reached;

[0097] Step 6: Use the output of step 5 as the final scheduling plan for the workflow task submitted by the user on the current cloud execution environment. Execute the task according to this scheduling plan and return the task execution result to the user.

[0098] Based on the same inventive concept, the present invention provides a computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the computer system implements the steps of real-time scheduling of workflow tasks in a cloud-edge collaborative environment based on deep reinforcement learning.

[0099] Based on the same inventive concept, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of real-time scheduling of workflow tasks in a cloud-edge collaborative environment based on deep reinforcement learning.

[0100] Those skilled in the art can understand that the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes several instructions to make a computer system (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in the embodiments of the present application. The storage medium includes: a U disk, a mobile hard disk, a read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk, and various storage media that can store computer programs.

Claims

1. A real-time scheduling method and system for cloud-edge collaborative workflow based on deep reinforcement learning, characterized in that: The steps include: Step 1: Extract workflow task attributes. Based on the workflow tasks uploaded by the user, the association relationship and DAG structure between the subtasks of the workflow task set are constructed based on the graph adjacency matrix, and the attributes of each subtask in the workflow task set are read; Step 2: Extract the cloud-edge server status attributes. Based on the rented cloud-edge collaborative computing environment, communicate with the service environment's automated monitoring system to obtain the server's current status attributes. Step 3: Task-Server Adaptability Awareness. Based on the subtask attributes of the workflow task set obtained in step 1, the numerical features are normalized using the Z-Score normalization method. The task association relationships are decomposed using the Laplace matrix. Based on the above two parts of workflow characteristics, the K-Means++ algorithm is used to perform n = 4 task clustering. The resulting clusters are divided into cloud environments or edge environments based on the minimum total execution time under the budget constraint. Step 4: Use LSTM neural network to predict server status online. Use an input sequence window with a capacity of W, input a real-time updated LSTM neural network model, and predict the status of each server at the next time step. Each sequence is stored in an experience replay buffer R DNN , using a dynamic real-time update strategy to asynchronously update the network model to achieve more accurate predictions; Step 5: Scheduling decision of reinforcement learning. Initialize the basic parameters of deep reinforcement learning, use the state obtained by sliding window sampling as the current state, perform action selection for t time steps after each training round initialization, and convert the four-tuple (s t ,a t ,r t ,s t+1 ) Store the value experience replay cache, randomly take samples from the cache to train and update the Actor and Critic networks, and repeat the above process until the maximum number of iterations or the termination condition is reached; Step 6: Use the output of step 5 as the final scheduling plan for the workflow tasks submitted by the user in the current cloud execution environment; the computer software and system execute the tasks according to this scheduling plan and return the task execution results to the user.

2. The method and system for real-time scheduling of cloud-edge collaborative workflows based on deep reinforcement learning according to claim 1 is characterized by: The attributes of the subtasks of the workflow task set in step 1 include: the number of instructions of the task, the estimated execution time of the task, and the estimated data output volume of the task.

3. The method and system for real-time scheduling of cloud-edge collaborative workflows based on deep reinforcement learning according to claim 1 is characterized by: The server attributes in step 1 include: the online status of the cloud edge server, the number of CPU cores of the server, the CPU frequency, memory size, bandwidth size, and response delay.

4. The method and system for real-time scheduling of cloud-edge collaborative workflows based on deep reinforcement learning according to claim 1 is characterized by: The step 3 further comprises: Step 31: For the task attribute values, perform data normalization using the following Z-Score normalization formula: Step 32: For the task association relationship, use the following formula to calculate the normalized Laplacian matrix and perform eigendecomposition as the low-dimensional embedding; L=D-A sym Step 33: For the processed data above, use the K-Means++ clustering method to perform clustering. Calculate the distance between the sample point and the cluster center according to the following formula, and select the nearest cluster center for each sample point: Step 34: Based on the state parameters of the cloud edge server, the characteristic values ​​of the server resources are calculated, and the optimal execution environment of each subtask of the workflow under the cost constraint is selected in cluster units.

5. The method and system for real-time workflow scheduling in a cloud-edge collaboration scenario based on deep reinforcement learning according to claim 1 is characterized by: The step 4 further comprises: Step 41: Initialize the parameters of the main LSTM neural network and the target LSTM neural network; Step 42: Each time step is 10 milliseconds long. Use a sliding window to sample the server status of the most recent W time windows and normalize them as neural network input. Step 43: For Server i , real-time sampling status s i,t =<cpu,mem,bandwidth> ; Step 44: For the real-time status of the sample i,t , normalized and stored in the experience replay cache R DNN , retain the latest N = 10000 data; Step 45: From R DNN Extract the most recent L states, input them into the main LSTM neural network model, and predict the state of the next time step Pass the state into deep reinforcement learning as the initial state s1 for action selection; Step 46: Sample real-time status i,t+Δt , the neural network loss L is calculated based on the following formula: Step 47: When the update time step or loss threshold L is reached Max When, from R DNN Randomly extract samples and calculate the average network loss L Batch And use the following formula to slowly update the network parameters; θ′=τ prediction θ+(1-τ prediction )θ′,τ∈(0,1) 6. The method and system for real-time workflow scheduling in a cloud-edge collaboration scenario based on deep reinforcement learning according to claim 1 is characterized by: The step 5 further comprises: Step 51: Randomly initialize the DDPG algorithm hyperparameters and network parameters, and randomly initialize an experience replay buffer R with a capacity of N; Step 52: For each time step, initialize the server and task state s1 according to online deep learning; Step 53: The Actor network is based on the current time step state s t Select action a t , the state changes to s t+1 , get reward t , the quaternion (s t ,a t ,r t ,s t+1 ) is stored in the experience replay buffer R; Step 54: Randomly extract B samples from the cache R, evaluate and update the Actor and Critic networks; Step 55: The critic network calculates the target Q value and uses the following formula to calculate the loss function, and obtains the new parameter θ by gradient descent. Q′ : Step 56: The Actor network optimizes the function by maximizing the Q value of the Critic network, calculates the optimization target according to the following formula, and obtains the new parameter θ by gradient ascent μ′ : Step 57: Update the parameters of the target network according to the following formula: i Q′ =tθ Q +(1-τ)θ Q′ ,τ∈(0,1) i μ′ =tθ μ +(1-τ)θ μ′ ,τ∈(0,1) 7. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is loaded into the processor, it implements the real-time workflow scheduling method in the cloud-edge collaboration scenario based on deep reinforcement learning as described in any one of claims 1-6.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it implements the real-time workflow scheduling method in the cloud-edge collaboration scenario based on deep reinforcement learning according to any one of claims 1-6.