End-side cloud collaborative configuration and scheduling method for video analysis task
By building an end-edge cloud collaborative architecture in video analysis tasks and introducing a reinforced learning network of the actor-critician system, optimizing video stream configuration and task scheduling, the problems of resource waste and unoptimized task results in the existing technology are solved, and efficient resource utilization and performance improvement are achieved.
Patent Information
- Application Number
- CN202510116051.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing video analysis tasks have problems such as waste of resources and unoptimized task results in cloud computing resource scheduling, especially in large-scale task deployment and task demand adjustment.
A method of collaborative configuration and scheduling for video analysis tasks is proposed. By building a video analysis architecture of end-edge cloud collaboration and introducing an actor-criticist system, the video stream configuration and task scheduling are optimized using reinforcement learning network to achieve efficient resource utilization and performance improvement.
It realizes efficient utilization of resources in a multi-device environment, improves the scalability and robustness of the system, significantly reduces energy consumption, and ensures that the performance of video analysis tasks can meet the task requirements.
Smart Images

Figure CN119966991A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to configuration adjustment and resource scheduling technology of the Internet application layer, specifically to an end-edge-cloud collaborative configuration and scheduling method for video analysis tasks. Background Art
[0002] With the development of the Internet of Things (IoT), a large number of surveillance cameras have been deployed in different fields, including security monitoring, crime prevention, smart business, etc. It is reported that by 2023, the number of surveillance cameras has reached 25 million, and the global surveillance camera market size is expected to reach US$45.93 billion by 2027. These cameras generate a large number of video streams and need to be analyzed in a timely manner. Requests to analyze video streams for specific purposes are usually called video analysis tasks. This type of task usually uses deep neural networks to process video streams, which has a large demand for computing resources. However, many terminal devices lack sufficient computing resources to run deep neural networks in real time. Therefore, it is a common solution to transfer video analysis tasks to edge computing or cloud computing servers, which makes video analysis tasks a "killer application" of cloud computing.
[0003] Existing cloud computing resources have a tendency to develop towards sky computing. By abstracting the heterogeneity of cloud computing resources and providing a common interface for computing resource calls to ensure that applications can be performed across clouds, the flexibility, scalability and disaster recovery capabilities of applications can be effectively improved. To achieve this goal, it is crucial to effectively schedule resources within the network. This not only involves solving the heterogeneity of computing resources, but also involves the diverse network environment of servers. Task performance preferences also vary depending on the task scenario and the preferences of the user who issues the task. For example, some tasks prefer high accuracy, while others may prefer lower processing latency rather than accuracy. Meeting these different task preferences is crucial to improving the quality of experience (QoE).
[0004] In addition, in video analysis tasks, video streams can be encoded and uploaded using various parameters (such as resolution, frame rate, etc.), and tasks can also be divided into multiple steps and scheduled to run on different computing nodes. Changing these configurations and scheduling strategies can lead to significantly different performance and affect the demand for computing and network resources. Therefore, it is necessary to jointly optimize the video stream configuration and task scheduling strategy so that the performance goals of each task are met and the use of each edge or cloud computing resource does not exceed its capacity.
[0005] Existing cloud computing solutions for video analysis dynamically adjust video configuration parameters by analyzing the performance requirements of different tasks (including key indicators such as accuracy, latency, and cost), effectively reducing the amount of data that needs to be transmitted and processed, thereby avoiding resource waste; at the same time, there are also research works dedicated to scheduling computing devices distributed in different geographical locations to meet the huge demand for computing resources for video analysis tasks, thereby achieving efficient resource utilization. However, the former work only considers the simple case of a single edge and cloud node, which is not suitable for large-scale task deployment; the latter does not consider adjusting the video configuration according to task requirements, which may result in the task result not being the optimal solution, resulting in a certain degree of resource waste. Summary of the invention
[0006] The present invention proposes an end-edge-cloud collaborative configuration and scheduling method for video analysis tasks, which performs configuration and scheduling optimization on the premise that application performance meets task requirements, so as to achieve the goals of fully utilizing device resources and improving QoE.
[0007] The device-edge-cloud collaborative configuration and scheduling method for video analysis tasks is implemented by following the steps below:
[0008] Step 1: Build a video analysis architecture that collaborates with the device, edge, and cloud, including device, edge, and cloud devices.
[0009] The end device is responsible for generating video streams according to the configuration, generating video analysis tasks at the same time, and submitting the video streams and video analysis tasks to the edge device; the edge device decides the configuration of the current video stream, schedules the tasks through the scheduler, processes part of the task according to the scheduling results, and submits the remaining task content to the cloud device; the video analysis model of the cloud device generates analysis results, and transmits the analysis results to the analyzer to evaluate the task execution, and feeds back the evaluation results to the scheduler of the edge device, and the scheduler optimizes the task scheduling decision.
[0010] Step 2: Introduce the actor-critic system into the edge-cloud collaborative video analysis architecture to obtain an edge-cloud collaborative scheduler based on a reinforcement learning network.
[0011] The actor-critic system is specifically as follows: the actor model is deployed on the edge node and acts as a scheduler to determine the configuration and scheduling strategy of the video stream in the next time slot in each time slot; the critic model is deployed on any node in the cloud during the training phase and acts as an analyzer to evaluate the task execution of each node at each moment, and feeds the evaluation results back to the actor model on each edge node for training.
[0012] The actor-critic neural network structure uses LSTM modules to capture the temporal information of historical decisions as feature vectors, and concatenates other dimensional vectors in the observation space, the number of tasks submitted to the edge nodes, and the preference of each task into the feature vector.
[0013] Step three, establish an optimization goal for the edge-cloud collaborative scheduler based on reinforcement learning network to maximize the accuracy a while minimizing the latency l and processing overhead o for executing each task i.
[0014] Assume there are M edge nodes, each of which manages n i There are N tasks in the system, and the running time of the whole system is T. In each time slot t, it is necessary to determine the resolution r, frame rate f, number of steps s to run on the edge node, cloud node c for processing the task, and allocated bandwidth b of each task video stream. The optimization objective is expressed as:
[0015]
[0016]
[0017] Among them, the subscript i,t indicates that the parameter belongs to task i in the tth time slot, the superscript E indicates an edge node, and the superscript C indicates a cloud node; represents the end-to-edge bandwidth, represents the edge-cloud bandwidth;
[0018] Constraints C1 to C4 ensure that the resolution, frame rate, and number of steps running on the edge nodes and the cloud nodes used to process the task are within the acceptable range. 10 This ensures that the device resources used by each node are within the limit.
[0019] The GPU memory h and memory m occupied by task i on edge nodes and cloud nodes are related to its resolution r, frame rate f and the number of steps s running on the edge node, expressed as:
[0020]
[0021] Step 4: Solve the optimization problem and decompose it into the decision on discrete quantity and the decision on bandwidth b.
[0022] (1) Decision-making on discrete quantities
[0023] The decision of discrete quantity makes the utility of each task the highest by determining the video stream configuration and task scheduling of each task. The formula is expressed as:
[0024]
[0025] The discrete quantity decision problem is modeled as a partially observable Markov decision process POMDP (partially observable Markov Decision Process) as follows:
[0026] Bandwidth, GPU memory, memory and GPU utilization are represented as B, H, M and G respectively;
[0027] State space: The optimal scheduling decision for each edge node is determined by the resource capacity of each computing node, the preference p of each task, and the scheduling strategies σ of other edge nodes. Therefore, the state space is defined as S' = {S n},in
[0028] S n ={(B E ,H E ,M E ,G E ,B C ,H C ,M C ,G C ,p,σ)}.
[0029] Observation space: Since the scheduling strategies of other edge nodes are unobservable, we can only observe the scheduling decision x and the corresponding reward r made by the edge node at the last moment to infer its scheduling strategy. Therefore, the observation space is defined as O = {O n}, where O n ={(B E ,H E ,M E ,G E ,B C ,H C ,M C ,G C ,p,x t-1 ,r t-1 )}.
[0030] Observation function: Y is the function S×O→[0,1]. Function Y=P(x k,t ,R k,t |σ k ) represents the condition σ in task k k Make a decision x k,t And get reward R k,t When all edge nodes make scheduling decisions at time t, the observed value will be transformed to satisfy o n,t+1 =∑s n Y(·|s n )P(s n |x t ).
[0031] Decision space: The decision space is the configuration of the video stream and the scheduling of tasks, that is, A = {x n}={(r n ,f n ,s n ,c n )}.
[0032] Reward space: If the task in time slot t is successfully completed, the reward is the same as in the problem statement. If some scheduling decisions fail because the resource allocation of a cloud node exceeds the capacity, that is, C7, C8, C9, C 10 If any condition in , a negative constant -C is set for the reward. Therefore, the reward space is expressed as R = {p i,1 a i -p i,2 l i -p i,3 o i}∪{-C}.
[0033] Applying the above modeling process to the actor-critic system in step 2 and letting the actor model make decisions in each time slot is the solution method for discrete quantities.
[0034] (2) Decision on bandwidth b
[0035] The decision for the continuous quantity, i.e., bandwidth b, is to minimize the delay by determining the bandwidth allocation for each task, which can be expressed as follows:
[0036]
[0037] The optimal bandwidth allocation for each task on the corresponding edge node is obtained by the Cauchy inequality. ω(s) is used to represent the ratio of the data size of the video stream after s steps of processing on the edge node to the original data size. The obtained bandwidth allocation for task i on edge node e is i The optimal bandwidth allocation on is as follows:
[0038]
[0039] Step 5: Based on the discrete quantity decision and bandwidth decision obtained by solving the optimization problem, the delay and overhead of task i in time slot t are calculated;
[0040] (1) The delay of task i in time slot t is expressed as:
[0041]
[0042] Among them, the end-to-edge transmission delay is the coding coefficient; edge-to-cloud transmission delay is the coding coefficient; represents the computational delay, ρ(s i,t ) represents the proportion of data processed on edge devices.
[0043] (2) The overhead of task i in time slot t is expressed as:
[0044]
[0045] in, represents the transmission unit price from edge node i to cloud node j, Indicates the rental unit price of the edge node, Indicates the unit price of renting a cloud node.
[0046] Step 6: The scheduler executes the corresponding video analysis tasks according to the corresponding steps based on the discrete decision quantity. After the current time slot step is completed, the remaining tasks are uploaded to the cloud node;
[0047] Step 7: The cloud node processes the remaining tasks and uses the task processing status to train the actor-critic system;
[0048] The training process of the actor-critic system is:
[0049] The critic model uses the average of all task feature vectors on each node as the input of the fully connected layer to generate the estimated value required to calculate the evaluation value, including modeling factors of overhead and delay, task accuracy, and task success rate. The corresponding delay and cost are calculated through modeling factors, and the estimated reward is calculated in combination with the accuracy, and the completion rate is multiplied by the estimated reward to obtain the corresponding evaluation value.
[0050] The actor model uses a fully connected layer to take each task feature vector as input and output a probability map for each dimension in the decision space.
[0051] The evaluation value obtained by the critic model is used to optimize the actor model and update the parameters of the actor model. At the same time, the execution results of the task after the actor model takes action and the actual delay and overhead obtained in step 5 will be used to optimize the critic model so that the evaluation value it generates can be closer to the actual benefit.
[0052] Step 8. After each iteration of the actor-critic system training process, the scheduler reconfigures the task generator, returns to step 3, and repeats the above steps until the actor-critic system converges.
[0053] The advantages and beneficial effects of the present invention are:
[0054] (1) Based on the actor-critic deep reinforcement learning architecture, this paper proposes a distributed task configuration and scheduling optimization algorithm, which achieves efficient resource utilization in a multi-device environment and improves the overall performance while ensuring the scalability of the entire system.
[0055] (2) The method of the present invention improves the scalability and robustness of the system, significantly reduces energy consumption, and ensures that the performance of the video analysis task can meet the requirements of the task. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a video analysis architecture model that combines end-edge and cloud collaboration;
[0057] Figure 2 This is a video analysis flow chart of device-edge-cloud collaboration;
[0058] Figure 3 This is a schematic diagram of the overall architecture of the actor-critic system;
[0059] Figure 4 It is a structural diagram of the reinforcement learning network. DETAILED DESCRIPTION
[0060] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0061] like Figure 1 As shown, the end-edge-cloud collaborative architecture for video analysis tasks analyzes different types of surveillance videos, mainly target detection, target tracking, and traffic analysis, etc. The proposed system has a three-layer structure of client, edge device, and cloud device. Video analysis tasks are mainly calculated through deep neural network models that require a large amount of computing and storage resources, and it is difficult to deploy them independently to resource-constrained edge servers for execution; and due to the huge amount of data in the video stream, transmitting all videos to the cloud platform for analysis will bring unnecessary transmission and computing costs. Therefore, the present invention divides the model used for the video analysis task, so that the edge device can offload subsequent steps to the cloud device after partially processing the data, so that the resources of the entire system can be flexibly allocated.
[0062] Cloud nodes will have most of the computing resources of the network, and these cloud nodes are built in different regions and have different rental fees. Therefore, the network environment and usage costs of different nodes are also different, which brings corresponding trade-off space for tasks with different preferences. For example, some tasks are more sensitive to system response delay, so the system tends to submit the task to a cloud node that is physically closer to obtain lower transmission delay; while some tasks prefer to save costs, so submitting the task to a node that is physically farther away but has lower costs is a better choice.
[0063] Making decisions on the cloud nodes called by the system based on task preference can effectively improve system efficiency and reduce costs. In the actual operation of the system, a more important issue is to make reasonable task allocation based on the resource capacity and utilization rate of each device to avoid problems such as the task load of some nodes exceeding the device resource capacity when there are too many tasks, and some devices being idle, resulting in resource waste.
[0064] The end-edge-cloud collaborative configuration and scheduling method for video analysis tasks proposed in the present invention has the following specific steps:
[0065] Step 1: Build a video analysis architecture that collaborates with the device, edge, and cloud, including device, edge, and cloud devices.
[0066] like Figure 2 As shown in the figure, the client device generates a video stream through the camera and submits the video analysis task to the nearest edge node; a scheduler is deployed at the edge node, which will adjust the video stream configuration on the terminal side in real time according to the resource situation and task requirements, determine the task steps to be performed on the edge device, and submit subsequent tasks to the cloud node; after the edge node completes its own task, it will upload the remaining data to the cloud node, and the video analysis model of the cloud node will generate the analysis result, and evaluate the execution of this task through the analyzer, and finally feed back the result to the scheduler of the edge device, and the scheduler will optimize its decision based on this result.
[0067] Step 2: Introduce the actor-critic system into the edge-cloud collaborative video analysis architecture to obtain an edge-cloud collaborative scheduler based on a reinforcement learning network.
[0068] The overall architecture of the actor-critic system is as follows Figure 3 As shown in the figure, specifically: the actor model is deployed on the edge node and acts as a scheduler to determine the configuration and scheduling strategy of the video stream in the next time slot in each time slot; the critic model is deployed in the analyzer module on any node in the cloud during the training phase, and acts as an analyzer to evaluate the task execution status of each node at each moment, and feeds the evaluation results back to the actor model on each edge node for training.
[0069] The replay buffer in the analyzer records resource utilization, task execution results, and actions taken by the actor model, which are used to optimize the critic model and generate corresponding evaluation values in each time slot for training the actor model. The present invention uses centralized critic and analyzer modules because the centralized critic module can capture information from all edge nodes in the network and make the training process easier to converge. The centralized analyzer and critic modules are only used in the training phase, so the system can work completely distributed in the inference phase, which does not conflict with the goal of designing a distributed scheduling system in the present invention.
[0070] The actor-critic neural network structure uses LSTM modules to capture the temporal information of historical decisions as feature vectors, and concatenates other dimensional vectors in the observation space, the number of tasks submitted to the edge nodes, and the preference of each task into the feature vector.
[0071] The specific reinforcement learning network structure is as follows Figure 4 As shown, since the observation space includes the historical strategies and rewards of all schedulers, it is very important to use the time information in the history. Therefore, the present invention uses LSTM modules to capture time information as feature vectors and embeds other dimensions in the observation space (the number of tasks submitted to the scheduler and the preference of each task) into the feature vectors. For the actor model, a fully connected layer is used to take the feature vector as input and output the probability map of each dimension in the decision space. The critic model uses embedding and LSTM modules similar to the actor model to process task information and generate feature vectors, and uses the average value of these vectors as the input of the fully connected layer to generate the estimated values required to calculate the evaluation value, including accuracy, success rate and other modeling factors, and calculates the corresponding delay and cost through the modeling formula of delay and overhead, and the accuracy will be used to calculate the estimated reward according to the method defined in the reward space. Finally, the success rate will be multiplied by the estimated reward to obtain the corresponding evaluation value.
[0072] Step three, establish an optimization goal for the edge-cloud collaborative scheduler based on reinforcement learning network to maximize the accuracy a while minimizing the latency l and processing overhead o for executing each task i.
[0073] Assume there are M edge nodes, each of which manages n i There are N tasks in the system, and the running time of the whole system is T. In each time slot t, it is necessary to determine the resolution r of each task video stream, the frame rate f, the number of steps s running on the edge node, the cloud node c used to process the task, and the allocated bandwidth b. The goal of the optimization system of the present invention is to maximize the accuracy a while minimizing the delay l and the processing overhead o. Each task has different preferences for these three factors, so the weighted linear combination of these three factors can be used to express the benefits of each task. The optimization goal is expressed as:
[0074]
[0075]
[0076]
[0077] Cloud Node; represents the end-to-edge bandwidth, represents the edge-cloud bandwidth;
[0078] Constraints C1 to C4 ensure that the resolution, frame rate, and number of steps running on the edge nodes and the cloud nodes used to process the task are within the acceptable range. 10 This ensures that the device resources used by each node are within the limit.
[0079] The video analysis task of the present invention is executed on the GPU, so the main resources considered in the present invention include GPU utilization g, GPU video memory h, memory m and network bandwidth b. Among these resources, the use of GPU video memory h and memory m is determined by the video configuration and scheduling decision. If the resources are insufficient, the task will fail. The usage of these resources can be determined in the training phase. The GPU video memory h and memory m occupied by task i on the edge node and cloud node are related to its resolution r, frame rate f and the number of steps s running on the edge node, which can be expressed as:
[0080]
[0081] As for the utilization of GPU cores, since GPU vendors do not provide APIs for fine-grained allocation of GPU cores, GPU cores are usually equally allocated to each task running on the GPU. Therefore, the utilization of the GPU can be modeled using the following formula, where c and e represent the steps required for the task on the edge node and cloud node, G represents the GPU capacity of the node, and Σ j [e j =e i ] and ∑ j [c j,t =c i,t ] indicates the number of tasks assigned to the node:
[0082]
[0083]
[0084] Step 4: Solve the optimization problem and decompose it into the decision on discrete quantity and the decision on bandwidth b.
[0085] (1) Decision-making on discrete quantities
[0086] The decision of discrete quantity makes the utility of each task the highest by determining the video stream configuration and task scheduling of each task. The formula is expressed as:
[0087]
[0088] For discrete quantity decision, since it is an integer programming problem with nonlinear objectives, it is NP-hard; at the same time, each decision maker can only collect partial information due to the limitations of the environment; in addition, the variable available resources and task preferences also make the situation more complicated. Therefore, the discrete quantity decision problem is modeled as a partially observable Markov decision process POMDP (partially observable Markov Decision Process). POMDP is a better modeling scheme because it allows decisions to be made under uncertainty and incomplete information. To this end, the present invention proposes an algorithm based on reinforcement learning to solve this problem. The process of the POMDP is described as follows:
[0089] Bandwidth, GPU memory, memory and GPU utilization are represented as B, H, M and G respectively;
[0090] State space: The optimal scheduling decision for each edge node is determined by the resource capacity of each computing node, the preference p of each task, and the scheduling strategies σ of other edge nodes. Therefore, the state space is defined as S' = {S n},in
[0091] S n ={(B E ,H E ,M E ,G E ,B C ,H C ,M C ,G C ,p,σ)}.
[0092] Observation space: Since the scheduling strategies of other edge nodes are unobservable, we can only observe the scheduling decision x and the corresponding reward r made by the edge node at the last moment to infer its scheduling strategy. Therefore, the observation space is defined as O = {O n}, where O n ={(B E ,H E ,M E ,G E ,B C ,H C ,M C ,G C ,p,x t-1 ,r t-1 )}.
[0093] Observation function: Y is the function S×O→[0,1]. Function Y=P(x k,t ,R k,t |σ k ) represents the condition σ in task kk Make a decision x k,t And get reward R k,t When all edge nodes make scheduling decisions at time t, the observed value will be transformed to satisfy o n,t+1 =∑s n Y(·|s n )P(s n |x t ).
[0094] Decision space: The decision space is the configuration of the video stream and the scheduling of tasks, that is, A = {x n}={(r n ,f n ,s n ,c n )}.
[0095] Reward space: If the task in time slot t is successfully completed, the reward is the same as in the problem statement. If some scheduling decisions fail because the resource allocation of a cloud node exceeds the capacity, that is, C7, C8, C9, C 10 If any condition in , a negative constant -C is set for the reward. Therefore, the reward space is expressed as R = {p i,1 a i -p i,2 l i -p i,3 o i}∪{-C}.
[0096] Applying the above modeling process to the actor-critic system in step 2 and letting the actor model make decisions in each time slot is the solution method for discrete quantities.
[0097] (2) Decision on bandwidth b
[0098] The decision for the continuous quantity, i.e., bandwidth b, is to minimize the delay by determining the bandwidth allocation for each task, which can be expressed as follows:
[0099]
[0100] The optimal bandwidth allocation for each task on the corresponding edge node is obtained by the Cauchy inequality. ω(s) is used to represent the ratio of the data size of the video stream after s steps of processing on the edge node to the original data size. The obtained bandwidth allocation for task i on edge node e is i The optimal bandwidth allocation on is as follows:
[0101]
[0102] Step 5: Based on the discrete quantity decision and bandwidth decision obtained by solving the optimization problem, the delay and overhead of task i in time slot t are calculated;
[0103] (1) The delay of task i in time slot t is expressed as:
[0104]
[0105] The allocation of bandwidth affects the transmission delay, which in turn affects the overall performance of the task. Therefore, the present invention models the transmission delay as a function determined by the bandwidth and the amount of data, and further depends on the resolution r and frame rate f of the video stream. According to previous studies, r 2 f models the data size. The amount of data that needs to be transmitted between the end and the edge also depends on the video encoding algorithm used on the end side, for which a coding coefficient C is required enc To express the ratio of the amount of data after encoding to the amount of data before encoding. Therefore, the end-to-edge transmission delay can be expressed using the following formula:
[0106]
[0107] The amount of data that needs to be transmitted between the edge and the cloud depends on the data size of the feature map obtained after the edge device processes the original data in s steps. A coefficient can also be used To express the ratio of the amount of data after processing to that before processing. Therefore, the edge-cloud transmission delay can be expressed by the following formula:
[0108]
[0109] The delay of each video clip includes transmission delay, propagation delay and computation delay. To represent the propagation delay between edge nodes and cloud nodes; ρ(s i,t ) represents the proportion of data processed on edge devices, which can be used To represent the computation delay.
[0110] (2) The overhead of each video segment consists of transmission overhead and computational overhead. Both types of overhead are linearly related to the amount of data. The overhead of task i in time slot t is expressed as:
[0111]
[0112] in, represents the transmission unit price from edge node i to cloud node j, Indicates the rental unit price of the edge node, Indicates the unit price of renting a cloud node.
[0113] Step 6: The scheduler executes the corresponding video analysis tasks according to the corresponding steps based on the discrete decision quantity. After the current time slot step is completed, the remaining tasks are uploaded to the cloud node;
[0114] Step 7: The cloud node processes the remaining tasks and uses the task processing status to train the actor-critic system;
[0115] The training process of the actor-critic system is:
[0116] The critic model uses the average of all task feature vectors on each node as the input of the fully connected layer to generate the estimated value required to calculate the evaluation value, including modeling factors of overhead and delay, task accuracy, and task success rate. The corresponding delay and cost are calculated through modeling factors, and the estimated reward is calculated in combination with the accuracy, and the completion rate is multiplied by the estimated reward to obtain the corresponding evaluation value.
[0117] The actor model uses a fully connected layer to take each task feature vector as input and output a probability map for each dimension in the decision space.
[0118] The evaluation value obtained by the critic model is used to optimize the actor model and update the parameters of the actor model. At the same time, the execution results of the task after the actor model takes action and the actual delay and overhead obtained in step 5 will be used to optimize the critic model so that the evaluation value it generates can be closer to the actual benefit.
[0119] Step 8. After each iteration of the actor-critic system training process, the scheduler reconfigures the task generator, returns to step 3, and repeats the above steps until the actor-critic system converges.
Claims
1. A device-edge-cloud collaborative configuration and scheduling method for video analysis tasks, characterized in that: The steps include: Step 1: Build a device-edge-cloud collaborative video analysis architecture, including device, edge device, and cloud device, and introduce the actor-critic system into it to obtain a device-edge-cloud collaborative scheduler based on reinforcement learning network; The actor-critic system is specifically as follows: the actor model is deployed on the edge node and acts as a scheduler to determine the configuration and scheduling strategy of the video stream in the next time slot in each time slot; the critic model is deployed on any node in the cloud during the training phase and acts as an analyzer to evaluate the task execution of each node at each moment, and feeds the evaluation results back to the actor model on each edge node for training; The actor-critic neural network structure uses LSTM modules to capture the temporal information of historical decisions as feature vectors, and concatenates other dimensional vectors in the observation space, the number of tasks submitted to the edge nodes, and the preference of each task into the feature vector; Step 2: Establish an optimization goal for the edge-cloud collaborative scheduler based on the reinforcement learning network to maximize the accuracy a while minimizing the latency l and processing overhead o for executing each task i; Assume there are M edge nodes, each of which manages n i There are N tasks in the system, and the running time of the whole system is T. In each time slot t, it is necessary to determine the resolution r, frame rate f, number of steps s to run on the edge node, cloud node c for processing the task, and allocated bandwidth b of each task video stream. The optimization objective is expressed as: Among them, the subscript i,t indicates that the parameter belongs to task i in the tth time slot, the superscript E indicates an edge node, and the superscript C indicates a cloud node; represents the end-to-edge bandwidth, represents the edge-cloud bandwidth; Constraints C1 to C4 ensure that the resolution, frame rate, and number of steps running on the edge nodes and the cloud nodes used to process the task are within the acceptable range. 10 This ensures that the device resources used by each node are within limits; The GPU memory h and memory m occupied by task i on edge nodes and cloud nodes are related to its resolution r, frame rate f and the number of steps s running on the edge node, expressed as: Step 3, solving the optimization problem, decomposing the optimization problem into a decision on discrete quantity and a decision on bandwidth; (1) Decision-making on discrete quantities The decision of discrete quantity makes the utility of each task the highest by determining the video stream configuration and task scheduling of each task. The formula is expressed as: s.t.C1,C2,C3,C4,C7,C8,C9,C 10 The discrete quantity decision problem is modeled as a partially observable Markov decision process POMDP, and the modeling process is actually applied to the actor-critic system, allowing the actor model to make a decision in each time slot, which is the solution method for discrete quantities. (2) Decision on bandwidth The decision for the continuous quantity, i.e., bandwidth b, is to minimize the delay by determining the bandwidth allocation for each task, which can be expressed as follows: stC5,C6 The optimal bandwidth allocation of each task on the corresponding edge node is obtained through the Cauchy inequality; Step 4: Based on the discrete quantity decision and bandwidth decision obtained by solving the optimization problem, the delay and overhead of task i in time slot t are calculated; (1) The delay of task i in time slot t is expressed as: Among them, the transmission delay between end and edge is the coding coefficient; edge-to-cloud transmission delay is the coding coefficient; represents the computational delay, ρ(s i,t ) represents the proportion of data processed on edge devices; (2) The overhead of task i in time slot t is expressed as: in, represents the transmission unit price from edge node i to cloud node j, Indicates the rental unit price of the edge node, Indicates the unit price of renting a cloud node; Step 5: The scheduler executes the corresponding video analysis tasks according to the corresponding steps based on the discrete decision quantity. After the current time slot step is completed, the remaining tasks are uploaded to the cloud node; Step 6: The cloud node processes the remaining tasks and uses the task processing status to train the actor-critic system; The training process of the actor-critic system is: The critic model uses the average of all task feature vectors on each node as the input of the fully connected layer to generate the estimated value required to calculate the evaluation value, including modeling factors of overhead and delay, task accuracy, and task success rate; the corresponding delay and cost are calculated through modeling factors, and the estimated reward is calculated in combination with the accuracy, and the completion rate is multiplied by the estimated reward to obtain the corresponding evaluation value; The actor model uses a fully connected layer to take each task feature vector as input and output a probability map for each dimension in the decision space; The evaluation value obtained by the critic model is used to optimize the actor model and update the parameters of the actor model. At the same time, the execution results of the task after the actor model takes action, as well as the actual delay and cost, will be used to optimize the critic model so that the evaluation value generated by it can be closer to the actual benefit. Step 7: After each iteration of the actor-critic system training process, the scheduler reconfigures the task generator, returns to step 2, and repeats the above steps until the actor-critic system converges.
2. The method for end-edge-cloud collaborative configuration and scheduling for video analysis tasks according to claim 1 is characterized in that: The end-edge-cloud collaborative video analysis architecture is specifically as follows: the end device is responsible for generating video streams according to the configuration, generating video analysis tasks at the same time, and submitting the video streams and video analysis tasks to the edge device; the edge device decides the configuration of the current video stream, schedules the tasks through the scheduler, processes part of the task according to the scheduling results, and submits the remaining task content to the cloud device; the video analysis model of the cloud device generates analysis results, and transmits the analysis results to the analyzer to evaluate the task execution status, and feeds back the evaluation results to the scheduler of the edge device, and the scheduler optimizes the task scheduling decision.
3. The method for end-edge-cloud collaborative configuration and scheduling for video analysis tasks according to claim 1 is characterized in that: The modeling process of partially observable Markov decision process POMDP is: Bandwidth, GPU memory, memory and GPU utilization are represented as B, H, M and G respectively; State space: The optimal scheduling decision for each edge node is determined by the resource capacity of each computing node, the preference p of each task, and the scheduling strategies σ of other edge nodes. Therefore, the state space is defined as S' = {S n }, where S n ={(B E ,H E ,M E ,G E ,B C ,H C ,M C ,G C ,p,σ)}; Observation space: Since the scheduling strategies of other edge nodes are unobservable, we can only observe the scheduling decision x and the corresponding reward r made by the edge node at the last moment to infer its scheduling strategy; Therefore, the observation space is defined as O = {O n }, where O n ={(B E ,H E ,M E ,G E ,B C ,H C ,M C ,G C ,p,x t-1 ,r t-1 )}; Observation function: Y is the function S×O→[0,1]; function Y=P(x k,t ,R k,t |σ k ) represents the condition σ in task k k Make a decision x k,t And get reward R k,t When all edge nodes make scheduling decisions at time t, the observed value will be transformed to satisfy Decision space: The decision space is the configuration of the video stream and the scheduling of tasks, that is, A = {x n }={(r n ,f n ,s n ,c n )}; Reward space: If the task in time slot t is successfully completed, the reward is the same as in the problem statement; if some scheduling decisions fail because the resource allocation of a cloud node exceeds the capacity, that is, C7, C8, C9, C 10 If any condition in , a negative constant -C is set for the reward; therefore, the reward space is expressed as R = {p i,1 a i -p i,2 l i -p i,3 o i }∪{-C}.
4. The method for end-edge-cloud collaborative configuration and scheduling for video analysis tasks according to claim 1 is characterized in that: The process of solving the optimal bandwidth allocation by using the Cauchy inequality is: We use ω(s) to represent the ratio of the data size of the video stream after s steps of processing on the edge node to the original data size. The obtained value for task i on the edge node e is i The optimal bandwidth allocation on is as follows:
Citation Information
Patent Citations
Methods and systems for purposeful computing
CN109101217A
Multi-dimensional resource intelligent joint optimization method and system for vehicle-mounted computing power network
CN113423091A
Multi-task edge calculation scheduling optimization method based on graph attention network
CN113946423A
End-side collaborative multi-unmanned aerial vehicle autonomous navigation method
CN114061589A
Real-time video analysis and processing method based on edge cloud collaboration
CN114697324A