A task scheduling method and device for video processing on the edge side

By constructing an adaptive cuckoo algorithm based on Q-learning in an edge computing environment, using Tent chaotic mapping and Q-learning optimization parameters, the problem of high response time delay in video processing task scheduling is solved, efficient and accurate task scheduling is achieved, and user response time delay and tail delay are reduced.

CN119166303BActive Publication Date: 2025-07-04INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411225022.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-03
Publication Date
2025-07-04
Estimated Expiration
2044-09-03

AI Technical Summary

Technical Problem

In the edge computing environment, the scheduling of video processing tasks has high computational density and high resource demand, resulting in high user response delay and low solution efficiency. The existing cuckoo algorithm has high sensitivity in parameter selection, making it difficult to effectively optimize task scheduling.

Method used

The adaptive cuckoo algorithm based on Q-learning is constructed, and the population is initialized through Tent chaotic mapping, combined with the Q-learning algorithm to dynamically tune the main parameters and adaptively update the secondary parameters, optimize the task scheduling model, and reduce the user response delay and tail delay.

Benefits of technology

It improves the solution efficiency and accuracy of task scheduling, reduces the average response delay and tail delay of users, and improves user experience and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119166303B_ABST
    Figure CN119166303B_ABST
Patent Text Reader

Abstract

The present invention proposes a task scheduling method for edge-side video processing, including: constructing a task scheduling model for an edge video processing system and an optimization index for the task scheduling model; generating an objective function and constraint conditions for the task scheduling model based on the optimization index; using a cuckoo search model, taking the task scheduling scheme corresponding to the objective function when the optimization index is the minimum value as the optimization scheme, and performing task scheduling operations on the edge video processing system with the optimization scheme. In view of the problem that the response latency of users needs to be reduced when performing video processing task scheduling in the edge environment, the present invention takes the average response latency and tail latency of users as the main optimization indexes, and uses two indexes of user fairness and task priority to reflect the influence of the submission order of task requests by different users and the execution order of different tasks of the same user on the task scheduling problem respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of edge computing, and particularly relates to a method and device for scheduling video processing tasks on an edge cluster. Background Art

[0002] With the development of the new generation of information technology, a large number of application scenarios with high real-time requirements have emerged. In these scenarios, video processing tasks using artificial intelligence technology account for an important proportion. Edge computing can sink the computing and storage resources required for tasks from the cloud to the edge side, effectively reducing data transmission latency. However, video processing tasks are characterized by high computational intensity and high resource requirements. When scheduling tasks in an edge environment with relatively limited resources, problems such as poor performance of the edge cluster and high user response latency will be faced. The cuckoo algorithm is inspired by the obligate parasitic breeding behavior of the cuckoo species. This algorithm combines the extensive exploration of global search and the fine adjustment of local search, so as to more effectively find high-quality solutions in the search space. However, when the cuckoo algorithm is used for task scheduling problems, its existing disadvantages include: (1) The cuckoo algorithm usually adopts the method of randomly initializing the population in the initialization stage of the algorithm, which may cause uneven population distribution, is not conducive to exploring the solution space and introducing unstable factors, etc. (2) When the cuckoo algorithm solves practical problems, it is sensitive to parameter selection. Inappropriate parameters will lead to low algorithm solving efficiency and poor accuracy, and parameters need to be set according to specific problems. The above problems will result in low solving efficiency and poor solving accuracy of the cuckoo algorithm when applied to edge environments for video processing task scheduling, thereby resulting in high user response latency. Therefore, there is currently a need for a reasonable and efficient video processing task scheduling algorithm to improve the solving efficiency and accuracy of the algorithm while reducing the user response latency, so as to improve the service quality of actual applications and let users obtain a better experience. Summary of the Invention

[0003] In view of the above problems, the present invention proposes a task scheduling method for edge-side video processing, including: constructing a task scheduling model for an edge video processing system and an optimization index for the task scheduling model; based on the optimization index, generating an objective function and constraint conditions for the task scheduling model; using a cuckoo search model, taking the task scheduling scheme corresponding to the objective function when the optimization index is the minimum value as the optimization scheme, and performing task scheduling operations on the edge video processing system with the optimization scheme.

[0004] Further, the optimization index includes: user fairness index UF inv , task priority index TP inv , user average response latency index RT avg and tail latency index TL 99th; Among them, a user sequence is constructed in ascending order according to the time when users submit tasks, and the response time of each user is used as the value of the element at the corresponding position in the user sequence. The number of inversions of the user sequence is calculated as the user fairness index UF. inv ; A task sequence is constructed in ascending order according to the time when each task is submitted, and the execution time of each task is used as the value of the element at the corresponding position in the task sequence. The number of inversions of the task sequence is calculated and summed as the task priority index TP. inv ; User average response delay index represents the response delay of user h, and x is the total number of users submitting requests within a specified time period; Tail latency index TL 99th is the part where the proportion of the response delay higher than the average response delay is higher than 99%.

[0005] Further, the objective function TS

[0006] min TS = ω1RT avg + ω2TL 99th + ω3UF inv + ω4TP inv

[0007] satisfies: and

[0008] Among them, ω1, ω2, ω3, and ω4 are weight values, and ω1 + ω2 + ω3 + ω4 = 1. represents task t i on edge device p j the amount of computing resource used, represents t i on p j the amount of storage resource used, represents edge device p j the available computing resource on it, represents edge device p j the available storage resource on it.

[0009] Further, in the iterative process of searching for the optimization solution using the cuckoo search model, the Q-learning algorithm is used to update the global step size coefficient α and the local distance scale coefficient β of the cuckoo search model, including: initializing the Q-value table, the state space and action space of the value set of the objective function; the Q-value represents the expected value of the cumulative reward when performing the action of selecting α and β in the given state s; in each iteration, the action of selecting the values of α and β is selected from the Q-value table according to the ε-greedy strategy, and the selected values of α and β are input into the cuckoo search model for solution, and the Q-value table is updated using the objective function TS and the reward function R.

[0010] The ε-greedy strategy.

[0011]

[0012] argmaxV(a) represents the optimal action known in the current state selected with a probability of (1 - ε); rand(a) represents an action randomly selected with a uniform distribution with a probability of ε.

[0013] Furthermore, the Tent chaotic map is used for the population initialization of the cuckoo search model:

[0014]

[0015] where z k is the random number at this iteration, and z k+1 is the random number generated in the next iteration after passing through the Tent chaotic map, and τ is the control parameter.

[0016] Furthermore, the discovery probability of the cuckoo search model

[0017]

[0018] where represents the average response delay of the user at the current iteration, represents the initial average response delay of the user at the current iteration, represents the 99th percentile delay of the user at the current iteration, represents the initial 99th percentile delay of the user at the current iteration; ω5 and ω6 are weight values.

[0019] Furthermore, the task scheduling model is represented as a three-dimensional tensor

[0020]

[0021] where x ijk ∈{0, 1}, when any task t i is assigned to the processing core c j in the edge computing device p k then x ijk = 1, otherwise x ijk = 0.

[0022] The present invention also provides a task scheduling device for edge-side video processing, including: a model construction module, configured to construct a task scheduling model for an edge video processing system and optimization metrics for the task scheduling model; an optimization objective module, configured to generate an objective function and constraint conditions for the task scheduling model based on the optimization metrics; and a solution selection module, configured to use a cuckoo search model, and take the task scheduling solution corresponding to the objective function when the optimization metrics are all at minimum values as the optimized solution, and perform a task scheduling operation on the edge video processing system with the optimized solution.

[0023] The present invention also provides a computer-readable storage medium storing computer-executable instructions, characterized in that when the computer-executable instructions are executed, the task scheduling method for edge-side video processing as described above is implemented.

[0024] The present invention also provides an electronic device including the task scheduling device for edge-side video processing as described above.

[0025] For the task scheduling method for edge-side video processing of the present invention, in view of the problem that the user response latency needs to be reduced when performing video processing task scheduling in an edge environment, the average user response latency and tail latency are taken as the main optimization metrics. In addition, considering the actual situation of multiple users submitting task requests, two metrics of user fairness and task priority are respectively used to reflect the influence of the order in which different users submit task requests and the execution order of different tasks of the same user on the task scheduling problem. Description of the Drawings

[0026] Figure 1 is a flowchart of the task scheduling method for edge-side video processing of the present invention.

[0027] Figure 2 is a schematic diagram of the task scheduling device for edge-side video processing of the present invention.

[0028] Figure 3 is a schematic diagram of the model construction module of the present invention.

[0029] Figure 4 is a schematic diagram of the solution selection module of the present invention.

[0030] Figure 5 is a schematic diagram of the model optimization module of the present invention.

[0031] Figure 6 is a schematic diagram of an electronic device of the present invention.

[0032] Figure 7 is a schematic diagram of the hardware structure of an electronic device of the present invention.

[0033] Among them, the reference signs are:

[0034] 100: Electronic device 10: Model construction module

[0035] 11: Task scheduling model construction module 12: Optimization metric determination module

[0036] 20: Optimization objective module 30: Solution selection module

[0037] 31: Model initialization module 32: Model optimization module

[0038] 321: Q-value function update module 322: State space partitioning module

[0039] 323: Action space partitioning module 324: Reward function setting module

[0040] 325: Action selection strategy module 326: Parameter update module

[0041] S1, S2, S3, S11, S12, S31, S32, S311, S312, S313, S314, S315, S321, S322, S323, S324, S325, S326: Steps Detailed implementation manners

[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific implementation methods described herein are only used to explain the present invention and are not used to limit the present invention.

[0043] When the present invention performs video processing task scheduling in an edge environment, the response latency of users needs to be reduced. The present invention takes the average response latency and tail latency of users as the main optimization metrics, models them in the form of a weighted sum, and on this basis, proposes an adaptive cuckoo algorithm based on Q-learning. Aiming at the limitations of the cuckoo algorithm in solving practical problems, such as high sensitivity to parameter selection, low algorithm solving efficiency, and poor accuracy, the present invention classifies the parameters required by the algorithm into two categories: main parameters and secondary parameters. The main parameters are dynamically optimized by using the Q-learning algorithm, and the secondary parameters adopt an adaptive update strategy according to the change of the optimization objective. This method can automatically optimize and update the algorithm parameters during the iteration process, improving the solving performance of the algorithm. In addition, in order to solve the problems such as uneven population distribution and introduction of unstable factors brought by the cuckoo algorithm in the initialization stage, the present invention changes the initialization method of the cuckoo population through Tent chaotic mapping, accelerating the convergence speed of the algorithm.

[0044] The technical solution of the present invention includes:

[0045] 1. Construction of task scheduling model in edge environment

[0046] In an edge environment, if the number of users submitting requests within a certain period is x. The set of the number of video processing tasks submitted by each user is Y = {y1, y2, …, y x}, y x represents the number of tasks submitted by the x-th user. The total number of video processing tasks received by the edge cluster within this period is n. Define the video processing task set as T = {t1, t2, …, t i , …, t n}, where t i represents the i-th task. In the edge cluster, the number of edge computing devices that can participate in task scheduling is m. Define the edge computing device set as P = {p1, p2, …, p j , …, p m}, where p j represents the j-th edge computing device. Define the set of cores in each edge computing device as C = {c1, c2, …, c k , …, c l}, where l is the number of cores in this edge computing device, and c k is the k-th core in this device. If the number of available cores in an edge computing device is k' (k' ≤ l), then each element in {c k'+1 , …, c l} is 0.

[0047] The scheduling of a single task can be described as the allocation relationship of task t i to core c j in edge computing device p k . The task scheduling scheme can be represented by a three-dimensional tensor X nml of order (n × m × l), as shown in formula (1):

[0048]

[0049] where, x ijk ∈ {0, 1}. If task t i is allocated to core c j in device p k , x ijk = 1; otherwise, x ijk = 0.

[0050] Since the tasks submitted by users may be executed on different edge devices, define the set of times required for each user's tasks to be submitted and executed on all devices as ET = {et1, et2, …, et j , …, et m}. Where et jRepresents the total completion time of the task allocated to the j-th device, where et j already includes the queuing time of the task. When the task of this user is not executed on this device, et j = 0. Then the total processing time t proc (including the task execution time and the task queuing time) can be calculated by formula (2):

[0051] t proc = max{et1, et2, …, et j , …, et m} (2)

[0052] In the video processing task scheduling problem for normal request status in the edge environment, it is mainly considered to optimize from the perspective of user response latency. Therefore, it is necessary to establish a communication model between the user and the edge device. According to Shannon's formula, the calculation formula for the task transmission rate is as shown in (3):

[0053]

[0054] Among them, R u represents the task transmission rate corresponding to user u, B w represents the channel bandwidth, τ u represents the bandwidth ratio allocated to user u. P uj represents the transmission power between the j-th edge device and user u, G is the channel gain, and σ 2 is the noise power.

[0055] According to the task transmission rate and the task data volume, the task transmission time can be calculated:

[0056]

[0057] In the above formula, t tran is the task transmission time, D i is the task data volume size.

[0058] The user response latency t resp is the sum of the total processing time of the user task and the task transmission time:

[0059] t resp = t proc + t tran (5)

[0060] The average response latency of the user can be calculated by formula (6):

[0061]

[0062] Among them, RT avg represents the average response latency of the user, Denote the response latency of user h, and x is the total number of users submitting requests during this time period.

[0063] To better improve the overall system performance and ensure the user experience, the present invention introduces tail latency as one of the optimization goals. Tail latency, also known as high-percentage latency, represents the part with a relatively small proportion when the response time is significantly higher than the mean. To reflect the degree, tail latency is usually expressed as a percentage. According to the industry-standard commonly used, the present invention uses the 99th percentile latency for description. Its calculation method is to sort the response latencies of each user from small to large and determine the position of the 99th percentile. The response latency corresponding to this position is the 99th percentile latency, denoted as TL. 99th 。

[0064] In reality, the time when different users submit task requests may be different, and each user may determine the submission order of tasks according to their own preferences. The present invention proposes two indicators, user fairness and task priority, which respectively reflect the order of different users submitting task requests and the impact of the execution order of different tasks of the same user on the task scheduling problem. In the user fairness indicator, the earlier a user submits a task, the more urgently the task needs to be executed; in the task priority indicator, each user can determine the submission order of tasks by themselves, and the earlier a task is submitted, the more urgently it needs to be executed. To build an optimization model, the present invention uses the concept of the inversion number to mathematically describe the above two indicators.

[0065] Construct a user sequence from small to large according to the order of user task submission time. Take the response time of each user as the value of the element at the corresponding position in this sequence. Then, calculating the inversion number of this sequence can represent user fairness, denoted as UF. inv 。UF inv The smaller the value of UF, the higher the user fairness. For each user, construct a sequence from small to large according to the order of task submission time, and take the execution time of each task as the value of the element at the corresponding position in this sequence. In this way, sequences with the same number as the number of users can be constructed. Calculate the inversion number of each sequence and sum them up, denoted as TP. inv 。TP inv It can reflect the task priority. The smaller its value, the more the expected execution order of each user's tasks can be satisfied.

[0066] In summary, in the video processing task scheduling problem for normal request status in the edge environment, the optimization metrics of the present invention include the average user response latency, tail latency, user fairness, and task priority. To prioritize real-time performance, the average user response latency and tail latency will be used as the main optimization metrics. The above four metrics are all minimum optimization metrics, that is, the optimization task is to make the metric values as small as possible. Therefore, the optimization goal of the present invention is to pursue the task scheduling scheme when minimizing the average user response latency and tail latency while ensuring user fairness and task priority as much as possible. This task scheduling problem belongs to a multi-objective optimization problem, and the objective function is presented in the form of the weighted sum of each metric:

[0067] min TS=ω1RT avg +ω2TL 99th +ω3UF inv +ω4TP inv (7)

[0068] In the above formula, TS is the objective function; ω1, ω2, ω3, and ω4 are weight values, reflecting the importance of different metrics, and the sum of each weight is equal to 1.

[0069] To prevent tasks from being assigned to edge devices with insufficient computing or storage resources during task scheduling, the constraint conditions are as follows:

[0070]

[0071] Among them, and respectively represent the usage amounts of computing and storage resources of task t i on edge device p j ; and respectively represent the available amounts of computing and storage resources on edge device p j .

[0072] 2. Initializing the algorithm using Tent chaotic mapping

[0073] As the main optimization algorithm of the present invention, each individual in the cuckoo population represents a solution to task scheduling. In the initialization stage of the algorithm, the quality of the initial population has a great influence on the optimization performance of the algorithm. The cuckoo algorithm usually uses the method of randomly initializing the population, which may cause uneven population distribution, is not conducive to exploring the solution space, and introduces unstable factors and other problems. To solve the above problems, the present invention will use Tent chaotic mapping to initialize the population.

[0074] Chaos is a kind of dynamic behavior that is irregular, unpredictable, and highly sensitive to initial conditions. A chaotic map is a type of mathematical map that exhibits chaotic behavior in a nonlinear system. Due to the high randomness and dispersion of chaotic maps, the advantages of using chaotic maps for population initialization include: increasing the diversity of individuals and avoiding the homogenization of individuals; providing a broader search space and accelerating the convergence speed of the algorithm towards excellent solutions.

[0075] The calculation formula of the Tent chaotic map is as follows:

[0076]

[0077] where τ is the control parameter, τ ∈ (0, 1); z k is the random number at the current iteration, and z k+1 is the random number generated for the next iteration after being processed by the chaotic map.

[0078] To better describe the process of using the Tent chaotic map for population initialization, the present invention first defines the solution space of the video processing task scheduling problem for edge environment facing normal requests. In the present invention, the horizontal axis and the vertical axis represent edge devices and cores respectively, and are represented in the form of natural numbers. Within this range, each coordinate point reflects a certain core of a certain edge device. For example, the point (1, 1) represents the first core of the first device. The scheduling of a single task is the allocation of the task to a certain core in a certain edge device, corresponding to a coordinate point; the task scheduling scheme is the allocation of all tasks, corresponding to the sequence formed by each coordinate point.

[0079] 3. Q - learning algorithm design

[0080] In the cuckoo algorithm, the formula for the global search strategy is as follows:

[0081]

[0082] The update of the position of the bird's nest represents the process of the algorithm exploring and optimizing in the solution space. In the above formula, and represent the positions of the i - th bird's nest at the t - th generation and the (t + 1) - th generation respectively, represents the position of the optimal bird's nest at the t - th generation, representing the optimal solution at the t - th iteration; α is the global step - size scale coefficient, aiming to scale the scale of the global step - size to an appropriate size; represents the dot - product of vectors, and Levy(s) is the random step - size obtained by Levy flight.

[0083] The formula for updating the position of the bird's nest in the local search strategy of the cuckoo algorithm is as follows:

[0084]

[0085] wherein, and respectively represent the positions of the i-th cuckoo nest at the t-th generation and the (t + 1)-th generation. β is the local distance scale coefficient, and its purpose is to scale the local movement distance of the cuckoo nest to an appropriate scale. and represent the positions of any other two cuckoo nests at the t-th generation, which are obtained by uniform distribution sampling.

[0086] In the cuckoo algorithm, the parameters α and β are used as the global step scale coefficient and the local distance scale coefficient respectively, and are used for global search and local search in the solution space. In order to more fully traverse the solution space and adjust the accuracy of the solution during the iteration process, the parameters α and β have different value ranges. Currently, the values of the parameters α and β are usually manually adjusted according to experience and the running situation of the algorithm. However, the cuckoo algorithm is sensitive to parameter settings, and inappropriate parameter selection will lead to low algorithm solving efficiency and poor accuracy. The above parameter tuning method is not only time-consuming and laborious, but also difficult to adapt the values of α and β to the running state of the algorithm, thus affecting the optimization performance of the algorithm. As a machine learning paradigm, reinforcement learning can learn online in an uncertain unknown environment, adjust strategies according to the feedback of the environment, and has good generalization ability. Therefore, the present invention introduces the Q-learning algorithm in reinforcement learning to dynamically find the optimal combination of the main parameters α and β by revealing the internal structure of the cuckoo population, so as to improve the solving efficiency and accuracy of the cuckoo algorithm.

[0087] The following will introduce the design process of the Q-learning algorithm from six parts: the update of the Q-value function, the division of the state space and the action space, the setting of the reward function, the ε-greedy action selection strategy, and the update steps of the main parameters α and β.

[0088] 1) Update of the Q-value function

[0089] The main goal of the Q-learning algorithm is to learn the Q-value function, which represents the expected value of the cumulative reward when performing a specific action in a given state. For the state-action pair (s, a), the Q-value function is expressed as Q(s, a). The update of the Q-value function uses the Bellman equation:

[0090] Q(s, a) new = Q(s, a) cur + μ[R + γ * maxQ(s', a') - Q(s, a) cur (13)

[0091] In the above formula, Q(s, a) new refers to the new Q-value of action a in state s, and Q(s, a) curis the current Q value. μ is the learning rate, which reflects the weight of new information in each update of the Q value; γ is the discount factor, used to measure the importance of future rewards. R is the immediate reward obtained after executing action a, and s' is the new state after executing action a. maxQ(s', a') is the maximum expected Q value after executing new action a' in the new state s'.

[0092] 2) Division of the state space

[0093] The set of values of the objective function constitutes the state space of the cuckoo search model. In the state space, the number of different states divided has a great influence on the search results. If the number of states divided is too large, a large amount of time will be spent optimizing in each iteration process; if the number of states divided is too small, the quality of the solution will be reduced. In the task scheduling problem of the present invention, the state space is divided into 15 intervals within 0 to 1, which are [0, 0.4), [0.4, 0.45), [0.45, 0.5), [0.5, 0.54), [0.54, 0.58), [0.58, 0.62), [0.62, 0.66), [0.66, 0.7), [0.7, 0.74), [0.74, 0.78), [0.78, 0.82), [0.82, 0.86), [0.86, 0.9), [0.9, 0.95) and [0.95, 1].

[0094] 3) Division of the action space

[0095] In the Q-learning algorithm, the agent always selects actions from the action space. In the cuckoo search model, the selection of actions is the selection of main parameters, including the global step size scale coefficient α and the local distance scale coefficient β. For the global step size scale coefficient α, in the present invention, it is divided into 18 intervals within the range of 0.5 to 5, and the length of each interval is 0.25; for the local distance scale coefficient β, in this paper, it is divided into 15 intervals within the range of 0.05 to 0.8, and the length of each interval is 0.05. The action space is composed of different intervals. When an interval is selected, a random value within the interval is selected as the value of the parameter.

[0096] 4) Setting of the reward function

[0097] In order to evaluate the advantages and disadvantages of the selection of the global step size scale coefficient α and the local distance scale coefficient β, an appropriate reward function needs to be selected. According to the objective function of the cuckoo search model, the definition of the reward function R is as follows:

[0098]

[0099] Among them, and represents the average response latency and the 99th percentile delay of users at the t-th generation, and represents the average response latency and the 99th percentile delay of users at the (t - 1)-th generation. ω5 and ω6 are weight values.

[0100] 5) ε-greedy action selection strategy

[0101] The present invention adopts the ε-greedy strategy to select an action from the action space:

[0102]

[0103] where argmaxV(a) represents the known optimal action in the current state, which is selected with a probability of (1 - ε), and this is the exploitation state at this time; rand(a) represents randomly selecting an action in a uniform distribution manner, which is selected with a probability of ε, and this is the exploration state at this time.

[0104] 6) Update steps of the main parameters α and β

[0105] In the cuckoo search model, an individual in the cuckoo population can be regarded as an agent in the Q-learning algorithm, which refers to a solution to task scheduling in the present invention. The steps to update the global step size scale coefficient α and the local distance scale coefficient β through the Q-learning algorithm are as follows: First, initialize the Q-table, state space, and action space. Then, in one iteration of the algorithm, select an action from the Q-table according to the ε-greedy strategy, that is, select α and β. After inputting them into the cuckoo algorithm for solution, update the Q-table according to the value of the objective function and the reward function. Repeat the above steps until the total number of iterations of the algorithm is satisfied.

[0106] 4. Adaptive update strategy

[0107] According to the different focuses and advantages of the global search strategy and the local search strategy in the cuckoo algorithm, the global search should be used as the main strategy in the early stage of the algorithm to avoid falling into local optimal solutions, and gradually increase the proportion of the local search strategy in the middle and late stages to further improve the quality of the solution. In the standard cuckoo algorithm, the discovery probability P a is a constant, which cannot reflect its role in balancing the global search and the local search, thus reducing the optimization performance of the algorithm.

[0108] To better balance the global search strategy and the local search strategy of the algorithm and improve the algorithm performance, the present invention defines the discovery probability P a as a secondary parameter and adopts an adaptive update strategy to adjust according to the change of the optimization objective in the task scheduling problem. The update formula of the discovery probability P a is as follows:

[0109]

[0110] Among them, and represent the average response latency of users and the initial average response latency of users at the current iteration, and represent the 99th percentile latency of users and the initial 99th percentile latency at the current iteration; ω7 and ω8 are weight values.

[0111] In the initial stage of the algorithm, due to and having larger values, the discovery probability P a has a smaller value, and the algorithm will mainly adopt the global search strategy of Levy flight; as the algorithm iterates continuously, and gradually decrease, and the value of P a will gradually increase, so the possibility of the local search strategy will gradually increase, ensuring a good balance between the global search strategy and the local search strategy. At the same time, if in a certain iteration, and do not update their values, the value of the discovery probability P a will remain unchanged, avoiding the problem that the value of P a changes continuously with the number of iterations, resulting in poor robustness of the algorithm. In addition, when updating the value of P a , the average response latency and tail latency of users are used as variables, which is more conducive to strengthening the ability of the algorithm to solve the main optimization objective.

[0112] Based on the above technical solution, in the first embodiment of the present invention, a task scheduling method for edge-side video processing is proposed, as Figure 1 shown, including:

[0113] Step S1, constructing a task scheduling model for the edge video processing system and the optimization metrics of this task scheduling model; specifically including:

[0114] Step S11, constructing a task scheduling model in the edge environment

[0115] In the edge environment, the number of users submitting requests within a certain time period is x. The set of the number of video processing tasks submitted by each user is Y = {y1, y2,..., y x}, and y x represents the number of tasks submitted by the xth user. The total number of video processing tasks received by the edge cluster within this time period is n. Define the video processing task set as T = {t1, t2,..., t i ,..., t n}, where t iDenote the \(i\)-th task. In the edge cluster, the number of edge computing devices that can participate in task scheduling is \(m\). Define the set of edge computing devices as \(P = \{p_1, p_2, \ldots, p_m\}\), where \(p_j\) represents the \(j\)-th edge computing device. Define the set of cores in each edge computing device as \(C = \{c_1, c_2, \ldots, c_l\}\), where \(l\) is the number of cores in this edge computing device, and \(c_k\) is the \(k\)-th core in this device. If the number of available cores in an edge computing device is \(k'\) (\(k' \leq l\)), then each element in \(\{c_{k'+1}, \ldots, c_l\}\) is \(0\). j , \ldots, p m \}, where \(p j represents the \(j\)-th edge computing device. Define the set of cores in each edge computing device as \(C = \{c_1, c_2, \ldots, c k , \ldots, c l \}, where \(l\) is the number of cores in this edge computing device, and \(c k is the \(k\)-th core in this device. If the number of available cores in an edge computing device is \(k'\) (\(k' \leq l\)), then each element in \(\{c k'+1 , \ldots, c l \} is \(0\).

[0116] The scheduling of a single task can be described as the allocation relationship of task \(t_i\) to core \(c_k\) in edge computing device \(p_j\). The task scheduling scheme can be represented by a three-dimensional tensor \(X\) of order \((n\times m\times l)\), as shown in formula (1). i to edge computing device \(p j and core \(c k in it. The task scheduling scheme can be represented by a three-dimensional tensor \(X\) of order \((n\times m\times l)\), as shown in formula (1). nml as shown in formula (1).

[0117] Step S12: Determine the optimization metrics of the task scheduling model

[0118] Since the tasks submitted by users may be executed on different edge devices, define the set of times required for all tasks of a certain user to be submitted and executed on all devices as \(ET = \{et_1, et_2, \ldots, et_n\}\). Among them, \(et_j\) represents the total completion time of the task allocated to the \(j\)-th device. At this time, \(et_j\) already includes the queuing time of the task. When the task of this user is not executed on this device, \(et_j = 0\). Then the total processing time \(t\) (including the task execution time and the task queuing time) of this user's task can be calculated by formula (2). j , \ldots, et m \}. Among them, \(et j represents the total completion time of the task allocated to the \(j\)-th device. At this time, \(et j already includes the queuing time of the task. When the task of this user is not executed on this device, \(et j = 0. Then the total processing time \(t proc (including the task execution time and the task queuing time) of this user's task can be calculated by formula (2).

[0119] In the task scheduling problem of video processing tasks for normal request status in the edge environment, it is mainly considered to optimize from the perspective of user response latency. Therefore, it is necessary to establish a communication model between users and edge devices. According to Shannon's formula, the calculation of the task transmission rate \(R\) is as shown in formula (3). u is as shown in formula (3).

[0120] According to the task transmission rate and the task data volume, the task transmission time \(t\) can be calculated. tran .

[0121] User response latency \(t\)resp is the sum of the total processing time of the user task and the task transmission time.

[0122] The average response delay RT of the user avg can be calculated by formula (6).

[0123] To better improve the overall system performance and ensure the user experience, the present invention introduces tail delay as one of the optimization goals. Tail delay, also known as high-percentage delay, represents the part with a relatively small proportion when the response time is significantly higher than the mean. To reflect the degree, tail delay is usually expressed as a percentage. According to the industry-standard commonly used, the present invention uses the 99th percentile delay for description. Its calculation method is to sort the response delays of each user from small to large and determine the position of the 99th percentile. The response delay corresponding to this position is the 99th percentile delay, denoted as TL. 99th .

[0124] In reality, the times when different users submit task requests may be different, and each user may determine the submission order of tasks according to their own preferences. The present invention proposes two indicators, user fairness and task priority, which respectively reflect the influence of the submission order of task requests by different users and the execution order of different tasks of the same user on the task scheduling problem. In the user fairness indicator, the earlier a user submits a task, the more urgently the task needs to be executed; in the task priority indicator, each user can determine the submission order of tasks by themselves, and the earlier a task is submitted, the more urgently it needs to be executed. To construct an optimization model, the present invention uses the concept of the inversion number to mathematically describe the above two indicators.

[0125] Construct a user sequence from small to large according to the order of the times when users submit tasks. Take the response time of each user as the value of the element at the corresponding position in this sequence. Then, calculating the inversion number of this sequence can represent user fairness, denoted as UF. inv . UF inv The smaller the value of UF, the higher the user fairness. For each user, construct a sequence from small to large according to the order of the submission times of each task. Take the execution time of each task as the value of the element at the corresponding position in this sequence. In this way, sequences with the same number as the number of users can be constructed. Calculate the inversion number of each sequence and sum them up, denoted as TP. inv . TP inv can reflect the task priority. The smaller its value, the more the expected execution order of tasks of each user can be satisfied.

[0126] Step S2: Based on this optimization index, generate the objective function and constraint conditions of the task scheduling model;

[0127] In the video processing task scheduling problem for normal request status in the edge environment, the optimization metrics of the present invention include the average user response latency, tail latency, user fairness, and task priority. To prioritize real-time performance, the average user response latency and tail latency will be used as the main optimization metrics. The above four metrics are all minimum optimization metrics, that is, the optimization task is to make the metric values as small as possible. Therefore, the optimization objective of the present invention is to pursue the task scheduling scheme with the minimum average user response latency and tail latency while ensuring user fairness and task priority as much as possible. This task scheduling problem belongs to a multi-objective optimization problem, and the objective function TS is presented in the form of the weighted sum of each metric, as shown in formula (7).

[0128] To prevent tasks from being assigned to edge devices with insufficient computing or storage resources during task scheduling, the constraint conditions are as shown in formulas (8) and (9).

[0129] Step S3: Use the cuckoo search model, take the task scheduling scheme corresponding to the objective function when the optimization metrics are all at their minimum values as the optimization scheme, and perform task scheduling operations on the edge video processing system with this optimization scheme. It includes:

[0130] Step S31: Initialize the cuckoo search model using Tent chaotic mapping;

[0131] The cuckoo algorithm is the main optimization algorithm of the present invention. Each individual in the cuckoo population represents a solution to task scheduling. In the initialization stage of the algorithm, the quality of the initial population has a greater impact on the optimization performance of the algorithm. The cuckoo algorithm usually uses the method of randomly initializing the population, which may cause uneven population distribution, be unfavorable for exploring the solution space, and introduce unstable factors and other problems. To solve the above problems, the present invention will use Tent chaotic mapping for population initialization.

[0132] Chaos is a kind of irregular, unpredictable and highly sensitive dynamic behavior to initial conditions. Chaotic mapping is a class of mathematical mappings that exhibit chaotic behavior in nonlinear systems. Due to the high randomness and dispersion of chaotic mapping, the advantages of using chaotic mapping for population initialization include: increasing the diversity of individuals, avoiding the homogenization of individuals; providing a broader search space and accelerating the convergence speed of the algorithm to excellent solutions.

[0133] The Tent chaotic mapping is as shown in formula (10).

[0134] To better describe the process of population initialization using the Tent chaotic map, the present invention first defines the solution space of the video processing task scheduling problem for edge environments facing normal requests. In the present invention, the horizontal and vertical axes represent edge devices and cores respectively, and are represented in the form of natural numbers. Within this range, each coordinate point reflects a certain core of a certain edge device. For example, the point (1,1) represents the first core of the first device. The scheduling of a single task is the allocation of the task to a certain core in a certain edge device, corresponding to a coordinate point; the task scheduling scheme is the allocation of all tasks, corresponding to the sequence formed by each coordinate point.

[0135] According to the above definitions, within the two-dimensional solution space composed of edge devices and cores, the steps of the Tent chaotic map initialization algorithm are as follows:

[0136] Step S311: Randomly generate a coordinate point as the iterative initial point, denoted as (x0, y0). Where x0, y0 ∈ (0, 1) and x0 ≠ y0;

[0137] Step S312: After setting the value of τ, substitute the horizontal and vertical coordinates of the initial point into z in formula (10) k , generate (x1, y1), and use it as the initial point for the next iteration;

[0138] Step S313: Repeat Step S312 until the number of generated coordinate points is equal to the number of tasks and form a sequence with each point;

[0139] Step S314: Map each point in the sequence to the solution space range to generate a new sequence, which is an individual within the cuckoo population and represents an allocation scheme for all tasks;

[0140] Step S315: Repeat Steps S311 to S314 until the number of generated sequences is equal to the number of individuals in the population, completing the population initialization of the cuckoo algorithm.

[0141] Step S32: Introduce the Q-learning algorithm in reinforcement learning to dynamically find the optimal combination of the main parameters α and β by revealing the internal structure of the cuckoo population, improving the solution efficiency and accuracy of the cuckoo search model; including six parts: the update of the Q-value function, the division of the state space and action space, the setting of the reward function, the ε-greedy action selection strategy, and the update steps of the main parameters α and β.

[0142] Step S321: Update of the Q-value function

[0143] The main goal of the Q-learning algorithm is to learn the Q-value function, which represents the expected value of the cumulative reward when performing a specific action in a given state. For the state-action pair (s,a), the Q-value function is denoted as Q(s,a). The update of the Q-value function uses the Bellman equation, as shown in formula (13).

[0144] Step S322: Divide the state space

[0145] The set of values of the objective function constitutes the state space of the cuckoo search model. In the state space, the number of different states divided has a great influence on the search results. If the number of states is divided too many, a large amount of time will be spent optimizing in each iteration process; if the number of states is divided too few, the quality of the solution will be reduced. In the task scheduling problem of the present invention, the state space is divided into 15 intervals within 0 to 1, which are [0,0.4), [0.4,0.45), [0.45,0.5), [0.5,0.54), [0.54,0.58), [0.58,0.62), [0.62,0.66), [0.66,0.7), [0.7,0.74), [0.74,0.78), [0.78,0.82), [0.82,0.86), [0.86,0.9), [0.9,0.95) and [0.95,1].

[0146] Step S323: Divide the action space

[0147] In the Q-learning algorithm, the agent always selects actions from the action space. In the cuckoo search model, the selection of actions is the selection of the main parameters, including the global step size scale coefficient α and the local distance scale coefficient β. For the global step size scale coefficient α, in the present invention, it is divided into 18 intervals within the range of 0.5 to 5, and the length of each interval is 0.25; for the local distance scale coefficient β, in this paper, it is divided into 15 intervals within the range of 0.05 to 0.8, and the length of each interval is 0.05. The action space is composed of different intervals. When an interval is selected, a random value within the interval is selected as the value of the parameter.

[0148] Step S324: Set the reward function

[0149] In order to evaluate the advantages and disadvantages of the selection of the global step size scale coefficient α and the local distance scale coefficient β, an appropriate reward function needs to be selected. According to the objective function of the cuckoo search model, the reward function R is as shown in formula (14).

[0150] Step S325: Set the ε-greedy action selection strategy; use formula (15) as the ε-greedy action selection strategy.

[0151] Step S326: Update the main parameters α and β

[0152] In the cuckoo search model, an individual in the cuckoo population can be regarded as an agent in the Q - learning algorithm, which refers to a solution for task scheduling in the present invention. The steps to update the global step - size scale coefficient α and the local distance scale coefficient β through the Q - learning algorithm are as follows: First, initialize the Q - table, state space, and action space. Then, in one iteration of the algorithm, select an action from the Q - table according to the ε - greedy strategy, that is, select α and β. After inputting them into the cuckoo algorithm for solution, update the Q - table according to the value of the objective function and the reward function. Repeat the above steps until the total number of iterations of the cuckoo search model is satisfied.

[0153] The specific solution process of the adaptive cuckoo algorithm based on Q - learning proposed by the present invention is as follows:

[0154]

[0155]

[0156] It should be noted that in various embodiments of the present invention, the magnitudes of the serial numbers of the above steps do not mean the sequence of execution. The execution sequence of each step should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0157] The following is a system embodiment corresponding to the above - mentioned method embodiment. This embodiment can be implemented in cooperation with the above - mentioned embodiment. The relevant technical details mentioned in the above - mentioned embodiment are still valid in this embodiment. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above - mentioned embodiment.

[0158] In the second embodiment of the present invention, a task - scheduling device for edge - side video processing is proposed. As Figure 2 shown, the task - scheduling device of the present invention includes:

[0159] A model - building module 10, configured to build a task - scheduling model for the edge video - processing system and determine the optimization index of the task - scheduling model; As Figure 3 shown, it includes:

[0160] A task - scheduling - model - building module 11, configured to build a task - scheduling model and label it in the form of a three - dimensional tensor

[0161]

[0162] An optimization - index - determining module 12, configured to determine the optimization index of the task - scheduling model

[0163] The optimization metrics include the user fairness metric UF inv 、the task priority metric TP inv 、the average user response latency metric RT avg and the tail latency metric TL 99th ; among them, a user sequence is constructed in ascending order of the time when users submit tasks, and the response time of each user is used as the value of the element at the corresponding position in this user sequence. The inversion number of this user sequence is calculated as the user fairness metric UF inv ; a task sequence is constructed in ascending order of the time when each task is submitted, and the execution time of each task is used as the value of the element at the corresponding position in this task sequence. The inversion number of this task sequence is calculated and summed as the task priority metric TP inv ; the average user response latency metric represents the response latency of user h, and x is the total number of users who submit requests within a specified time period; the tail latency metric TL 99th is the part where the proportion of the response latency higher than the average response latency is higher than 99%.

[0164] The optimization objective module 20 is used to determine the objective function and constraint conditions of this task scheduling model based on this optimization metric;

[0165] This objective function TS

[0166] min TS = ω1RT avg + ω2TL 99th + ω3UF inv + ω4TP inv

[0167] Satisfy the constraint conditions: and

[0168] where, ω1, ω2, ω3 and ω4 are weight values, and ω1 + ω2 + ω3 + ω4 = 1, represents the computing resource usage of task t i on edge device p j ; represents t i on p j 's storage resource usage, represents the available computing resources on edge device p j ; represents the available storage resources on edge device p j ;

[0169] The solution selection module 30 is used to use the cuckoo search model, and take the task scheduling solution corresponding to the objective function when the optimization metrics are all minimum values as the optimization solution, and perform task scheduling operations on the edge video processing system with this optimization solution. As Figure 4 shown, it includes:

[0170] The model initialization module 31 is used to initialize the cuckoo search model with Tent chaotic mapping;

[0171] The cuckoo algorithm is the main optimization algorithm of the present invention. Each individual in the cuckoo population represents a solution to task scheduling. In the initialization stage of the algorithm, the quality of the initial population has a greater impact on the optimization performance of the algorithm. The cuckoo algorithm usually adopts the method of randomly initializing the population, which may cause uneven population distribution, is not conducive to solution space exploration, and introduces unstable factors and other problems. To solve the above problems, the present invention will use Tent chaotic mapping for population initialization.

[0172] Chaos is a kind of irregular, unpredictable and highly sensitive dynamic behavior to initial conditions. Chaotic mapping is a class of mathematical mappings that exhibit chaotic behavior in nonlinear systems. Due to the high randomness and dispersion of chaotic mapping, the benefits of using chaotic mapping for population initialization include: increasing the diversity of individuals, avoiding the homogenization of individuals; providing a broader search space and accelerating the convergence speed of the algorithm to excellent solutions.

[0173] The Tent chaotic mapping is shown in formula (10).

[0174] To better describe the process of using Tent chaotic mapping for population initialization, the present invention first defines the solution space of the video processing task scheduling problem for normal requests in the edge environment. In the present invention, the horizontal axis and the vertical axis respectively represent edge devices and cores, and are represented in the form of natural numbers. Within this range, each coordinate point reflects a certain core of a certain edge device. For example, the point (1,1) represents the first core of the first device. The scheduling of a single task is the allocation of the task to a certain core in a certain edge device, corresponding to a coordinate point; the task scheduling solution is the allocation of all tasks, corresponding to the sequence formed by each coordinate point.

[0175] The model optimization module 32 is used to introduce the Q-learning algorithm in reinforcement learning, and dynamically find the optimal combination of the main parameters α and β by revealing the internal structure of the cuckoo population, so as to improve the solution efficiency and accuracy of the cuckoo search model. As Figure 5 shown, it includes a Q-value function update module 321, a state space partitioning module 322, an action space partitioning module 323, a reward function setting 324, an action selection strategy module 325, and a parameter update module 326;

[0176] The Q-value function update module 321 is used for Q-value function update;

[0177] The main objective of the Q-learning algorithm is to learn the Q-value function, which represents the expected value of the cumulative reward when performing a specific action in a given state. For the state-action pair (s, a), the Q-value function is denoted as Q(s, a). The update of the Q-value function uses the Bellman equation, as shown in Equation (13).

[0178] The state space partitioning module 322 is used to partition the state space

[0179] The set of values of the objective function constitutes the state space of the cuckoo search model. In the state space, the number of different states partitioned has a great influence on the search results. If the number of state partitions is too large, a large amount of time will be spent optimizing in each iteration process; if the number of state partitions is too small, the quality of the solution will be reduced. In the task scheduling problem of the present invention, the state space is partitioned into 15 intervals within 0 to 1, namely [0, 0.4), [0.4, 0.45), [0.45, 0.5), [0.5, 0.54), [0.54, 0.58), [0.58, 0.62), [0.62, 0.66), [0.66, 0.7), [0.7, 0.74), [0.74, 0.78), [0.78, 0.82), [0.82, 0.86), [0.86, 0.9), [0.9, 0.95) and [0.95, 1].

[0180] The action space partitioning module 323 is used to partition the action space;

[0181] In the Q-learning algorithm, the agent always selects an action from the action space. In the cuckoo search model, the selection of an action is the selection of the main parameters, including the global step size scale coefficient α and the local distance scale coefficient β. For the global step size scale coefficient α, in the present invention, it is partitioned into 18 intervals within the range of 0.5 to 5, and the length of each interval is 0.25; for the local distance scale coefficient β, in this paper, it is partitioned into 15 intervals within the range of 0.05 to 0.8, and the length of each interval is 0.05. The action space is composed of different intervals. When an interval is selected, a random value within the interval is selected as the value of the parameter.

[0182] The reward function setting 324 is used to set the reward function;

[0183] In order to evaluate the advantages and disadvantages of the selection of the global step size scale coefficient α and the local distance scale coefficient β, an appropriate reward function needs to be selected. According to the objective function of the cuckoo search model, the reward function R is as shown in Equation (14).

[0184] The action selection strategy module 325 is used to set the ε-greedy action selection strategy; the present invention uses formula (15) as the ε-greedy action selection strategy.

[0185] The parameter update module 326 is used to update the main parameters α and β

[0186] In the cuckoo search model, an individual in the cuckoo population can be regarded as an agent in the Q-learning algorithm, which refers to a solution for task scheduling in the present invention. The steps for updating the global step size scale coefficient α and the local distance scale coefficient β through the Q-learning algorithm are as follows: First, initialize the Q-table, state space, and action space, and then select an action from the Q-table according to the ε-greedy strategy in one iteration of the algorithm, that is, select α and β. After inputting them into the cuckoo algorithm for solution, update the Q-table according to the value of the objective function and the reward function. Repeat the above steps until the total number of iterations of the cuckoo search model is satisfied.

[0187] In the third embodiment of the present invention, a computer-readable storage medium is proposed. For the task scheduling device for video processing on the edge side of the present invention, when its functions are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. Therefore, in the third embodiment of the present invention, a computer-readable storage medium is provided for storing a computer program for executing a task scheduling method for video processing on the edge side. It should be understood that the computer-readable storage medium in the embodiments of the present invention can be a volatile memory and / or a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).

[0188] Figure 6 is a schematic diagram of an electronic device of the present invention. As Figure 6As shown in the figure, in the fourth embodiment of the present invention, an electronic device 100 is proposed, which includes the task scheduling device for edge-side video processing as described above. Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by instructing relevant hardware (such as a processor, FPGA, ASIC, etc.) through a program. All or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module in the above embodiments can be implemented in the form of hardware, for example, by an integrated circuit to implement its corresponding function, or can be implemented in the form of a software function module, for example, by a processor executing a program / instruction stored in a memory to implement its corresponding function. The embodiments of the present invention are not limited to any specific form of combination of hardware and software.

[0189] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or combine certain components, or have a different component arrangement.

[0190] The electronic device of the present invention can be any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by a processor of any device with data processing capabilities reading the corresponding computer program instructions in a non-volatile memory into the memory and running them. Figure 7 It is a schematic diagram of the hardware structure of an electronic device of the present invention. As Figure 7 shown, from the hardware level, it is a hardware structure diagram of any device with data processing capabilities where the task scheduling device for edge-side video processing of the present invention is located. In addition to Figure 7 the processor, memory, network interface, and non-volatile memory shown, the any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated here.

[0191] When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more collections of available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium may be a solid-state drive.

[0192] The task scheduling method for edge-side video processing of the present invention aims at the problem that the response latency of users needs to be reduced when scheduling video processing tasks in the edge environment, and takes the average response latency and tail latency of users as the main optimization indicators. In addition, considering the actual situation of multiple users submitting task requests, two indicators of user fairness and task priority are respectively used to reflect the order of different users submitting task requests and the impact of the execution order of different tasks of the same user on the task scheduling problem. The present invention models the above indicators in the form of a weighted sum, and on this basis, proposes an adaptive cuckoo algorithm based on Q-learning. Aiming at the limitation that the cuckoo algorithm is highly sensitive to parameter selection when solving practical problems, resulting in low algorithm solving efficiency and poor accuracy, the present invention classifies the parameters required by the algorithm into two categories: main parameters and secondary parameters. The main parameters are dynamically optimized using the Q-learning algorithm, and the secondary parameters adopt an adaptive update strategy according to the change of the optimization target. This method can automatically optimize and update the algorithm parameters during the iteration process, improving the solving performance of the algorithm. In addition, the algorithm changes the initialization method of the cuckoo population through Tent chaotic mapping, accelerating the convergence speed of the algorithm. Simulation experiments show that compared with the other three comparison algorithms, the algorithm of the present invention not only improves the solving efficiency, but also effectively reduces the average response latency and tail latency of users on the premise of ensuring user fairness and the reasonable execution order of tasks.

[0193] The above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the present invention. The patent protection scope of the present invention shall be defined by the claims.

Claims

1. A task scheduling method for video processing on the edge side, characterized in that Including: Construct a task scheduling model for an edge video processing system and optimization metrics for the task scheduling model; The optimization metrics include: User Fairness metric UF inv , Task Priority metric TP inv , Average User Response Latency metric RT avg and Tail Latency metric TL 99th ; among them, a user sequence is constructed in ascending order according to the time when users submit tasks, and the response time of each user is used as the value of the element at the corresponding position in the user sequence, and the number of inversions of the user sequence is calculated as the User Fairness metric UF inv ; a task sequence is constructed in ascending order according to the time when each task is submitted, and the execution time of each task is used as the value of the element at the corresponding position in the task sequence, and the number of inversions of the task sequence is calculated and summed as the Task Priority metric TP inv ; Average User Response Latency metric , represents the response latency of user h, and x is the total number of users who submit requests within a specified time period; Tail Latency metric TL 99th is the part where the proportion of the response latency higher than the average response latency is higher than 99%; Generate an objective function and constraint conditions for the task scheduling model based on the optimization metrics; Using the cuckoo search model, the task scheduling scheme corresponding to the objective function when all the optimization metrics are at their minimum values is taken as the optimization scheme, and the edge video processing system is subjected to task scheduling operations with this optimization scheme; among them, during the iterative process of searching for this optimization scheme using the cuckoo search model, the Q-learning algorithm is used to update the global step size scale coefficient and the local distance scale coefficient of the cuckoo search model. and the local distance scale coefficient , including: initializing the Q-value table, the state space and the action space of the value set of the objective function; the Q-value represents the expected value of the cumulative reward when performing the actions of selecting and under the given state ; in each iteration, according to the policy, select the actions with the selected and values from the Q-value table , input the selected and values into the cuckoo search model for solution, and update the Q-value table with the objective function and the reward function ; and perform the next iteration; This Strategy: denotes the optimal action known in the current state selected with probability; denotes randomly selecting an action in a uniform distribution manner with This reward function : Represents the average response latency of users at the n-th generation, the 99th percentile latency of users at the n-th generation, represents the average response latency of users at the m-th generation, represents the 99th percentile latency of users at the m-th generation, and are weight values.

2. The task scheduling method according to claim 1, wherein Obtain the objective function when the optimized metrics are all at their minimum values : Satisfy: and ; Among them, , , and are weight values, and + + + = 1, represents the amount of computing resources used by task on the edge device . represents the amount of storage resources used on . represents the available amount of computing resources on the edge device . represents the available amount of storage resources on the edge device .

3. The task scheduling method according to claim 1, wherein Use Tent chaotic mapping for population initialization of the cuckoo search model: Among them, is the random number for this iteration, is the random number generated in the next iteration after Tent chaotic mapping, is the control parameter.

4. The task scheduling method according to claim 1, characterized in that The discovery probability of the cuckoo search model Among them, represents the average response latency of the user at the current iteration, represents the initial average response latency of the user at the current iteration, represents the 99th percentile latency of the user at the current iteration, represents the initial 99th percentile latency of the user at the current iteration; and are weight values.

5. The task scheduling method according to claim 1, characterized in that The task scheduling model is represented as a three-dimensional tensor Among them, When any task is assigned to the processing core in the edge computing device then otherwise , Total number of video processing tasks, represents the number of edge devices, represents the number of cores in the edge device, represents the th core in the edge device.

6. A task scheduling device for video processing on the edge side, characterized in that, Including: A model construction module for constructing a task scheduling model for an edge video processing system and optimization metrics for the task scheduling model; The optimization metrics include: User fairness metric UF inv , task priority metric TP inv , average user response latency metric RT avg and tail latency metric TL 99th ; among them, a user sequence is constructed in ascending order according to the time when users submit tasks, and the response time of each user is used as the value of the element at the corresponding position in the user sequence, and the inversion number of the user sequence is calculated as the user fairness metric UF inv ; a task sequence is constructed in ascending order according to the time when each task is submitted, and the execution time of each task is used as the value of the element at the corresponding position in the task sequence, and the inversion number of the task sequence is calculated and summed as the task priority metric TP inv ; average user response latency metric , represents the response latency of user h, and x is the total number of users who submit requests within a specified time period; the tail latency metric TL 99th is the part where the proportion of the response latency higher than the average response latency is higher than 99%; An optimization objective module for generating an objective function and constraint conditions for the task scheduling model based on the optimization metrics; A solution selection module is used to use the cuckoo search model. The task scheduling solution corresponding to the objective function when all the optimization metrics are at their minimum values is taken as the optimization solution, and the edge video processing system is subjected to task scheduling operations using this optimization solution. During the iterative process of searching for the optimization solution using the cuckoo search model, the Q-learning algorithm is used to update the global step size scale coefficient and the local distance scale coefficient of the cuckoo search model. and the local distance scale coefficient , including: initializing the Q-value table, the state space and the action space of the value set of the objective function; the Q-value represents the expected value of the cumulative reward when performing the actions of selection and under a given state ; in each iteration, according to the policy, select the actions with the selected and values from the Q-value table , input the selected and values into the cuckoo search model for solution, and update the Q-value table with the objective function and the reward function ; and perform the next iteration; This Strategy: denotes the optimal action known in the current state selected with probability; denotes a randomly selected action with a uniform distribution with The reward function : Represents the average response latency of users at the nth generation, the 99th percentile delay of users at the nth generation, represents the average response latency of users at the mth generation, represents the 99th percentile delay of users at the mth generation, and are weight values.

7. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed, the task scheduling method for edge-side video processing according to any one of claims 1 to 5 is implemented.

8. An electronic device, including the task scheduling device for edge-side video processing according to claim 6.

Citation Information

Patent Citations

  • A data resource pre-scheduling method based on edge computing

    CN109819030A