Quality of service aware multi-render task scheduling method for interactive applications
Patent Information
- Application Number
- CN202111625247.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-12-28
AI Technical Summary
[0006]在现有技术中,渲染任务调度方案主要存在以下缺陷:1)共享GPU的调度方法适用于一般的应用程序,并不针对交互式应用,因此,这些方法只考虑资源利用率,而没有考虑交互式应用的性能;2)现有方法只决策调度任务,而不涉及分辨率的决策;3)针对边缘辅助或云辅助的交互式应用,现有方法从渲染机制、渲染指令压缩、和压缩参数选择等方面提高用户体验质量或降低延迟,但是没有通过多个渲染任务的调度来优化性能
[0017]与现有技术相比,本发明的优点在于,考虑了交互式应用的性能,既决策调度任务又决策分辨率的选取;可以在低负载、计算资源充裕的情况下,显著提高任务的分辨率和帧率,提升任务性能;考虑了最低帧率需求,通过合理建模和方法设计,提升了满足最低帧率需求的几率;通过设计效用函数,充分权衡性能之间的冲突,使得多种性能尽可能地同时满足需求,提升了满足所有服务质量的几率。
Smart Images

Figure CN116360929B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud computing technology, and more specifically, to a service quality-aware multi-rendering task scheduling method for interactive applications. Background Technology
[0002] Leveraging emerging edge computing and 5G networks, 3D rendering tasks for interactive applications (such as virtual reality and cloud gaming) can be offloaded to edge servers. To improve resource utilization, multiple rendering tasks run on the same GPU server and compete for computing resources. Each task has corresponding performance requirements, i.e., quality of service (QoS) requirements. Given a set of rendering tasks running on a single server and sharing GPU resources, and a set of available resolutions, and knowing the QoS requirements for each rendering task, the scheduler needs to make two decisions when server computing resources are idle: 1) which task to schedule; 2) which resolution to use for rendering. The scheduler's optimization goal is: on the one hand, to meet the QoS requirements of all tasks; on the other hand, the scheduler needs to make full use of resources and optimize the performance of all tasks to maximize user satisfaction. For example, for the QoS and performance of rendering tasks, performance factors affecting the user experience of interactive applications, such as resolution, frame rate, and latency, need to be considered.
[0003] Edge computing takes the form of localized cloud. Because edge servers are closer to users, response times for user requests can be reduced. Cloud-based interactive applications, such as virtual reality and cloud gaming, utilize cloud resources to handle compute-intensive tasks, thus avoiding the need for high-end hardware (often expensive and energy-intensive) on user devices, making clients lightweight. However, cloud-based interactive applications require high-throughput and low-latency network connections. If the distance between the user and the data center is large, the user's low-latency requirements will be difficult to meet. One solution to overcome this problem is to leverage emerging mobile edge computing. Specifically, edge-assisted interactive applications offload compute-intensive 3D rendering tasks to mobile edge computing systems and transmit the edge-rendered video stream to the end user via a 5G connection. Because the edge server is closer to the end user, this approach can significantly reduce latency.
[0004] like Figure 1As shown, an edge-assisted interactive application consists of three parts distributed across different locations: application logic running on a cloud server, a rendering engine running on an edge server, and display and control components running on the user device. These three parts interact with each other. Specifically, the rendering engine receives rendering instructions from the application logic, executes the instructions, and transmits the rendered frames to the user device. The user device generates control instructions and sends them back to the cloud server, which then updates the application logic upon receiving them. The application logic can also be offloaded to the edge server, and the rendering task scheduling problem remains the same in this case.
[0005] For a rendering task, once its resource requirements and quality of service (QoS) are set, the provider allocates it to an edge server. To improve resource utilization, multiple rendering tasks run simultaneously on a single edge server and share the same processor, with each task providing rendering services for a specific interactive application. Rendering tasks compete for computing resources. The execution of each task is scheduled by a scheduler to achieve a preset QoS. Therefore, the technical problem to be solved is: how to schedule rendering tasks so that all tasks achieve good performance while simultaneously meeting QoS requirements.
[0006] In existing technologies, rendering task scheduling schemes mainly suffer from the following drawbacks: 1) Shared GPU scheduling methods are suitable for general applications but not for interactive applications. Therefore, these methods only consider resource utilization and do not consider the performance of interactive applications; 2) Existing methods only make decisions on scheduling tasks and do not involve resolution decisions; 3) For edge-assisted or cloud-assisted interactive applications, existing methods improve user experience quality or reduce latency from aspects such as rendering mechanisms, rendering instruction compression, and compression parameter selection, but do not optimize performance through the scheduling of multiple rendering tasks.
[0007] For example, a simple method for scheduling multiple rendering tasks is to schedule them in a round-robin fashion, using a fixed resolution. Since the tasks all execute at the same frequency, this results in all tasks receiving a fixed frame rate. First, different tasks may have different frame rate requirements, so the frame rate requirements of some tasks cannot be met. Second, this method does not fully utilize computational resources. Therefore, this method cannot achieve the optimization goal of simultaneously deciding on the scheduling tasks and the resolution used. Furthermore, increasing the resolution increases both processing and transmission time, and excessively increasing the resolution may violate latency requirements. Therefore, the scheduling algorithm must consider the performance trade-offs between each pair and attempt to make the optimal choice. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of the prior art and propose a scheduling method for multiple rendering tasks for edge-assisted or cloud-assisted interactive applications. This method simultaneously decides the scheduling tasks and the resolution used, so that the service quality of each task meets the requirements, while maximizing the performance of multiple applications.
[0009] The technical solution of this invention is to provide a service quality-aware multi-rendering task scheduling method for interactive applications. The method includes the following steps:
[0010] Modeling multiple rendering tasks as a task scheduling problem, which involves selecting the task to be rendered and its resolution to meet quality of service requirements and maximize the minimum utility among all tasks, is expressed as:
[0011]
[0012]
[0013]
[0014]
[0015] in, It is the defined utility function. and They represent tasks s respectively i The required resolution, maximum tolerable frame interval, and maximum tolerable latency, m i Indicates task s i The number of instructions executed, r ij h represents the resolution selected to execute the j-th instruction. ij and d ij This indicates the frame interval and delay obtained after the instruction is executed;
[0016] The task scheduling problem is solved through a multi-round interaction between a resolution adjustment algorithm and a frame rate fair scheduling algorithm, where the resolution adjustment algorithm is used to select the resolution for the task, and the frame rate fair scheduling algorithm is used to determine the task to be processed.
[0017] Compared with existing technologies, the advantages of this invention are that it considers the performance of interactive applications, making decisions on both task scheduling and resolution selection; it can significantly improve task resolution and frame rate under low load and with sufficient computing resources, thereby improving task performance; it considers the minimum frame rate requirement and increases the probability of meeting the minimum frame rate requirement through reasonable modeling and method design; and it fully balances the conflicts between performance by designing a utility function, so that multiple performance requirements can be met simultaneously as much as possible, thereby increasing the probability of meeting all quality of service requirements.
[0018] Other features and advantages of the invention will become clear from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.
[0020] Figure 1 This is a schematic diagram of the system architecture for edge 3D rendering according to an embodiment of the present invention;
[0021] Figure 2 This is a schematic diagram illustrating the use of a heuristic algorithm to solve for scheduling sequences according to an embodiment of the present invention;
[0022] Figure 3 This is a schematic diagram of the empirical cumulative distribution function of processing time samples rendered at various resolutions according to an embodiment of the present invention.
[0023] Figure 4 It is a histogram of delay limits according to an embodiment of the present invention;
[0024] Figure 5 It is the minimum utility value including a penalty term under various computational loads according to an embodiment of the present invention, wherein Figure 5 (a) Corresponds to low load. Figure 5 (b) Corresponds to high load;
[0025] Figure 6 According to an embodiment of the present invention, the minimum utility value without penalty terms under various computational loads is, wherein Figure 6 (a) Corresponds to low load. Figure 6 (b) Corresponds to high load;
[0026] Figure 7 This is a schematic diagram illustrating the penalty for frame interval variation under various computational loads according to an embodiment of the present invention;
[0027] Figure 8 This is a schematic diagram showing the percentage of instances that meet the quality of service under various computing loads according to an embodiment of the present invention.
[0028] Figure 9 According to an embodiment of the present invention, when each method is combined with a resolution adjustment algorithm, the minimum utility value including the penalty term is obtained under various computational loads, wherein Figure 9 (a) Corresponds to low load. Figure 9 (b) Corresponds to high load;
[0029] Figure 10This is a schematic diagram illustrating the utility gain resulting from combining each method of the present invention with a resolution adjustment algorithm according to an embodiment of the present invention;
[0030] Figure 11 This is a schematic diagram showing the percentage of instances that satisfy the Quality of Service (QoS) when combined with and without the resolution adjustment algorithm according to an embodiment of the present invention. Detailed Implementation
[0031] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the invention.
[0032] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.
[0033] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0034] In all the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0035] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0036] This invention proposes a quality-of-service (QoS)-aware multi-rendering task scheduling method. This method makes real-time decisions on which resolution to use for which task, ensuring that user QoS requirements are met while maximizing user performance. This invention models the multi-task scheduling problem as a QoS-constrained minimum-maximum utility problem and proposes an efficient scheduling method to solve it. The provided scheduling method consists of a resolution adjustment algorithm and a frame rate fair scheduling algorithm. The resolution adjustment algorithm intelligently selects the resolution for each task, while the frame rate fair scheduling algorithm determines which task to process.
[0037] For clarity, the following sections will detail service quality; task scheduling modeling, including the service quality-constrained maximum and minimum utility problem, utility function, and penalties for frame interval changes; and overall scheduling methods, including resolution adjustment methods and frame rate fair scheduling methods. The resolution adjustment method and the frame rate fair scheduling method require solving two sub-problems: the weighted maximum and minimum frame rate problem and the scheduling sequence problem. The weighted maximum and minimum frame rate problem and its solution methods are also detailed below.
[0038] I. Service Quality
[0039] The following description uses the example of considering three performance parameters for interactive applications: resolution, frame rate, and latency. Resolution and frame rate are two key issues for video applications. High resolution provides better image quality, while high frame rates result in smooth user interaction. Latency is crucial for the user experience of interactive applications. Low latency ensures responsiveness to user interactions. The following discussion only considers latency related to rendering scheduling, i.e., the delay between the arrival of a rendering instruction in the system and the user receiving the corresponding rendered frame. There is a trade-off between resolution and latency. High resolution increases rendering and transmission time, thus increasing latency. There is also a trade-off between resolution and frame rate. High resolution reduces scheduling frequency, thus reducing the frame rate. The quality of service (QoS) target mentioned in this invention refers to the minimum requirements of these three performance parameters.
[0040] II. Modeling the Task Scheduling Problem
[0041] For a given task, a utility function can be used to quantify its performance. Since there are multiple tasks to optimize, we will use maximizing the minimum utility as an example, a concept known as minimax modeling. The basic principle of minimax modeling is to strive to meet the requirements of all users and provide additional performance (e.g., better frame quality) when computational resources are sufficient. Given n tasks on a server, where each task either has one instruction or no instruction once the server is idle, the task scheduling problem is to select an instruction to render (corresponding to a specific task) and a resolution to use, such that all quality-of-service (resolution, frame rate, and latency) requirements are met, and the minimum utility among all tasks is maximized. This is a quality-of-service-constrained minimax utility problem.
[0042] 1) The problem of maximum and minimum utility with limited service quality
[0043] For a task, given its required resolution r min And the obtained resolution r, its performance is modeled as Here, u(x) is a concave, non-decreasing utility function. If x is greater than 1, u(x) is positive; otherwise, it is negative. Similarly, for the maximum tolerable delay d...max Model its performance as However, for the required minimum frame rate f min Use its reciprocal, i.e., the frame interval. To measure performance, the frame interval is the time elapsed between two consecutive frames, while the frame rate is a statistical measure over a period of time. The frame interval has stricter constraints because a constant frame interval ensures that the frames received by the user are evenly distributed over time, while a constant frame rate does not. Therefore, for frame rate requirements, its performance is modeled as... Where h is the measured frame interval. In one embodiment, the overall performance is defined as a weighted sum, expressed as:
[0044]
[0045] Where, θ r θ h and θ d The weights used to balance each performance level are all non-negative and sum to 1.
[0046] Given n tasks, use {s1, s2, ..., sn} n Let m represent the process. Each task receives a series of instructions, some of which are executed and others are discarded. For those instructions that are executed, the resulting rendered frames are encoded and transmitted over the network. However, congestion can occur in both stages, potentially leading to frame drops. Therefore, the performance is evaluated using frames received by the user. Let m i Indicates task s i The number of instructions executed, and the rendered frames for these instructions successfully received by the corresponding users. To execute the j-th instruction, a resolution r is selected. ij After the instruction is executed, the frame interval h can be obtained. ij and delay d ij The instantaneous utility can be obtained through formula (1), i.e., u(r) ij ,h ij ,d ij In one embodiment, the utility of a task is defined as the average utility of all its executed instructions, denoted by u. i This means, that is:
[0047]
[0048] However, this utility function does not account for fluctuations in frame intervals, which could negatively impact the user experience. To address this, in another embodiment, a penalty term v is added to the utility function. i The final utility is defined as follows:
[0049]
[0050] Where φ is the weighting parameter. It should be noted that u i and All can be used as utility functions, unless otherwise specified. In the following description, the default is to use... Let's take an example to illustrate.
[0051] Specifically, the multi-rendering task scheduling problem is modeled as follows:
[0052]
[0053]
[0054]
[0055]
[0056] in and They represent tasks s respectively i The required resolution, maximum tolerable frame interval, and maximum tolerable latency are defined. These three constraints in the modeling ensure average performance for each type. For the modeled problem and a scheduling policy, if all three constraints are satisfied, the quality of service is said to be satisfied; otherwise, it is said to be violated.
[0057] 2) Definition of utility function u(x)
[0058] To effectively constrain the three average performance characteristics mentioned above, in one embodiment, a utility function is defined, generally expressed as:
[0059]
[0060] Where α is a non-negative constant used to adjust the penalty. When x>1, u(x) is in (0,1]; when x=1, u(x)=0; when x<1, u(x) is in [-∞,-α).
[0061] The advantages of this utility function compared to the traditional log(x) are mainly reflected in the following aspects: First, its upper bound is 1, rather than increasing to infinity. This helps to balance the trade-offs among the three performance characteristics. Second, the utility function has two segments, where the parameters of the negative segment can be adjusted so that any violation of performance characteristics is sufficiently penalized, thus avoiding the problems that arise from using the natural logarithm as the utility function, where excessively improving one performance characteristic degrades another. In other words, by setting the parameter α, it is possible to ensure that u(r,h,d)<0 holds true when any type of performance characteristic is violated.
[0062] Preferably, there are three functions. and Set α separately, and use α separately. r α h and α d The specific settings are as follows:
[0063] α r =σ(θ) r ),α h =σ(θ) h ), and α d =σ(θ) d (6)
[0064] in:
[0065]
[0066] 3) Penalty for frame interval changes
[0067] Variations in frame intervals also have a significant impact on user experience. For a series of frames received by a user, smaller fluctuations in frame intervals result in a better user experience. Sudden increases in frame intervals can disrupt user interaction. For example, relative standard deviation can be used to model variations in frame intervals. Relative standard deviation is the ratio of the standard deviation to the mean.
[0068] For task s i Suppose the user receives a series of frames, the number of which is m. i Let h ij This represents the frame interval at the instant after the user receives the j-th frame. Then, task s... i The average frame interval is
[0069]
[0070] The average of the squared frame intervals is
[0071]
[0072] Let v i Indicates task s i The relative standard deviation of the frame interval is then:
[0073]
[0074] III. Task Scheduling Algorithms
[0075] In one embodiment, an algorithm is proposed to solve the multi-task scheduling problem, consisting of a resolution adjustment algorithm and a frame rate fair scheduling algorithm. The resolution adjustment algorithm intelligently selects a resolution for each task, while the frame rate fair scheduling algorithm determines which task to process. The two interact with each other. Once the server is idle, the frame rate fair scheduler selects a task and executes its instructions using the resolution determined by the resolution adjustment algorithm. After multiple rounds of scheduling, the utility is evaluated and fed back to the resolution adjustment algorithm, which then decides how to update the resolution arrangement.
[0076] IV. Resolution Adjustment Algorithm
[0077] For each task, there are a total of k resolutions available. Low resolutions may not meet the resolution requirements, while high resolutions may increase processing and transmission times, leading to increased latency and potentially violating latency constraints. Therefore, when there is only one task, its utility is a concave function of resolution. However, when multiple tasks compete with each other, how the resolution decisions of each task affect the final utility (the minimum among all tasks) is very complex.
[0078] Given n tasks and k resolutions, each task chooses one resolution, resulting in n k There are a number of permutations and combinations, each representing a resolution arrangement. A simple approach is to try them all and choose the best one; however, this method requires O(n log n) time complexity. k Furthermore, trying too many poor arrangements degrades overall performance. To address these issues, in one embodiment, the resolution is gradually increased from an initial arrangement until overall utility no longer improves, thereby minimizing poor arrangements.
[0079] For such an algorithm, there are three important design problems: the initial resolution, how to improve a resolution arrangement, and the convergence condition. These three design problems will be discussed in detail below. First, the initial resolution for each task should be the resolution required by the user, thus satisfying the resolution requirement. When computational resources are sufficient, improving the resolution of some tasks may improve overall performance. Among all tasks, the task with the lowest utility should be prioritized, as it limits overall performance. Therefore, a resolution arrangement is improved by increasing the resolution of the worst-performing task by one level. By continuously increasing the resolution until the overall utility decreases due to insufficient computational resources, the algorithm can be considered converged. However, randomness in the operating system may cause performance fluctuations, misleading the algorithm to converge prematurely. To tolerate this uncertainty, convergence is only triggered when utility decreases significantly. It should be noted that attempts should continue when utility remains unchanged, because multiple tasks may simultaneously have the lowest utility; in such cases, increasing a single resolution may not improve the lowest performance, while increasing all resolutions may.
[0080] Another key issue is when to adjust the resolution. To evaluate the performance of a resolution arrangement as accurately as possible, it must be held for a sufficiently long time before being changed. Furthermore, video coding complicates the issue. Rendered frames are typically encoded in groups of pictures (GOPs). Each GOP consists of one independently coded intra-frame and many inter-coded frames encoded based on previous frames within the same GOP. A GOP should use the same resolution; otherwise, common video coding techniques (such as H.264, MPEG-4, etc.) cannot support it. Therefore, for each task, its resolution can only be adjusted when a GOP is completed. Considering both of these aspects, in one embodiment, initiating a resolution adjustment requires simultaneously meeting the following two conditions: 1) sufficient time has elapsed since the most recent adjustment; 2) the task to be adjusted has completed encoding a GOP. In the following simulation experiments, the GOP length is 64 frames, and the holding time for a resolution arrangement is set to the time required for all tasks to render a total of 1024 frames.
[0081] Specifically, the resolution adjustment algorithm includes the following steps:
[0082] Step S11: Let vector r represent the current resolution arrangement set, where each element represents the resolution of a task, and initialize it with the initial arrangement r. min (The resolution requested by the user). Let r * This represents the optimal resolution arrangement found so far, and its corresponding overall utility (in terms of...). (Indicates) the highest.
[0083] Step S12, if among all tasks, there is one task s k If both of the following conditions are met, then task s k A resolution adjustment was initiated: ① Sufficient time has passed since the last adjustment; ② Tasks s k The encoding of a group of images has been completed.
[0084] Step S13: Calculate the average utility of each task since the most recent resolution adjustment (the most recent change in r) according to formula (3). Will The minimum value in the total utility Overall utility Used to evaluate the performance of the current resolution arrangement r.
[0085] Step S14: If the algorithm has converged, arrange r according to the currently found optimal resolution. * Tasks k Adjust the resolution to the optimal value. Return to step S12.
[0086] Step S15, if the resolution adjustment task is started... k The utility is not the lowest, so we return to step S12. The reason for not increasing its resolution is that this would not improve the lowest utility.
[0087] Step S16, record optimal performance information: if Compare Large, then Become Optimal resolution arrangement r * It becomes r.
[0088] Step S17: If the following three conditions are met simultaneously, then by completing task s... k To improve r by increasing the resolution by one level: ① Compared to No significant decrease (i.e.) Where ξ≥0, the parameter ξ prevents the algorithm from converging prematurely due to fluctuations in utility; ② The average performance obtained using the current resolution arrangement r can satisfy the three constraints of formula (4b-4d); ③ Task s k There is still room for improvement in resolution. Return to step S12.
[0089] Step S18: If the three conditions in step S17 cannot be satisfied simultaneously, the algorithm converges, and r is arranged according to the currently found optimal resolution. * Tasks k Adjust the resolution to the optimal value. Return to step S12.
[0090] In summary, the proposed resolution adjustment method can find the resolution arrangement that maximizes overall utility from an exponentially large number of resolution arrangements in a short period of time.
[0091] V. Frame Rate Fair Scheduling Algorithm
[0092] Given n tasks {s1, s2, ..., sn} n Tasks can be scheduled in a round-robin fashion, resulting in similar rendering frequencies and frame rates for each task. However, this approach fails to achieve optimal fairness when tasks have different frame rate / frame interval requirements, and some tasks with higher requirements may violate frame interval constraints. Instead, a weighted maximum-minimum fairness in terms of frame rate is achieved by optimizing the scheduling frequency of each task; this is known as the weighted maximum-minimum frame rate problem. If the frame rate of a task cannot be increased without reducing the frame rate of another task that already has a lower frame rate, then the frame rate allocation is maximum-minimum fair.
[0093] The goal of this invention is to achieve fairness in long-term frame rates, but the scheduling frequency of a task is only planned for the next cycle, which determines the short-term frame rate. The relationship between the short-term scheduling frequency and the long-term frame rate can be derived. For task s... i , let f i This indicates the expected scheduling frequency in the next cycle, the duration of which will be determined later. Assume the average frame rate (i.e., long-term frame rate) achieved by the task so far is... So, what is the long-term frame rate (denoted by x) that the task will achieve after the next cycle? i (Indicates) to become:
[0094]
[0095] Where β is a constant parameter. Therefore, the long-term frame rate ratio is: The weighted maximum-minimum frame rate problem is to find the weighted maximum-minimum fairness vector x on a feasible set, which will be analyzed below. The weighted maximum-minimum fairness is defined as follows: given some positive weights... A vector x is a weighted maximum-minimum fair vector on a feasible set if and only if an additional component x is added to the vector x. s It will definitely reduce some other component x t , making
[0096] After obtaining the weighted maximum-minimum fairness vector x, the short-term frame rate vector f can be obtained through formula (11). To ensure that if task s... i short-term frame rate f i If the value is not zero, it will be scheduled at least once, and the duration of the next cycle will be set to [value missing]. Therefore, task si The number of scheduling times in the next cycle (using Q) i (represented as):
[0097]
[0098] Once the number of scheduling attempts for each task is obtained, tasks can be scheduled cyclically based on these quantities. However, this scheduling method may lead to large fluctuations in frame intervals, resulting in performance degradation. To smooth out changes in frame intervals, the order of task scheduling needs to be optimized, called the scheduling sequence. This can be modeled as a vector, where the j-th element represents task s. k This indicates that task s k This will be processed in step j. For example, given three tasks s1, s2, and s3, assuming that in a plan, the three tasks have 1, 2, and 3 instructions to execute respectively, the vectors (s1, s2, s3, s2, s3, s3) and (s1, s3, s2, s3, s2, s3) are two scheduling sequences. Assume that the time to process one instruction is the same for each task. Then, for task s3, there are two frame interval samples under the first scheduling sequence: 1 and 0; while under the second scheduling sequence, there are two frame interval samples: 1 and 1. The second scheduling sequence is better than the first in terms of frame interval variation. In one embodiment, the scheduling sequence problem is defined as: given a set of instructions to be executed in one cycle, optimize the scheduling sequence of these instructions such that the penalty for frame interval variation is minimized.
[0099] The weighted maximum / minimum frame rate problem and the scheduling sequence problem are two core problems in scheduling. Once these two problems are solved, the scheduling algorithm becomes easier to design.
[0100] Specifically, the frame rate fair scheduling algorithm includes the following steps:
[0101] Step S21, for each task s i , making D i This represents the deficit value and is initialized to 0. The iterator next of the scheduling sequence vector is initialized to 1. The total number of instructions planned to be executed in the current cycle, m, is initialized to 0.
[0102] Step S22: When the processor is idle and there are instructions in the queue of at least one task, perform the following steps S23, S24 and S25 to schedule a task.
[0103] Step S23: If the iterator next is greater than m, then one traversal of the scheduling sequence vector is completed, and a new cycle begins. The following steps are performed to obtain a new scheduling sequence vector:
[0104] Step S23.1: Solve the weighted maximum and minimum frame rate problem to obtain a scheduling frequency vector f;
[0105] Step S23.2, based on f, for each task s i The scheduling count Q for the next cycle can be obtained using formula (12). i In its deficit value D i Accumulated Q i And round down to the nearest integer D i Get m i This refers to the number of instructions planned to be executed. The total number of instructions planned to be executed in the current cycle, m = ∑ i m i ;
[0106] Step S23.3, transform vector m (whose i-th element is m) i Using as input, solve the scheduling sequence problem to obtain a scheduling sequence vector S;
[0107] Step S23.4: Reset the iterator next of vector S to 0.
[0108] Step S24: Traverse vector S using the iterator next, obtaining the corresponding task s for each element traversed. idx The iterator next increments by 1 until task s. idx The queue contains instructions and its deficit value D idx The iteration ends only when the value is not less than 1.
[0109] Step S25, schedule task s idx The corresponding deficit value D idx Decrease by 1. Return to step S22.
[0110] In summary, the proposed frame rate fairness scheduling method is used to decide the next scheduling task. By optimizing the scheduling frequency and the scheduling sequence, it can optimize the performance of all rendering tasks. Optimizing the scheduling frequency can solve for the short-term scheduling frequency of each task over a period of time, achieving weighted maximum and minimum fairness from the perspective of frame rate for each task. Optimizing the scheduling sequence can sort multiple instructions belonging to multiple tasks to be executed over a period of time, with the aim of smoothing out changes in the frame interval of tasks.
[0111] VI. Modeling and Algorithm for the Weighted Maximum and Minimum Frame Rate Problem
[0112] We model the weighted maximum-minimum frame rate problem. First, we model the feasible set of x. The short-term frame rate f... i It is non-negative and must not exceed the instruction arrival rate (denoted by λ). i (represented by f), that is, 0≤f i ≤λ i Using formula (11), we can obtain:
[0113]
[0114] The arrival rate of instructions can be estimated through statistical analysis. The time consumed in executing all instructions is ∑ i t i f i , where t i It is task s i The average time to process each instruction, f i This represents the short-term frame rate. Let the duration of one cycle be 1 second, then we get:
[0115] ∑ i=1…n t i f i ≤1 (14)
[0116] Using formula (11), we can obtain:
[0117]
[0118] Therefore, the feasible set of x is:
[0119]
[0120]
[0121] It is a compact convex set. The goal is to obtain a weighted maximum-minimum fair allocation x on the set (16), where x i The weight is According to the maximum-minimum fairness theory, x exists and can be found using the water-filling algorithm.
[0122] To illustrate the algorithm more clearly, the problem above is transformed into a concise form. Let y i =t i x i The problem then becomes equivalent to finding the weighted maximum minimum fair allocation y within the following feasible set.
[0123]
[0124] Among them, capacity lower limit upper limit y i The weight is
[0125] Specifically, optimizing the scheduling frequency algorithm includes the following steps:
[0126] Step S31, the input to the algorithm includes f min λ and t are respectively composed of λi and t i The vector is composed of these inputs. The following parameters are obtained from these inputs: capacity C, lower bound L. i , upper limit U i y i weight W i .
[0127] Step S32, the ratio of each task (i.e., y) i Initialize it to its corresponding lower bound L i Let C′ represent the remaining capacity, initialized to... Let I represent the set of tasks for which the ratio can be further improved, initialized as {1,2,…,n}.
[0128] Step S33: Find the set with the smallest weighted ratio (i.e., y) in set I. i / W i A subset of tasks (denoted by V1) is denoted as...
[0129] Step S34: Increase the weighting ratio of tasks in I′ until one of the following occurs: ① the remaining capacity is exhausted; ② the weighting ratio reaches the second minimum value in I. In case ①, the weighting ratio of the tasks can reach a value of... In case ②, the weighted ratio that the task can achieve is the second smallest weighted ratio in I, denoted by V2. When I′=I, V2 does not exist. Therefore, the target weighted ratio (denoted by V) that the task can achieve is the minimum value achievable in both cases.
[0130] Step S35, the weighting ratio y for each task in I′ i Become VW i With upper limit U i The minimum value between.
[0131] Step S36, update the remaining capacity C′ to And remove from I the value that reaches its upper limit U. i The task is to repeat steps S33-S36 until the remaining capacity C′ is not greater than 0 or the set I is empty.
[0132] Step S37, for each task, obtain y i Then, we can get x. i =y i / t i And f is obtained according to formula (11) i .
[0133] Step S38: Return the scheduling frequency vector f.
[0134] VII. Modeling and Algorithms for Scheduling Sequence Problems
[0135] The modeling of the scheduling sequence problem is as follows: Given m instructions to be executed in one cycle, where m instructions belong to task s i The instructions include m i The duration of one cycle is T = ∑ i m i t i , where t i It is task s i The average time to process each instruction. The sorting problem is to schedule m instructions within a period [0, T] such that the penalty for frame interval variation is minimized. The objective is:
[0136] m i nim i ze max i=1,...,n v i (18)
[0137] Where v i The task s is defined in formula (10). i The relative standard deviation of the frame interval.
[0138] In the task scheduling modeling described in Part II above, the frame interval was measured at the user end. However, in the scheduling sequence problem, the execution of instructions is planned rather than implemented, so the frame interval cannot be obtained in advance. Instead, the frame interval is approximated using the interval between the planned start times of two instructions.
[0139] This problem is a combinatorial optimization problem. In one embodiment, a heuristic algorithm is proposed to solve it. The main idea of the algorithm is to distribute m as uniformly as possible within a period. i One instruction to smooth out changes in frame intervals. Specifically, for task s i Try to arrange its m i The start time of each instruction is such that the interval between two consecutive starts is approximately equal to...
[0140]
[0141] Let o ij Indicates task s i The j-th instruction, τ ij Indicate its start time. Then, use τ. ij -τi ,j-1 To approximate the frame interval h ij For the instruction o ij After scheduling its start time τ ij Then, its end time can be approximated as τ. ij +t i Therefore, [τ] ij,τ ij +t i To become a busy interval, it must be removed from the interval [0, T]. Note that the intervals mentioned in this section are left-closed and right-open. After instructions are scheduled, the interval [0, T] becomes an ordered set of free intervals. These free intervals are separated by a series of scheduled busy intervals. This represents the set of free intervals, where and These are the start and end times of the kth free interval, respectively.
[0142] Given a set of free intervals E, for instruction o ij Determine the start time τ ij First, ensure that the previous instruction is followed. i,j-1 At least g has elapsed since startup. i ,Right now:
[0143] τ ij ≥τi ,j-1 +g i (20)
[0144] Next, try scheduling instructions in idle periods rather than busy periods to minimize changes to already scheduled instructions. In the following two cases, instruction o... ij It can be arranged in the free space Internal: 1) When τ i,j-1 +g i When included in this free interval, instructions can be scheduled within that interval, and τ ij Let it be τ i,j-1 +g i ;2) When the start time of the free interval is later than τ i,j-1 +g i At the same time, instructions can also be assigned to this interval, and τ ij Set as Therefore, it can be represented as:
[0145]
[0146] For a given instruction, there may be several available free intervals, with a selection count of O(m). An interesting question is which one to choose. Different choices may yield different target values (18). Generally, the earliest interval yields the highest target value because it minimizes the change in the current task's frame interval. Therefore, the earliest interval can be directly selected. Alternatively, each interval can be tried, and the interval that best yields the target value can be selected, but this is time-consuming. Simulation experiments compared these two approaches, revealing that their performance is very similar.
[0147] A new instruction arrangement may affect existing arrangements. Suppose an instruction o... ij Inserted into free space Within this range, its busy zone is [τ] ij ,τ ij +t i If the busy interval exceeds the idle interval, that is... The new busy zone will inevitably overlap with other existing busy zones, therefore some existing arrangements must be adjusted to avoid this overlap. Specifically, for example... Figure 2 As shown in (a), the start time is later than The interval (busy interval and idle interval) offset Δ, where:
[0148]
[0149] Therefore, we have:
[0150]
[0151] Similarly, offset free interval As shown below:
[0152]
[0153] Once an instruction is scheduled, the set of free intervals E must be updated accordingly. In addition to the offsets mentioned above, the scheduled free intervals... It becomes:
[0154]
[0155] like Figure 2 As shown in (b), if the new busy interval [τ] ij ,τ ij +t i Completely contained within the free interval, i.e. The free space is then divided into and Two parts; otherwise, such as Figure 2 As shown in (a), the original free interval is truncated and becomes
[0156] Specifically, optimizing the scheduling sequence algorithm includes the following steps:
[0157] Step S41, the algorithm input includes m, t, and τ0, which are respectively determined by m i t i and τ i0 The vector formed by τ i0 It is task s iThe start time (relative to the current time) of the most recently executed instruction (the instruction preceding the first scheduled instruction) is a negative value. From these inputs, the duration of a cycle can be obtained as T = ∑ i m i t i Add the interval [0,T] to the set of free intervals E, and calculate the value of each task s using formula (19). i Target average frame interval g i .
[0158] Step S42, according to the number of instructions in the task (i.e., m) i The tasks are sorted in descending order of their parameters. Each task will then be processed sequentially according to this order. This is because the more instructions a task has, the higher its target average frame interval g becomes. i The lower the value, the more unstable the frame interval.
[0159] Step S43: Based on the task sequence in step S42, for each task s in the sequence... i Determine its m in sequence i The start time of each instruction.
[0160] Step S44, for task s i Each instruction o ij The earliest available free interval that can be used to schedule the instruction is obtained from the free interval set E, denoted as .
[0161] Step S45, calculate instruction o using formula (21) ij start time τ ij Let τ represent the expression by τ. i The start time matrix is composed of j. τ is updated using formula (23), that is, the start time of the previously scheduled instructions is updated.
[0162] Step S46: Offset the free interval according to formula (24), remove the new busy interval according to formula (25), that is, update the free interval set E.
[0163] Step S47: After all instructions for all tasks have been scheduled, obtain the scheduling sequence S based on the latest τ and return it.
[0164] To further verify the effectiveness of the present invention, a simulation experiment was conducted. To make the simulation experiment more realistic, tracking data was used to generate the simulation environment. The specific settings are as follows.
[0165] 1) Collection of tracking data
[0166] The game engine Unity was used to render 3D animated scenes at different resolutions and track processing times. It was run on an Nvidia GeForce GTX 1060. A higher-performance GPU wasn't used because Unity couldn't provide sufficient accuracy to track very short processing times. Instead, a mid-performance GPU was used to collect data, and then the processing time was reduced by a factor to simulate rendering on a high-performance GPU. The factor used was 0.3.
[0167] Ideally, each frame should be rendered at multiple resolutions, and the timing of each rendering operation should be recorded. However, it has been found that this approach makes it difficult to obtain the accurate time for each rendering operation due to Unity's numerous built-in caching and threading mechanisms. Instead, based on experience, the rendering time for a frame is found to be approximately linearly related to the resolution. Therefore, the scene is first rendered at a selected resolution, called the baseline resolution, and then the rendering time for other resolutions is obtained by scaling its rendering time. In the following simulation experiment, the candidate resolution set is {1920×1080, 2560×1440, 3072×1728, 3840×2160}, with 2560×1440 selected as the baseline resolution. The scaling factors for the four resolutions are set to 0.73, 1.0, 1.37, and 2.18, respectively. The scaling factors are obtained empirically. The scene is first rendered independently at different resolutions, and the resulting data is called raw data. Then, the median rendering time for each resolution is divided by the median of the baseline resolution to obtain the corresponding scaling factor. Figure 3 The empirical cumulative distribution function (CDF) for processing time is shown, with generated data displayed as solid lines and original data as dashed lines. It can be seen that for each resolution other than the baseline resolution (2560×1440), the generated data and the original data have similar statistical characteristics.
[0168] 2. Service quality requirements
[0169] Trade-offs should be considered when determining service quality requirements. Here, a simple approach is introduced. First, the trade-off between modeling resolution and latency is expressed as:
[0170] d max ≥t p (r min )+t x (r min (26)
[0171] Where t p (r) and t x (r) represents the average processing time and transmission time when using resolution r, respectively. Otherwise, the latency limit dmax This will not be satisfied. Modeling Where b is the number of bits per pixel, c is the compression ratio, and B is the bandwidth. Second, the trade-off between frame rate and latency is modeled as follows:
[0172]
[0173] In this conservative estimate, for the sake of simplicity, the parallelism of processing and transmission is ignored.
[0174] The quality of service requirements are generated as follows. First, a bandwidth value B is randomly selected from the corresponding candidate set that follows a uniform distribution, and the frame rate requirement f is... min Resolution requirement r min These candidate sets are shown in Table 1. Second, do not include d. max Instead of setting it to a single value (in milliseconds), set it to:
[0175]
[0176] d max The higher the value, the greater the chance of improving resolution. Figure 4 The statistical data of latency limits in the simulation experiment are shown. There are a total of 1349 samples, with a median of 40 milliseconds and an average of 46.3 milliseconds. Third, the service quality requirements that meet the above conditions (26) and (27) are screened. The settings of parameters b and c are shown in Table 1. Finally, h max Set to 1 / f min Second.
[0177] Table 1 shows the parameter settings for the simulation environment.
[0178]
[0179]
[0180] In Table 1, Unif{a,b} represents a discrete uniform distribution between a and b.
[0181] 3. Task allocation
[0182] The task allocation is simulated to generate a group of tasks, called a task group, allocated to a single server. The computational load of a task is defined as the proportion of time per second spent processing it. For a task, given f... min and r min Let L represent the computational load, then we have:
[0183] L = f min ·t p (r min (29)
[0184] Load is essentially the minimum computing power required for a preset quality of service. Assume n tasks are allocated to a server, each task having a bandwidth of B. i Service quality requirements and load L i Then there are two constraints on bandwidth and load:
[0185]
[0186] Among them B max and L max These are bandwidth constraints (set to 100 megabits per second in the simulation) and load constraints (set to 1 in the simulation). Other resources can be modeled in the same way. For simplicity, we assume isomorphic settings, so other resource constraints can be translated into constraints on the number of tasks. For example, we can limit the number of tasks to between 2 and 4, because scheduling a single task is negligible, while the total load of 5 tasks far exceeds the load constraint.
[0187] The task groups are generated as follows: First, generate the possible quality of service requirements (along with bandwidth). Second, generate combinations of these quality of service requirements and filter those that satisfy formula (30). Here, each combination corresponds to a task group. Third, select some of these task groups for simulation experiments based on their total load. Specifically, given a set of loads, i.e., {30%,...,100%}, for each load, randomly select task groups with loads close to it (allowing ±1.5% fluctuation). Here, 30% is the lowest possible load.
[0188] 4. Generate instructions
[0189] For a given task, a series of instructions are generated, with an arrival rate higher than the required frame rate. Given the required frame rate f... min The arrival interval obeys the following Seconds and The data is uniformly distributed across seconds. Command arrival was simulated over 120 seconds. The processing time for each command was read sequentially from the trace data stream, which is a concatenation of multiple randomly selected segments (each segment containing 5000 samples) of the original trace data, generated according to the method described above. The random combination of trace data segments allows for testing of more variations.
[0190] In addition, dynamic bandwidth was simulated. Three sets of bandwidth tracking data were collected using a continuous speed test tool, each set representing one of the candidate bandwidths (10, 20, and 30 megabits per second). For each set of data, it was first divided into 50 segments, each consisting of a bandwidth sample lasting 120 seconds. Then, an offset was added to each data segment so that its average value equaled the specified bandwidth (10, 20, and 30 megabits per second). The instantaneous bandwidth at the time of transmission for each frame was read sequentially from randomly selected data segments.
[0191] Furthermore, video encoding is simplified into a short process. The encoding time for each frame is randomly selected based on a uniform distribution between 2 and 6 milliseconds. The length of a group of images is 64 frames. The compression ratio is 1:x, where x is an integer value. For intra-coded frames, x follows a uniform distribution on [400, 600], and for inter-coded frames, x follows a uniform distribution on [800, 1200]. Other parameter settings for the simulation environment are shown in Table 1.
[0192] 5) Performance Evaluation
[0193] The method of this invention is evaluated through simulation experiments: Frame Rate Fair Scheduling (FRF-RA) combined with resolution adjustment algorithm. Unless otherwise specified, the relevant parameter settings are shown in Table 2.
[0194] The method proposed in this invention is compared with the following classic scheduling methods:
[0195] 1) Circular scheduling (RR): It iterates through tasks in a circular manner.
[0196] 2) First-Come, First-Served (FCFS): It tends to run the earliest arriving instruction first.
[0197] 3) Shortest Remaining Time First (SRTF): This prioritizes scheduling the most urgent tasks. For a task, its urgency is assessed using a deadline, which is the maximum tolerable frame interval for the task since the most recent scheduling.
[0198] In the following evaluations, each value is averaged from 50 different simulation instances. Each instance uses a task assignment, a processing time trace data segment, and a bandwidth trace data segment, all of which are independently and randomly selected.
[0199] Table 2 Algorithm Parameter Settings
[0200]
[0201] 6. Frame Rate Fair Scheduling Algorithm
[0202] First, evaluate the frame rate fair scheduling without incorporating resolution adjustment algorithms.
[0203] 1) Increased utility
[0204] The objective of this invention is to minimize the utility including the penalty term, i.e. Figure 5 This demonstrates the value of this utility under various computational loads. Figure 5 (a) Corresponds to low load. Figure 5 (b) Under high load. Frame Rate Fair Scheduling (FRF) can be observed to achieve the highest utility under almost all loads, while Shortest Remaining Time First Scheduling (SRTF) performs slightly worse than FRF. Round Robin Scheduling (RR) and First-Come, First-Served Scheduling (FCFS) perform poorly, especially under high load. Figure 6 This demonstrates the minimum utility without penalty terms, i.e., min i=1…n u i ,That Figure 6 (a) Corresponds to low load. Figure 6 (b) Corresponding to high load. It can be seen that Frame Rate Fair Scheduling (FRF) achieved the highest value under all load conditions. Furthermore, Figure 7 This demonstrates the penalty for frame interval variations, namely max. i=1…n v i It can be seen that the method of this invention (FRF) has almost the same performance as the shortest remaining time first scheduling (SRTF), and both are better than first-come-first-served scheduling (FCFS), but worse than round-robin scheduling (RR).
[0205] 2) Increased percentage of instances meeting service quality requirements
[0206] For a given problem instance, if all constraints (4b)-(4d) in modeling (4) are satisfied, it is said to satisfy Quality of Service (QoS); otherwise, it is said to violate QoS. Evaluate the percentage of instances that satisfy QoS (QoS-SAT) out of 50 simulation instances.
[0207] It should be noted that each method can satisfy the resolution constraint (4b), but may violate other constraints. Figure 8 The percentage of instances meeting Quality of Service (QoS) is shown. It can be seen that the method of this invention (FRF) achieves the highest value. In contrast, all other methods have a large number of instances violating QoS. Frame Rate Fair Scheduling (FRF) ensures that all instances meet QoS requirements even under 90% high load.
[0208] Furthermore, simulation results show that when the latency limit is set according to formula (28), all methods can satisfy the latency constraint (4d). In this case, all instances that violate the quality of service also violate the frame interval constraint (4c).
[0209] 7. Resolution Adjustment Algorithm
[0210] The resolution adjustment algorithm can be combined with any scheduling algorithm. The impact of the resolution adjustment algorithm on all methods was evaluated.
[0211] 1) Increased utility
[0212] The resolution adjustment algorithm was combined with each method, and the impact on utility was evaluated. For example... Figure 9 As shown, the method of the present invention (FRF-RA) exhibits the best performance under all loads. Figure 10 The utility gain of each method compared to the original method is shown. It can be observed that at loads below 50%, the utility of almost every method increases. Roughly speaking, the lower the load, the greater the gain due to more idle computational power.
[0213] 2) Impact on the percentage of instances that meet service quality standards
[0214] Next, we will evaluate the impact of the resolution adjustment algorithm on meeting the quality of service requirements. Figure 11 The percentage of instances that satisfy Quality of Service (QoS) is shown for each method, with and without the resolution adjustment algorithm. It can be seen that the resolution adjustment algorithm slightly reduces the performance in satisfying QoS under high load. This decrease comes at the cost of improved utility under low load, such as... Figure 10 As shown. This is mainly due to the attempts at resolution arrangement in the resolution adjustment algorithm.
[0215] In summary, compared with the prior art, the present invention has at least the following technical advantages:
[0216] 1) Compared with the scheduling methods that share the GPU in the existing technology, the present invention is designed for interactive applications and takes into account the performance of interactive applications, including resolution, frame rate and latency.
[0217] 2) Compared with the scheduling methods of shared GPUs in the prior art, the present invention makes decisions on both the scheduling task and the selection of resolution.
[0218] 3) For edge-assisted or cloud-assisted interactive applications, existing technologies improve user experience quality or reduce latency through rendering mechanisms, rendering command compression, and compression parameter selection, but do not optimize performance through the scheduling of multiple rendering tasks. This invention achieves performance optimization through frame-by-frame scheduling of multiple rendering tasks.
[0219] 4) Compared with polling-based scheduling methods that use fixed resolution, this invention can significantly improve task resolution and frame rate under low load and with ample computing resources, thereby enhancing task performance.
[0220] 5) Compared with polling-based scheduling methods that use fixed resolution, this invention takes into account the minimum frame rate requirement and improves the probability of meeting the minimum frame rate requirement through reasonable modeling and method design.
[0221] 6) Compared with polling-based scheduling methods that use fixed resolution, this invention, by designing a utility function u(x) and an effective method, fully balances the conflicts between performance, so that multiple performance aspects (resolution, frame rate / frame interval, latency) can meet the requirements simultaneously as much as possible, thereby increasing the probability of meeting all quality of service requirements.
[0222] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0223] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0224] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.
Claims
1. A service quality-aware multi-rendering task scheduling method for interactive applications, comprising the following steps: Modeling multiple rendering tasks as a task scheduling problem, which involves selecting the task to be rendered and its resolution to meet quality of service requirements and maximize the minimum utility among all tasks, is expressed as: The task scheduling problem is solved through a multi-round interaction between a resolution adjustment algorithm and a frame rate fair scheduling algorithm, where the resolution adjustment algorithm is used to select the resolution for the task and the frame rate fair scheduling algorithm is used to determine the task to be processed. in, It is the defined utility function. , and Representing tasks The required resolution, maximum tolerable frame interval, and maximum tolerable latency, Indicates task The number of instructions executed. In order to execute the first The instruction specifies the selected resolution. and This indicates the frame interval and delay obtained after the instruction is executed; The resolution adjustment algorithm performs the following steps: Step S10: Let vector This represents the current set of resolution arrangements, where each element represents the resolution of a task, initialized to the resolution requested by the user. ,make This represents the optimal resolution arrangement found so far, and its corresponding overall utility. Highest; Step S20: If among all tasks, there is one task If both of the following conditions are met, then the task... A resolution adjustment was initiated: the last resolution adjustment occurred after a set time threshold had elapsed; task Encoding of a group of images has been completed; Step S30: Calculate the average utility of each task since the most recent resolution adjustment. , Will The minimum value in the total utility Used to evaluate the current resolution arrangement Performance; Step S40: If the algorithm meets the set convergence conditions, arrange the algorithm according to the currently found optimal resolution. The task Adjust the resolution to the optimal value; Step S50: If the resolution adjustment task is initiated. If the utility is not the lowest, return to step S20; Step S60: Record optimal performance information: if Compare Large, then Become Optimal resolution arrangement Become ; Step S70: If the following three conditions are met simultaneously, the task will be completed. Increase the resolution by one level to improve : ,in Use the current resolution arrangement The obtained average performance satisfies all constraints in the task scheduling problem; task There is room for improvement in resolution; Step S80: If the three conditions cannot be met simultaneously, the algorithm is considered converged, and the algorithm is arranged according to the currently found optimal resolution. The task Adjust the resolution to the optimal value; The frame rate fair scheduling algorithm performs the following steps: Step S100, for each task ,make Represents its deficit value, initialized to 0, and is the iterator of the scheduling sequence vector. Initialize to 1, representing the total number of instructions planned to be executed in the current cycle. Initialize to 0; Step S200: When the processor is idle and there are instructions in the queue of at least one task, execute the following steps S300, S400 and S500 to schedule a task. Step S300, if the iterator Greater than This completes one traversal of the scheduling sequence vector and begins a new cycle. The following steps are then performed to obtain a new scheduling sequence vector: Solving the weighted maximum and minimum frame rate problem yields a scheduling frequency vector. ; based on For each task Calculate the number of scheduling cycles in the next cycle. In its deficit value Accumulation Round down get This gives the number of instructions planned to be executed, and the total number of instructions planned to be executed in the current period. ; vector The first element As input, the scheduling sequence problem is solved to obtain a scheduling sequence vector. ; Reset Vector iterator =0; Step S400, through iterator Traversing vectors For each element traversed, its corresponding task is obtained. , iterator Increment by 1 until the task is completed. The queue contains instructions and its deficit value The iteration ends when the value is not less than 1. Step S500, Schedule the task , The corresponding deficit value Decrease by 1 and return to step S200.
2. The method according to claim 1, wherein, For a given task, the utility function is set as follows: in, Overall performance is defined as: It is the minimum required resolution. That is the obtained resolution. This is the maximum tolerable delay. It is the required minimum frame rate. , It is the measured frame interval. It is a penalty term for changes in frame interval. These are weight parameters. , and It is the weight of the corresponding item.
3. The method according to claim 1, wherein, The weighted maximum and minimum frame rate problem is modeled according to the following steps: Modeling Feasible set, short frame rate It is non-negative and not greater than the instruction arrival rate. , represented as It satisfies the following formula: in, This indicates the long-term frame rate that the task will achieve after the next cycle. The instruction arrival rate was estimated through statistical analysis, and the time consumed in executing all instructions was... ,in It is a task The average time to process each instruction, with a cycle duration of 1 second, yields: Based on formula ,get: therefore, The feasible set is: , The goal of the weighted maximum minimum frame rate problem is to obtain a weighted maximum minimum fair allocation. ,in The weight is ; make Then the problem is equivalent to finding a weighted maximum-minimum fair allocation in the following feasible set. : Among them, capacity lower limit upper limit , The weight is , It is a constant parameter. That is the average frame rate.
4. The method according to claim 3, wherein, The scheduling sequence is used to optimize the order of task scheduling. The scheduling sequence problem is modeled according to the following steps: Given an execution within a cycle One instruction, of which belongs to the task The instructions include One, the duration of one cycle is ,in It is a task Average time to process each instruction; In the cycle Internal arrangements An instruction that minimizes the penalty for frame interval changes is expressed as: in It is a task The relative standard deviation of the frame interval is used, and the frame interval is approximated by the interval between the start times planned by the two instructions.
5. The method according to claim 2, wherein, , and Using the same utility function definition, it can be expressed as: And for three functions , and Set separately , respectively represented as , and And there are: in: in, It is a non-negative constant used to adjust the penalty.
6. The method according to claim 4, wherein, For the task The one instruction The start time is arranged according to the following formula. : in, This indicates the interval between the start times of two consecutive instructions. Once the instructions are scheduled, the interval... It becomes an ordered set of free intervals, separated by a series of scheduled busy intervals, using... This indicates the set of free intervals. and They are the first The start and end times of each free interval; For an instruction Inserted into free space Internal situation, busy area is If the busy interval exceeds the idle interval, that is... Then the start time will be later than interval offset , is represented as:
7. A computer-readable storage medium having a computer program stored thereon, wherein, When executed by a processor, the program implements the steps of the method according to any one of claims 1 to 6.
8. A computer device comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Task scheduling method and device suitable for distributed rendering
CN112015533A
Computational task offloading for virtualized graphics
US10423463B1