Parallel task scheduling optimization method and device for heterogeneous multi-core real-time system
By optimizing the DAG task scheduling method in a heterogeneous multi-credit real-time system, selecting the first instance for scheduling, and sorting the instance's cutoff time and weight, the problem of waste of computing and storage space is solved, and efficient task scheduling and energy-saving effects are achieved.
Patent Information
- Application Number
- CN202510291070.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has problems of wasted computing and storage space in parallel DAG task scheduling, resulting in inefficiency.
By selecting the first instance for each DAG task for scheduling in a heterogeneous multi-credit real-time system, prioritizing the instance's cutoff time and weights, optimizing task allocation using ready queues and to-release queues, and further optimizing processor frequency to reduce computing and storage requirements.
The amount of data and calculation during the optimization period is reduced, the scheduling efficiency is improved, and energy-saving effects are achieved without affecting real-time.
Smart Images

Figure CN120295759A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to multi-core system scheduling, and more specifically, relates to a parallel task scheduling optimization method and device for heterogeneous multi-core real-time systems. Background Art
[0002] Heterogeneous multi-core systems have developed rapidly and have quite extensive applications in actual production and life. The parallel programming model is the basis for utilizing the performance of these architectures. With the increase in the number of processor cores, the parallelism between tasks continues to increase. Parallel tasks are usually modeled as directed acyclic graphs (DAGs). Each task corresponds to a DAG, which is called a DAG task. Each DAG task contains subtasks (nodes) and edges (communication costs).
[0003] In a real-time system, some DAG tasks are periodically executed. Completing a DAG task once represents completing an instance of the DAG task. For example, the current task instance of the DAG task is executed within the current task cycle. The absolute deadline of the current task instance is the end time of the current cycle, and the end time of the current cycle is used as the start time of the next cycle, that is, the release time of the next instance. The next task instance of the DAG task is executed within the next task cycle. The time period from when a DAG instance is released to when it is executed and completed is called the scheduling length. Therefore, the current instance must be executed within its absolute deadline; otherwise, the next task instance cannot be executed normally. That is, the scheduling length of each instance cannot exceed its task cycle; otherwise, the system will report an error.
[0004] Currently, for the scheduling of parallel DAG tasks, in order to ensure that each task instance can be completed before its deadline, usually all parallel DAG task instances within a scheduling cycle are merged into one DAG task instance, thus turning the multi-DGA problem into a problem of solving the scheduling of a single DAG. However, this approach also has obvious drawbacks. The data volume of the merged DAG task instance is large, and when performing scheduling optimization, it will inevitably be accompanied by a large amount of calculation and waste of storage space, resulting in low efficiency of the entire algorithm. Summary of the Invention
[0005] In view of the above defects or improvement requirements of the prior art, the present invention provides a parallel task scheduling optimization method and device for heterogeneous multi-core real-time systems, aiming to improve the efficiency of parallel task scheduling optimization.
[0006] To achieve the above object, in a first aspect, the present invention provides a parallel task scheduling optimization method for heterogeneous multi-core real-time systems, which includes:
[0007] S1. Determine the number of instances of each DAG task within the scheduling period and the deadline of each instance. For each DAG task, select the first instance and place it in the task set;
[0008] S2. Sort the instances in the task set according to the deadlines of the instances. The earlier the deadline, the higher the priority of the instance. Each instance corresponds to a ready queue that can reflect the priority of its unscheduled nodes;
[0009] S3. Take the node τ with the highest priority in the ready queue of the target instance in the current task set. top , if a new instance is added to the task set, the target instance is the instance with the highest priority in the current task set. If no new instance is added to the task set, the current target instance is the instance with a priority second only to the previous instance selected as the target instance in the current task set; Find the processor π that can make the node τ top have the earliest completion time, top record the start execution time AST(τ top ) of π top for τ top ;
[0010] S4. Determine whether the completion time of τ top is later than the deadline of the task instance to which this node belongs. If so, end the scheduling and report an error. Otherwise, execute S5;
[0011] S5. Determine whether there is an instance in the queue to be released whose release time is earlier than AST(τ tpp ). If there is, remove this instance from the queue to be released and add it to the task set, then jump to S2. Otherwise, execute S6;
[0012] S6. Put the scheduling information of τ top into the scheduling sequence. Determine whether τ top is the last node of the corresponding instance. If not, delete τ top from the ready queue of the corresponding instance and jump to S3. If so, delete the instance to which τ top belongs from the task set and add the next unscheduled instance of the task to which this instance belongs to the queue to be released. Determine whether the current task set is empty. If so, execute S7. Otherwise, jump to S3; The scheduling information includes the instance to which the node belongs, the allocated processor, and the execution information of the processor;
[0013] S7. Output the scheduling sequence.
[0014] Optionally, the nodes of each instance are arranged in the order of node weights to form a ready queue. The higher the weight, the higher the priority; The calculation formula for the weight R(τ i ) of any node τ i is:
[0015]
[0016] Wherein, P is the set of system processors, and p x is any one of the processors, and Rank(τ i , p x ) represents the sorting value of node τ i under processor p x .
[0017] Optionally, if node τ top is the last node of the instance to which it belongs, it is also placed in queue Q;
[0018] After S7, the method further includes:
[0019] S8. Optimize the scheduling sequence, and the optimization steps include:
[0020] S81. Prioritize all the nodes in queue Q according to the completion time of the nodes. The earlier the completion time, the higher the priority;
[0021] S82. Obtain the node with the highest priority in the current queue Q, and calculate the optimizable duration OT of the instance to which the node belongs. OT is the deadline of the instance to which it belongs minus the completion time of the node;
[0022] S83. Judge whether the duration OT is 0. If so, jump to S86; otherwise, execute S84;
[0023] S84. Allocate the duration OT to all the nodes in the instance to which the node belongs according to the node weights, and update the execution duration of each node that obtains the allocation to the sum of its default execution duration and the allocated duration;
[0024] S85. For each node with the updated execution duration, judge whether there are adjustable nodes. The processor of the adjustable node can find a frequency that satisfies the constraint. The constraint is that the frequency is less than the current frequency and the execution duration of the processor at this frequency does not exceed the updated execution duration. If there are no adjustable nodes, jump to S86. If there are adjustable nodes, update the frequencies of all the processors corresponding to the adjustable nodes to the minimum frequency that satisfies the constraint, and recalculate the completion time of each node in the scheduling sequence. If there is a node whose completion time is later than the deadline of the instance to which it belongs, cancel the current frequency update; otherwise, retain the frequency update;
[0025] S86. Delete the node with the highest priority in the current queue Q;
[0026] S87. Judge whether the current queue Q is empty. If so, jump to S88; otherwise, jump to S82;
[0027] S88; Output the optimized scheduling sequence.
[0028] Optionally, in S84, the step of allocating the duration OT to all nodes in the instance to which the node belongs according to the node weight includes:
[0029] Using all nodes in the instance to which the node belongs as the nodes to be allocated, calculate the normalized weight of each node to be allocated;
[0030] Taking the product result of the normalized weight of the node to be allocated and the duration OT as the duration allocated to the corresponding node to be allocated.
[0031] Optionally, the scheduling period belongs to the least common multiple of the periods of each DAG task.
[0032] Optionally, each instance in the task set is sorted in ascending order according to its deadline, and the priority of the instances arranged from start to end gradually decreases.
[0033] Optionally, the nodes in the ready queue are sorted in descending order according to their weights, and the priority of the nodes arranged from start to end gradually decreases.
[0034] In a second aspect, the present invention provides an electronic device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any item of the first aspect are implemented.
[0035] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any item of the first aspect are implemented.
[0036] In a fourth aspect, the present invention provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the method described in any item of the first aspect are implemented.
[0037] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the present invention mainly has the following beneficial effects:
[0038] 1. In the present invention, through S1 to S7, the instance nodes in the task set are used as the objects for scheduling and allocation. The nodes that have completed the allocation will be removed from the task set. During the scheduling and allocation loop, new instances will be added to the task set through the queue to be released. However, it can be ensured that the number of instances in the task set does not exceed the number of DAG tasks, that is, it is ensured that each DGA task will have at most one instance in the task set. Compared with the traditional method of integrating all instances of all DAG tasks to form a larger DAG instance for scheduling and allocation, each DAG task in the task set of the present invention only contains one instance, and it is not necessary to place all instances of the DAG tasks in the task set for decision-making, which greatly reduces the amount of data involved during optimization, the storage space occupied by the algorithm is reduced, and the computational amount is reduced. Therefore, the optimization efficiency is improved.
[0039] 2. Further, through S81 to S88, further optimization of the scheduling sequence can be achieved. This kind of optimization will not change the real-time performance of the scheduling structure. By further optimizing the execution frequency of the processor, the processor can reduce the execution frequency as much as possible on the premise of ensuring the normal scheduling of tasks, and thus the effect of energy saving can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of the steps of the parallel task scheduling optimization method in an embodiment of the present invention;
[0041] Figure 2 is a flowchart of the steps for optimizing the scheduling sequence in an embodiment of the present invention;
[0042] Figure 3 is a structure diagram of 3 real-time DAG tasks;
[0043] Figure 4 is a table of processor parameters of 3 different types;
[0044] Figure 5 is the execution time of 3 real-time DAG tasks under three different processors;
[0045] Figure 6 is a timing diagram of the scheduling result of scheduling 3 real-time DAG tasks using the method of the present invention;
[0046] Figure 7 is the ratio of the final energy consumption savings after optimizing the scheduling sequence compared with before optimization. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0048] Embodiment 1
[0049] The present invention provides an optimization method for parallel task scheduling in a heterogeneous multi-core real-time system, as Figure 1 shown in the flowchart of the steps of the parallel task scheduling optimization method in an embodiment of the present invention. The steps will be described in detail below.
[0050] S1. Determine the number of instances of each DAG task within the scheduling period and the deadline of each instance, and select the first instance of each DAG task and place it in the task set.
[0051] Specifically, the heterogeneous multi-core real-time system also performs periodic scheduling on parallel DAG tasks according to a certain scheduling period. Therefore, the scheduling period needs to be set in advance as required. Preferably, the scheduling period can be set as the super period of each DAG task to be scheduled, and the super period of multiple tasks is the least common multiple of the task periods of all tasks.
[0052] After determining the scheduling period, calculate the number of instances of each DAG task within one scheduling period, that is, calculate the maximum number of times each DAG task can be cyclically executed within one scheduling period. If a DAG task can be cyclically executed N times within one scheduling period, then this DAG task has N instances within one scheduling period, and each cycle represents the execution of one instance of this DAG task. Since the task periods of each DAG task are usually different, the number of instances of different DAG tasks within one scheduling period is usually different.
[0053] For example, as Figure 3 shown, G 1 、G 2 、G 3 are three different parallel DAG tasks, whose task periods are 90, 60, and 180 respectively. The scheduling period can be set as 180. Through calculation, it can be obtained that within this scheduling period, G 1 、G 2 、G 3 need to execute 2, 3, and 1 instances respectively.
[0054] After determining the number of instances of each DAG task, select the first instance of each DAG task and place it in the task set to form an initial task set. At this time, each DAG task in the task set contains only one instance, and it is not necessary to place all instances of the DAG task in the task set for decision-making, which greatly reduces the amount of data involved during optimization, reduces the storage space occupied by the algorithm, and reduces the computational complexity. Therefore, the optimization efficiency is improved.
[0055] S2. Sort the instances in the task set according to the deadline of the instances. The earlier the deadline, the higher the priority of the instance. Each instance corresponds to a ready queue that can reflect the priority of its unscheduled nodes.
[0056] Specifically, each instance has a corresponding deadline, which is the end time of the task cycle in which the instance is located. For example, if the task cycle of a DAG task is 60, and at t = 0, the first instance of this task is released, then this first instance must be executed and completed before t = 60, and t = 60 is the deadline of the first instance. Similarly, its second instance will be released by the system at t = 60, and t = 60 is also the release time of the second instance, and it must be completed before t = 120, and t = 60 is the deadline of the second instance, and so on.
[0057] Specifically, the rule for sorting the priorities of the instances is: the earlier the deadline, the higher the priority of the instance. For an instance with an earlier deadline, if it does not obtain the processor resource for a long time, it is more likely to miss the deadline and cause the system to malfunction. Therefore, the present invention will preferentially allocate scheduling resources to it.
[0058] In actual operation, the instances in the task set can be sorted in ascending order according to their deadlines. For the instances arranged from beginning to end, their priorities gradually decrease. Therefore, when selecting the instance with the highest priority subsequently, directly select the instance located at the head. In this way, the target instance can be quickly selected.
[0059] Specifically, for each instance in the task set, there is a ready queue that can reflect the priority of its unscheduled nodes, so that the unscheduled nodes in it can be traversed one by one based on this ready queue subsequently and scheduling resources can be allocated. The sorting rule of this ready queue can be flexibly set according to the actual situation.
[0060] In a specific embodiment, the nodes of each instance are arranged in the order of node weights to form a ready queue. The higher the weight, the higher the priority. In a DAG instance, the higher the weight of a node, the more remaining workload there is for the instance to which the node belongs. If the processor does not process this node in time, it is very likely that the corresponding instance will miss the deadline, resulting in system errors. Therefore, sorting the nodes in the ready queue according to the weights of the nodes can reduce the system error rate.
[0061] Any node τ i The weight R(τ i ) is calculated by the formula:
[0062]
[0063] In the formula, P is the set of system processors, p x is any one of them, Rank(τ i ,p x ) represents the sorting value of node τ i under the processor p x .
[0064] The sorting value Rank(τ i under the processor p x actually reflects the remaining workload within the corresponding instance period when executing the subtask of node τ i on this processor p x . x i i i
[0065] The sorting value Rank(τ i under the processor p x can be specifically calculated by the following formula: i ,p x )
[0066]
[0067] In the formula, Succ(τ i ) represents the set of successor nodes of node τ i , τ j is a node in the set of successor nodes Succ(τ i ), W(τ i ,p x ) represents the default execution duration of node τ i on the processor p x , C(τ i ,τ j ,p x ,p y ) from the location of the processor py node τ on j to the node τ on the processor p N node τ on i the communication cost between.
[0068] Through the above calculation method, the weight of each node can be calculated, and the unscheduled nodes in each instance are sorted by priority based on the weight to form a ready queue for the corresponding instance.
[0069] In actual operation, the nodes in the instance can be sorted in descending order according to their weights. For the nodes arranged from beginning to end, their priorities gradually decrease. Therefore, when selecting the node with the highest priority subsequently, directly select the node located at the head, so that the target node can be quickly selected.
[0070] S3. Select the node τ with the highest priority in the ready queue of the target instance in the current task set top , if a new instance is added to the task set, the target instance is the instance with the highest priority in the current task set; if no new instance is added to the task set, the current target instance is the instance with the second highest priority in the current task set after the previous instance selected as the target instance; find the processor π that can make the node τ top have the earliest completion time top , record the start execution time AST(τ top ) and the completion time AFT(τ top ) of the processor π for the node τ top . top
[0071] Specifically, this step can be decomposed into sub-steps S31 to S33.
[0072] S31. Select the target instance from the current task set.
[0073] Specifically, in the entire loop process, when certain conditions are met, new instances will be added from the pending release queue to the task set. Therefore, each time the target instance is selected from the task set, the task set may have new instances added compared to the previous time, or no new instances added. Therefore, different choices need to be made in two cases.
[0074] Case 1. New instances are added to the task set.
[0075] In this case, new instances are added from the pending release queue to the task set. Since different instances correspond to different deadlines, when new instances are added, the priorities of the instances in the task set may change. At this time, the instance with the highest priority in the current task set needs to be used as the target instance again.
[0076] It should be noted that as long as a new instance is added to the task set, even if the original instances are deleted, it is considered that a new instance has been added to the task set.
[0077] It should be noted that when selecting the target instance for the first time, it is defaulted that the task set changes from the original empty set to a non-empty set, which also belongs to the case of adding a new instance to the task set.
[0078] Case 2: No new instance is added to the task set.
[0079] In this case, the priorities of the instances in the task set remain unchanged. Therefore, only the instance with a priority just lower than the previous selected target instance needs to be selected from the current task set as the current target instance. That is, when no new instance is added to the task set, the target instance is selected cyclically according to the priorities of the instances in the task set for resource allocation. In this way, not only can the fairness between different tasks be guaranteed, but also the high-priority task nodes can be processed preferentially, making the possibility of missing the deadline smaller to a certain extent.
[0080] S32. Select the highest-priority node τ from the ready queue of the target instance top .
[0081] Specifically, when the nodes in the instance are sorted in descending order according to their weights, only the head node in the current ready queue needs to be taken, which is the highest-priority node v in the ready queue top .
[0082] S32. Allocate a processor for node τ top and record the processing information.
[0083] Specifically, find the processor π that can make node τ tpp have the earliest completion time tpp .
[0084] Among them, the formula for the earliest completion time EFT(τ i , p x ) of any node τ i under any processor p x is as follows;
[0085] EET(τ i , p x ) = EST(τ i , p x ) + W(τ i , p x )
[0086] In the formula, EST(τ i , p x ) represents the earliest start time of node τ i on processor p xThe earliest start execution time on, W(τ i , p x ) represents the node τ i On the processor p x The default execution duration.
[0087] EST(τ i , p x ) is calculated as:
[0088]
[0089] In the formula, Avail(p x ) represents the earliest available time of the processor p x , Pred(τ i ) represents the set composed of the predecessor nodes of the node τ i , AFT(τ j ) represents the completion time of the node τ j , C(τ j , τ i , p y , p x ) represents from the node τ y on the processor p j to the node τ x on the processor p i The communication time between (if p y = p x , the communication time is 0, otherwise it is the communication cost between the node τ j and the node τ i divided by the bandwidth between p y and p x ).
[0090] Based on the above formula, allocate the processor π top that can make the node τ top have the earliest completion time, and its earliest completion time is denoted as AFT(τ top ), then:
[0091]
[0092] After allocating the processor π top for the node τ top , the start execution time AST(τ top ) of the node τ top on the processor π top ) can be calculated:
[0093] AST(τ top ) = AFT(τ top ) - W(τ top , πtop );
[0094] Wherein, W(τ top , π top ) is the default execution duration of node τ top on processor π top .
[0095] Through the above process, a processor π can be found that enables node τ top to have the earliest completion time, record the start execution time AST(τ top ) and the completion time AFT(τ top ) of processor π for node τ top . top ) top .
[0096] S4. Determine whether the completion time of the current node τ top is later than the deadline of the task instance to which the node belongs. If so, end the scheduling and report an error. Otherwise, normally execute S5.
[0097] Specifically, since the scheduling must ensure that any instance needs to be executed before its deadline, therefore, after calculating the earliest completion time AFT(τ top ) of node τ top , if AFT(τ top ) is later than the deadline of the task instance to which the node belongs, if so, it means that the instance cannot be executed before its deadline, and the current scheduling strategy is not feasible. Clear the scheduling sequence and output an empty sequence as an error signal. If not, it means that the current resource strategy is feasible, and continue with the next operation.
[0098] S5. Determine whether there is an instance in the queue to be released whose release time is earlier than AST(τ top ). If so, remove the task instance from the queue to be released and add it to the task set, then jump to S2. Otherwise, execute S6.
[0099] S6. Put the scheduling information of τ top into the scheduling sequence, determine whether τ top is the last node of the corresponding instance. If not, delete τ top from the ready queue of the corresponding instance and jump to S3. If so, delete the instance to which τ top belongs from the task set and add the next unscheduled instance of the task to which the instance belongs to the queue to be released. Determine whether the current task set is empty. If so, execute S7. Otherwise, jump to S3.
[0100] Specifically, the initial state of the queue to be released is empty. However, in subsequent steps, when certain conditions are met, instances are added to the queue to be released. Therefore, the purpose of S5 is to identify whether there are instances in the queue to be released whose release time is earlier than time AST(τ top ). If there are, it means that the instance can be executed before τ top , and it needs to be preferentially allocated resources. Therefore, the instance is added to the task set, and the process jumps to S2 to reselect an instance node with a higher priority for resource allocation. Only when there are no instances in the queue to be released whose release time is earlier than AST(τ top ) does S6 execute, allocate resources for the current node τ top and store its scheduling information in the scheduling sequence.
[0101] In S6, it is also necessary to determine whether τ top is the end node (exit node) of the instance to which it belongs. If so, it means that the current instance has completed all scheduling allocations, and the instance needs to be removed from the task set. Moreover, if there are still unscheduled instances in the DAG task to which the instance belongs, the next unscheduled instance is added to the queue to be released and will be added to the task set from the queue to be released at an appropriate time. It can be seen that when an instance is removed from the task set, it is possible for a new instance to be added to the queue to be released. That is, taking the task set and the queue to be released as a large set, each DAG task will have at most one instance in the large set, and it is not necessary to place all instances of the DAG task in the task set for decision-making. This greatly reduces the amount of data involved during optimization, reduces the storage space occupied by the algorithm, reduces the computational amount, and thus improves the optimization efficiency.
[0102] During specific operations, deleting the instance to which τ top belongs from the task set and adding the next unscheduled instance of the task to which the instance belongs to the queue to be released actually involves a judgment process, that is, judging whether there are unscheduled instances in the task to which the instance belongs. If there are, the next unscheduled instance is added to the queue to be released; if not, it is not added.
[0103] S7. Output the scheduling sequence.
[0104] Specifically, when the task set is empty, that is, after all instances have completed scheduling allocations, the scheduling sequence is output. The scheduling sequence stores the scheduling information of each node of each instance, including the instance to which the node belongs, the allocated processor, and the execution information of the processor. The execution information includes the start execution time AST(τ top ), the default execution duration, and the completion time AFT(τ tpp ).
[0105] In the present invention, through the above S1 to S7, the instance nodes in the task set are used as the objects for scheduling and allocation. The nodes that have completed the allocation will be removed from the task set. During the process of the scheduling and allocation loop, new instances will be added to the task set through the queue to be released. However, it can be ensured that the number of instances in the task set does not exceed the number of DAG tasks, that is, it is ensured that each DGA task will have at most one instance in the task set. Compared with the traditional method of integrating all instances of all DAG tasks to form a larger DAG instance for scheduling and allocation, each DAG task in the task set of the present invention only contains one instance, and it is not necessary to place all instances of the DAG task in the task set for decision-making, which greatly reduces the amount of data involved during optimization, reduces the storage space occupied by the algorithm, and reduces the computational amount. Therefore, the optimization efficiency is improved.
[0106] Further, the present invention can also optimize the resource allocation for the scheduling sequence output by S7 to achieve energy-saving efficiency while improving the efficiency.
[0107] That is, in S6, if the node τ top is the last node of the instance to which it belongs, it will also be put into the queue Q. After S7, the method further includes:
[0108] S8. Optimize the scheduling sequence.
[0109] As Figure 2 shown, the optimization steps include S81 to S88.
[0110] S81. Sort all the nodes in the queue Q according to the completion time of the nodes. The earlier the completion time, the higher the priority.
[0111] Specifically, the nodes in the queue Q are the last nodes of all instances, and they are sorted according to the completion time. The earlier the completion time, the higher the priority. For example, the nodes in the queue Q can be sorted in ascending order according to their completion time. The nodes arranged from the beginning to the end have gradually decreasing priorities. Therefore, when selecting the node with the highest priority later, directly select the node at the head. In this way, the target node can be quickly selected.
[0112] Among them, the completion time of each node is the completion time AFT(τ top ) calculated in S3 for this node.
[0113] S82. Obtain the node with the highest priority in the current queue Q, and calculate the optimizable duration OT of the instance to which the node belongs. OT is the deadline of the instance to which it belongs minus the completion time of this node.
[0114] This step can be divided into the following processes.
[0115] S821. Determine the instance to which the node belongs.
[0116] Specifically, first clarify the instance to which the node belongs. When storing the node in the queue Q, the instance number to which it belongs can also be recorded. Alternatively, the node τ can be calculated by i The instance number to which it belongs is calculated as follows:
[0117] Use τ i The completion time AFT(τ i ) divided by τ i The task period of the DAG to which it belongs and take the integer part to obtain τ i The instance number of the DAG instance to which it belongs (that is, the instance number of the DAG).
[0118] S821. Calculate the optimizable time OT of the instance.
[0119] Specifically, the duration OT may be optimized to be equal to the deadline of the corresponding instance minus the completion time of the node.
[0120] S83, determine whether the time length OT is 0, if so, jump to S86, otherwise, execute S84.
[0121] Specifically, if there is a duration that can be optimized, then the subsequent optimization steps are entered; otherwise, the process jumps to S86 for the next round of judgment.
[0122] S84: Allocate the duration OT to all nodes in the instance to which the node belongs according to the node weight, and update the execution duration of each node to the sum of its default execution time and the allocated duration.
[0123] Specifically, the optimizable time OT is allocated to each node in the instance to increase the execution time of the processor while ensuring normal scheduling.
[0124] The principle of allocating the OT time that can be optimized is that the higher the weight of the node, the longer the allocated time. The weight calculation method of each node is referred to above.
[0125] In specific operations, the normalized weight of each node in the instance can be calculated first, such as the node τ i The normalized weight λ(τ i ) can be expressed as:
[0126]
[0127] Where G is the node set of the instance.
[0128] Then, with the normalized weight λ(τ i ) multiplied by the time OT, as the node τ iThe allocated duration is updated to the sum of the default execution duration and the allocated duration for each node, wherein the default execution duration is the execution duration of the node task executed by the processor used when scheduling and allocating in S1 to S7.
[0129] It should be noted here that for each node, the processor generally has multiple optional frequencies. The higher the frequency, the shorter the execution time for the node. Usually, the frequency of the processor will be given when scheduling and allocating. The purpose of this embodiment is to further optimize the frequency of the processor after obtaining the scheduling sequence to achieve energy saving.
[0130] S85. For each node that updates the execution duration, determine whether there is an adjustable point. The processor of the adjustable point can find a frequency that meets the constraint. The constraint is that the frequency is less than the current frequency and the execution duration of the processor at this frequency does not exceed the updated execution duration. If there is no adjustable point, jump to S86. If there is an adjustable point, update the frequencies of the processors corresponding to all adjustable points to the minimum frequency that meets the constraint, and recalculate the completion time of each node in the scheduling sequence. If there is a node whose completion time is later than the deadline of the instance to which it belongs, cancel this frequency update. Otherwise, retain the frequency update.
[0131] Specifically, after increasing the execution time of the processor for the node, it is also necessary to determine whether there is a frequency in the discrete optional frequencies of the processor that can achieve the updated execution time of the node. If there is a frequency in the optional frequencies that is smaller than the current frequency of the processor, so that the execution time of the processor at this frequency does not exceed the updated execution time, the execution frequency of the processor for the node is updated to the minimum frequency that can be found that meets the above conditions, so as to ensure that the execution frequency of the processor is reduced as much as possible without exceeding the updated execution time, thereby achieving energy saving.
[0132] After the frequency is updated, the execution time of the node will change. At this time, it is necessary to recalculate the completion time of each node in the scheduling sequence to determine whether there is an error, that is, whether the node completion time is later than the deadline of the instance to which it belongs. If not, it means that it can be scheduled normally after the update, and this frequency update is retained. Otherwise, it means that it cannot be scheduled normally after the update, and this frequency update is canceled.
[0133] S86. Delete the node with the highest priority in the current queue Q;
[0134] S87: Determine whether the current queue Q is empty. If so, jump to S88; otherwise, jump to S82;
[0135] S88: Output the optimized scheduling sequence.
[0136] Specifically, compared with the scheduling sequence output by S7, the optimized scheduling sequence updates the execution information of the processor, including the start time, completion time of the node, and the frequency of the processor.
[0137] Through the above S81 - S88, further optimization of the scheduling sequence can be achieved. This kind of optimization will not change the real-time performance of the scheduling structure. By further optimizing the execution frequency of the processor, the processor can reduce the execution frequency as much as possible on the premise of ensuring the normal scheduling of tasks, thereby achieving the effect of energy saving.
[0138] The following uses a specific experiment to illustrate the effect of this solution.
[0139] As Figure 3 shown, G 1 、G 2 、G 3 are three periodic real-time tasks, whose periods are 90, 60, and 180 respectively. According to this method, the period of the set composed of these three tasks is 180, and the number of instances to be executed within this period is 2, 3, and 1 respectively.
[0140] Figure 4 For different types of processor parameters, among them, C ef 、m y 、 are respectively the static power of the processor, independent power, effective switching capacitance, dynamic power exponent, critical frequency, and maximum frequency.
[0141] Figure 5 are the execution times of each node of all tasks on different processors.
[0142] According to the scheduling method of the present invention, the timing diagram of the final scheduling result is as Figure 6 shown. It can be seen from this that by using the scheduling optimization method proposed by the present invention, the normal scheduling of parallel tasks by the heterogeneous multi-core system can be achieved.
[0143] As Figure 7 shown is the ratio of the final energy consumption saved after the scheduling sequence is optimized compared with before optimization. It can be seen from this that after the scheduling sequence is optimized, the effect of energy saving of 20 - 24% can be achieved. At the same time, when the number of processors is the same, the energy saving effect will decrease with the increase of the task set size.
[0144] Embodiment 2
[0145] The present invention also relates to an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0146] The electronic device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory can be used to store computer programs and / or modules. The processor runs or executes the computer programs and / or modules stored in the memory, and calls the data stored in the memory to perform various functions of the electronic device.
[0147] Embodiment 3
[0148] The present invention also relates to a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0149] Specifically, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0150] Embodiment 4
[0151] The embodiment of the present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the method in the above embodiment of the present invention.
[0152] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification. It should be noted that the "in one embodiment", "for example", "again for example", etc. of the present invention are intended to illustrate the present invention, rather than to limit the present invention.
[0153] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention.
Claims
1. A parallel task scheduling optimization method for heterogeneous multi-core real-time systems, characterized in that Including: S1. Determine the number of instances of each DAG task within the scheduling period and the deadline of each instance, and for each DAG task, select the first instance and place it in the task set; S2. Sort the instances in the task set according to the deadlines of the instances. The earlier the deadline, the higher the priority of the instance. Each instance corresponds to a ready queue that can reflect the priority of its unscheduled nodes; S3. Take the node τ with the highest priority in the ready queue of the target instance in the current task set top , if a new instance is added to the task set, the target instance is the instance with the highest priority in the current task set; if no new instance is added to the task set, the current target instance is the instance in the current task set whose priority is second only to the previous instance selected as the target instance; Find a processor π that can make node τ top have the earliest completion time, record π top 's start execution time AST(τ tpp ) for τ top ; top ) S4. Determine τ top whether the completion time is later than the deadline of the task instance to which the node belongs. If so, end the scheduling and report an error; otherwise, execute S5; S5. Determine whether there is an instance in the release queue whose release time is earlier than AST(τ top ). If so, remove the instance from the release queue and add it to the task set, then jump to S2; otherwise, execute S6; S6. Put the scheduling information of τ top into the scheduling sequence, and determine whether τ top is the last node of its corresponding instance. If not, delete τ top from the ready queue of the corresponding instance, and jump to S3. If so, delete the instance to which τ top belongs from the task set, add the next unscheduled instance of the task to which this instance belongs to the queue to be released, and determine whether the current task set is empty. If so, execute S7. Otherwise, jump to S3; the scheduling information includes the instance to which the node belongs, the allocated processor, and the execution information of the processor; S7. Output the scheduling sequence.
2. The parallel task scheduling optimization method according to claim 1, wherein The nodes of each instance are arranged in the order of the node weights to form a ready queue. The higher the weight, the higher the priority. For any node τ i The weight R(τ i ) is calculated by the following formula: Where P is the set of system processors, and p x is any one of the processors, and Rank(τ i , p x ) represents the sorting value of node τ i under processor p x .
3. The parallel task scheduling optimization method according to claim 1, wherein In S6, if node τ top is the last node of the instance to which it belongs, it is also placed in queue Q; After S7, the method further includes: S8. Optimize the scheduling sequence. The optimization steps include: S81. Sort all the nodes in queue Q according to the completion time of the nodes. The earlier the completion time, the higher the priority; S82. Obtain the node with the highest priority in the current queue Q, and calculate the optimizable duration OT of the instance to which the node belongs. OT is the deadline of the instance to which the node belongs minus the completion time of the node; S83. Determine whether the duration OT is 0. If so, jump to S86; otherwise, execute S84; S84. Allocate the duration OT to all the nodes in the instance to which the node belongs according to the node weights, and update the execution duration of each node that obtains the allocation to the sum of its default execution duration and the allocated duration; S85. For each node with the updated execution duration, determine whether there is an adjustable node. The processor of the adjustable node can find a frequency that meets the constraint. The constraint is that the frequency is less than the current frequency and the execution duration of the processor at this frequency does not exceed the updated execution duration. If there is no adjustable node, jump to S86. If there is an adjustable node, update the frequencies of all the processors corresponding to the adjustable nodes to the minimum frequency that meets the constraint, and recalculate the completion times of all the nodes in the scheduling sequence. If there is a node whose completion time is later than the deadline of the instance to which it belongs, cancel the current frequency update; otherwise, retain the frequency update; S86. Delete the node with the highest priority in the current queue Q; S87. Determine whether the current queue Q is empty. If so, jump to S88; otherwise, jump to S82; S88. Output the optimized scheduling sequence.
4. The parallel task scheduling optimization method according to claim 3, wherein In S84, the step of allocating the duration OT to all the nodes in the instance to which the node belongs according to the node weights includes: Using all the nodes in the instance to which the node belongs as the nodes to be allocated, and calculate the normalized weight of each node to be allocated; Using the product result of the normalized weight of the node to be allocated and the duration OT as the duration allocated to the corresponding node to be allocated.
5. The parallel task scheduling optimization method according to claim 1, wherein The scheduling period to which it belongs is the least common multiple of the periods of each DAG task.
6. The parallel task scheduling optimization method according to claim 1, wherein The instances in the task set are sorted in ascending order according to their deadlines. For the instances arranged from the beginning to the end, their priorities gradually decrease.
7. The parallel task scheduling optimization method according to claim 1, wherein The nodes in the ready queue are sorted in descending order according to their weights. For the nodes arranged from the beginning to the end, their priorities gradually decrease.
8. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, it implements the steps of the method described in any one of claims 1 to 7.