A Data Flow Role Load-Aware Scheduling Method for Heterogeneous Environments

By proposing a load-aware scheduling method in a heterogeneous environment, and optimizing task scheduling using data flow graph model and LAQPSA algorithm, the problem of uneven utilization of computing resources in a heterogeneous environment is solved, and scheduling efficiency and processor utilization are improved.

CN115904746BActive Publication Date: 2025-05-27UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211150949.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-05-27
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

The data stream role scheduling algorithm in heterogeneous environments fails to make full use of computing resource differences, resulting in uneven utilization of computing resource and low scheduling length ratio efficiency.

Method used

A load-aware scheduling method for data flow roles in heterogeneous environments is proposed. By inputting the data flow graph model, load factors are initialized, perception queues are established, and scheduling is used to enqueue, optimizing the priority sorting and allocation of tasks.

Benefits of technology

Through the load-aware scheduling mechanism, the utilization rate of heterogeneous processors is improved, the scheduling span value is reduced, the scheduling acceleration ratio is improved, the scheduling needs in heterogeneous environments are met, and the scheduling performance of data stream roles is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115904746B_ABST
    Figure CN115904746B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for load-aware scheduling of data flow roles in a heterogeneous environment. First, a data flow graph model representation is input, then the load factor is initialized, a perception queue is established, and the LAQPSA algorithm is used for scheduling and enqueueing. Finally, the scheduling result is output. The role division mechanism proposed by the method of the present invention further describes the differences of heterogeneous resources, defines the communication overhead of roles through the method of enqueueing predecessor roles, better defines the differences of computing platforms using the extreme value description method, makes it easier to mine the parallelism in the data flow graph by establishing a perception queue, and the initial load factor and fluctuation value maintain the scheduling balance of the computing platform on the basis of exerting the processing capabilities of heterogeneous platforms, further reducing the overall scheduling cross value and increasing the scheduling speedup ratio, meeting the scheduling requirements in a heterogeneous environment, and at the same time effectively improving the scheduling performance of data flow roles and mining the concurrent characteristics of the data flow graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of optimizing data flow role scheduling in heterogeneous multi-core systems, and particularly relates to a data flow role load-aware scheduling method for heterogeneous environments. Background Art

[0002] A data flow is a form of transfer of different data forms between adjacent or non-adjacent work tasks. At the same time, a data flow is also a solution for expressing an application program. Different from the writing specification of traditional control flow programs, a data flow expresses an application program as a data flow graph; each work task in the graph represents a data flow role, and there is a data dependency between roles. Only after all the predecessor tasks of a role are completed can the data flow role be executed concurrently.

[0003] A heterogeneous environment refers to the computing resources for scheduling and executing data flow roles. Compared with the computing capabilities of computing cores in a homogeneous system, the computing capabilities of computing cores in a heterogeneous environment are asymmetric, and the computing capabilities of different cores are not the same. For example: a computing platform that combines hardware structures with different units such as CPUs, GPUs, FPGAs, and ASICs. The differences in the internal structure of a heterogeneous system can minimize the impact of the failure of Moore's Law on the scale of integrated circuits to the greatest extent, and can also improve the computing speed of processor cores. At the same time, heterogeneous systems also bring problems such as uneven utilization of computing resources and low scheduling length ratio efficiency in the scheduling field. The main reasons are as follows: in the data flow role priority calculation stage, the drawbacks of global scheduling and local scheduling are very prominent during the scheduling process, and starvation is likely to occur in heterogeneous computing resources.

[0004] Therefore, in view of the fact that the data flow role scheduling algorithm in a heterogeneous environment fails to fully consider the differences in computing resources, especially in the trend of the rapid development of integrated circuit structures, future processors will develop towards large-scale heterogeneity. In response to this trend, the scheduling problem becomes increasingly important for effectively organizing computing capabilities and improving role execution efficiency. In the research on role scheduling, it mainly focuses on list-based heuristic scheduling algorithms, random search scheduling algorithms, and scheduling algorithms based on clustering and replication.

[0005] The random search-based scheduling algorithms mainly include genetic algorithms and ant colony algorithms. Genetic algorithms mainly design chromosome gene coding according to the task stage, and perform gene coding interaction between individuals through the mutation operator, genetic operator, and crossover operator designed for the chromosome. Finally, the gene coding method of different individuals is determined according to the fitness function and the number of iterations. The ant colony algorithm selects the optimal path from the starting point to the end point based on the pheromone remaining on the best path by ants. However, due to the excessive number of iterations, both of them will cause the problem of too high overall complexity, and the random search characteristic causes the return function to fall into premature convergence or local optimality.

[0006] The above methods are all the mainstream scheduling algorithms in the current scheduling direction. However, most of the scheduling algorithms do not consider the heterogeneity factors sufficiently, and cannot give full play to the scheduling convenience brought by the powerful computing capabilities of heterogeneous resources. In actual estimation, there is an urgent need for a high-efficiency data flow role scheduling algorithm in a heterogeneous environment. Summary of the Invention

[0007] To solve the above technical problems, a data flow role load-aware scheduling method for a heterogeneous environment is provided.

[0008] The technical solution adopted by the present invention is a data flow role load-aware scheduling method for a heterogeneous environment, and the specific steps are as follows:

[0009] Step 1: Input the data flow graph model representation;

[0010] For the scheduling problem in a heterogeneous multi-core system, it is first necessary to perform mathematical modeling on the tasks to be solved. The task model is described using a DAG graph, that is, the task model is characterized using a quadruple G = {V, E, C, W}.

[0011] Among them, V = {v i | i = 1, 2, 3…, n} represents the task set in the task graph to be solved, v i represents the task numbered i in the task set, and n = |V| represents the number of task sets. E = {(v i , v j ) | 1 ≤ i ≤ n, 1 ≤ j ≤ n} represents that there is an association relationship between any two tasks in the task set, that is, there is a directed edge relationship between task v i and task v j . C = {c ij | 1 ≤ i ≤ n, 1 ≤ j ≤ n} represents the communication overhead between any two tasks with an association relationship in the task set, c ij represents the communication overhead between task i and task j, task i is the predecessor task of task j, and task j is the successor task of task i. W = {w ij | 1 ≤ i ≤ n, 1 ≤ j ≤ m} represents the computing overhead of task i on processor core p j , and m is the total number of processor cores. The priority level of each task in the DAG graph is higher than that of all its successor tasks. Only when all the predecessor parent tasks of the task are scheduled can the successor task be scheduled.

[0012] Step 2: Initialize the load factor and establish a perception queue;

[0013] To improve the robustness of the perception queue to the single-entry task set, during the task priority division stage, a corresponding number of perception queues (Q m));

[0014] Calculate the initial load factor of the processor core, the initial factor ζ 0 Is the ratio of the number of the same tasks processed by each processor core per unit time. The computing capabilities of the cores in the same processor are approximately the same; that is, the initial factor ζ 0 Among the components Q 1 :Q 2 :...:Q j Indicates the number of the same amount of load tasks processed by each processor core per unit time. The load factor is a parameter that perceptually reflects the load situation of the queue in real time. By repeatedly calculating the load factor during each scheduling process, the processor core can be prevented from starving, thereby improving the utilization rate of heterogeneous processors.

[0015] When calculating the priority level of each task (Level(v i ), i = 1, 2,... n), considering the differences in the computing performance of heterogeneous processors and the impact of communication overhead on priority sorting, the task priority is jointly determined by the priority level of the maximum predecessor task, the maximum computing overhead of the task, and the maximum communication overhead between predecessors and successors. For the entry task without a predecessor task, the task priority is determined by the maximum computing overhead and the maximum communication overhead of the successor. The initial load factor ζ of the processor core 0 The calculation formula is shown in Equation (1):

[0016] ζ 0 =Q 1 :Q 2 :...:Q j , j ≤ m (1)

[0017] According to the initial load factor, the queue load threshold η can be calculated. The calculation formula is shown in Equation (2):

[0018] η = max(Q i ) - min(Q j ) (2)

[0019] Among them, Q i and Q j Are the queue with the largest initial load factor capacity and the queue with the smallest initial load factor capacity, respectively.

[0020] The calculation of the priority level Level of the task starts from the entry task. The calculation formula is shown in Equation (3):

[0021]

[0022] Among them, c ji Is the communication overhead between the predecessor task v j and the task v i cik For task v i and the subsequent task v k the communication overhead between them, and w(i, p) is the computational overhead of task v i on processor p. For the entry task without a predecessor task, its priority level is calculated as shown in Equation (4):

[0023]

[0024] where c k is the communication overhead between the entry task v entry and the subsequent task v k and w(entry, p) is the computational overhead of the entry task on processor p.

[0025] Step 3: Schedule and enqueue the LAQPSA algorithm;

[0026] The task priority division algorithm sorts the tasks in terms of three aspects, greatly reducing the impact of processor computational differences and the communication overheads of predecessors and successors on the division results;

[0027] It is agreed that tasks are assigned to the processor core queues in the order of priority. During the enqueue process of tasks, the task start time (EST), task end time (EFT), processor core available time (Avail), and scheduling fluctuation value σ are as shown in Equations (5) to (8):

[0028] σ(i) = ζ max -ζ min (5)

[0029]

[0030] EFT(v i ) = EST(v i ) + w(i, j) (7)

[0031] Avail(j) = EFT(v i ) (8)

[0032] where ζ max and ζ min represent the load factor, λ represents the communication coefficient, c ki represents the communication overhead between task v k and task v i , and w(i, j) is the computational overhead of task v i on processor j.

[0033] In the queue selection algorithm, the tasks waiting to enter the corresponding queue are called waiting tasks. The selection of the enqueue number is mainly determined by four parameters: the scheduling fluctuation value (σ) of the current scheduling, the earliest start time (EST) of the task, the earliest finish time (EFT) of the task, and the available time (Avail) of the processor core. The specific steps are as follows:

[0034] First, obtain ζ, the load factor after the previous scheduling, before enqueueing max and ζ min . Calculate the scheduling fluctuation value at the current scheduling, and compare the initial load factor load threshold η with the scheduling fluctuation value

[0035] Case 1: If the fluctuation value σ is less than η, at this time, as defined in the above formulas (6) and (7), the waiting task is scheduled to the processor core queue with the smallest finish time; if there is a waiting task predecessor task in the target queue, λ = 0; if there is no waiting task predecessor task in the target queue, then λ = 1. At the same time, update the available time of the processor core as described in formula (7), and the Q value of the target processor core queue capacity increases by 1

[0036] Case 2: If the fluctuation value σ is not less than η, the waiting task is scheduled to the processor core queue corresponding to ζ min ; if the processor core queue corresponding to ζ min exists and is not unique, the waiting task is assigned to the processor core with the smallest finish time. When there is also a waiting task predecessor task in the target queue, λ = 0; if there is no waiting task predecessor task in the target queue, then λ = 1. Finally, update the available time of the processor core and update the Q value of the target processor core queue capacity

[0037] Step 4: Output the scheduling result

[0038] By randomly generating DAG graphs with different numbers of tasks as the experimental task set, using two parameters, the schedule length (ScheduleLength, SL) and the task speedup ratio (Speedup Ratio, SR), as evaluation metrics, and then conducting experimental comparisons with the HEFT algorithm and the PQDSA algorithm, the expressions for the schedule length SL and the speedup ratio SR are as follows

[0039] SL = min{EFT(v exit )} (9)

[0040]

[0041] where EFT(v exit ) represents the finish time of the exit task in the task graph. The smaller the SL value of the scheduling result, the better the scheduling efficiency of the algorithm. The speedup ratio represents the ratio of the finish time of a task when executed on a single processor core to the schedule length, T exitE represents the end time of the export task, P represents all processor cores. The larger the SR value generated by the scheduling algorithm, the more efficient the scheduling performance of the algorithm.

[0042] Advantages of the present invention: The method of the present invention first represents the input data flow graph model, then initializes the load factor, establishes a perception queue, schedules and enqueues using the LAQPSA algorithm, and finally outputs the scheduling result. The role division mechanism proposed by the method of the present invention further describes the differences of heterogeneous resources. The communication overhead of roles is defined by the method of the predecessor role enqueueing, and the differences of the computing platform are better defined by the extreme value description method. Establishing a perception queue makes it easier to mine the parallelism in the data flow graph. The initial load factor and the fluctuation value maintain the scheduling balance of the computing platform while exerting the processing capabilities of the heterogeneous platform, further reducing the overall scheduling cross value and increasing the scheduling speedup ratio, meeting the scheduling requirements in the heterogeneous environment, and at the same time effectively improving the scheduling performance of the data flow roles and mining the concurrent characteristics of the data flow graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flowchart of a data flow role load perception scheduling method in a heterogeneous environment according to the present invention.

[0044] Figure 2 It is a DAG model task graph of the method of the present invention.

[0045] Figure 3 It is a scheduling model in an embodiment of the present invention.

[0046] Figure 4 It is a statistical chart of the average scheduling length of the algorithm under different task numbers in an embodiment of the present invention.

[0047] Figure 5 It is a statistical chart of the average scheduling length of the algorithm under different processor queue numbers in an embodiment of the present invention.

[0048] Figure 6 It is a statistical chart of the speedup ratio of the algorithm under different task numbers in an embodiment of the present invention.

[0049] Figure 7 It is a statistical chart of the speedup ratio of the algorithm under different processor queue numbers in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0050] The method of the present invention will be further described below with reference to the accompanying drawings.

[0051] As Figure 1 shown, a flowchart of a data flow role load perception scheduling method in a heterogeneous environment according to the present invention is as follows:

[0052] First, some parameters to be used are defined, as shown in Table 1:

[0053] Table 1

[0054] Parameter Definition <![CDATA[ζ 0 > Initial load factor of the sensing queue Q Queue load capacity Precount Number of predecessor tasks of the task <![CDATA[ζ max > Maximum component of the load factor <![CDATA[ζ min > Minimum component of the load factor EST Task start time EFT Task end time Avail Available time of the processor core η Load threshold σ Fluctuation value

[0055] Step 1: Input the data flow graph model representation;

[0056] As Figure 2 shown, for the scheduling problem under heterogeneous multi-core systems, it is first necessary to perform mathematical modeling on the tasks to be solved. The task model is described using a DAG graph, that is, the task model is characterized using the quadruple G = {V, E, C, W}. Figure 2 The computational overhead of each task in the task graph shown above on the processor is shown in Table 2:

[0057] Table 2

[0058] <![CDATA[v i > <![CDATA[p 1 > <![CDATA[p 2 > <![CDATA[p 3 > <![CDATA[p 4 > <![CDATA[v 1 > 3 4 6 12 <![CDATA[v 2 > 4 7 13 8 <![CDATA[v 3 > 2 12 9 5 <![CDATA[v 4 > 23 17 6 9 <![CDATA[v 5 > 12 16 4 8 <![CDATA[v 6 > 14 18 12 12 <![CDATA[v 7 > 22 12 25 15 <![CDATA[v 8 > 14 27 19 18 <![CDATA[v 9 > 9 32 34 12 <![CDATA[v 10 > 11 13 12 6

[0059] Among them, V = {v i | i = 1, 2, 3…, n} represents the set of tasks in the task graph to be solved, v i represents the task numbered i in the task set, and n = |V| represents the number of task sets. E = {(v i , v j ) | 1 ≤ i ≤ n, 1 ≤ j ≤ n} represents that there is an association relationship between any two tasks in the task set, that is, there is a directed edge relationship between task v i and task v j . C = {c ij | 1 ≤ i ≤ n, 1 ≤ j ≤ n} represents the communication overhead between any two tasks with an association relationship in the task set, c ij represents the communication overhead between task i and task j. Task i is the predecessor task of task j, and task j is the successor task of task i. W = {w ij | 1 ≤ i ≤ n, 1 ≤ j ≤ m} represents the computational overhead of task i on processor core p j , and m is the total number of processor cores. The priority level of each task in each DAG graph is higher than that of all its successors. Only when all the predecessor parent tasks of a task are scheduled can the successor task be scheduled.

[0060] Step 2: Initialize the load factor and establish a perception queue;

[0061] To improve the robustness of the perception queue to the single-entry task set, during the task priority division stage, establish the corresponding number of perception queues (Q m ) according to the number of processor cores;

[0062] Calculate the initial load factor of the processor core, and the initial factor ζ 0It is the ratio of the number of the same tasks processed by each processor core per unit time. The computing capabilities of the cores in the same processor are approximately the same. That is, the initial factor ζ 0 each component Q 1 :Q 2 :...:Q j represents the number of tasks with the same workload processed by each processor core per unit time. The load factor is a parameter that perceives the load situation of the queue in real time. By repeatedly calculating the load factor in each scheduling process, the processor core can be prevented from being in a starvation state, thereby improving the utilization rate of heterogeneous processors.

[0063] When calculating the priority level of each task (Level(v i ), i = 1, 2,... n), considering the differences in the computing performance of heterogeneous processors and the impact of communication overhead on priority sorting, the task priority is jointly determined by the priority level of the maximum predecessor task, the maximum computing overhead of the task, and the maximum communication overhead between predecessors and successors. For the entry task without a predecessor task, the task priority is determined by the maximum computing overhead and the maximum communication overhead of the successor. The initial load factor ζ 0 The calculation formula is shown in Equation (1):

[0064] ζ 0 =Q 1 :Q 2 :...:Q j , j ≤ m (1)

[0065] According to the initial load factor, the queue load threshold η can be calculated. The calculation formula is shown in Equation (2):

[0066] η = max(Q i ) - min(Q j ) (2)

[0067] Among them, Q i and Q j are the queue with the largest initial load factor capacity and the queue with the smallest initial load factor capacity respectively.

[0068] The calculation of the priority level Level of the task starts from the entry task. The calculation formula is shown in Equation (3):

[0069]

[0070] Among them, c ji is the communication overhead between the predecessor task v j and the task v i , c ik is the communication overhead between the task v i and the successor task v k and w(i, p) is the task vi The computational overhead on processor p. For an entry task without a predecessor task, its priority level is calculated as shown in Equation (4):

[0071]

[0072] where c k is the communication overhead between the entry task v entry and its successor task v k , and w(entry, p) is the computational overhead of the entry task on processor p.

[0073] Step 3: Schedule and enqueue the LAQPSA algorithm;

[0074] As Figure 3 shown, in the scheduling model of the embodiment of the present invention, the task priority division algorithm comprehensively sorts the tasks in three aspects, greatly reducing the influence of processor calculation differences and the communication overhead of predecessors and successors on the division result;

[0075] In the priority division stage, considering the topological structure of the overall tasks and the association relationship between tasks, the order of tasks at the same level may be disordered, and tasks must wait for the division of the upper-level tasks to end before they can be sorted by priority. In the processor core queue selection stage, tasks are assigned to appropriate processor core queues according to the task priority levels, and the tasks entering the processor core queues are sequentially scheduled to run on the corresponding processor cores of the queues, that is, the tasks entering the processor core queues are calculated on the processing cores.

[0076] It is agreed that tasks are assigned to processor core queues in the order of priority. During the process of task enqueueing, the task start time (EST), task end time (EFT), processor core available time (Avail), and scheduling fluctuation value σ are as shown in Equations (5) to (8):

[0077] σ(i) = ζ max -ζ min (5)

[0078]

[0079] EFT(v i ) = EST(v i ) + w(i, j) (7)

[0080] Avail(j) = EFT(v i ) (8)

[0081] where ζ max and ζ min represent the load factor, λ represents the communication coefficient, and c ki represents task vk The communication overhead with task v i w(i, j), and the computing overhead of task v on processor j i on processor j

[0082] In the queue selection algorithm, the tasks waiting to enter the corresponding queue are called waiting tasks. The selection of the enqueue number is mainly determined by four parameters: the scheduling fluctuation value (σ) of the current scheduling, the earliest start time (EST) of the task, the earliest finish time (EFT) of the task, and the available time (Avail) of the processor core. The specific steps are as follows:

[0083] First, obtain ζ of the load factor after the previous scheduling before enqueueing max and ζ min . Calculate the scheduling fluctuation value at the current scheduling, and compare the initial load factor load threshold η with the scheduling fluctuation value

[0084] Case 1: If the fluctuation value σ is less than η, at this time, as defined in the above formulas (6) and (7), the waiting task is scheduled to the processor core queue with the smallest finish time; if there is a waiting task predecessor task in the target queue, λ = 0; if there is no waiting task predecessor task in the target queue, then λ = 1. At the same time, update the available time of the processor core as described in formula (7), and the capacity Q value of the target processor core queue increases by 1

[0085] Case 2: If the fluctuation value σ is not less than η, the waiting task is scheduled to the processor core queue corresponding to ζ min ; if the processor core queue corresponding to ζ min exists and is not unique, the waiting task is assigned to the processor core with the smallest finish time. When there is also a waiting task predecessor task in the target queue, λ = 0; if there is no waiting task predecessor task in the target queue, then λ = 1. Finally, update the available time of the processor core and update the capacity Q value of the target processor core queue

[0086] Step 4: Output the scheduling result

[0087] By randomly generating DAG graphs with different numbers of tasks as the experimental task set, using two parameters, the schedule length (SL) and the task speedup ratio (SR), as evaluation metrics, and then conducting experimental comparisons with the HEFT algorithm and the PQDSA algorithm, the expressions for the schedule length SL and the speedup ratio SR are as follows

[0088] SL = min{EFT(v exit )} (9)

[0089]

[0090] where EFT(v exit) represents the end time of the exit task node in the task graph. The smaller the SL value of the scheduling result, the better the scheduling efficiency of the algorithm. The speedup ratio represents the ratio of the end time of the task executed on a single processor core to the scheduling length. exit It represents the end time of the export task, P represents all processor cores, and the larger the SR value generated by the scheduling algorithm, the more efficient the algorithm scheduling performance is.

[0091] like Figure 4 As shown, it is a statistical diagram of the average scheduling length of the algorithm under different numbers of tasks in an embodiment of the present invention. Figure 5 This is a statistical diagram of the average scheduling length of the algorithm under different numbers of processor queues in an embodiment of the present invention. Figure 6 This is a statistical diagram of algorithm acceleration ratios under different numbers of tasks in an embodiment of the present invention. Figure 7 This is a statistical diagram of algorithm acceleration ratios under different numbers of processor queues in an embodiment of the present invention. Figure 4 and Figure 5 The experimental results show that the proposed LAQPSA algorithm has a certain degree of performance improvement in scheduling span as the number of tasks increases and the processor queue increases. Figure 6 and Figure 7 The experimental results show that in terms of acceleration performance, the proposed algorithm also has different degrees of optimization effect compared with the compared classic algorithm.

[0092] In summary, it can be seen that for the scheduling algorithms on existing heterogeneous multi-core systems, there is a problem of over-normalization of role division; as the trend of computing unit heterogeneity continues to increase, and the quantitative analysis of communication overhead becomes increasingly complex, the role division method using the average value method is more and more obvious for the resource gap between different platforms. The role division mechanism proposed in the present invention not only makes a further description of the differences in heterogeneous resources, but also defines the communication overhead of the role through the predecessor role enqueue method; finally, considering the different computing capabilities of different computing resources, the introduction of the load prior step is conducive to simplifying the complexity of role enqueue, and the scheduling fluctuation value ensures the load balance of the overall scheduling platform. In general, the extreme value description method is used to better define the differences in computing platforms based on the load-aware scheduling partitioning algorithm, and the established perception queue is easier to mine parallelism in the data flow graph; the initial load factor and fluctuation value maintain the scheduling balance of the computing platform on the basis of exerting the processing capabilities of the heterogeneous platform, thereby further reducing the overall scheduling span value and improving the scheduling acceleration ratio.

Claims

1. A data flow role load-aware scheduling method for heterogeneous environments, the specific steps are as follows: Step 1: Input the data flow graph model representation; Perform mathematical modeling on the task to be solved. The task model is described using a DAG graph, that is, the task model is characterized using the quadruple G = {V, E, C, W}; Where, V = {v i | i = 1, 2, 3…, n} represents the set of tasks in the task graph to be solved, where v i represents the task numbered i in the task set, n = |V| represents the number of task sets, E = {(v i , v j ) | 1 ≤ i ≤ n, 1 ≤ j ≤ n} represents that there is an association relationship between any two tasks in the task set, that is, there is a directed edge relationship between task v i and task v j , C = {c ij | 1 ≤ i ≤ n, 1 ≤ j ≤ n} represents the communication overhead between any two tasks with an association relationship in the task set, c ij represents the communication overhead between task i and task j. Task i is the predecessor task of task j, and task j is the successor task of task i. W = {w ij | 1 ≤ i ≤ n, 1 ≤ j ≤ m} represents the computing overhead of task i on processor core p j , and m is the total number of processor cores; Step 2: Initialize the load factor and establish a perception queue; In the task priority division stage, a corresponding number of perception queues (Q m ) are established according to the number of processor cores; Calculate the initial load factor of the processor core, the initial factor ζ 0 It is the ratio of the number of the same tasks processed by each processor core per unit time, the initial factor ζ 0 Each component Q in 1 :Q 2 :...:Q j Indicates the number of the same amount of load tasks processed by each processor core per unit time; Calculate the priority level of each task (Level(v i ), i = 1, 2, … n), and the initial load factor ζ of the processor core 0 As shown in Equation (1): ζ 0 = Q 1 :Q 2 :...:Q j , j ≤ m (1) Calculate the queue load threshold η, as shown in Equation (2): η=max(Q i )-min(Q j ) (2) where Q i and Q j are the queue with the largest initial load factor capacity and the queue with the smallest initial load factor capacity, respectively; The calculation of the priority level Level of the task starts from the entry task, as shown in Equation (3): Among them, c ji is the communication overhead between the predecessor task v j and the task v i ; c ik is the communication overhead between the task v i and the successor task v k ; w(i, p) is the computing overhead of the task v i on the processor p. The priority calculation of the entry task without a predecessor task is shown in Equation (4): Among them, c k is the communication overhead between the entry task v entry and the successor task v k , and w(entry, p) is the computational overhead of the entry task on the processor p; Step 3: Schedule and enqueue using the LAQPSA algorithm; It is agreed that tasks are assigned to the processor core queue in the order of priority. During the enqueue process of tasks, the task start time (EST), task end time (EFT), processor core available time (Avail), and scheduling fluctuation value σ are as shown in Equations (5) - (8): σ(i) = ζ max -ζ min (5) EFT(v i ) = EST(v i ) + w(i, j) (7) Avail(j) = EFT(v i ) (8) Among them, ζ max and ζ min represent the load factor, λ represents the communication coefficient, c ki represents the communication overhead between task v k and task v i , and w(i, j) represents the computing overhead of task v i on processor j; The selection of the enqueue number is mainly determined by the four parameters of the current scheduling fluctuation value (σ), task start time (EST), task end time (EFT), and processor core available time (Avail). The specific steps are as follows: Obtain ζ of the load factor after the previous scheduling before enqueueing max and ζ min , calculate the scheduling fluctuation value during the current scheduling, and compare the initial load factor load threshold η with the scheduling fluctuation value; Case 1: If the fluctuation value σ is less than η, as defined in the above Equations (6) and (7), wait for the task to be scheduled to the processor core queue with the smallest end time; there are waiting task predecessor tasks in the target queue, λ = 0; there are no waiting task predecessor tasks in the target queue, then λ = 1; at the same time, update the processor core available time as described in Equation (7), and the capacity Q value of the target processor core queue increases by 1; Case 2: If the fluctuation value σ is not less than η, wait for the task to be scheduled to ζ min on the corresponding processor core queue; if ζ min the corresponding processor core queue exists and is not unique, wait for the task to be assigned to the processor core with the smallest end time; when there are also waiting task predecessor tasks in the target queue, λ = 0; if there are no waiting task predecessor tasks in the target queue, then λ = 1; finally, update the available time of the processor core and update the Q value of the capacity of the target processor core queue Step 4: Output the scheduling result; By randomly generating DAG graphs with different numbers of tasks as the experimental task set, using two parameters, the scheduling length (ScheduleLength, SL) and the task speedup ratio (Speedup Ratio, SR), as evaluation indicators, and then conducting experimental comparisons with the HEFT algorithm and the PQDSA algorithm, the expressions for the scheduling length SL and the speedup ratio SR are as follows: SL = min{EFT(v exit )} (9) Among them, EFT(v exit ) represents the end time of the exit task in the task graph. The smaller the SL value of the scheduling result, the better the scheduling efficiency of the algorithm. The speedup ratio represents the ratio of the end time of a task executed on a single processor core to the scheduling length. T exit represents the end time of the exit task, P represents all the processor cores, and the larger the SR value generated by the scheduling algorithm, the more efficient the scheduling performance of the algorithm.

Citation Information

Patent Citations

  • Dependent task scheduling method of heterogeneous multi-core processor

    CN103473134A

  • Task scheduling method and device based on heterogeneous computing

    CN112328380A