Data scheduling method and device in many-core system
By evaluating operator parameters and path redundancy in a many-core system and adjusting data transmission priorities, the problem of inter-core data contention is resolved and the execution efficiency of applications is improved.
Patent Information
- Application Number
- CN202510978920.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-30
AI Technical Summary
In a many-core system, data contention is prone to occur when data interacts between cores, which affects the execution speed of applications.
The task redundancy evaluation scheduling unit evaluates operator parameters and path redundancy, adjusts data transmission priority to reduce the transmission priority of operators with large path redundancy, and ensures that data of non-redundant operators pass first.
It improves the execution efficiency of applications, avoids delays caused by data competition, and increases the overall computing speed.
Smart Images

Figure CN120723408A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and specifically provides a data scheduling method and device in a many-core system. Background Art
[0002] With advancements in process technology, the number of transistors integrated into a single chip has skyrocketed, often integrating dozens or even hundreds of cores on a single chip, connected via an on-chip network to form a single computing system. An application is often split into multiple operators for computation, with one or more operators assigned to a core simultaneously. These operators may have data dependencies.
[0003] However, when different cores interact with each other, data contention may occur, causing some operators to obtain data later and thus affecting the execution speed of the entire application. Summary of the Invention
[0004] The present invention aims to address the above-mentioned deficiencies in the prior art and provides a highly practical data scheduling method in a multi-core system.
[0005] A further technical task of the present invention is to provide a data scheduling device in a multi-core system that is rationally designed, safe and applicable.
[0006] The technical solution adopted by the present invention to solve its technical problem is:
[0007] A data scheduling method in a many-core system includes a task redundancy evaluation and scheduling unit performing task redundancy evaluation, wherein the task redundancy evaluation and scheduling unit includes an evaluation unit, a network configuration unit, and a contention data scheduling unit;
[0008] The host sends the layout coordinates of the on-chip network operators, operator parameter sizes, and the dependencies of each operator for this application to the evaluation unit of the many-core system;
[0009] The network configuration unit configures the on-chip network according to the operator layout coordinates and sends the corresponding operator instructions and execution data to the corresponding network nodes. At the same time, the evaluation unit evaluates the parameters of each operator.
[0010] The contention data scheduling unit calculates the path of each task branch based on the specific location of the operator actually mapped on the network and the operator's evaluation value, determines the path redundancy of each branch, and controls the operator data priority of the data contention node.
[0011] Furthermore, the sub-paths are divided into A, B, C, D, E, F, G and H. A, E and F form one branch, and B, C, D and G form another branch, and finally the result of H is returned.
[0012] Subsequent tasks in the same branch need to wait for the previous tasks to complete and transmit the calculation results before they can continue to run. Therefore, the H operator needs to wait for all previous tasks to complete before it can continue to execute.
[0013] Furthermore, the operator layout coordinates are:
[0014] Node A(0,0), B(0,1), C(1,0), D(1,1), E(2,0), F(2,1), G(1,2), H(2,2).
[0015] Furthermore, when the evaluation unit evaluates the parameters of each operator, the evaluation value includes a communication evaluation value and a calculation evaluation value. The communication evaluation value is 1 for each hop on the on-chip network, and the calculation evaluation value is accumulated by 1 for each KB of the operator calculation parameter.
[0016] Furthermore, the evaluation unit evaluates the task according to the size of the task parameters and stores the evaluation results in an internal table. Each time a new task comes in, it will be dynamically updated.
[0017] Furthermore, the contention data scheduling unit will determine the redundancy of task branches based on the network configuration status transmitted by the network configuration unit and the task evaluation value transmitted by the evaluation unit;
[0018] According to the network configuration, the node with coordinates (1, 0) is determined to be the point where competition occurs. Starting from the starting point of the task, that is, the time point when A and B start the operation, the evaluation cumulative value reaching the competition point is calculated. The evaluation cumulative value = calculation evaluation value + communication evaluation value.
[0019] Furthermore, the cumulative evaluation value of path A when it reaches the contention point is P, and the cumulative evaluation value of path B when it reaches the contention point is Q. If P = Q, the evaluation values of the two paths are equal, indicating that there is data contention between the two paths, that is, the data has time overlap at the contention point, and both operators may send data to the contention point at the same time.
[0020] Furthermore, the evaluation value after the contention point is calculated, starting from the end point. After the contention point, the cumulative evaluation value of path B is Q, and the evaluation value of path A is P. If P is greater than Q, it means that the evaluation value of path A is greater than that of path B, that is, path A requires more time. Therefore, path B is judged as a redundant path with a redundancy of 1.
[0021] Path A is determined to be a non-redundant path. Therefore, when path A competes with path B at a contention point, the non-redundant path has the highest priority and data passes through first.
[0022] A data scheduling device in a many-core system includes: at least one memory and at least one processor;
[0023] The at least one memory is configured to store a machine-readable program;
[0024] The at least one processor is configured to call the machine-readable program to execute the data scheduling method in the many-core system.
[0025] Compared with the prior art, the data scheduling method and device in the many-core system of the present invention have the following outstanding beneficial effects:
[0026] This method can evaluate the operator calculation delay based on the task mapping results, and based on the operator's path redundancy, when a conflict occurs on the overlapping path of the on-chip network, reduce the data transmission priority of the operator with large path redundancy, and give priority to the operator data with low redundancy or non-redundancy. This can ensure the execution progress of subsequent operators and thus improve the execution efficiency of the application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 It is a schematic diagram of a framework of a data scheduling method in a many-core system;
[0029] Figure 2 It is a schematic diagram of an evaluation unit table entry in a data scheduling method in a many-core system;
[0030] Figure 3 The present invention is a flowchart of a competitive data scheduling unit in a data scheduling method in a many-core system. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention will be further described in detail below in conjunction with specific embodiments. Obviously, the embodiments described are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0032] A best embodiment is given below:
[0033] like Figure 1As shown, the data scheduling method in the many-core system in this embodiment includes a task redundancy evaluation and scheduling unit performing task redundancy evaluation, and the task redundancy evaluation and scheduling unit includes an evaluation unit, a network configuration unit and a contention data scheduling unit;
[0034] The host sends the layout coordinates, parameter sizes, and dependencies of the on-chip network operators for this application to the evaluation unit of the many-core system. Operators are divided into A / B / C / D / E / F / G / H, where A / E / F / form one branch and B / C / D / G form another. The result is finally returned to H. Subsequent tasks in the same branch must wait for the previous task to complete and transmit the calculation results before continuing to run. Therefore, the H operator must wait for all previous tasks to complete before continuing to execute.
[0035] The network configuration unit configures the on-chip network according to the operator layout coordinates and sends the corresponding operator instructions and execution data to the corresponding network nodes. A is the (0, 0) node, B (0, 1), C (1, 0), D (1, 1), E (2, 0), F (2, 1), G (1, 2), and H (2, 2).
[0036] At the same time, the evaluation unit will evaluate the parameters of each operator. The evaluation value includes the communication evaluation value and the calculation evaluation value. The communication evaluation value is 1 for each hop in the on-chip network; the calculation evaluation value is accumulated by 1 for each KB of the operator calculation parameter.
[0037] The contention data scheduling unit calculates the path of each task branch based on the specific location of the operator actually mapped on the network and the operator's evaluation value, determines the path redundancy of each branch, and controls the operator data priority of the data contention node.
[0038] like Figure 2 As shown, the evaluation unit evaluates tasks based on their parameters and stores the results in an internal table. This table is dynamically updated each time a new task arrives. The values evaluated for each task are A=6; B=5; C=4; D=7; E=9; F=2; G=1; and H=1.
[0039] like Figure 3 As shown, the contention data scheduling unit will determine the redundancy of task branches based on the network configuration status transmitted by the network configuration unit and the task evaluation value transmitted by the evaluation unit.
[0040] According to the network configuration, the node with coordinates (1,0) is determined to be the point where competition occurs. Starting from the starting point of the task, that is, the time point when A and B start the operation, the evaluation cumulative value reaching the competition point is calculated. The evaluation cumulative value = calculation evaluation value + communication evaluation value.
[0041] The cumulative evaluation value of path A when it reaches the contention point is 11(10+1), and the cumulative evaluation value of path B when it reaches the contention point is 11(5+2+4). The evaluation values of the two paths are equal, indicating that there is data contention between the two paths, that is, the data overlaps in time at the contention point, and both operators may send data to the contention point at the same time.
[0042] Next, we calculate the evaluation value after the contention point, starting from the end point. After the contention point, the cumulative evaluation value of path B is 11 (1+7+1+1), and the evaluation value of path A is 12 (1+9+1+1). Path A has a larger evaluation value than path B, meaning it takes longer. Therefore, to avoid delaying the overall application execution, path B is determined to be a redundant path with a redundancy of 1.
[0043] Path A is determined to be a non-redundant path. Therefore, when path A competes with path B at a contention point, the non-redundant path (path A) has the highest priority and data passes through first.
[0044] By using the above method, tasks with longer computing time after the competition point of the application are given priority in obtaining data, which can improve the overall execution speed of the application.
[0045] Based on the above method, the data scheduling device in the many-core system in this embodiment includes: at least one memory and at least one processor;
[0046] The at least one memory is configured to store a machine-readable program;
[0047] The at least one processor is configured to call the machine-readable program to execute the data scheduling method in the many-core system.
[0048] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.
[0049] The memory can be used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, the memory can also include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state memory devices.
[0050] The above-mentioned specific implementation methods are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above-mentioned specific implementation methods. Any technical solutions that conform to the above-mentioned specific implementation methods of the present invention and any appropriate changes or substitutions made thereto by ordinary technicians in the relevant technical field shall fall within the patent protection scope of the present invention.
[0051] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A data scheduling method in a many-core system, characterized in that: It includes a task redundancy evaluation and scheduling unit for performing task redundancy evaluation, and the task redundancy evaluation and scheduling unit includes an evaluation unit, a network configuration unit and a contention data scheduling unit; The host sends the layout coordinates of the on-chip network operators, operator parameter sizes, and the dependencies of each operator for this application to the evaluation unit of the many-core system; The network configuration unit configures the on-chip network according to the operator layout coordinates and sends the corresponding operator instructions and execution data to the corresponding network nodes. At the same time, the evaluation unit evaluates the parameters of each operator. The contention data scheduling unit calculates the path of each task branch based on the specific location of the operator actually mapped on the network and the operator's evaluation value, determines the path redundancy of each branch, and controls the operator data priority of the data contention node.
2. The data scheduling method in a many-core system according to claim 1, characterized in that: The operators are divided into A, B, C, D, E, F, G and H. A, E, F form one branch, B, C, D and G form another branch, and finally the result is returned to H. Subsequent tasks in the same branch need to wait for the previous tasks to complete and transmit the calculation results before they can continue to run. Therefore, the H operator needs to wait for all previous tasks to complete before it can continue to execute.
3. The data scheduling method in a many-core system according to claim 2, characterized in that: The operator layout coordinates are: Node A(0,0), B(0,1), C(1,0), D(1,1), E(2,0), F(2,1), G(1,2), H(2,2).
4. The data scheduling method in a many-core system according to claim 3, characterized in that: The evaluation unit evaluates each operator based on its parameters. The evaluation value includes a communication evaluation value and a computational evaluation value. The communication evaluation value is 1 for each hop on the on-chip network, and the computational evaluation value is accumulated by 1 for each KB of the operator's operation parameters.
5. The data scheduling method in a many-core system according to claim 4, characterized in that: The evaluation unit evaluates the task according to the size of the task parameters and stores the evaluation results in the internal table. Every time a new task comes in, it will be dynamically updated.
6. The data scheduling method in a many-core system according to claim 5, characterized in that: The contention data scheduling unit will determine the redundancy of task branches based on the network configuration status transmitted by the network configuration unit and the task evaluation value transmitted by the evaluation unit; According to the network configuration, the node with coordinates (1, 0) is determined to be the point where competition occurs. Starting from the starting point of the task, that is, the time point when A and B start the operation, the evaluation cumulative value reaching the competition point is calculated. The evaluation cumulative value = calculation evaluation value + communication evaluation value.
7. The data scheduling method in a many-core system according to claim 6, characterized in that: The cumulative evaluation value of path A when it reaches the contention point is P, and the cumulative evaluation value of path B when it reaches the contention point is Q. If P = Q, the evaluation values of the two paths are equal, indicating that there is data contention between the two paths, that is, the data has time overlap at the contention point, and both operators may send data to the contention point at the same time.
8. The data scheduling method in a many-core system according to claim 6, characterized in that: Calculate the evaluation value after the contention point, starting from the end point. After the contention point, the cumulative evaluation value of path B is Q, and the evaluation value of path A is P. If P is greater than Q, it means that the evaluation value of path A is greater than that of path B, that is, path A requires more time. Therefore, path B is judged as a redundant path with a redundancy of 1. Path A is determined to be a non-redundant path. Therefore, when path A competes with path B at a contention point, the non-redundant path has the highest priority and data passes through first.
9. A data scheduling device in a many-core system, characterized in that: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 8.