A control method and controller for software-defined hardware

By designing software-defined hardware control methods and controllers on TSP architecture chips, task mapping, resource segmentation and adaptive routing are implemented, and the problems of low resource utilization and large memory usage in the existing technology are solved, and the chip performance and resource utilization are improved.

CN114356837BActive Publication Date: 2025-05-06GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111474044.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-05-06
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

In the prior art, there are few software-defined hardware controllers designed for tensor stream processor (TSP) architecture, resulting in low resource utilization and large amounts of memory occupies, reducing chip utilization.

Method used

A software-defined hardware control method and controller are proposed to be applied to TSP architecture chips. The method includes task mapping and resource segmentation, updating the network transmission link through an adaptive routing algorithm, using improved A* algorithm and odd_even algorithm to avoid deadlocks, and optimizing resource allocation and transmission paths.

Benefits of technology

By reasonably arranging the mapping and blocking resources between tasks and resources on the chip, the utilization rate of resources is improved, storage resources is saved, transmission links are shortened, and running time and chip performance are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114356837B_ABST
    Figure CN114356837B_ABST
Patent Text Reader

Abstract

The present invention provides a control method and controller for software-defined hardware, which are applied to TSP architecture chips, wherein the control method includes the following steps: S1 selects idle IP cores on the TSP chip for task mapping for tasks to be processed, and divides the remaining resources into the largest remaining resource block as the area for the next task mapping; S2 when the task mapping area in the TSP on-chip network is updated, the transmission links between nodes in the network are updated based on an adaptive routing algorithm. The control scheme design of task mapping and adaptive routing proposed by the present invention is helpful to improve the resource utilization of the TSP chip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software-defined hardware control, and in particular to a control method and controller for software-defined hardware. Background Art

[0002] The essence of software-defined hardware (SDH) is to realize virtualization, flexibility, diversity and customization through software programming on the basis of digitalization and standardization of hardware resources, provide customized special intelligent and customized services to the outside world, and realize the deep integration of application software and hardware. Its core is API (Application Programming Interface). API removes the coupling relationship between software and hardware, promotes the development of application software in the direction of personalization, hardware resources in the direction of standardization, and system functions in the direction of intelligence. Software-defined hardware programs will create a scalable hardware / software architecture, which, unlike ASIC (Application Specific Integrated Circuit), allows applications to modify hardware configurations at runtime.

[0003] The tensor streaming processor (TSP) architecture developed by Groq also provides the basic idea of ​​"software-defined hardware", that is, the control and scheduling of all operations in the chip are completed by software, thereby reducing the corresponding hardware overhead and using the saved part for computing and on-chip storage.

[0004] In the prior art, there are few designs of software-defined hardware controllers based on the TSP architecture. Therefore, there is an urgent need to propose a software-defined hardware control method and controller based on the TSP architecture. Summary of the invention

[0005] In view of the above problems, the present invention aims to provide a control method and controller for software-defined hardware.

[0006] The purpose of the present invention is achieved by the following technical solutions:

[0007] In a first aspect, the present invention provides a method for controlling software-defined hardware, which is applied to a TSP architecture chip. The method comprises the following steps:

[0008] S1 selects idle IP cores on the TSP chip for task mapping according to the task to be processed, and divides the remaining resources into the largest remaining resource block as the area for the next task mapping;

[0009] S2 When the task mapping area in the TSP on-chip network is updated, the transmission links between nodes in the network are updated based on the adaptive routing algorithm.

[0010] In one implementation, step S1 specifically includes:

[0011] S11 obtains the topological structure information of the idle IP cores in the current TSP on-chip network;

[0012] S12 adopts edge mapping, selecting mapping locations along the edge of the internal network for task mapping;

[0013] S13 adopts the maximum remaining resource partitioning method to determine the optimal partitioning scheme to divide the maximum remaining resource block as the area for the next task mapping.

[0014] In one implementation, step S12 specifically includes:

[0015] On a piece of free rectangular IP core resources, first select free IP cores from the upper left corner for task mapping, and the order is: upper left, upper right, lower right, lower left;

[0016] In the process of task mapping, it is preferred to allocate two new tasks so that their mapping positions in the TSP on-chip network are not adjacent.

[0017] In one implementation, step S13 specifically includes:

[0018] After the task is mapped, a rectangle containing the largest number of IP cores is segmented from the graph composed of the remaining IP cores to serve as the area for the next task mapping.

[0019] In one implementation, step S2 specifically includes:

[0020] The S21 node dynamically updates the real-time status of the links around it and maintains it in a real-time status table.

[0021] S22 uses an improved A* algorithm combined with the odd_even algorithm to find the shortest path and avoid chip deadlock.

[0022] In one implementation, step S21 specifically includes:

[0023] When a fault occurs in the TSP network-on-chip or a task mapping area is updated, each node in the TSP network-on-chip updates the real-time status of the links around itself and maintains a real-time status table of the links around the node; wherein the real-time status table is continuously updated;

[0024] When a blockage or failure occurs in one direction of a node, the transmission between nodes is re-evaluated, problem information is obtained through the real-time status table of the node, and steps are jumped in advance and the preset escape virtual channel is used to avoid the problem area, so as to ensure the minimum cost overhead of transmission on the chip.

[0025] In one implementation, step S22 specifically includes:

[0026] Based on the acquired TSP network-on-chip topology and link consumption information between nodes, an improved A* algorithm is used to calculate the shortest path between the source node and the destination node corresponding to each task, and routing and forwarding are performed according to the calculated shortest path; in the TSP network-on-chip topology, occupied nodes and their corresponding links are regarded as blocked areas.

[0027] In one implementation, step S22 specifically includes: using the improved A* algorithm to implement dynamic bidirectional search. Searching from the starting point and the target point at the same time, and taking the node reached by the other party at the last moment as the end node to guide the search direction. In the search process, a dynamic step length is introduced to perform the search.

[0028] In one implementation, step S22 specifically includes: when using the improved A* algorithm to calculate the shortest path between the source node and the destination node corresponding to the task, combining the odd_even algorithm to avoid deadlock:

[0029] If the Y coordinate of the node is an odd number, the column is called an odd column; if the Y coordinate of the column is an even number, the column is called an even column; E / S / W / N represent the four directions of the on-chip network, namely, the southeast, northwest, and northwest. NW represents a turn from north to west, SW represents a turn from south to west, EN represents a turn from east to north, and ES represents a turn from east to south.

[0030] Deadlock can be avoided by using the following constraints:

[0031] NW and SW turns are prohibited at odd-numbered nodes, that is, the turning destination of odd-numbered nodes is prohibited to be westward;

[0032] It is forbidden to make EN and ES turns at even-numbered columns, that is, the starting direction of even-numbered columns is forbidden to be eastward;

[0033] 180 degree turns are prohibited.

[0034] In a second aspect, the present invention shows a software-defined hardware controller, which is applied to a TSP architecture chip. The controller is used to implement the software-defined hardware control method shown in any one of the implementations of the first aspect above.

[0035] The beneficial effects of the present invention are:

[0036] The present invention uses resources in blocks by reasonably arranging the mapping between tasks and resources in the chip, so that resources can be better utilized. It is not necessary to keep multiple programs on the chip, saving storage resources on the chip. At the same time, the method of mapping tasks in blocks is adopted, which greatly shortens the transmission link, improves the running time, and also improves the performance of the chip.

[0037] The present invention adopts an adaptive routing algorithm, changes deterministic routing into adaptive routing, improves the A* algorithm, and combines XY routing to avoid the occurrence of deadlock. At the same time, a path with the shortest communication distance is selected to reduce transmission time and ensure full utilization of resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The present invention is further described using the accompanying drawings, but the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative work.

[0039] Figure 1 A schematic diagram of a method for controlling software-defined hardware according to an embodiment of the present invention;

[0040] Figure 2 For the present invention Figure 1 Schematic diagram of TSP chip architecture in real-time example;

[0041] Figure 3 For the present invention Figure 1 Schematic diagram of edge mapping in an embodiment;

[0042] Figure 4 For the present invention Figure 1 Schematic diagram of maximum remaining resource partitioning in the embodiment;

[0043] Figure 5 For the present invention Figure 1 Schematic diagram of maximum remaining resource partitioning in the embodiment;

[0044] Figure 6 For the present invention Figure 1 Schematic diagram of maximum remaining resource partitioning in the embodiment;

[0045] Figure 7 For the present invention Figure 1 Flowchart of the improved A* algorithm in the embodiment. DETAILED DESCRIPTION

[0046] The present invention is further described in conjunction with the following application scenarios.

[0047] See also Figure 1 The embodiment shows a method for controlling software-defined hardware, which is applied to a tensor stream processor (TSP) chip based on a TSP architecture. The method comprises the following steps:

[0048] S1 selects idle IP cores on the TSP chip for task mapping according to the task to be processed, and divides the remaining resources into the largest remaining resource block as the area for the next task mapping;

[0049] S2 When the task mapping area in the TSP on-chip network is updated, the transmission links between nodes in the network are updated based on the adaptive routing algorithm.

[0050] In one implementation, the TSP chip includes functional units such as a memory module, a switch execution module, and a vector switch unit.

[0051] Unlike traditional 2D network structure processor chips, in traditional multi-core processor chips, each slice is independent and is a heterogeneous collection of functional units, but is globally isomorphic. The TSP architecture is the opposite, with local functional isomorphism but global heterogeneity.

[0052] In one scenario, the TSP chip architecture is as follows Figure 2 As shown, the memory module (MEM) consists of 44 parallel static random access memory slices and provides the memory concurrency required to fully utilize 32 streams in each direction. Each slice provides 13-bit physical addressing of 16-byte memory words, and each byte is mapped to a channel, for a total of 220 megabytes of on-chip static random access memory.

[0053] The Switch Execution Module (SXM) is similar to a network interface used for communication between cores.

[0054] The structure of TSP is bilaterally symmetrical, with both ends containing memory modules (MEM) and switch execution modules (SXM). The unit in the middle is the vector switch unit (VXM).

[0055] The data transmission on the chip is jointly executed by the memory module (MEM) and the switch execution module (SXM). In the horizontal direction, data can flow in both directions on 320 channels through the movement of the memory module (MEM); in the vertical direction, data can only be transmitted in the same super channel by rearranging the vector elements through the switch execution module (SXM). The 20 super channels set up in the TSP chip are equivalent to dividing the entire chip into 20 group computing groups.

[0056] In the prior art, the task mapping design for TSP chips has improved the resource utilization, but has the technical problem of occupying a large amount of memory and reducing the chip utilization.

[0057] In one implementation, step S1 specifically includes:

[0058] S11 obtains the topological structure information of the idle IP cores in the current TSP on-chip network;

[0059] S12 adopts edge mapping method, selecting mapping locations along the edge of the internal network for task mapping.

[0060] S13 adopts the maximum remaining resource partitioning method to determine the optimal partitioning scheme to divide the maximum remaining resource block as the area for the next task mapping.

[0061] In one scenario, based on the TSP chip architecture proposed above, the entire TSP chip is divided into 20 groups, and the scheduling selection of tasks can be carried out within the group, which greatly shortens the transmission link, improves the running time, and improves the performance of the chip. Among them, the scheduling of tasks can be carried out within the group, or several adjacent groups can be combined to form a new group. The mapping process of new tasks is called dynamic allocation of system resources. The (SDH) controller selects an idle IP core in the current topology to complete the task, thereby forming a new topology. When there are many idle IP cores, there are multiple options, but any random selection may cause uneven distribution of cores, and the path of the transmission task is not the shortest. When other tasks are output in the (SDH) controller, it will interfere with the IP core. Therefore, the above proposal adopts edge mapping to map, which can ensure the integrity of resources.

[0062] In one implementation, step S12 specifically includes:

[0063] On a piece of free rectangular IP core resources, first select free IP cores from the upper left corner for task mapping, and the order is: upper left, upper right, lower right, lower left;

[0064] In the process of task mapping, it is preferred to allocate two new tasks so that their mapping positions in the TSP on-chip network are not adjacent.

[0065] See also Figure 3 , which shows a schematic diagram of edge mapping of resources on a TSP, in which each horizontal row represents a super channel. A super channel can form a computing group by itself, or it can form a computing group together with its nearby super channels. Edge mapping means that on a piece of idle rectangular resources, first select idle IP cores from the upper left corner for task mapping, and the order is: upper left, upper right, lower right, and lower left. In the process of task mapping, try to ensure that two new tasks are not adjacent. When a new task appears, according to the size of the task, the IP core in the upper left corner will be filled first, followed by the upper right, lower right, and lower left in turn. Task mapping by the above-mentioned edge mapping method can effectively improve the performance of the chip.

[0066] In one implementation, step S13 specifically includes:

[0067] After the task is mapped, a rectangle containing the largest number of IP cores is segmented from the graph composed of the remaining IP cores to serve as the area for the next task mapping.

[0068] Considering the uncertainty of the number of IP cores occupied by tasks on the TSP chip and the irregularity of the topology, for a randomly specified topology, the same resources have different partitioning methods. Therefore, according to the actual situation, the maximum remaining resource partitioning method is proposed to divide the remaining resources into the maximum remaining resource blocks, see Figure 4-Figure 6 : After the task mapping, a rectangle is cut out from the graph composed of the remaining IP cores, and the way of cutting is selected that contains the most IP cores in the rectangle (such as Figure 4 ) as the interval for the next task mapping. The maximum remaining resource split at this time is considered to be the optimal split solution.

[0069] In one implementation, step S2 specifically includes:

[0070] The S21 node dynamically updates the real-time status of the links around it and maintains it in a real-time status table.

[0071] S22 uses an improved A* algorithm combined with the odd_even algorithm to find the shortest path and avoid chip deadlock.

[0072] In one implementation, step S21 specifically includes:

[0073] When a fault occurs in the TSP network-on-chip or a task mapping area is updated, each node in the TSP network-on-chip updates the real-time status of the links around itself and maintains a real-time status table of the links around the node; wherein the real-time status table is continuously updated;

[0074] When a blockage or failure occurs in one direction of a node, the transmission between nodes is re-evaluated, problem information is obtained through the real-time status table of the node, and steps are jumped in advance and the preset escape virtual channel is used to avoid the problem area, so as to ensure the minimum cost overhead of transmission on the chip.

[0075] The technical problems of the technical solution of using XY dimensional order routing algorithm to realize data transmission on chip in the prior art are as follows: XY routing algorithm belongs to deterministic routing, which is simple and efficient. In the transmission of small amounts of data, this algorithm has the highest cost performance. However, when the data increases sharply, if a large number of packets need to be transmitted horizontally (along the X-axis) or vertically (along the Y-axis) at the same time, it may cause serious blockage of certain rows and columns in the on-chip network. Since it cannot respond to dynamic changes in network status, when network congestion increases, resources are consumed too much, more time will be spent, and the requirements of low latency cannot be met. The performance will also decline rapidly, and there is also a deadlock problem.

[0076] In one implementation, step S22 specifically includes:

[0077] Based on the acquired TSP network-on-chip topology structure and the link consumption information between nodes, the improved bidirectional A* algorithm is used to calculate the shortest path between the source node and the destination node corresponding to each task, and routing and forwarding are performed according to the calculated shortest path; in the TSP network-on-chip topology structure, occupied nodes and their corresponding links are regarded as blocked areas.

[0078] The A* algorithm combines the advantages of Dijkstra and BFS, combining the information blocks of the Dijkstra algorithm (nodes close to the initial point) and the BFS algorithm (nodes close to the target point). Since the TSP chip only allows four directions of movement: up, down, left, and right, the Manhattan distance is introduced as a heuristic function to guide its search direction. Improvements are made to the A* algorithm. The traditional A* algorithm simply moves from the starting point to the end point. The two-way A* algorithm searches from the starting point and the target point at the same time on the basis of the A* algorithm. When one party detects a node that the other party has checked, the search ends. In terms of search time, the two-way A* algorithm will save more time.

[0079] In order to consider both search accuracy and search time, a dynamic step size is introduced. First, a large step size T1 and a small step size T2 are set. Based on the detection of occupied nodes in the TSP chip, the large step size T is used to expand the node first, and the current node is connected to the node to be expanded. If the connection contains an occupied node, the length of the current node to the expanded node is adjusted to a small step size T2. In this step, a large step size search is used for nodes far from the occupied nodes to reduce the search nodes of the bidirectional A* algorithm and reduce the search time of the algorithm; a small step size search is used for nodes close to the occupied nodes to increase the search nodes of the bidirectional A* algorithm, improve the search accuracy of the algorithm, and avoid the reuse of occupied nodes.

[0080] When using the improved A* algorithm to calculate the shortest path between the source node and the destination node corresponding to the task, the odd_even algorithm is combined to avoid deadlock.

[0081] If the Y coordinate of the node is an odd number, the column is called an odd column; if the Y coordinate of the column is an even number, the column is called an even column; E / S / W / N represent east, south, west, and north respectively, NW represents a turn from north to west, SW represents a turn from south to west, EN represents a turn from east to north, and ES represents a turn from east to south;

[0082] Deadlock can be avoided by using the following constraints:

[0083] NW and SW turns are prohibited at odd-numbered nodes, that is, the turning destination of odd-numbered nodes is prohibited to be westward;

[0084] It is forbidden to make EN and ES turns at even-numbered columns, that is, the starting direction of even-numbered columns is forbidden to be eastward;

[0085] 180 degree turns are prohibited.

[0086] For the TSP architecture, the following constraints are added: when data flows through the MEM, it can be transmitted horizontally; when data flows through the SVM, the data can be transmitted vertically between channels but is restricted to one super channel.

[0087] By optimizing the traditional XY dimensional order routing algorithm, an adaptive routing algorithm is implemented, which changes the deterministic routing into adaptive routing to avoid deadlock. At the same time, a path with the shortest communication distance is selected to reduce the transmission time and ensure full utilization of resources.

[0088] For the TSP chip, its internal hardware topology and the number of functional units will not change under normal circumstances, so the transmission cost between nodes is calculated in advance and the calculation results are maintained in a table. The controller refers to the topology and link consumption in the table, and can regard the occupied nodes and their surrounding links as blocked areas. The improved A* algorithm is used to calculate the shortest path between the source node and the destination node for routing forwarding. When a fault occurs in the network or the task mapping area changes, each node must maintain a real-time status table of the surrounding links. The real-time status table is continuously updated to ensure that the packets passing through the node can be smoothly forwarded under the control of the algorithm. When congestion or failure occurs in one direction, the controller re-evaluates the transmission between nodes, skips in advance, and uses the set escape virtual channel to avoid the problem area, so as to ensure the minimum cost overhead of transmission on the chip.

[0089] At the same time, the present invention also proposes a software-defined hardware controller, which is applied to a TSP chip of a tensor stream processor architecture. The controller can implement the above-mentioned Figure 1 The functions of a software-defined hardware control method shown in the embodiment can realize different implementation methods of the above method in real time. For details, please refer to the above description of the method, which will not be described in detail here.

[0090] It should be noted that, in the description of the present invention, the directions or positional relationships indicated by terms such as “longitudinal”, “lateral”, “up”, “down”, “front”, “back”, “left”, “right”, “vertical”, “horizontal”, “east”, “south”, “west” and “north” are based on the directions or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that they must be constructed and operated in a specific direction, and therefore should not be understood as limiting the present invention.

[0091] Through the description of the above implementation mode, it can be clearly understood by those skilled in the art that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code or any appropriate combination thereof. For hardware implementation, the processor can be implemented in one or more of the following units: application specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing device (DSPD), programmable logic device (PLD), field programmable gate array (FPGA), processor, controller, microcontroller, microprocessor, other electronic units designed to implement the functions described herein or their combination. For software implementation, part or all of the process of the embodiment can be completed by instructing the relevant hardware through a computer program. When implemented, the above program can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein the communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium that can be accessed by a computer. Computer-readable media can include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention, rather than to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should analyze that the technical solution of the present invention can be modified or replaced by equivalents without departing from the essence and scope of the technical solution of the present invention.

Claims

1. A control method for software-defined hardware, applied to a TSP architecture chip, characterized in that: The method comprises the following steps: S1 selects idle IP cores on the TSP chip for task mapping according to the task to be processed, and divides the remaining resources into the largest remaining resource block as the area for the next task mapping; S2 When the task mapping area in the TSP on-chip network is updated, the transmission links between nodes in the network are updated based on the adaptive routing algorithm, specifically including: The S21 node dynamically updates the real-time status of the links around it and maintains it as a real-time status table; S22 uses an improved bidirectional A* algorithm combined with the odd_even algorithm to find the shortest path and avoid chip deadlock; Wherein, step S22 comprises: Based on the obtained TSP on-chip network topology and the link consumption information between nodes, the improved bidirectional A* algorithm is used to calculate the shortest path between the source node and the destination node corresponding to each task, and routing forwarding is performed according to the calculated shortest path; in the TSP on-chip network topology, the occupied nodes and their corresponding links are regarded as blocked areas; The improved bidirectional A* algorithm is based on the A* algorithm. It searches both the starting point and the target point at the same time, and takes the node reached by the other party at the last moment as the end node to guide the search direction. In the search process, a dynamic step size is introduced for search. First, a large step size T1 and a small step size T2 are set. Based on the detection of occupied nodes in the TSP chip, the large step size T1 is used to expand the node first, connecting the current node with the node to be expanded. If the connection contains an occupied node, the length of the current node to the expanded node is adjusted to a small step size T2. When one party detects a node that the other party has checked, the search ends. When using the improved bidirectional A* algorithm to calculate the shortest path between the source node and the destination node corresponding to the task, the odd_even algorithm is combined to avoid deadlock; specifically, the following are included: If the Y coordinate of the node is an odd number, the column corresponding to the Y coordinate is called an odd column; if the Y coordinate is an even number, the column corresponding to the Y coordinate is called an even column; E / S / W / N represent east, south, west and north respectively, NW represents a turn from north to west, SW represents a turn from south to west, EN represents a turn from east to north, and ES represents a turn from east to south; Deadlock can be avoided by using the following constraints: NW and SW turns are prohibited at odd-numbered nodes, that is, the turning destination of odd-numbered nodes is prohibited to be westward; It is forbidden to make EN and ES turns at even-numbered columns, that is, the starting direction of even-numbered columns is forbidden to be eastward; 180 degree turns are prohibited.

2. A method for controlling software-defined hardware according to claim 1, characterized in that: Step S1 specifically includes: S11 obtains the topological structure information of the idle IP cores in the current TSP on-chip network; S12 adopts edge mapping, selecting mapping locations along the edge of the internal network for task mapping; S13 adopts the maximum remaining resource partitioning method to determine the optimal partitioning scheme to divide the maximum remaining resource block as the area for the next task mapping.

3. The method for controlling software-defined hardware according to claim 1, characterized in that: Step S12 specifically includes: On a piece of free rectangular IP core resources, first select free IP cores from the upper left corner for task mapping, and the order is: upper left, upper right, lower right, lower left; In the process of task mapping, it is preferred to allocate two new tasks so that their mapping positions in the TSP on-chip network are not adjacent.

4. The method for controlling software-defined hardware according to claim 1, characterized in that: Step S13 specifically includes: After the task is mapped, a rectangle containing the largest number of IP cores is segmented from the graph composed of the remaining IP cores to serve as the area for the next task mapping.

5. The method for controlling software-defined hardware according to claim 1, characterized in that: Step S21 specifically includes: When a fault occurs in the TSP network-on-chip or a task mapping area is updated, each node in the TSP network-on-chip updates the real-time status of the links around itself and maintains a real-time status table of the links around the node; wherein the real-time status table is continuously updated; When a blockage or failure occurs in one direction of a node, the transmission between nodes is re-evaluated, problem information is obtained through the real-time status table of the node, and steps are jumped in advance and the preset escape virtual channel is used to avoid the problem area, so as to ensure the minimum cost overhead of transmission on the chip.

6. A software-defined hardware controller, applied to a TSP architecture chip, characterized in that: The controller is used to implement the software-defined hardware control method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Heterogeneous multi-core processing system based on on-chip network

    CN104794100A

  • Placement and scheduling of radio signal processing dataflow operations

    CN111247513A