Network-on-chip routing method and device, electronic equipment and storage medium
By determining the faulty node in the on-chip network and configuring the target edge node as a logical reconstruction node, dynamically reconstructing the on-chip network architecture, the problems of low data processing efficiency and high failure rate in the prior art are solved, and efficient data transmission and system fault tolerance are achieved.
Patent Information
- Application Number
- CN202510152468.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-10
AI Technical Summary
Existing bus-based systems have limitations in efficient data processing and are difficult to meet the complex needs of modern high-performance applications. Especially when the number of multi-core processor cores increases, NoCs face more complex traffic management challenges, and chip miniaturization leads to increased manufacturing defects and failure rates.
A routing method for on-chip network is provided. By determining the faulty node and sending reconstruction instructions to the target edge node, the target edge node is configured as a logical reconstruction node of the faulty node, thereby realizing dynamic reconstruction of the on-chip network architecture and improving the load balancing and fault tolerance of the system.
It realizes rapid path adjustment in the event of node failure, improves the stability and reliability of the system, reduces the complexity of the system design and hardware overhead, avoids packet transmission delay and lock problems, and improves data transmission efficiency and network fault tolerance.
Smart Images

Figure CN120128526A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of chip technology, and in particular, to a routing method, apparatus, electronic device, and storage medium for a network on chip. Background Art
[0002] In the context of the rapid development of digitalization and intelligence, technological progress has continuously driven the rapid growth of computing demands. Especially in the fields of data centers, artificial intelligence, and big data analysis, the growth rate of existing demands has far exceeded the processing limits of traditional technologies. Traditional system-on-a-chip (SoC) based on a bus architecture has shown obvious limitations in efficient data processing and is difficult to meet the complex requirements of modern high-performance applications. In view of this, the communication interconnection method based on a network on chip (NoC), especially in an SoC equipped with a large number of processor cores, is gradually becoming a key technological trend to promote the development of future data processing systems.
[0003] As the number of multi-core processor cores increases, the NoC faces more complex traffic management challenges, requiring the data transmission strategy to achieve a uniform distribution of data flows in the network, avoid local overload and resource idleness, improve network efficiency and response speed; at the same time, avoid the increase in manufacturing defects and failure rates caused by chip miniaturization. Summary of the Invention
[0004] The present disclosure provides a routing method, apparatus, electronic device, and storage medium for a network on chip to at least solve the above technical problems existing in the prior art.
[0005] In a first aspect, an embodiment of the present disclosure provides a routing method for a network on chip, where the network on chip has multiple routing nodes, and the method is applied to any one of the first routing nodes in the network on chip. The method includes:
[0006] Determine a faulty node, and determine a target edge node from the edge nodes related to the faulty node; the edge node is a node directly connected to the faulty node;
[0007] Send a reconstruction instruction to the target edge node; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the faulty node.
[0008] In a second aspect, an embodiment of the present disclosure provides a routing method for a network on chip, where the network on chip has multiple routing nodes, and the method is applied to a target edge node. The method includes:
[0009] Receive a reconstruction instruction from a first routing node in the network on chip; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of a faulty node;
[0010] Perform logical configuration according to the reconstruction instruction.
[0011] In a third aspect, an embodiment of the present disclosure provides a routing device for a network-on-chip. The network-on-chip has multiple routing nodes. The device is applied to any first routing node in the network-on-chip. The device includes:
[0012] A first processing module, configured to determine a faulty node and determine a target edge node from the edge nodes associated with the faulty node; the edge node is a node directly connected to the faulty node;
[0013] A first communication module, configured to send a reconstruction instruction to the target edge node if the target edge node is not itself; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the faulty node.
[0014] In a fourth aspect, an embodiment of the present disclosure provides a routing device for a network-on-chip. The network-on-chip has multiple routing nodes. The device is applied to a target edge node. The device includes:
[0015] A first communication module, configured to receive a reconstruction instruction from a first routing node in the network-on-chip;
[0016] A second processing module, configured to be configured as a logical reconstruction node of the faulty node according to the reconstruction instruction.
[0017] In a fifth aspect, an embodiment of the present disclosure provides an electronic device for a network-on-chip, including:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the routing method of the network-on-chip on the side of the first routing node; or so that the at least one processor can execute the routing method of the network-on-chip on the side of the target edge node.
[0021] In a sixth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions. The computer instructions are used to cause a computer to execute the routing method of the network-on-chip on the side of the first routing node; or the computer instructions are used to cause a computer to execute the routing method of the network-on-chip on the side of the target edge node.
[0022] Embodiments of the present disclosure provide a routing method, apparatus, electronic device, and storage medium for a network-on-chip. The network-on-chip has multiple routing nodes. The method is applied to any first routing node in the network-on-chip, and the method includes: determining a faulty node, and determining a target edge node from the edge nodes associated with the faulty node; the edge node is a node directly connected to the faulty node; sending a reconstruction instruction to the target edge node; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the faulty node. Correspondingly, the target edge node receives a reconstruction instruction from the first routing node in the network-on-chip; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the faulty node; and perform logical configuration according to the reconstruction instruction. In this way, when a faulty node appears in the network-on-chip, the target edge node of the faulty node is configured as a logical reconstruction node to perform the reconstruction of the network-on-chip architecture, realizing the load balancing and fault tolerance of the network-on-chip system architecture, and at the same time avoiding problems such as high system design complexity, large hardware overhead, high message transmission delay, and livelock.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0024] Figure 1 It is a flowchart of a routing method for a network-on-chip provided by an embodiment of the present disclosure;
[0025] Figure 2 It is a schematic diagram of a network-on-chip topology architecture provided by an embodiment of the present disclosure;
[0026] Figure 3 It is a flowchart of another routing method for a network-on-chip provided by an embodiment of the present disclosure;
[0027] Figure 4 It is a schematic diagram of a fault handling of a network-on-chip provided by an application embodiment of the present disclosure;
[0028] Figure 5 It is a schematic diagram of the structure of a routing node provided by an application embodiment of the present disclosure;
[0029] Figure 6 It is a schematic diagram of a bypass connection provided by an application embodiment of the present disclosure;
[0030] Figure 7 It is a schematic diagram of the structure of a routing apparatus for a network-on-chip provided by an embodiment of the present disclosure;
[0031] Figure 8Schematic structural diagram of another on-chip network routing device provided by an embodiment of the present disclosure;
[0032] Figure 9 Schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0033] To make the objectives, features, and advantages of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present disclosure.
[0034] Figure 1 Schematic flowchart of a routing method for an on-chip network provided by an embodiment of the present disclosure. The on-chip network has multiple routing nodes, and the method is applied to any first routing node in the on-chip network. As Figure 1 shown, the method includes:
[0035] Step 101: Determine a faulty node, and determine a target edge node from the edge nodes related to the faulty node; the edge node is a node directly connected to the faulty node;
[0036] Step 102: Send a reconstruction instruction to the target edge node; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the faulty node.
[0037] In some embodiments, the on-chip network is an on-chip network with a mesh structure. The on-chip network has multiple routing nodes, and each routing node can implement Figure 1 the routing method shown. That is to say, when any routing node discovers a faulty node, it can be used as the first routing node and execute Figure 1 the routing method shown.
[0038] The mesh-structured on-chip network in a fault-free or initial situation may include N*M routing nodes, where N and M are positive integers, and N may be equal to or different from M; N and M respectively represent the number of routing nodes in each row and each column. For example, as Figure 2 shown, an on-chip network with a 5*5 (i.e., N and M are equal, both are 5) mesh structure is provided. Each routing node has a corresponding routing identifier. As Figure 2 shown, R00, R01... R44 are the routing identifiers of 25 routing nodes respectively.
[0039] Here, a faulty node refers to a routing node in the on-chip network that fails to work properly or malfunctions due to hardware failures, communication problems, or other reasons. A faulty node may be unable to forward data or handle communication tasks. Therefore, the on-chip network needs to bypass these faulty nodes through dynamic routing methods to ensure that data transmission is not affected and the normal operation of the entire network is guaranteed.
[0040] Through the method provided by the embodiments of the present disclosure, for the on-chip network with a Mesh structure combined with the dynamic routing method, when a routing node fails, the routing path of the on-chip network can be quickly adjusted, ensuring the stability and reliability of the on-chip network, thereby improving the fault tolerance of the system, enabling communication to continue when some routing nodes fail, and avoiding the risk of single-point failures. At the same time, dynamically optimizing the communication path can improve the data packet transmission efficiency and has strong scalability.
[0041] It should be noted that in the on-chip network, a node can refer to a connection point in the on-chip network, which can be any computing or data processing unit in the on-chip network. A node is usually responsible for sending, receiving, and routing data. Nodes can include but are not limited to the following types:
[0042] Computing nodes, such as: processor cores, processing units, responsible for executing computing tasks.
[0043] Storage nodes, such as: memory modules, responsible for storing and providing data.
[0044] Routing nodes, also known as routers, are responsible for routing or forwarding data, and they can decide how to forward data packets according to routing algorithms and target addresses; Figure 2 Only the connections between routing nodes are shown in, but it should be understood that routing nodes can actually be connected to or have computing nodes and / or storage nodes attached, rather than only being connected to routing nodes as Figure 2 shown.
[0045] In some embodiments, determining the target edge node from the edge nodes related to the faulty node includes:
[0046] Determining the Manhattan distance between the faulty node and the edge of the on-chip network;
[0047] Selecting the target edge with the smallest Manhattan distance according to the Manhattan distance;
[0048] Taking the edge node of the faulty node in the direction of the target edge as the target edge node.
[0049] Here, the on-chip network provided by the embodiments of the present disclosure supports data transmission functions such as control, response, and listening. In the state without faulty nodes, routing nodes (such as Figure 2The packets of the 25 routing nodes shown can pass through Figure 2 the solid - line links in it. If a faulty node is encountered, this on - chip network also supports re - configuring the on - chip network topology according to the fault information and actual transmission requirements to ensure the continuous and stable operation of the on - chip network and efficient data transmission.
[0050] To implement the re - configuration of the on - chip network topology, a target edge node can be determined from the edge nodes related to the faulty node, and the target edge node can be configured as the logical reconstruction node of the faulty node, thus realizing the re - configuration of the on - chip network topology.
[0051] Here, a method for determining the target edge node is provided to select a suitable target edge node. Specifically, when a certain routing node in the on - chip network fails (denoted as the faulty node), after the first routing node discovers it, the distance between the faulty node and the edges of the on - chip network in the four directions of east (E), west (W), south (S), and north (N) is calculated to identify the nearest edge, and the edge node of the faulty node in the direction of the minimum - distance edge (i.e., the target edge direction) is used as the target edge node, that is, this direction is used as the direction of network architecture re - configuration.
[0052] In an example, as Figure 2 the on - chip network topology architecture shown, assuming the faulty nodes are R11, R12, R13, R21, R22, R23, R31, R32, or R33, and these routing nodes each have four edge nodes. If these nodes are faulty nodes, the distance between them and the edges of the on - chip network in the four directions of E, W, S, and N can be calculated to identify the nearest edge;
[0053] For example, assuming the faulty node is R11, and its edge nodes are R01, R10, R21, and R12, the distance between it and the edges of the on - chip network in the four directions of E, W, S, and N can be calculated to identify the nearest edge.
[0054] In another example, as Figure 2 the on - chip network topology architecture shown, assuming the faulty nodes are R10, R20, R30, R01, R02, R03, R41, R42, R43, R14, R24, or R34, and these routing nodes each have three edge nodes. If these routing nodes are faulty nodes, since they each have only three edge nodes, they can calculate the distance between them and the edges of the on - chip network in three directions (for example, for R10, R20, R30, it is the three directions of E, S, and N; for R01, R02, R03, it is the three directions of E, W, and S; for R41, R42, R43, it is the three directions of E, W, and N; for R14, R24, R34, it is the three directions of W, S, and N) to identify the nearest edge;
[0055] For example, assume that the faulty node is R20, and its edge nodes are R10, R21, and R30. Then, the distances to the edges of the on-chip network in the E, S, and N directions can be calculated to identify the nearest edge.
[0056] In another example, as Figure 2 shown in the on-chip network topology architecture, assume that the faulty nodes are R00, R04, R40, or R44. These routing nodes each have two edge nodes. If these nodes are faulty, since they each have only two edge nodes, they can calculate the distances to the edges of the on-chip network in two directions respectively to identify the nearest edge;
[0057] For example, assume that the faulty node is R00, and its edge nodes are R01 and R10. Then, the distances to the edges of the on-chip network in the E and S directions can be calculated to identify the nearest edge;
[0058] Assume that the faulty node is R04. Then, the distances to the edges of the on-chip network in the W and S directions can be calculated to identify the nearest edge;
[0059] Assume that the faulty node is R40. Then, the distances to the edges of the on-chip network in the N and E directions can be calculated to identify the nearest edge;
[0060] Assume that the faulty node is R44. Then, the distances to the edges of the on-chip network in the N and W directions can be calculated to identify the nearest edge.
[0061] Combined with the above examples, it is not difficult to understand that to calculate the distances to the edges in which directions, it is necessary to specifically combine the edge nodes it has. Furthermore, the target edge node can be determined according to the edge nodes of the faulty node in a certain direction.
[0062] In this way, by dynamically selecting the target edge node to achieve on-chip network reconfiguration, not only the robustness of the network is enhanced, but also the efficiency of data transmission in the case of node failure is guaranteed.
[0063] In some embodiments, determining the Manhattan distance between the faulty node and the edge of the on-chip network includes:
[0064] Determining at least one edge node of the faulty node;
[0065] Determining the Manhattan distance between the faulty node and the edge in the direction where each edge node in the on-chip network is located.
[0066] Here, the Manhattan Distance refers to the distance between two points on a grid-like plane, which is the total distance in the horizontal and vertical directions, rather than the diagonal distance.
[0067] Considering that not all routing nodes have four edge nodes, such as R20, R00, etc. in the above example, therefore, it is possible to first determine the edge nodes of the faulty node, and then determine the Manhattan distance between the faulty node and the edge in the direction of the edge nodes of the faulty node in the on-chip network.
[0068] In this way, by determining at least one edge node of the faulty node, calculating the Manhattan distance between the faulty node and the edge in the direction of each of its edge nodes, and selecting the target edge node, the routing strategy of the network can be optimized, the fault tolerance and stability of the system can be improved, the data transmission delay can be reduced, and the path can be quickly adjusted when the network fails to ensure communication efficiency and reliability.
[0069] In some embodiments, if there are at least two target edge directions with the smallest Manhattan distance, determining the target edge node from the edge nodes related to the faulty node includes:
[0070] Randomly select the target edge node from the edge nodes of the faulty node in the at least two target edge directions with the smallest Manhattan distance.
[0071] Here, considering that there is a situation where there are at least two smallest Manhattan distances at the same time, if this situation occurs, one of the directions with the smallest distance can be randomly selected as the direction for network architecture reconfiguration, that is, the target edge node is randomly selected from the edge nodes of the faulty node in the at least two target edge directions with the smallest Manhattan distance.
[0072] It can be understood that if there are multiple directions with the smallest and same Manhattan distance from the faulty node, one of these directions can be randomly selected, that is, first determine which directions have the smallest Manhattan distance, and then randomly select one edge node from the edge nodes corresponding to the directions with the smallest distance as the target edge node.
[0073] In this way, in the case of multiple optional directions, a path can be randomly selected to avoid overly single selection, thus providing a more flexible network path selection.
[0074] Combined with Figure 2 The following example is provided.
[0075] In one example, assume that the faulty node is R11, and its edge nodes are R01, R10, R21, R12. Then, the distance to the edges in the four directions of east (E), west (W), south (S), and north (N) of the on-chip network can be calculated to identify the nearest edge; assume that after calculation, the distances to the edges in the west direction and the north direction are the smallest and equal. Then, one can be randomly selected from the edge nodes in the west direction and the north direction (i.e., R01 and R10) as the target edge node.
[0076] In another example, assume that the faulty node is R22, and its edge nodes are R12, R21, R32, and R23. Then, the distances to the edges of the on-chip network in the east, west, south, and north directions can be calculated to identify the nearest edge. If it is calculated that the distances to the edges in all four directions are equal, one can be randomly selected from the four edge nodes R12, R21, R32, and R23 as the target edge node.
[0077] In yet another example, assume that the faulty node is R20, and its edge nodes are R10, R21, and R30. Then, the distances to the edges of the on-chip network in the west, south, and north directions can be calculated to identify the nearest edge. If it is calculated that the distances to the edges in the north and south directions are the smallest and equal, one can be randomly selected from the edge nodes in the north and south directions (i.e., R10 and R30) as the target edge node.
[0078] In some embodiments, determining the faulty node includes at least one of the following:
[0079] If it is determined that the periodic signal of the second routing node is not received, determine that the second routing node is a faulty node; the second routing node is an edge node adjacent to the first routing node;
[0080] Send a first signal to the second routing node. If the feedback signal of the second routing node is not received, determine that the second routing node is a faulty node.
[0081] Here, the manner in which each routing node discovers the faulty node can be diverse;
[0082] For example, periodic signals can be sent between directly connected routing nodes. This periodic signal is used to inform the connected routing nodes (i.e., edge nodes) that they are currently working properly, that is, to ensure that the connection between nodes is normal. If a certain routing node does not receive the periodic signal that its directly connected routing node should have sent regularly, this may indicate that the node has failed, that is, a faulty node may be discovered.
[0083] For another example, each routing node can actively detect whether its directly connected routing node is correct. For example, if the first routing node sends a specific signal (i.e., the first signal) to its directly connected routing node (i.e., the second routing node), and no feedback signal is received from the second routing node, it can also be determined that the second routing node has failed. Among them, the first signal can be a detection signal used to detect whether the node is correct, and the feedback signal is usually a response to the sent signal used to confirm whether the node is working properly.
[0084] Thus, by monitoring the periodic signal and / or feedback signal, faulty nodes can be detected in a timely and accurate manner, thereby improving the stability and reliability of the network.
[0085] It should be noted that in practical applications, there may also exist or be designed other methods for determining faulty nodes, and the above examples do not limit the methods for determining faulty nodes.
[0086] In some embodiments, if the target edge node is itself, the method further includes:
[0087] A logical reconstruction node configured as the faulty node.
[0088] Here, considering the situation where the first routing node discovers that itself is the target edge node, that is, the first routing node that discovers the faulty node and the target edge node of the faulty node are the same routing node. If this situation occurs, the first routing node configures itself as the logical reconstruction node of the faulty node.
[0089] Here, configuring as the logical reconstruction node of the faulty node may include: adjusting its own routing identifier according to the faulty node.
[0090] The routing identifier may refer to the identifier (ID) of the routing node in the on-chip network. For example, assume that the routing identifier of the faulty node is R21, and the determined target edge node is R20. The target edge node R20 adjusts its own routing identifier to R21.
[0091] In some embodiments, if the target edge node is itself, the method further includes:
[0092] Determine the communication requirements of the data packet to be forwarded;
[0093] Establish a bypass connection according to the connection relationship of the faulty node and the communication requirements of the data packet to be forwarded;
[0094] Transmit the data packet using the established bypass connection.
[0095] Here, if the target edge node is itself and there is a data packet to be forwarded, clarify its communication requirements according to the data packet to be transmitted, such as the target address. Due to the existence of the faulty node, it is necessary to establish a suitable bypass connection according to the connection relationship between the faulty node and other routing nodes and in combination with the specific requirements of the data packet. After establishing the bypass connection, the established bypass connection can be used to continue transmitting the data packet subsequently, ensuring that the data can reach the target smoothly without being blocked by the faulty node.
[0096] The above operations can also be understood as part of the logical reconstruction node configured as the faulty node.
[0097] In this way, by bypassing the faulty node, communication interruption or data loss caused by the failure of the faulty node is avoided. Moreover, the bypass connection can optimize the data transmission path, reduce latency, and ensure that data reaches the target at the fastest speed. This method improves the fault tolerance of the network, enabling the network system to still operate efficiently in the face of node failures, thus ensuring the stability and continuity of overall communication.
[0098] In some embodiments, if the target edge node is itself, the method further includes:
[0099] Determining a reconstruction direction according to the positional relationship between the faulty node and itself;
[0100] Determining the routing nodes to be updated according to the reconstruction direction;
[0101] Sending an update signal to the routing nodes to be updated, where the update signal is used to instruct the routing nodes to be updated to update their own routing identifiers.
[0102] Here, the routing identifier may refer to the identifier (ID) of the routing node in the on-chip network.
[0103] Since the central area of the on-chip network usually has a high load, while the edge area has a low load. Therefore, the fault handling strategy tends to make the data transmission process detour to the edge area, avoid congestion in the high-load central area, and effectively improve the utilization efficiency of the network edge nodes.
[0104] On this basis, to achieve the reconfiguration of the on-chip network architecture, the direction corresponding to the minimum distance from the faulty node to the edge (denoted as the dis side) is selected as the reconstruction direction. Therefore, it is necessary to process the IDs of the normal working routing nodes in the selected direction. The processing rules are as follows:
[0105] When dis is west (W), that is, the reconstruction direction of the on-chip network is W, the routing node IDs in the same row on the W side of the faulty node are incremented by 1, that is, the number representing the column number is incremented by 1;
[0106] When dis is east (E), that is, the reconstruction direction of the on-chip network is E, the routing node IDs in the same row on the E side of the faulty node are decremented by 1, that is, the number representing the column number is decremented by 1;
[0107] When dis is south (S) and north (N) respectively, the processing strategies are the same as those in the E and W directions respectively, and the routing node IDs in the same column are processed respectively. If the routing identifier rule is as Figure 2 shown, when dis is south (S), that is, the reconstruction direction of the on-chip network is S, the routing node IDs in the same row on the S side of the faulty node are decremented by 10, that is, the number representing the row number is decremented by 1;
[0108] When dis is north (N), that is, the reconfiguration direction of the on-chip network is N, the routing node id of the same row on the N side of the faulty node increases by 10, that is, the number representing the row number is increased by 1.
[0109] Specifically, if the target edge node is itself, and is configured as a logical reconstruction node of the faulty node, it can determine the routing node to be updated according to the reconstruction direction after determining the reconstruction direction, and send an update signal to the routing node to be updated, and the update signal is used to instruct the routing node to be updated to update its own routing identifier. After receiving the update signal, the routing node to be updated updates its own routing identifier, and if it also has a routing node connected in the reconstruction direction, it continues to send update signals to the routing node, and so on, to complete the update of the routing identifier.
[0110] In some embodiments, if the faulty node is a destination node of the data packet, the method further includes:
[0111] A transmission feedback signal is sent to the originating node of the data packet, wherein the transmission feedback signal is used to inform the destination node of the data packet of a failure.
[0112] Here, it is considered that there is a situation where the faulty node is the destination node of the data packet, that is, during the data transmission process, the destination node (i.e., the destination) of the data packet fails and cannot receive the data packet. At this time, a feedback signal can be sent to the starting node (i.e., the source node) of the data packet to notify the source node that the destination node has a fault.
[0113] The feedback signal is used to inform the starting node (ie, the source node) that the destination node cannot receive data, so that the source node knows that the data cannot be transmitted as expected, and can take other measures, which are not limited here.
[0114] In this way, the fault condition can be notified through the feedback signal, helping the starting node (ie, the source node) to make an appropriate response and ensure that data transmission can be carried out.
[0115] In some embodiments, the method further comprises:
[0116] A reconstruction result signal is received from the logic reconstruction node, where the reconstruction signal is used to indicate a result of the reconstruction operation.
[0117] Here, after the logical reconstruction node configuration is completed, that is, after the reconfiguration is completed and the path caused by the faulty node on the chip network is repaired, a reconstruction result signal can be sent, which is used to indicate the execution result of the reconstruction operation, that is, to tell the first routing node whether the reconstruction is successful, or whether there is a problem or failure during the reconstruction process.
[0118] In this way, the first routing node can be aware of the result of the logical reconstruction operation to make the next decision or adjustment. For example, if the reconstruction fails, it can reselect other target edge nodes for reconstruction, further ensuring the continuous and stable operation of the on-chip network and efficient data transmission.
[0119] In some embodiments, the on-chip network is a two-dimensional Mesh on-chip network or a three-dimensional Mesh on-chip network.
[0120] Here, in the two-dimensional Mesh on-chip network, the routing nodes communicate through two-dimensional grid-like connections, specifically connected to each other through horizontal and vertical connections.
[0121] The three-dimensional Mesh on-chip network is different from the two-dimensional Mesh on-chip network, mainly reflected in the dimension of its topological structure. The three-dimensional Mesh structure adds "vertical" connections on the basis of the two-dimensional Mesh, enabling the routing nodes to not only connect on the plane but also perform vertical communication through multi-layer stacking.
[0122] As Figure 2 shown, a two-dimensional Mesh on-chip network is provided. In actual application, the method provided by the embodiments of the present disclosure can also be applied to a three-dimensional Mesh on-chip network, or at least can be applied to any plane of the three-dimensional Mesh on-chip network, that is, each plane on the three-dimensional Mesh on-chip network is understood as a two-dimensional Mesh on-chip network, and the method provided by the embodiments of the present disclosure is applied.
[0123] Figure 3 It is a schematic flowchart of a routing method for another on-chip network provided by the embodiments of the present disclosure; the on-chip network has multiple routing nodes, and the method is applied to a target edge node. As Figure 3 shown, the method includes:
[0124] Step 301, receiving a reconstruction instruction from the first routing node in the on-chip network; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of a faulty node;
[0125] Step 302, performing logical configuration according to the reconstruction instruction.
[0126] In some embodiments, the target edge node has a crossbar switch; the performing logical configuration according to the reconstruction instruction includes:
[0127] Determining the communication requirements of the data packet to be forwarded;
[0128] Establishing a bypass connection according to the connection relationship of the faulty node and the communication requirements of the data packet to be forwarded; the bypass connection is used to enable the data packet to be transmitted without passing through the crossbar switch;
[0129] Transmit the data packet using the bypass connection.
[0130] Here, if there is a data packet to be forwarded, clarify its communication requirements according to the data packet to be transmitted, such as the destination address. Due to the existence of the faulty node, it is necessary to select a bypass connection based on the connection relationship between the faulty node and other routing nodes and in combination with the specific requirements of the data packet, and establish a new connection in the way of bypassing the faulty node to transmit data. After establishing the bypass connection, use the established bypass connection to continue transmitting the data packet to ensure that the data can reach the destination smoothly without being blocked by the faulty node.
[0131] Specifically, the routing node has a crossbar switch (Switch). When it is configured as the target edge node, the bypass link is started so that it no longer passes through the crossbar switch of the node; that is, the target edge node establishes a bypass connection, and the original transmission path that needs to pass through the target edge node and perform routing judgment and other operations does not pass through the crossbar switch but can be directly transmitted through the bypass connection, so that the data can be directly transmitted to the corresponding routing node.
[0132] Provide an example, such as Figure 4 shown, provide a schematic diagram of the fault handling of the network-on-chip; as Figure 4 shown, the black-marked routing node R21 is the faulty node, the gray-marked routing node R20 is the logical reconstruction node, and the remaining routing nodes are normal working nodes.
[0133] After the routing node R21 fails (as Figure 4 shown in the left figure), the routing node R21 cannot communicate with the neighboring routing nodes R11, R20, R22, and R31 (as Figure 4 shown in the middle figure).
[0134] The network-on-chip reconstructs the system architecture according to the routing method provided in the embodiments of the present disclosure. The original routing node 20 is configured as the routing node 21, and the result is as Figure 4 shown in the right figure. Assume that when the routing node R30 needs to transmit data to the routing node R10, the routing node R20 starts the bypass link, that is, establishes a bypass connection so that it no longer passes through the crossbar switch of the node, and the routing node 30 directly transmits the data to the routing node R10.
[0135] In this way, by bypassing the faulty node, communication interruption or data loss caused by the failure of the faulty node is avoided. Moreover, the bypass connection can optimize the data transmission path, reduce latency, and ensure that the data reaches the destination at the fastest speed. This method improves the fault tolerance of the network, enabling the network system to still operate efficiently in the face of node failures, thus ensuring the stability and continuity of the overall communication.
[0136] In some embodiments, the logical configuration according to the reconstruction instruction includes:
[0137] Adjust its own routing identifier according to the faulty node.
[0138] Here, the routing identifier may refer to the identifier (ID) of a routing node in the on-chip network. For example, assume that the routing identifier of the faulty node is R21, and the target edge node is R20. The target edge node R20 adjusts its own routing identifier to R21.
[0139] In some embodiments, the method further includes:
[0140] Determine the reconstruction direction according to the positional relationship between the faulty node and itself;
[0141] Determine the routing nodes to be updated according to the reconstruction direction;
[0142] Send an update signal to the routing nodes to be updated, where the update signal is used to instruct the routing nodes to be updated to update their routing identifiers.
[0143] Here, the routing identifier may refer to the identifier (ID) of a routing node in the on-chip network.
[0144] Since the central area of the on-chip network usually has a high load, while the edge area has a low load. Therefore, the fault handling strategy tends to make the data transmission process detour to the edge area, avoid congestion in the high-load central area, and effectively improve the utilization efficiency of the network edge nodes.
[0145] On this basis, to achieve the reconfiguration of the on-chip network architecture, the direction corresponding to the minimum distance from the faulty node to the edge (denoted as the dis side) is selected as the reconstruction direction. Therefore, it is necessary to process the IDs of the normally operating routing nodes in the selected direction. The processing rules are as follows:
[0146] When dis is west (W), that is, the reconstruction direction of the on-chip network is W, the routing node IDs in the same row on the W side of the faulty node are incremented by 1, that is, the number representing the column number is incremented by 1;
[0147] When dis is east (E), that is, the reconstruction direction of the on-chip network is E, the routing node IDs in the same row on the E side of the faulty node are decremented by 1, that is, the number representing the column number is decremented by 1;
[0148] When dis is south (S) and north (N) respectively, the processing strategies are the same as those in the E and W directions respectively, and the routing node IDs in the same column are processed respectively. If the rules of the routing identifier are as Figure 2 shown, when dis is south (S), that is, the reconstruction direction of the on-chip network is S, the routing node IDs in the same row on the S side of the faulty node are decremented by 10, that is, the number representing the row number is decremented by 1;
[0149] When dis is north (N), that is, when the reconfiguration direction of the on-chip network is N, the routing node id in the same row on the N side of the faulty node is increased by 10, that is, the number representing the row number is incremented by 1.
[0150] Specifically, after determining the reconfiguration direction, the to-be-updated routing node is determined according to the reconfiguration direction, and an update signal is sent to the to-be-updated routing node, where the update signal is used to instruct the to-be-updated routing node to update its own routing identifier. After receiving the update signal, the to-be-updated routing node updates its own routing identifier, and at the same time, if it also has a routing node connected in the reconfiguration direction, it continues to send the update signal to this routing node, and so on, to complete the update of the routing identifier.
[0151] In some embodiments, the method further includes:
[0152] Sending a reconfiguration result signal to the first routing node, where the reconfiguration signal is used to indicate the result of the reconfiguration operation.
[0153] Here, after the configuration of the to-be-logically reconfigured node is completed, that is, after the reconfiguration is completed and problems such as paths caused by repairing the faulty nodes on the on-chip network are solved, a reconfiguration result signal can be sent, and this signal is used to indicate the execution result of the reconfiguration operation. That is, it tells the first routing node whether the reconfiguration is successful, or whether there are problems or failures during the reconfiguration process.
[0154] In this way, the first routing node can know the result of the logical reconfiguration operation, so as to make the next decision or adjustment, further ensuring the continuous and stable operation of the on-chip network and efficient data transmission.
[0155] In some embodiments, the on-chip network is a two-dimensional Mesh on-chip network, or a two-dimensional Mesh on-chip network.
[0156] Through the method provided by the embodiments of the present disclosure, adopting an on-chip network routing algorithm with load balancing and fault tolerance capabilities is crucial for ensuring the stable operation of the system and data integrity. Through the above method, the routing can be automatically adjusted when nodes or links fail, and the complex network topology and data traffic distribution can be managed.
[0157] The embodiments of the present disclosure provide a reconfiguration strategy. Specifically, aiming at the problems brought by permanent faults to the data transmission performance of the on-chip network, a reconfigurable on-chip network based on a logical Mesh structure is designed to ensure that the continuity of network communication can still be maintained when node failures occur.
[0158] Specifically, when a node in the network-on-chip fails, the system will identify the nearest edge by calculating the distances between the failed node and the edges of the network in the four directions of east, west, south, and north. For example, when the first routing node discovers a failed node, it can calculate the distances between the failed node and the edges in the four directions. Then, a direction is randomly selected from the directions with the minimum distances as the direction for reconfiguring the network-on-chip architecture. In this way, dynamic reconfiguration is achieved, which not only enhances the robustness of the network but also ensures the efficiency of data transmission in the case of node failures.
[0159] Specifically, the reconstruction strategy includes determining the Manhattan distance between the failed node and the edge of the network-on-chip, and selecting the target edge with the minimum Manhattan distance according to the Manhattan distance. A calculation example is provided, and the following calculations can be specifically performed:
[0160] distance E = cols - (id mod rows)
[0161] distance S = rows - (id / rows)
[0162] distance W = id mod rows
[0163] distance N = id / rows
[0164] dis ~ {min(distance E , distance S , distance W , distance N )}
[0165] Among them, distance E , distance S , distance W and distance N respectively represent the distances between the failed node and the edges of the network-on-chip in the four directions of east (E), south (S), west (W), and north (N);
[0166] rows represents the total number of rows of the network-on-chip;
[0167] cols represents the total number of columns of the network-on-chip;
[0168] mod represents the modulo operation;
[0169] The id represents the row number or column number of the faulty node. For example, when calculating the edge distance from E and W, the id is the column number of the routing node; when calculating the edge distance from S and N, the id is the row number of the routing node. Combining Figure 2 with the routing node identifier shown, R01, where 0 represents the row number and 1 represents the column number; that is, the first digit represents the row number of the routing node, and the second digit represents the column number of the routing node.
[0170] min represents taking the minimum value, and dis represents the direction corresponding to the minimum distance from the edge.
[0171] In this way, through the above calculations, the distances of the faulty node from the edges in the four directions of the on-chip network, east (E), south (S), west (W), and north (N), can be obtained, and then the direction corresponding to the minimum distance from the edge can be selected.
[0172] A routing node provided by an embodiment of the present disclosure, such as Figure 5 shown, specifically provides a routing node (router) applicable to the reconfigurable architecture of the on-chip network, which may include:
[0173] A routing calculation unit (RC, Routing Computation), a buffer, a virtual channel allocator (VA, Virtual Channel Allocator), a crossbar arbiter (SA, Switch Allocator), and a crossbar (Switch).
[0174] Among them, the routing calculation unit is responsible for determining the path of the data packet from one node to another node. The routing calculation unit can calculate the transmission path of the data packet according to the network topology, routing policy, and the current network state (such as the availability of the link) to select the best path for each data packet.
[0175] The buffer is used to temporarily store the data packets to be sent or received to avoid data loss. In the on-chip network, the data may encounter congestion or delay, and the buffer can temporarily save the data until it can be smoothly transmitted. It can balance the transmission rates of each routing node in the on-chip network and avoid losing data packets due to sudden traffic.
[0176] The virtual channel allocator is responsible for allocating virtual channels and managing the use of each channel. In order to effectively utilize the network bandwidth, the on-chip network usually uses the virtual channels (Virtual Channels) technology to simultaneously transmit different logical data streams on the same physical channel. The task of the virtual channel allocator is to allocate appropriate virtual channels to different data streams according to the routing information and the current network load situation, so as to avoid blocking and improve parallelism.
[0177] A crossbar arbiter is responsible for coordinating the data streams of multiple requests, determining which data streams are transmitted through the crossbar; if there are multiple data packets to be forwarded simultaneously, it is also responsible for allocating the order of transmission, etc. Routing nodes in a network-on-chip usually have multiple input and output ports, and the crossbar is responsible for sending the data stream from the input port to the correct output port through a "crosspoint". The crossbar arbiter then needs to determine the scheduling of the data stream according to network status, data priority, and other factors to ensure that data streams do not conflict or deadlock.
[0178] The crossbar, inside the routing node, is responsible for actual data transmission. The crossbar connects the input ports and output ports in a cross-connected manner, and data packets can be actually transmitted from one port to another through the crossbar. The crossbar is usually implemented as a multi-port matrix structure, and each port can independently receive and send data. In the figure, E, W, S, and N respectively represent the ports in the east, west, south, and north directions, and L represents the connection port to the attached storage node and computing node.
[0179] Based on this infrastructure, a specific connection strategy is implemented. Through this strategy, the input ports in four different directions are connected to the corresponding output ports through internal wiring.
[0180] To enhance the fault tolerance of the network, the routing node also includes or integrates:
[0181] A fault tolerance (FT) module, which is responsible for real-time monitoring of the network status, capturing and feedback any possible fault signals;
[0182] A bypass selection unit (BS), which is used to dynamically adjust the bypass connection based on the fault signal.
[0183] Through this strategy and the structure of the routing node, it is realized that when a fault occurs, the data packet can bypass the affected routing node and be directly delivered to the target location according to the previous data routing strategy.
[0184] As Figure 6 shown, according to the strategy provided by the embodiments of the present disclosure, four bypass connection modes can be implemented to ensure that data packets can effectively bypass the faulty router in the central area. The specific connection methods include:
[0185] A north-to-east (N-E) bypass connection can be established between the northeast (NE) ports;
[0186] A south-to-west (S-W) bypass connection can be established between the southwest (SW) ports;
[0187] A south - east (S - E) bypass connection can be established between the south - east (SE) ports;
[0188] A north - west (N - W) bypass connection can be established between the north - west (NW) ports.
[0189] When a faulty node appears in the network - on - chip, the bypass selection unit can dynamically establish corresponding bypass connections according to the specific communication requirements of data packets, thus directly bypassing the damaged router. This flexible bypass selection mechanism greatly enhances the adaptability and communication efficiency of the network in the face of various traffic pattern changes and changing network topologies.
[0190] Combined with the above embodiments, the embodiments of the present disclosure provide a specific router architecture (i.e., routing node architecture). Self - testing, fault feedback, and reconfigurable bypass control and other technologies are introduced into the routing node to achieve the optimized design of the router micro - architecture. Through this design, the router can detect network faults, establish a bypass according to the fault information, and select the optimal bypass connection to bypass the faulty router.
[0191] In addition, the network - on - chip provided by the embodiments of the present disclosure adopts the on - chip network architecture reconfiguration technology of logical Mesh. By calculating the Manhattan distance between the faulty node and the four - week of the network - on - chip, the faulty node is logically migrated to the edge area with the shortest distance, effectively reducing network congestion in the central area and over - high load of adjacent routing nodes. When a fault occurs in the network - on - chip, it can maintain the logical integrity of the Mesh topology in the central area by constructing a bypass link and ensure that the faulty node is always logically located at the network edge. This method not only enhances the fault - tolerance and load - balancing capabilities of the network - on - chip, but also reduces network latency and effectively prevents the occurrence of deadlocks.
[0192] Figure 7 It is a schematic structural diagram of a routing device for a network - on - chip provided by the embodiments of the present disclosure; as Figure 7 shown, the network - on - chip has multiple routing nodes, and the device is applied to any one of the first routing nodes in the network - on - chip. The device includes:
[0193] A first processing module, configured to determine a faulty node and determine a target edge node from the edge nodes related to the faulty node; the edge node is a node directly connected to the faulty node;
[0194] A first communication module, configured to send a reconstruction instruction to the target edge node if the target edge node is not itself; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the faulty node.
[0195] In some embodiments, the first processing module is configured to determine the Manhattan distance between the faulty node and the edge of the network - on - chip;
[0196] Select a target edge with the smallest Manhattan distance according to the Manhattan distance.
[0197] Use the edge node of the failed node in the direction of the target edge as the target edge node.
[0198] In some embodiments, the first processing module is configured to determine at least one edge node of the failed node.
[0199] Determine the Manhattan distance between the failed node and the edge in the direction of each edge node in the on-chip network.
[0200] In some embodiments, if there are at least two target edge directions with the smallest Manhattan distance, the first processing module randomly selects a target edge node from the edge nodes of the failed node in the at least two target edge directions with the smallest Manhattan distance.
[0201] In some embodiments, determining the failed node includes at least one of the following:
[0202] If it is determined that the periodic signal of the second routing node is not received, determine that the second routing node is a failed node; the second routing node is an edge node adjacent to the first routing node.
[0203] Send a first signal to the second routing node. If the feedback signal of the second routing node is not received, determine that the second routing node is a failed node.
[0204] In some embodiments, if the target edge node is itself, the first processing module is further configured to be a logical reconstruction node of the failed node.
[0205] In some embodiments, if the target edge node is itself, the first processing module is further configured to determine a reconstruction direction according to the positional relationship between the failed node and itself.
[0206] Determine a routing node to be updated according to the reconstruction direction.
[0207] Send an update signal to the routing node to be updated, where the update signal is used to instruct the routing node to be updated to update its own routing identifier.
[0208] In some embodiments, if the target edge node is itself, the first processing module is further configured to determine the communication requirements of the data packet to be forwarded.
[0209] Establish a bypass connection according to the connection relationship of the failed node and the communication requirements of the data packet to be forwarded.
[0210] Transmit the data packet by using the established bypass connection.
[0211] In some embodiments, if the failed node is the destination node of the data packet, the first communication module is further configured to send a transmission feedback signal to the source node of the data packet, where the transmission feedback signal is used to inform that the destination node of the data packet has failed.
[0212] In some embodiments, the first communication module is further configured to receive a reconstruction result signal from the logic reconstruction node, where the reconstruction signal is used to indicate the result of the reconstruction operation.
[0213] In some embodiments, the network-on-chip is a two-dimensional Mesh network-on-chip or a two-dimensional Mesh network-on-chip.
[0214] It can be understood that when the routing device of the network-on-chip provided in the above embodiments implements the routing method of the corresponding network-on-chip on the first routing node side, the above processing can be allocated to different program modules as needed to complete all or part of the above-described processing. In addition, the device provided in the above embodiments and the embodiments of the corresponding method belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.
[0215] Figure 8 FIG. is a schematic structural diagram of another routing device of a network-on-chip provided by an embodiment of the present disclosure; as Figure 8 shown, the network-on-chip has multiple routing nodes, and the device is applied to a target edge node, and the device includes:
[0216] A second communication module, configured to receive a reconstruction instruction from a first routing node in the network-on-chip;
[0217] A second processing module, configured to be configured as a logic reconstruction node for the failed node according to the reconstruction instruction.
[0218] In some embodiments, the target edge node has a crossbar switch; the second processing module is configured to determine the communication requirements of the data packet to be forwarded;
[0219] Establish a bypass connection according to the connection relationship of the failed node and the communication requirements of the data packet to be forwarded; the bypass connection is used to enable the data packet to be transmitted without passing through the crossbar switch;
[0220] Transmit the data packet by using the bypass connection.
[0221] In some embodiments, the second processing module is configured to adjust its own routing identifier according to the failed node.
[0222] In some embodiments, the second processing module is further configured to determine a reconstruction direction according to the positional relationship between the faulty node and itself;
[0223] Determine the routing nodes to be updated according to the reconstruction direction;
[0224] The second communication module is further configured to send an update signal to the routing nodes to be updated, and the update signal is used to instruct the routing nodes to be updated to update the routing identifier.
[0225] In some embodiments, the second communication module is further configured to send a reconstruction result signal to the first routing node, and the reconstruction signal is used to indicate the result of the reconstruction operation.
[0226] In some embodiments, the network-on-chip is a two-dimensional Mesh network-on-chip or a two-dimensional Mesh network-on-chip.
[0227] It can be understood that when the routing device of the network-on-chip provided in the above embodiments implements the routing method of the corresponding network-on-chip on the target edge node side, the above processing can be allocated to different program modules as needed to complete all or part of the processing described above. In addition, the device provided in the above embodiments and the embodiments of the corresponding method belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.
[0228] The embodiments of the present disclosure provide a computer-readable storage medium storing executable instructions, where the executable instructions, when executed by a processor, will trigger the processor to execute the routing method on the first routing node side or the routing method on the target edge node side provided by the embodiments of the present disclosure.
[0229] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device and a readable storage medium.
[0230] Figure 9 It is a schematic structural diagram of an electronic device of a network-on-chip provided by an embodiment of the present disclosure; as Figure 9 shown, the electronic device 90 of the network-on-chip includes: a processor 901 and a memory 902 communicatively connected to the processor 901; the memory 902 stores instructions executable by the processor 901;
[0231] If the electronic device is a first routing node, the instructions are executed by the processor 901 to enable the processor 901 to execute:
[0232] Determine a faulty node, and determine a target edge node from the edge nodes related to the faulty node; the edge node is a node directly connected to the faulty node;
[0233] Send a reconstruction instruction to the target edge node; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the faulty node.
[0234] If the electronic device is the target edge node, the instruction is executed by the processor 901, so that the processor 901 can execute:
[0235] Receive a reconstruction instruction from the first routing node in the on-chip network; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the faulty node;
[0236] Perform logical configuration according to the reconstruction instruction.
[0237] Of course, the electronic device provided in the above embodiment and the embodiment of the corresponding method belong to the same concept. The processor 901 can also execute the routing method of the on-chip network in any of the above items. The specific implementation process can be found in the method embodiment and will not be elaborated here.
[0238] In practical applications, the electronic device 90 may further include: at least one network interface 903. Each component in the electronic device 90 is coupled together through a bus system 904. It can be understood that the bus system 904 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 904 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 5 all kinds of buses are labeled as the bus system 904. Among them, the number of the processors 901 can be at least one, and the number of the memories 902 can be at least one. The network interface 903 is used for the communication between the electronic device 90 and other devices in a wired or wireless manner.
[0239] The memory 902 in the embodiment of the present disclosure is used to store various types of data to support the operation of the electronic device 90.
[0240] The method disclosed in the above embodiments of the present disclosure can be applied to or implemented by the processor 901. The processor 901 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 901 or by instructions in the form of software. The above-mentioned processor 901 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 901 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present disclosure, it can be directly embodied as being executed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory 902. The processor 901 reads the information in the memory 902 and combines its hardware to complete the steps of the foregoing method.
[0241] In some embodiments, the electronic device 90 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components, and is used to execute the foregoing method.
[0242] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved. This is not limited herein.
[0243] In the above description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0244] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. The terms used in this disclosure are for the purpose of describing embodiments of this disclosure only and are not intended to limit this disclosure.
[0245] It should be understood that in various embodiments of this disclosure, the magnitude of the sequence number of each implementation process does not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this disclosure.
[0246] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this disclosure, "a plurality" means two or more unless otherwise specifically defined.
[0247] As described above, the above are only specific implementation manners of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by this disclosure can easily think of changes or substitutions, which should all be covered by the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.
Claims
1. A routing method for a network on chip, characterized in that: The network on chip has a plurality of routing nodes, and the method is applied to any first routing node in the network on chip, and the method includes: Determine a faulty node, and determine a target edge node from edge nodes related to the faulty node; the edge node is a node directly connected to the faulty node; A reconstruction instruction is sent to the target edge node; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the faulty node.
2. The method according to claim 1, characterized in that: The determining of the target edge node from the edge nodes related to the faulty node includes: Determining a Manhattan distance between the faulty node and an edge of the network on chip; According to the Manhattan distance, select the target edge with the smallest Manhattan distance; The edge node of the faulty node in the target edge direction is used as the target edge node.
3. The method according to claim 2, characterized in that Determining the Manhattan distance between the faulty node and the edge of the on-chip network includes: Determining at least one edge node of the faulty node; The Manhattan distance between the faulty node and the edge in the direction of each edge node in the on-chip network is determined.
4. The method according to claim 3, characterized in that If there are at least two target edge directions with the smallest Manhattan distances, determining the target edge node from the edge nodes related to the faulty node includes: A target edge node is randomly selected from the edge nodes in the target edge direction of the faulty node with the at least two smallest Manhattan distances.
5. The method according to claim 1, characterized in that The determining of the faulty node comprises at least one of the following: If it is determined that the periodic signal of the second routing node is not received, it is determined that the second routing node is a faulty node; the second routing node is an edge node adjacent to the first routing node; A first signal is sent to the second routing node, and if no feedback signal from the second routing node is received, the second routing node is determined to be a faulty node.
6. The method according to claim 1, characterized in that If the target edge node is itself, the method further includes: The node is configured as a logical reconstruction node of the failed node.
7. The method according to claim 1, characterized in that If the target edge node is itself, the method further includes: Determine a reconstruction direction according to a position relationship between the faulty node and itself; Determining a routing node to be updated according to the reconstruction direction; An update signal is sent to the routing node to be updated, where the update signal is used to instruct the routing node to be updated to update its own routing identifier.
8. The method according to claim 1, characterized in that If the target edge node is itself, the method further includes: determining the communication requirements of the data packets to be forwarded; Establishing a bypass connection according to the connection relationship of the faulty node and the communication requirements of the data packet to be forwarded; The data packet is transmitted using the established bypass connection.
9. The method according to claim 1, characterized in that: If the faulty node is a destination node of the data packet, the method further includes: A transmission feedback signal is sent to the originating node of the data packet, wherein the transmission feedback signal is used to inform the destination node of the data packet of a failure.
10. The method according to claim 1, characterized in that The method further comprises: A reconstruction result signal is received from the logical reconstruction node, where the reconstruction signal is used to indicate a result of a reconstruction operation.
11. The method according to claim 1, characterized in that: The on-chip network is a two-dimensional mesh on-chip network, or a two-dimensional mesh on-chip network.
12. A routing method for a network on chip, characterized in that: The network on chip has a plurality of routing nodes, and the method is applied to a target edge node, and the method comprises: Receiving a reconstruction instruction from a first routing node in the network on chip; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the failed node; Logic configuration is performed according to the reconstruction instruction.
13. The method according to claim 12, characterized in that The target edge node has a crossbar switch; the logical configuration according to the reconstruction instruction includes: determining the communication requirements of the data packets to be forwarded; Establishing a bypass connection according to the connection relationship of the faulty node and the communication requirements of the data packet to be forwarded; the bypass connection is used to prevent the transmission of the data packet from passing through the cross switch; The data packet is transmitted using the bypass connection.
14. The method according to claim 12, characterized in that The performing logic configuration according to the reconstruction instruction includes: The self-routing identifier is adjusted according to the faulty node.
15. The method according to claim 12 or 14, characterized in that The method further comprises: Determine a reconstruction direction according to a position relationship between the faulty node and itself; Determining a routing node to be updated according to the reconstruction direction; An update signal is sent to the routing node to be updated, where the update signal is used to instruct the routing node to be updated to update the routing identifier.
16. The method according to claim 12, characterized in that The method further comprises: A reconstruction result signal is sent to the first routing node, where the reconstruction signal is used to indicate a result of the reconstruction operation.
17. A routing device for a network on chip, characterized in that: The network on chip has a plurality of routing nodes, and the device is applied to any first routing node in the network on chip, and the device includes: A first processing module is used to determine a faulty node and determine a target edge node from edge nodes related to the faulty node; the edge node is a node directly connected to the faulty node; The first communication module is used to send a reconstruction instruction to the target edge node if the target edge node is not itself; the reconstruction instruction is used to instruct the target edge node to be configured as a logical reconstruction node of the faulty node.
18. A routing device for a network on chip, characterized in that: The network on chip has a plurality of routing nodes, and the device is applied to a target edge node, and the device includes: A first communication module, configured to receive a reconstruction instruction from a first routing node in the network on chip; The second processing module is used to configure a logical reconstruction node of the faulty node according to the reconstruction instruction.
19. An electronic device with a network on chip, characterized in that: include: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1 to 11; or, execute the method of any one of claims 12 to 16.
20. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 11; or, the computer instructions are used to cause a computer to execute the method according to any one of claims 12 to 16.