Fault link processing method and device, electronic equipment and storage medium

By passing fault information and data packets between nodes in the on-chip network, the resource occupation problem caused by node perception of all links is solved, and efficient fault awareness and low-latency transmission are achieved.

CN120200964APending Publication Date: 2025-06-24SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371058.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Each node in the on-chip network needs to perceive all links in the entire network, occupying a large amount of computing resources and network bandwidth, resulting in large system delays.

Method used

The node receives the fault information of adjacent fault nodes, and when receiving the data packet, it determines the next hop node based on the fault information and routing information, attaches the fault information to the data packet, and sends it to the next hop node to realize the transmission of fault information.

Benefits of technology

Each node only needs to detect failures with adjacent nodes, which reduces link awareness requirements for the entire network, saves computing resources and network bandwidth, and reduces system latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200964A_ABST
    Figure CN120200964A_ABST
Patent Text Reader

Abstract

The invention discloses a fault link processing method and device, electronic equipment and a storage medium, and relates to the technical field of computers, and the method comprises the steps: receiving fault information sent by a first fault node, the first fault node is an adjacent node of a node, the first fault node is adjacent to a second fault node, and the second fault node is adjacent to the first fault node; a link between the first fault node and the second fault node is a fault link; when a first data packet sent by a previous node is received, a next hop node corresponding to the first data packet is determined according to the fault information and the first routing information, and the first data packet carries the first routing information; and adding the fault information to the first data packet to obtain a second data packet, and sending the second data packet to the next hop node. According to the method and the device, the problem that each node of the network-on-chip needs to perceive all links in the whole network-on-chip, a large amount of computing resources and network bandwidth are occupied, and the whole system-on-chip is relatively long in delay can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, electronic device, and storage medium for processing a faulty link. Background Art

[0002] A Network on Chip (NoC) is a multi-core interconnection bus architecture and a main component of multi-core technology. Compared with traditional bus-based systems, a Network on Chip can provide richer routing paths, making data transmission in the Network on Chip more efficient and the system more scalable. Due to its relatively complex interconnection structure and links, various faults may occur in the Network on Chip, and the most common fault is a link fault.

[0003] In the face of link faults in the Network on Chip, the Network on Chip can adopt a dedicated monitoring module, enabling each node to sense all links in the entire Network on Chip to ensure that each node can promptly sense link faults and adopt corresponding strategies to handle the faults. This approach will consume a large amount of computing resources and network bandwidth, resulting in a relatively large delay in the entire on-chip system. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, and storage medium for processing a faulty link, so as to at least solve the problem in the related art that each node in the Network on Chip needs to sense all links in the entire Network on Chip, consuming a large amount of computing resources and network bandwidth, resulting in a relatively large delay in the entire on-chip system.

[0005] This application provides a method for processing a faulty link, which is applied to a node in the Network on Chip. The method includes: receiving fault information sent by a first faulty node, where the first faulty node is an adjacent node of the node, the first faulty node is adjacent to a second faulty node, and the link between the first faulty node and the second faulty node is a faulty link; when receiving a first data packet sent by the previous node, determining a next-hop node corresponding to the first data packet according to the fault information and the first routing information, where the first data packet carries the first routing information; attaching the fault information to the first data packet to obtain a second data packet, and sending the second data packet to the next-hop node.

[0006] The present application also provides a processing device for a faulty link. The device includes: a transceiver module, configured to receive fault information sent by a first faulty node, where the first faulty node is an adjacent node of a node, the first faulty node is adjacent to a second faulty node, and the link between the first faulty node and the second faulty node is a faulty link; a processing module, configured to, when receiving a first data packet sent by a previous node, determine a next-hop node corresponding to the first data packet according to the fault information and the first routing information, where the first data packet carries the first routing information; and the transceiver module, further configured to attach the fault information to the first data packet to obtain a second data packet, and send the second data packet to the next-hop node.

[0007] The present application also provides an electronic device, including: a memory, configured to store a computer program; and a processor, configured to implement the steps of any one of the above-mentioned processing methods for a faulty link when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned processing methods for a faulty link are implemented.

[0009] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned processing methods for a faulty link are implemented.

[0010] Through the present application, it is possible to receive fault information sent by a first faulty node, and when receiving a first data packet sent by a previous node, attach the fault information to the first data packet to obtain a second data packet, and send the second data packet to the next-hop node.

[0011] Since each node only needs to detect faults with adjacent nodes, the total number of faulty links that each node needs to detect is very small, and there is no need to sense all the links in the entire on-chip network. At the same time, after attaching the fault information to the first data packet, as the subsequent routing of the data packet can be transmitted to more nodes, ultimately each node can sense this fault information. Furthermore, it is possible to achieve that all nodes can sense the faulty links in the entire on-chip network without monitoring all the links, without occupying a large amount of computing resources and network bandwidth, and thus reducing the latency of the entire on-chip system. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0013] Figure 1Topological diagram of the fault link processing system provided by the embodiments of the present application;

[0014] Figure 2 Flowchart of the fault link processing method provided by the embodiments of the present application;

[0015] Figure 3 Schematic diagram of a method for representing nodes related to on-chip network link faults provided by the embodiments of the present application;

[0016] Figure 4 Schematic diagram of the data packet structure of the second data packet provided by the embodiments of the present application;

[0017] Figure 5 Flowchart of another fault link processing method provided by the embodiments of the present application;

[0018] Figure 6 Schematic diagram of the first on-chip network link fault provided by the embodiments of the present application

[0019] Figure 7 Schematic diagram of the second on-chip network link fault provided by the embodiments of the present application;

[0020] Figure 8 Schematic diagram of the third on-chip network link fault provided by the embodiments of the present application;

[0021] Figure 9 Schematic diagram of the fourth on-chip network link fault provided by the embodiments of the present application;

[0022] Figure 10 Structural block diagram of the fault link processing device provided by the embodiments of the present application;

[0023] Figure 11 Schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0025] It should be noted that in the description of this application, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0026] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further describes this application in detail with reference to the accompanying drawings and specific embodiments.

[0027] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the processing method of the faulty link depends, the specific application environment architecture or specific hardware architecture is described herein.

[0028] The embodiment of the present invention is applied to the scenario where a faulty link appears in a network-on-chip.

[0029] In the face of link failures in a network-on-chip, usually a detour strategy is adopted to bypass the faulty link for data packet transmission. Since the detour strategy cannot calculate the minimum routing path of the data packet, when the data packet is transmitted in the network-on-chip, the number of hops increases, resulting in a large delay, low transmission efficiency, and poor performance of the entire network-on-chip.

[0030] In addition, to solve the problem that each node in the network-on-chip needs to sense all the links in the entire network-on-chip, which will occupy a large amount of computing resources and network bandwidth, resulting in a large delay in the entire on-chip system.

[0031] To solve the above problems, the embodiment of this application provides a method for processing a faulty link. The method includes: receiving fault information sent by a first faulty node, where the first faulty node is an adjacent node of a node, the first faulty node is adjacent to a second faulty node, and the link between the first faulty node and the second faulty node is a faulty link; when receiving a first data packet sent by the previous node, determining the next-hop node corresponding to the first data packet according to the fault information and the first routing information, where the first data packet carries the first routing information; attaching the fault information to the first data packet to obtain a second data packet, and sending the second data packet to the next-hop node. In this way, it can at least solve the problem in the related art that each node in the network-on-chip needs to sense all the links in the entire network-on-chip, occupying a large amount of computing resources and network bandwidth, resulting in a large delay in the entire on-chip system.

[0032] The following takes Figure 1 the processing system 100 of the faulty link shown as an example to describe the method provided by the embodiment of this application.Figure 1 This is only a schematic diagram and does not constitute a limitation on the applicable scenarios of the technical solution provided by this application.

[0033] As Figure 1 shown, Figure 1 This is a topology diagram of the fault link processing system provided by an embodiment of this application. Figure 1 In this, the fault link processing system 100 may include a node 101, a previous node 102, a next-hop node 103, and a first fault node 104.

[0034] The node 101, the previous node 102, the next-hop node 103, or the first fault node 104 in this embodiment may be any device with computing and communication capabilities. For example, the node 101, the previous node 102, the next-hop node 103, or the first fault node 104 may be a network-on-chip routing node, and each routing node is connected to a processor core, a memory, and a peripheral interface below.

[0035] Figure 1 The shown fault link processing system 100 is only for illustration and is not used to limit the technical solution of this application. Those skilled in the art should understand that in the specific implementation process, the fault link processing system 100 may also include other nodes in the network-on-chip, without limitation.

[0036] According to an embodiment of the present invention, an embodiment of a method for processing a fault link is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0037] In an embodiment of the present invention, a method for processing a fault link is provided, which can be used for the above-mentioned node 101, Figure 2 This is a flowchart of the method for processing a fault link provided by an embodiment of this application. As Figure 2 shown, this process includes the following steps:

[0038] S201: Receive fault information sent by the first fault node.

[0039] The first fault node is an adjacent node of the node. The first fault node is adjacent to the second fault node, and the link between the first fault node and the second fault node is a fault link.

[0040] The fault information is the information of the faulty link between the first fault node and the second fault node. The fault information includes the first node coordinates of the first fault node, the second node coordinates of the second fault node, and the first transfer node coordinates of each first transfer node in the first transfer node set corresponding to the faulty link; the first transfer node set includes multiple nodes adjacent to the first fault node except the second fault node, and multiple nodes adjacent to the second fault node except the first fault node.

[0041] Exemplarily, fault detection lines are provided at the end nodes of each link in the on-chip network, and online signals are periodically exchanged. With the timing period being t l and the time threshold being t m as an example, when the transmission of the tail flit of the previous data packet on a link ends, timing starts. After t l / 2, the node that sends the tail flit sends an online signal. The node that receives the tail flit receives the online signal and feeds back its online signal to the other party after t l / 2, and the online signals are cyclically exchanged with the period t l The two end nodes determine that the link between the two end points has a fault, that is, the link fault is sensed, and the fault information is transmitted to its adjacent nodes.

[0042] For example, when node 101 is a transfer node, if the first fault node does not receive the online signal returned by the second fault node after exceeding the time threshold t m , the first fault node determines that the link between the first fault node and the second fault node has a fault, and sends the fault information to the transfer node adjacent to the first fault node. The transfer node receives the fault information sent by the first fault node.

[0043] Among them, the time threshold t m can be set to the maximum allowable delay of the on-chip network. The timing period t l can be set to N clock cycles corresponding to the on-chip network. For example, N is 100.

[0044] It can be understood that the time threshold and the timing period can be set according to actual needs without limitation.

[0045] In the embodiments of the present invention, in order to accurately describe the position of the faulty link, as Figure 3 shown, Figure 3 is a schematic diagram of a method for representing nodes related to on-chip network link faults provided by an embodiment of the present application. In Figure 3 , Figure 3It is an 8×8 2D Mesh network-on-chip. In this network-on-chip, two links have failed, namely the vertical failed link (3, 3)-(3, 4) and the horizontal failed link (5, 1)-(6, 1). The nodes (3, 3) and (5, 1) with smaller ordinate or abscissa in the two failed links are called E00, and the nodes (3, 4) and (6, 1) at the other end coordinates are called E10. The adjacent node sets of the end nodes E00 and E10 of the failed link are {N ij}, i = 0, 1, j = 1, 2, 3, 4. Starting from the western node and counterclockwise, each node is numbered. After excluding the end nodes of the link failure, there will be omissions in the numbering (for example, omitting N 04 and N 12 ). Therefore, based on the coordinate relationship, if the coordinates of the E00 node are known, the coordinates of E10 and {N ij} can be obtained. Except for the above nodes, the remaining nodes in the network-on-chip form a node set {O}.

[0046] Optionally, when two links of the E00 node have failed, label overlaps may occur when using the above method. Since the result calculated based on the coordinate relationship is unique, N ij is any one of the overlapping labels.

[0047] Optionally, the source node corresponding to the first data packet can be a node in {O}. When the destination node is E00 or E10, a node in {N ij} can be selected as the relay node.

[0048] It can be understood that after the relay node receives the failure information sent by the first failure node, the router in the relay node reconfigures the port weights, updates the routing table information, and enables the relay function to perform relay routing on the data packets passing through the current node (i.e., the relay node) and requiring to pass through the failed link, and forwards the data packets.

[0049] It can be understood that if the source node does not receive the response information from the destination node within the time threshold t m , it is considered that the transmitted data packet has been lost and the data packet needs to be retransmitted.

[0050] S202: When receiving the first data packet sent by the previous node, determine the next-hop node corresponding to the first data packet according to the failure information and the first routing information.

[0051] Among them, the first data packet is also called a package, which can be divided into a head flit, a body flit, and a tail flit. The head flit contains the first routing information, the source node address, and the destination node address; the body flit contains the data body; the tail flit indicates the end of the data packet, and the tail flit contains a check code, which is used to check whether the first data packet is complete.

[0052] In some alternative embodiments, when receiving a first data packet sent by a previous node, based on the routing rules corresponding to the system-on-chip and the first routing information, calculate the first routing path for the next two hops of the first data packet when there is no fault in the on-chip network; based on the fault information, determine whether the first routing path passes through a faulty link; if so, obtain the routing table information; determine a target preset path from at least one preset path, and determine the node adjacent to the node in the target preset path as the next-hop node.

[0053] Among them, the routing table information includes at least one preset path, and each preset path passes through a node. For example, for a single-link fault of a node, the at least one preset path included in the routing table information has four paths. Among them, two paths can be paths that U-shaped bypass from both sides of the end node, and the other two paths can be paths that U-shaped bypass from the adjacent nodes of the end node. Optionally, the at least one preset path can be recorded in the routing table information or can be set in the router in advance.

[0054] In one example, a node determines a target preset path from at least one preset path, including: randomly selecting any one preset path from at least one preset path and determining the randomly selected one preset path as the target preset path; or, obtaining the transmission frequency of each preset path in at least one preset path and determining the preset path with the lowest transmission frequency as the target preset path.

[0055] Optionally, if the first routing route does not pass through a faulty link, the node determines the node adjacent to the node in the first routing path as the next-hop node.

[0056] S203: Append the fault information to the first data packet to obtain a second data packet, and send the second data packet to the next-hop node.

[0057] In one example, the relay node appends the fault information to the tail flit of the first data packet to obtain a second data packet, and sends the second data packet to the next-hop node.

[0058] In one example, as Figure 4 shown, Figure 4 is a schematic diagram of the data packet structure of the second data packet provided by the embodiment of the present application. In Figure 4 , the header flit of the second data packet includes the source node coordinates, the destination node coordinates, the relay node coordinates, and the routing information; the body flit of the second data packet includes data; the tail flit of the second data packet includes the faulty node coordinates related to the faulty link, the adjacent node set {N} coordinates, and the check code.

[0059] Optionally, the node can also record the destination node of the data packet with fault information attached, and at the same time record the source node of the data packet header flit with the transit node information set, to ensure that the fault information is sent to the destination node and the transit node at most once and all nodes in {O} receive the fault information. Furthermore, the node stops attaching the fault information after determining that all nodes in {O} perceive the fault information based on the record. Similarly, if the faulty link function is restored, the node can also use the same method to send link function recovery information until all nodes in {O} perceive that the faulty link has been restored, which will not be elaborated here.

[0060] In one example, the failure of a failed link is divided into a vertical link failure and a horizontal failed link in the system on chip.

[0061] For example, when the node is a transit node, for a longitudinal link failure, when N ij When the horizontal coordinate is an odd number, the transit node sends fault information to the node with an odd horizontal coordinate in {O} and records the node with an odd horizontal coordinate that has sensed the fault; when N ij When the horizontal coordinate is an even number, the fault information is sent to the node with an even horizontal coordinate in {O} and the source node with an even horizontal coordinate that has sensed the fault is recorded. ij When the vertical coordinate is an odd number, send fault information to the nodes with odd vertical coordinates in {O} and record the source nodes with odd vertical coordinates that have sensed the fault; when N ij When the ordinate is an even number, the fault information is sent to the node with an even ordinate in {O} and the source node with an odd ordinate that has sensed the fault is recorded.

[0062] It can be understood that the transmission node sends fault information according to the odd and even classification of the node coordinates, which can reduce information redundancy and improve the efficiency of fault information transmission.

[0063] It can be understood that before the transit node receives the fault information, the corresponding bits of the transit node coordinates included in the data packet header flit are 0; after receiving the fault information, the transit node coordinates are added to the data packet header flit; after receiving the link fault resolution information, the corresponding bits of the transit node coordinates are all returned to 0.

[0064] Based on the above Figure 2 According to the method, a node can receive fault information sent by a first faulty node, and when receiving a first data packet sent by a previous node; the fault information is attached to the first data packet to obtain a second data packet, and the second data packet is sent to a next hop node.

[0065] Since each node only needs to detect faults between adjacent nodes, the total number of faulty links that each node needs to detect is small, and there is no need to sense all the links in the entire on-chip network. At the same time, after attaching the fault information to the first data packet, as the packet is routed later, it can be transmitted to more nodes, and finally each node can sense this fault information. Thus, all nodes can sense the faulty links in the entire on-chip network without monitoring all the links, without occupying a large amount of computing resources and network bandwidth, and thus reducing the latency of the entire on-chip system.

[0066] Further, after obtaining at least one fault information of all nodes in the on-chip network, the transmission path of subsequent data packets can be planned based on the at least one fault information. Specifically, as Figure 5 shown, Figure 5 FIG. is a flowchart of another method for processing faulty links provided by an embodiment of the present application. When the node is a source node, the process includes the following steps:

[0067] S501: Obtain at least one fault information of at least some nodes in the on-chip network, and obtain a third data packet.

[0068] Wherein, the third data packet carries second routing information, and the second routing information carries the first destination node address of the first destination node and the first source node address of the first source node.

[0069] The at least one fault information corresponds to a set of faulty nodes and each faulty node address.

[0070] S502: Determine whether the first destination node is a third faulty node according to the at least one fault information.

[0071] The third faulty node is any one of the faulty nodes in the set of faulty nodes corresponding to the at least one fault information. The link between the third faulty node and the fourth faulty node is the target faulty link.

[0072] In one example, the node determines whether the first destination node address is one of the addresses of each faulty node in the set of faulty nodes according to each faulty node address in the set of faulty nodes and the first destination node address; if so, it is determined that the first destination node is a third faulty node; if not, it is determined that the first destination node is not a third faulty node.

[0073] In some alternative embodiments, if the first destination node is not a third faulty node, the third data packet is transmitted to the first destination node according to the second routing information and the routing rules corresponding to the on-chip system.

[0074] S503: If so, determine the target relay node corresponding to the third data packet according to the target fault information corresponding to the third faulty node.

[0075] The target fault information includes the coordinates of each transit node in the set of transit nodes corresponding to the target fault link; the set of transit nodes includes multiple transit nodes adjacent to the third fault node except the fourth fault node, and multiple transit nodes adjacent to the fourth fault node except the third fault node.

[0076] In some alternative embodiments, the node obtains the number of transmissions of each packet transmitted by each transit node in the set of transit nodes; determines a weight value corresponding to each transit node in the set of transit nodes according to the number of transmissions corresponding to each transit node; and determines the target transit node from the set of transit nodes according to at least one weight value.

[0077] Wherein, each weight value corresponds to a transit node one by one.

[0078] In one example, the node selects the maximum weight value from at least one weight value; and determines the second transit node corresponding to the maximum weight value as the target transit node.

[0079] Optionally, each node weight can also be associated with the positional relationship between the coordinates of the destination node and the coordinates of the source node. Embodiments of the present application can also determine the weight of each node (the weight is represented by w) corresponding to each transit node in the set of transit nodes {N ij} based on the positional relationship between the coordinates of the destination node and the coordinates of the source node.

[0080] Exemplary 1, taking Figure 6 as an example, Figure 6 is a schematic diagram of the first on-chip network link fault provided by the embodiment of the present application; if the source node is located southwest of the longitudinal fault link, the source node is O1, the destination node is E10, and the selectable transit nodes are N11 or N14. Since using N14 as the transit node increases the number of hops, the node weight w11>w14 is set.

[0081] If the source node is located southeast of the longitudinal fault link, the source node is O2, the destination node is E10, and the selectable transit nodes are N13 or N14. Since using N14 as the transit node increases the number of hops, the node weight w13>w14 is set.

[0082] Exemplary 2, taking Figure 7 as an example, Figure 7 is a schematic diagram of the second on-chip network link fault provided by the embodiment of the present application; if the source node is located northwest of the longitudinal link fault link, the source node is O3, the destination node is E00, and the selectable transit nodes are N01, N02. The node weight w01>w02 is set.

[0083] If the source node is located northeast of the longitudinal link failure link, the source node is O4, the destination node is E00, the selectable relay nodes are N03 and N02, and the node weights are set as w03 > w02.

[0084] Exemplary 3, taking Figure 8 as an example, Figure 8 is a schematic diagram of the third on-chip network link failure provided by the embodiment of the present application; if the source node is located southwest of the horizontal failure link, the source node is O5, the target node is E10, the selectable relay nodes are N12 and N13, and since using N13 as the relay node increases the number of hops, the node weights are set as w12 > w13.

[0085] If the source node is located northwest of the horizontal failure link, the source node is O6, the target node is E10, the selectable relay nodes are N14 and N13, and since using N13 as the relay node increases the number of hops, the node weights are set as w14 > w13.

[0086] Exemplary 4, taking Figure 9 as an example, Figure 9 is a schematic diagram of the fourth on-chip network link failure provided by the embodiment of the present application; when the source node is located southeast of the horizontal link failure, the source node is O7, the destination node is E00, the selectable relay nodes are N02 and N01, and the node weights are set as w02 > w01.

[0087] When the source node is located northeast of the horizontal link failure, the source node is O8, the destination node is E00, the selectable relay nodes are N04 and N01, and the node weights are set as w04 > w01.

[0088] It can be understood that since the target relay node of the third data packet can be determined based on each node weight, the relay node with a larger weight can be set as the target relay node to balance the load of the relay nodes.

[0089] S504: Add the address of the target relay node to the third data packet to obtain a fourth data packet.

[0090] S505: Transmit the fourth data packet to the target relay node according to the address of the target relay node.

[0091] It can be understood that after obtaining at least one failure information of all nodes in the on-chip network, the transmission path of subsequent data packets can be planned based on the at least one failure information, effectively avoiding the failure link, ensuring that the data packets can be transmitted normally to the greatest extent, and reducing the delay.

[0092] In this embodiment, a processing device for a faulty link is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0093] This embodiment provides a processing device for a faulty link, as Figure 10 shown Figure 10 is a structural block diagram of the processing device for a faulty link provided by an embodiment of the present application; it is applied to a node in a network-on-chip. The device includes: a transceiver module 1001, configured to receive fault information sent by a first faulty node, where the first faulty node is an adjacent node of the node, the first faulty node is adjacent to a second faulty node, and the link between the first faulty node and the second faulty node is a faulty link.

[0094] A processing module 1002, configured to determine a next-hop node corresponding to the first data packet according to the fault information and the first routing information when receiving a first data packet sent by the previous node, where the first data packet carries the first routing information.

[0095] The transceiver module 1001 is further configured to attach the fault information to the first data packet to obtain a second data packet, and send the second data packet to the next-hop node.

[0096] In some alternative implementation manners, the processing module 1002 is specifically configured to, when receiving a first data packet sent by the previous node, calculate a first routing path for the next two hops of the first data packet when there is no fault in the network-on-chip based on the routing rules corresponding to the system-on-chip and the first routing information; based on the fault information, determine whether the first routing path passes through a faulty link; if so, obtain routing table information, where the routing table information includes at least one preset path, and each preset path passes through a node; determine a target preset path from the at least one preset path, and determine the node adjacent to the node in the target preset path as the next-hop node.

[0097] In some alternative implementation manners, the processing module 1002 is further specifically configured to, if the first routing route does not pass through a faulty link, determine the node adjacent to the node in the first routing path as the next-hop node.

[0098] In some alternative embodiments, the transceiver module 1001 is further configured to obtain at least one fault information of at least some nodes in the network-on-chip, and obtain a third data packet, where the third data packet carries second routing information, and the second routing information carries a first destination node address of a first destination node and a first source node address of a first source node; the processing module 1002 is further configured to determine, according to the at least one fault information, whether the first destination node is a third faulty node, where the third faulty node is any faulty node in a set of faulty nodes corresponding to the at least one fault information; if so, determine a target relay node corresponding to the third data packet according to target fault information corresponding to the third faulty node; the transceiver module 1001 is further configured to add an address of the target relay node to the third data packet to obtain a fourth data packet; and transmit the fourth data packet to the target relay node according to the address of the target relay node.

[0099] In some alternative embodiments, the processing module 1002 is further configured to, if the first destination node is not a third faulty node, transmit the third data packet to the first destination node according to the second routing information and a routing rule corresponding to the system-on-chip.

[0100] In some alternative embodiments, a link between the third faulty node and a fourth faulty node is a target faulty link; the target fault information includes each relay node coordinate of each relay node in a set of relay nodes corresponding to the target faulty link; the set of relay nodes includes multiple relay nodes adjacent to the third faulty node except the fourth faulty node, and multiple relay nodes adjacent to the fourth faulty node except the third faulty node; the processing module 1002 is specifically further configured to obtain each transmission count of each relay node in the set of relay nodes for transmitting data packets; determine a weight value corresponding to each relay node in the set of relay nodes according to each transmission count corresponding to each relay node, where each weight value corresponds to a relay node; and determine a target relay node from the set of relay nodes according to at least one weight value.

[0101] In some alternative embodiments, the processing module 1002 is specifically further configured to select a maximum weight value from the at least one weight value; and determine a second relay node corresponding to the maximum weight value as the target relay node.

[0102] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.

[0103] For the description of the features corresponding to the embodiments of the processing device for the faulty link, reference can be made to the relevant description of the embodiments of the method for processing the faulty link, which will not be elaborated here one by one.

[0104] Embodiments of the present application further provide an electronic device, such as Figure 11 shown Figure 11 is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present application; the electronic device includes a memory and a processor, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described method embodiments for processing a faulty link.

[0105] Embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any of the above-described method embodiments for processing a faulty link when running.

[0106] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs.

[0107] Embodiments of the present application further provide a computer program product, the above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described method embodiments for processing a faulty link are implemented.

[0108] Embodiments of the present application further provide another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-described method embodiments for processing a faulty link are implemented.

[0109] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0110] The above has introduced in detail a method, apparatus, electronic device, and storage medium for processing a faulty link provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for processing a faulty link, characterized in that: Applied to a node in a network on chip, the method comprises: receiving fault information sent by a first faulty node, where the first faulty node is an adjacent node of the node, the first faulty node is adjacent to a second faulty node, and a link between the first faulty node and the second faulty node is a faulty link; when receiving a first data packet sent by a previous node, determining a next hop node corresponding to the first data packet according to the fault information and first routing information, where the first data packet carries the first routing information; The fault information is appended to the first data packet to obtain a second data packet, and the second data packet is sent to the next hop node.

2. The method for processing a faulty link according to claim 1, characterized in that: The method of, when receiving a first data packet sent by a previous node, determining a next hop node corresponding to the first data packet according to the fault information and the first routing information, comprises: When receiving the first data packet sent by the previous node, based on the routing rule corresponding to the system on chip and the first routing information, a first routing path of the first data packet for subsequent two hops when the network on chip is fault-free is calculated; Based on the fault information, determining whether the first routing path passes through the faulty link; If yes, obtain routing table information, the routing table information includes at least one preset path, each of the preset paths passes through the node; A target preset path is determined from the at least one preset path, and a node adjacent to the node in the target preset path is determined as the next hop node.

3. The method for processing a faulty link according to claim 2, characterized in that: The method further comprises: If the first routing route does not pass through the faulty link, a node adjacent to the node in the first routing path is determined as the next hop node.

4. The method for processing a faulty link according to claim 3, characterized in that: The method further comprises: Acquire at least one fault information of at least part of the nodes in the on-chip network, and acquire a third data packet, wherein the third data packet carries second routing information, and the second routing information carries a first destination node address of a first destination node and a first source node address of a first source node; Determine, according to the at least one fault information, whether the first destination node is a third fault node, where the third fault node is any one fault node in the set of fault nodes corresponding to the at least one fault information; If yes, determining the target transfer node corresponding to the third data packet according to the target fault information corresponding to the third fault node; Adding the address of the target transfer node to the third data packet to obtain a fourth data packet; The fourth data packet is transmitted to the target transfer node according to the address of the target transfer node.

5. The method for processing a faulty link according to claim 4, characterized in that: The method further comprises: If the first destination node is not the third faulty node, the third data packet is transmitted to the first destination node according to the second routing information and the routing rule corresponding to the system on chip.

6. The method for processing a faulty link according to claim 4, characterized in that: The link between the third fault node and the fourth fault node is a target fault link; the target fault information includes each transfer node coordinate of each transfer node in the transfer node set corresponding to the target fault link; the transfer node set includes a plurality of transfer nodes adjacent to the third fault node except the fourth fault node, and a plurality of transfer nodes adjacent to the fourth fault node except the third fault node; The step of determining the target transfer node corresponding to the third data packet according to the target fault information corresponding to the third fault node includes: Obtaining each transmission number of each data packet transmitted by each transfer node in the transfer node set; Determine a weight value corresponding to each of the transfer nodes in the transfer node set according to each of the transmission times corresponding to each of the transfer nodes, wherein each of the weight values ​​corresponds to a transfer node one-to-one; The target transfer node is determined from the transfer node set according to the at least one weight value.

7. The method according to claim 6, characterized in that The step of determining a target transfer node from the transfer node set according to the at least one weight value comprises: selecting a maximum weight value from the at least one weight value; The second transfer node corresponding to the maximum weight value is determined as the target transfer node.

8. A device for processing a faulty link, characterized in that: Applied to a node in a network on chip, the device comprises: a transceiver module, configured to receive fault information sent by a first faulty node, wherein the first faulty node is an adjacent node of the node, the first faulty node is adjacent to a second faulty node, and a link between the first faulty node and the second faulty node is a faulty link; a processing module, configured to determine, when receiving a first data packet sent by a previous node, a next hop node corresponding to the first data packet according to the fault information and first routing information, wherein the first data packet carries the first routing information; The transceiver module is further configured to attach the fault information to the first data packet to obtain a second data packet, and send the second data packet to the next hop node.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for processing a failed link as claimed in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for processing a failed link according to any one of claims 1 to 7.