Arbitration method and related equipment for on-chip network transmission information
By determining the request weight value based on the location and historical information of the target node in the on-chip network, sorting and responding to the request, the problem of traffic load imbalance is solved, and the fault tolerance of the network and the fairness of resource allocation are improved.
Patent Information
- Application Number
- CN202510741445.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-06-05
AI Technical Summary
There is a problem of traffic load imbalance in on-chip networks, and the existing technology has not been effectively solved.
During the arbitration cycle, the weight value of the request to be sent is determined based on the location and historical information of the target node, and the weight value is sorted and responded in the order of the weight value from high to low, and the request is processed using a cross switch.
It effectively avoids traffic load imbalance in the on-chip network, improves the fault tolerance of transmission information and the fairness of resource allocation.
Smart Images

Figure CN120263759B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an arbitration method for network-on-chip (NOC) transmission information and related equipment. Background Art
[0002] Network-on-Chip (NoC) is a new communication architecture used in System-on-Chip (SoC). It connects multiple processor cores, memory, and other functional modules within the chip in a networked manner, solving the bottleneck problem of on-chip communication.
[0003] Typically, the nodes in a NoC are arranged in a grid pattern, and an arbitrator is installed on each node in the NoC. The main function of the arbitrator is to solve the problem of multiple requests competing for the same resource at the same time. The arbitration mechanism set up in the arbitrator can provide a fairer response opportunity for each request. However, in related technologies, the arbitrator in the NoC is designed based on the information of the node itself, without considering the changes in the overall traffic in the NoC. This can lead to an imbalance in the traffic load of the NoC. Currently, there is no effective solution to this technical problem. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide an arbitration method and related equipment for information transmission in an on-chip network, so as to solve the technical problem of unbalanced traffic load in the on-chip network in the related art.
[0005] In order to solve the above technical problems, the present invention provides an arbitration method for information transmission in an on-chip network, which is applied to a target node in an on-chip network arranged in a grid shape, wherein the target node is any node in the on-chip network, comprising:
[0006] Determine the requests to be sent by the target node in each input direction in the current arbitration cycle to obtain a target set;
[0007] Determine a weight value corresponding to each pending request in the target set based on a failure status of a node closest to the target node in a message transmission path of the target request, the number of free spaces in the flit size, the number of historical requests in the same transmission direction, and a historical weight value; the target request is any pending request in the target set; and the historical weight value is a weight value corresponding to each pending request sent by the target node in a previous arbitration cycle;
[0008] The requests to be sent in the target set are sorted in descending order of weight value, so as to respond to the sorted requests to be sent by using the crossbar switch on the target node.
[0009] In a specific embodiment of the present application, determining the requests to be sent by the target node in each input direction in the current arbitration cycle to obtain the target set includes:
[0010] Taking the position of the target node in the on-chip network as a reference point, determine the requests to be sent by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction and the local incoming direction in the current arbitration cycle, and determine the requests to be sent that are retained by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction and the local incoming direction in the previous arbitration cycle, to obtain the target set.
[0011] In a specific embodiment of the present application, it also includes:
[0012] The data packet header flit of each to-be-sent request in the target set is parsed to determine the output direction of each to-be-sent request in the target set relative to the target node.
[0013] In a specific embodiment of the present application, parsing the data packet header flit of each to-be-sent request in the target set includes:
[0014] The router in the target node is used to parse the header fragments of each data packet of the to-be-sent request in the target set.
[0015] In a specific embodiment of the present application, determining the requests to be sent by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the current arbitration cycle, and determining the requests to be sent that are retained by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the previous arbitration cycle, to obtain the target set, includes:
[0016] Determine pending requests cached by the target node in the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area in a current arbitration cycle, and determine pending requests cached by the target node in the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area in a previous arbitration cycle, to obtain a first subset, a second subset, a third subset, a fourth subset, and a fifth subset;
[0017] The target set includes the first subset, the second subset, the third subset, the fourth subset, and the fifth subset;
[0018] The first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area are cache areas corresponding to the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction, respectively;
[0019] The arbitration cycle is: starting from processing the pending requests cached and retained in the first cache area, the second cache area, the third cache area, the fourth cache area and the fifth cache area, and the time required to complete the processing of the pending requests cached and retained in the first cache area, the second cache area, the third cache area, the fourth cache area and the fifth cache area.
[0020] In a specific embodiment of the present application, sorting the requests to be sent in the target set in descending order of weight values so as to respond to the sorted requests to be sent using the crossbar switch on the target node includes:
[0021] Prioritizing responding to pending requests with local destination addresses in the first subset, the second subset, the third subset, the fourth subset, and the fifth subset, and sorting the pending requests in the first subset, the second subset, the third subset, the fourth subset, and the fifth subset, excluding the pending requests with local destination addresses, in descending order of weight to obtain a first request sequence, a second request sequence, a third request sequence, a fourth request sequence, and a fifth request sequence;
[0022] The first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence are sent to the crossbar switch on the target node, so that the crossbar switch is used to respond to each of the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence.
[0023] In a specific embodiment of the present application, the sending of the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence to the crossbar switch on the target node, so as to use the crossbar switch to respond to each of the to-be-sent requests in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, includes:
[0024] If there is only one request to be sent in each of the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, sending the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence to the crossbar switch on the target node, so that the crossbar switch responds to the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence at the same time;
[0025] If there is more than one request to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence are sent to the cross switch on the target node, so that the cross switch responds to the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence based on a first-in-first-out principle.
[0026] In a specific embodiment of the present application, determining the weight value corresponding to each to-be-sent request in the target set according to the failure status of the node closest to the target node in the message transmission path of the target request and the number of free spaces in the flit size, the number of historical requests in the same transmission direction, and the historical weight value includes:
[0027] Determining weight values corresponding to the other requests to be sent in the target set excluding the requests to be sent whose destination addresses are local according to the weight setting model;
[0028] Wherein, the expression of the weight setting model is:
[0029] ;
[0030] Where, Indicates a weight value corresponding to a target cache request in a target cache area, where the target cache area is any one of the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area; and the target cache request is any pending request in the target cache area except pending requests with a local destination address. is the fault state of the peripheral node, the peripheral node is the node closest to the target node in the message transmission path of the target cache request; when the peripheral node fails, , when the surrounding nodes are not faulty, ; Indicates the number of free spaces of the microchip size of the peripheral node in the peripheral cache area; the direction of the peripheral cache area relative to the peripheral node is: the corresponding direction of the target cache area relative to the target node; represents a value reassigned based on the weight values of all pending requests of the target node in the previous arbitration cycle and in the target cache area, excluding pending requests with local destination addresses; Indicates the preset coefficient; Indicates the number of pending requests in the target cache area with the same output direction as the target cache request in the previous round of arbitration cycle. If the target node does not have a pending request in the target cache area with the same output direction as the target cache request in the previous round of arbitration cycle, let .
[0031] In a specific embodiment of the present application, it also includes:
[0032] In the previous arbitration cycle and in the target cache area, obtaining weight values of other pending requests on the target node except for pending requests with local destination addresses, and sorting the pending requests in descending order of weight values to obtain a first sequence;
[0033] If there are pending requests sent in a target output direction in the first sequence, determining the pending requests sent in the target output direction in the first sequence to obtain a target screening set; the target output direction is the same as the output direction of the target cache request;
[0034] Determine the to-be-sent request corresponding to the largest weight value in the target screening set to obtain the target screening request, and filter out the to-be-sent requests that appear for the first time in different output directions from the first sequence to obtain a second sequence;
[0035] According to the order of arrangement of the target screening request in the second sequence Assign a value.
[0036] In a specific embodiment of the present application, it also includes:
[0037] If there is no request to be sent by the target output direction in the first sequence, Assign a value.
[0038] In a specific embodiment of the present application, it also includes:
[0039] A target detection signal is sent to the peripheral node, and whether the peripheral node fails is determined according to a feedback signal returned by the peripheral node.
[0040] In a specific embodiment of the present application, the sending of the target detection signal to the surrounding node and determining whether the surrounding node has a fault according to the feedback signal returned by the surrounding node includes:
[0041] Sending the target detection signal to the surrounding nodes, and determining whether the surrounding nodes can return a feedback signal corresponding to the target detection signal within a preset time;
[0042] If yes, it is determined that the peripheral node has not failed;
[0043] If not, it is determined that the peripheral node fails.
[0044] In a specific embodiment of the present application, it also includes:
[0045] When it is determined that the peripheral node fails, the destination node corresponding to the target cache request is determined, and the routing transmission path between the target node and the destination node corresponding to the target cache request is recalculated according to a routing algorithm.
[0046] In a specific embodiment of the present application, before recalculating the routing transmission path between the target node and the target node corresponding to the target cache request according to the routing algorithm, the method further includes:
[0047] The target cache request is retained in the cache area where the target cache request is located, and the target cache request is added to the next round of arbitration cycle.
[0048] In a specific embodiment of the present application, it also includes:
[0049] The credit information of the peripheral node is acquired, and the number of free spaces of the size of the microchip of the peripheral node in the peripheral cache area is determined according to the credit information of the peripheral node.
[0050] In a specific embodiment of the present application, it also includes:
[0051] Pre-establishing a first table, a second table, and a third table;
[0052] Using the first table to record the fault status of the surrounding nodes;
[0053] Using the second table to record the number of free spaces of the size of the flit of the peripheral nodes in the peripheral cache area;
[0054] The third table is used to record the weight values corresponding to the other pending requests of the target node in the previous arbitration cycle and in each cache area except the pending requests with local destination addresses.
[0055] In order to solve the above technical problems, the present invention further provides an arbitration device for information transmission in a network on chip, which is applied to a target node in a network on chip arranged in a grid shape, wherein the target node is any node in the network on chip, and includes:
[0056] a request acquisition module, configured to acquire requests to be sent by the target node in each input direction in the current arbitration cycle, and obtain a target set;
[0057] A weight calculation module is configured to determine a weight value corresponding to each pending request in the target set based on the failure status of the node closest to the target node in the message transmission path of the target request, the number of free spaces in the flit size, the number of historical requests in the same transmission direction, and a historical weight value; the target request is any pending request in the target set; and the historical weight value is the weight value corresponding to each pending request sent by the target node in the previous arbitration cycle;
[0058] The request response module is used to sort the requests to be sent in the target set in descending order of weight value, so as to respond to the sorted requests to be sent by using the crossbar switch on the target node.
[0059] In order to solve the above technical problems, the present invention further provides an arbitration device for transmitting information on a network on a chip, comprising:
[0060] memory for storing computer programs;
[0061] The processor is configured to execute the computer program to implement the steps of the arbitration method for transmitting information in an on-chip network as disclosed above.
[0062] In order to solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the arbitration method for on-chip network transmission information disclosed above are implemented.
[0063] In order to solve the above technical problem, the present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the arbitration method for on-chip network transmission information as disclosed above.
[0064] Beneficial effect: In an arbitration method for information transmission in an on-chip network provided by the present invention, the target node first determines the requests to be sent by the target node in each input direction in the current arbitration cycle to obtain a target set; then, the weight value corresponding to each request to be sent in the target set is determined according to the fault status of the node closest to the target node in the message transmission path of the target request and the number of free spaces in the micro-chip size, the number of historical requests in the same transmission direction, and the historical weight value; wherein the target request is any request to be sent in the target set; the historical weight value is the weight value corresponding to when the target node sent each request to be sent in the previous arbitration cycle; finally, the requests to be sent in the target set are sorted in descending order according to the weight value, so as to use the cross switch on the target node to respond to the sorted requests to be sent.
[0065] Compared with the related art, in the present invention, the weight value corresponding to each request to be sent in the target set is determined according to the fault status of the node closest to the target node in the message transmission path of the target request, the number of free spaces of the microchip size, the number of historical requests in the same transmission direction, and the historical weight value, and the various requests to be sent in the target set are sorted in descending order of weight value, so as to utilize the cross switch on the target node to respond to the sorted requests to be sent. Under this setting, when the target node transmits information in the on-chip network, it is equivalent to considering the changes in the overall traffic load in the on-chip network, thereby avoiding the problem of unbalanced traffic load in the on-chip network. Moreover, when the link transmission status of the request to be sent is also included in the arbitration strategy of the target node, the fault tolerance of the on-chip network when transmitting information can also be improved.
[0066] Correspondingly, the arbitration device, equipment, computer-readable storage medium, and program product for network-on-chip information transmission provided by the present invention also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0068] Figure 1 A flowchart of an arbitration method for network-on-chip transmission information provided by an embodiment of the present invention;
[0069] Figure 2 A schematic diagram of the structure of a network on chip provided by an embodiment of the present invention;
[0070] Figure 3 A schematic diagram illustrating how a target node determines the weights of each pending request in a target set based on the failure status of the node closest to the target node in the target request message transmission path, the number of free spaces in the flit size, the number of historical requests in the same transmission direction, and the historical weights.
[0071] Figure 4 A schematic diagram of sending the pending requests in each cache area of the target node to the crossbar switch;
[0072] Figure 5 A schematic diagram of a target node calculating a weight value corresponding to each request to be sent in a target set by calling the data in the first table, the second table, and the third table according to the weight setting model;
[0073] Figure 6 A structural diagram of an arbitration device for network-on-chip information transmission provided by an embodiment of the present invention;
[0074] Figure 7 This is a structural diagram of an arbitration device for transmitting information on a network on a chip provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0075] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0076] The terms "including" and "having," as used in the present description and accompanying drawings, and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements and may include steps or elements that are not listed.
[0077] In order to enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0078] See Figure 1 , Figure 1 This is a flowchart of a method for arbitrating information transmitted in a network on chip (NOC) according to an embodiment of the present invention. The method is applied to a target node in a NOC arranged in a grid shape, where the target node is any node in the NOC, and includes:
[0079] Step S11: determining the requests to be sent by the target node in each input direction in the current arbitration cycle to obtain a target set;
[0080] Step S12: Determine the weight value corresponding to each pending request in the target set based on the failure status of the node closest to the target node in the message transmission path of the target request, the number of free spaces in the flit size, the number of historical requests in the same transmission direction, and the historical weight value; the target request is any pending request in the target set; the historical weight value is the weight value corresponding to each pending request when the target node sent it in the previous arbitration cycle;
[0081] Step S13: sorting the requests to be sent in the target set in descending order of weight value, so as to use the crossbar switch on the target node to respond to the sorted requests to be sent.
[0082] In this embodiment, a method for arbitrating information transmitted on a network on a chip is provided, which can avoid the problem of unbalanced traffic load on the network on a chip. Figure 2 , Figure 2 A schematic diagram of the structure of a network on chip provided by an embodiment of the present invention. Figure 2 All nodes in the NoC are arranged in a 3x3 grid, with each node capable of transmitting information to adjacent connected nodes. If node 5 is the destination node, then nodes 2, 8, 4, and 6 are the transmission nodes for the destination node in the north, south, west, and east directions, respectively. Node 5 itself is the transmission node for node 5 in the local direction.
[0083] In practical applications, a target node may receive pending requests from multiple input directions at the same time. However, the target node cannot respond to all of these pending requests at the same time. In this case, it is necessary to determine the priority of each of these pending requests and use the crossbar switch on the target node to respond to these requests in sequence. Therefore, in this embodiment, the target node first determines the pending requests to be sent in each input direction during the current arbitration cycle to obtain a target set.
[0084] After obtaining the target set, in order to determine the priority of each request to be sent in the target set, it is necessary to determine the weight value corresponding to each request to be sent in the target set based on the failure status of the node closest to the target node in the message transmission path of the target request and the number of free spaces in the micro-slice size, the number of historical requests in the same transmission direction, and the historical weight value.
[0085] The target request is any pending request in the target set; the historical weight value is the weight value corresponding to each pending request sent by the target node in the previous arbitration cycle; and the number of free spaces of the micro-chip size refers to: after the target node's input buffer area divides its cache space by micro-chip size, a space will be divided on the input buffer interval of the target node. These spaces will wait for the reception of micro-chips, and the number of spaces not occupied by micro-chips.
[0086] See Figure 3 , Figure 3 This is a diagram showing how the target node determines the weights of each pending request in the target set based on the failure status of the node closest to the target node in the target request message transmission path, the number of free spaces in the flit size, the number of historical requests in the same transmission direction, and the historical weight values. Figure 3 In the figure, the left arrows outside the target node represent the pending requests received by the target node in the four incoming directions of east, west, south, and north, and the right arrows outside the target node represent the requests sent by the target node in the four outgoing directions of east, west, south, and north.
[0087] When determining the weight values corresponding to each request to be sent in the target set, the purpose of introducing the fault status of the node closest to the target node in the message transmission path of the target request is to determine whether there is a fault or abnormality in the node closest to the target node in the message transmission path of the target node. The purpose of introducing the number of free spaces in the micro-slice size of the node closest to the target node in the message transmission path of the target request is to judge the overall traffic load in the on-chip network. The purpose of introducing the number of historical requests and historical weight values in the same transmission direction is to prevent certain nodes from obtaining transmission resources multiple times when transmitting information in the on-chip network. Under this setting method, it is equivalent to incorporating the above four influencing factors into the arbitration strategy of the target node. In this way, when the target node transmits information, it will not only use the information of the node itself as a reference to design the arbitration mechanism, but will comprehensively consider the influence of the above four factors, and then incorporate the changes in the overall traffic load in the on-chip network and fault information into the arbitration design of the target node.
[0088] After determining the weights corresponding to the requests in the target set, the requests are sorted in descending order of weight and sent to the crossbar switch on the target node. This allows the target node to respond to multiple requests for the same resource in a more fair and reasonable manner.
[0089] Obviously, when determining the weight value corresponding to each request to be sent in the target set, the failure status of the node closest to the target node in the message transmission path of the target request and the number of free spaces in the micro-chip size, the number of historical requests in the same transmission direction, and the historical weight value are taken into consideration. This is equivalent to incorporating the changes in the overall traffic load in the on-chip network and the fault information into the arbitration design of the target node. In this way, the arbitration strategy of the target node can reasonably allocate resources and effectively disperse traffic, thereby avoiding the problem of unbalanced traffic load in the on-chip network and improving the fault tolerance of the on-chip network when transmitting information.
[0090] Based on the above embodiment, this embodiment further illustrates and optimizes the technical solution. As a preferred implementation method, the above step of determining the requests to be sent by the target node in each input direction in the current arbitration cycle to obtain the target set includes:
[0091] Taking the position of the target node in the on-chip network as a reference point, the requests to be sent by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction and the local incoming direction in the current arbitration cycle are determined, and the requests to be sent that are retained by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction and the local incoming direction in the previous arbitration cycle are determined to obtain a target set.
[0092] In this embodiment, the process of a target node acquiring a target set is specifically described. When acquiring the target set, the target node first uses its location in the NoC as a reference point. It then determines the target node's pending requests in the east, west, south, north, and local directions during the current arbitration cycle. It also determines the target node's pending requests in the east, west, south, north, and local directions during the previous arbitration cycle, thereby obtaining the target set.
[0093] It should be noted that the pending requests that are retained by the target node in the east incoming direction, west incoming direction, south incoming direction, north incoming direction and local incoming direction in the previous round of arbitration cycle include both the pending requests that are retained by the target node in the east incoming direction, west incoming direction, south incoming direction, north incoming direction and local incoming direction due to congestion of downstream nodes in the previous round of arbitration cycle, and the pending requests that are retained by the target node in the east incoming direction, west incoming direction, south incoming direction, north incoming direction and local incoming direction due to failure of downstream nodes in the previous round of arbitration cycle.
[0094] Obviously, through the technical solution provided by this embodiment, the target node can accurately obtain the target set.
[0095] As a preferred embodiment, the arbitration method for on-chip network transmission information further includes:
[0096] The data packet header flit of each to-be-sent request in the target set is parsed to determine the output direction of each to-be-sent request in the target set relative to the target node.
[0097] In this embodiment, the output direction of each pending request in the target set relative to the target node is determined by parsing the header flit of each request. Because a request's header flit contains information such as the source address, destination address, and length of the packet, the output direction of each request in the target set relative to the target node can be determined by parsing the flit of each request.
[0098] Obviously, through the technical solution provided by this embodiment, the output direction of each request to be sent in the target set relative to the target node can be accurately determined.
[0099] As a preferred embodiment, the above step of parsing the header flit of each data packet of the request to be sent in the target set includes:
[0100] The router in the target node is used to parse the header fragments of each data packet of the request to be sent in the target set.
[0101] It is understood that since the router is a key component responsible for forwarding data packets in each node of the on-chip network, in this embodiment, the router in the target node can be used to parse the data packet flit of each pending request in the target set. Upon obtaining the data packet header flit of each pending request in the target set, the router in the target node can parse the data packet header flit of each pending request to determine the destination address of each pending request, thereby determining the output direction of each pending request in the target set relative to the target node.
[0102] Obviously, through the technical solution provided by this embodiment, the output direction of each to-be-sent request in the target set relative to the target node can be determined efficiently and reliably.
[0103] As a preferred embodiment, the above steps of: determining the requests to be sent by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the current arbitration cycle, and determining the requests to be sent that were retained by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the previous arbitration cycle, to obtain the target set, include:
[0104] Determine pending requests cached by the target node in the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area in a current arbitration cycle, and determine pending requests cached by the target node in the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area in a previous arbitration cycle, to obtain a first subset, a second subset, a third subset, a fourth subset, and a fifth subset;
[0105] The target set includes a first subset, a second subset, a third subset, a fourth subset, and a fifth subset;
[0106] The first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area are cache areas corresponding to the target node in the east, west, south, north, and local directions, respectively.
[0107] The arbitration cycle is: the processing of the pending requests cached and retained in the first cache area, the second cache area, the third cache area, the fourth cache area and the fifth cache area is taken as the starting point, and the time required to complete the processing of the pending requests cached and retained in the first cache area, the second cache area, the third cache area, the fourth cache area and the fifth cache area.
[0108] In practical applications, the target node is typically configured with five cache areas: the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area. These areas correspond to the target node's cache areas for incoming traffic from the east, west, south, north, and local directions, respectively.
[0109] When the target node receives pending requests from various input directions, it caches these requests in the first, second, third, fourth, and fifth cache areas, respectively, according to the incoming directions of the pending requests. Under this configuration, when the target node obtains the target set, it obtains the pending requests in the first, second, third, fourth, and fifth cache areas, respectively, during the current arbitration cycle, and also obtains the pending requests that were retained in the first, second, third, fourth, and fifth cache areas, respectively, during the previous arbitration cycle, thereby obtaining the first, second, third, fourth, and fifth subsets.
[0110] That is, the first subset is the requests to be sent obtained by the target node from the first cache area in the current arbitration cycle and the previous arbitration cycle, the second subset is the requests to be sent obtained by the target node from the second cache area in the current arbitration cycle and the previous arbitration cycle, the third subset is the requests to be sent obtained by the target node from the third cache area in the current arbitration cycle and the previous arbitration cycle, the fourth subset is the requests to be sent obtained by the target node from the fourth cache area in the current arbitration cycle and the previous arbitration cycle, and the fifth subset is the requests to be sent obtained by the target node from the fifth cache area in the current arbitration cycle and the previous arbitration cycle.
[0111] It should be noted that the arbitration cycle described in the present invention refers to: the processing of the pending requests cached and retained in the first cache area, the second cache area, the third cache area, the fourth cache area and the fifth cache area is taken as the starting point, and the time required to complete the processing of the pending requests cached and retained in the first cache area, the second cache area, the third cache area, the fourth cache area and the fifth cache area.
[0112] Under this configuration, the duration of each arbitration cycle may be different. For example, in the first arbitration cycle, the target node obtains only one pending request in each of the eastbound, westbound, and southbound directions. Furthermore, no pending requests remain in any of the directions during this arbitration cycle. Therefore, the duration of the first arbitration cycle is the time required for the target node to process the three requests. In the second arbitration cycle, the target node obtains one pending request in each of the eastbound, westbound, southbound, northbound, and localbound directions. Furthermore, the target node obtains one pending request in each of the eastbound and westbound directions that remained in the first arbitration cycle. Therefore, the duration of the second arbitration cycle is the time required for the target node to process all seven requests.
[0113] Obviously, the technical solution provided by this embodiment can ensure the accuracy and reliability of the target set acquisition results.
[0114] As a preferred embodiment, the above step of sorting the requests to be sent in the target set in descending order of weight value, and responding to the sorted requests to be sent by using the crossbar switch on the target node, includes:
[0115] Prioritizing responses to pending requests with local destination addresses in the first subset, the second subset, the third subset, the fourth subset, and the fifth subset, and sorting the pending requests in the first subset, the second subset, the third subset, the fourth subset, and the fifth subset, excluding the pending requests with local destination addresses, in descending order of weight to obtain a first request sequence, a second request sequence, a third request sequence, a fourth request sequence, and a fifth request sequence;
[0116] The first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence are sent to a crossbar switch on a target node, so that the crossbar switch is used to respond to each of the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence.
[0117] In this embodiment, when using the cross switch on the target node to respond to the sorted requests to be sent, since the requests to be sent with a local destination address in the target set require less resource overhead and a shorter delay time, the cross switch on the target node will give priority to responding to the requests to be sent with a local destination address in the first subset, second subset, third subset, fourth subset and fifth subset when responding to the requests to be sent.
[0118] Then, the target node will sort the other pending requests in the first subset, the second subset, the third subset, the fourth subset and the fifth subset in order of weight value from high to low, excluding the pending requests with local destination addresses, to obtain a first request sequence, a second request sequence, a third request sequence, a fourth request sequence and a fifth request sequence, and send the first request sequence, the second request sequence, the third request sequence, the fourth request sequence and the fifth request sequence to the cross switch on the target node, so as to use the cross switch to respond to each pending request in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence and the fifth request sequence.
[0119] Obviously, the technical solution provided by this embodiment can reduce the resource overhead required by the target node.
[0120] As a preferred embodiment, the step of sending the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence to the crossbar switch on the target node, so as to use the crossbar switch to respond to each of the to-be-sent requests in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, includes:
[0121] If there is only one request to be sent in each of the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, sending the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence to the crossbar switch on the target node, so that the crossbar switch responds to the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence at the same time;
[0122] If there is more than one request to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence and the fifth request sequence, the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence and the fifth request sequence are sent to the cross switch on the target node, so that the cross switch responds to the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence and the fifth request sequence based on the first-in-first-out principle.
[0123] See Figure 4 , Figure 4 This diagram illustrates how pending requests in each cache area of a target node are sent to a crossbar switch. Since a crossbar switch is a multi-port switching network used to connect multiple input ports and multiple output ports in an on-chip network, the input ports of the target node's crossbar switch can be connected to the first, second, third, fourth, and fifth cache areas of the target node, respectively. Furthermore, the crossbar switch has five output ports: output port 1, output port 2, output port 3, output port 4, and output port 5. These five output ports can be used to independently arbitrate pending requests in the first, second, third, fourth, and fifth cache areas of the target node, respectively.
[0124] In this embodiment, if there is only one pending request in each of the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, the target node will send the pending requests in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence to the crossbar switch on the target node. Upon receiving the pending requests in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, the crossbar switch on the target node will simultaneously respond to the pending requests in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence using output port 1, output port 2, output port 3, output port 4, and output port 5 on the crossbar switch, respectively.
[0125] If there is more than one request to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence and the fifth request sequence, then after the target node sends the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence and the fifth request sequence to the cross switch on the target node, the cross switch will use its output port 1, output port 2, output port 3, output port 4 and output port 5 to respond to the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence and the fifth request sequence respectively based on the first-in-first-out principle.
[0126] Obviously, the technical solution provided by this embodiment can enable the target node to respond to multiple requests for the same resource in a fairer and more reasonable manner.
[0127] As a preferred embodiment, the step of determining the weight value corresponding to each to-be-sent request in the target set based on the failure status of the node closest to the target node in the message transmission path of the target request, the amount of free space in the flit size, the number of historical requests in the same transmission direction, and the historical weight value includes:
[0128] Determine, according to the weight setting model, weight values corresponding to the other pending requests in the target set excluding the pending requests with local destination addresses;
[0129] Among them, the expression of the weight setting model is:
[0130] ;
[0131] Where, Indicates the weight value corresponding to the target cache request in the target cache area, where the target cache area is any one of the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area; and the target cache request is any pending request in the target cache area except those with a local destination address. is the fault state of the peripheral node. The peripheral node is the node closest to the target node in the message transmission path of the target cache request. When a peripheral node fails, , when the surrounding nodes are not faulty, ; The number of free spaces of the microchip size of the surrounding nodes in the surrounding cache area is represented by: the direction of the surrounding cache area relative to the surrounding nodes is the corresponding direction of the target cache area relative to the target node; The value reassigned based on the weights of all pending requests in the target cache area during the previous arbitration cycle, excluding those with local destination addresses. Indicates the preset coefficient; Indicates the number of pending requests in the target cache area with the same output direction as the target cache request in the previous round of arbitration cycle. If the target node does not have a pending request with the same output direction as the target cache request in the target cache area in the previous round of arbitration cycle, let .
[0132] In this embodiment, in order to accurately quantify the weight values corresponding to each pending request in the target set, a weight setting model is used to determine the weight values corresponding to other pending requests in the target set except for the pending requests with local destination addresses.
[0133] Here, we first explain the target cache area, target cache request, surrounding nodes, and surrounding cache areas in detail. The target cache area refers to any of the first, second, third, fourth, and fifth cache areas. A target cache request refers to any request in the target cache area, excluding pending requests with local destination addresses. A surrounding node refers to the node closest to the target node in the message transmission path of the target cache request. See [Note: The following sentences appear to be unrelated and should be omitted:] Figure 2 Assuming node 5 is the target node, if node 5 has a pending request A in its third cache area (that is, node 5 receives a pending request A in the southbound incoming direction), then when calculating the weight value corresponding to pending request A, the first step is to parse pending request A. After parsing, it is found that the transmission path corresponding to pending request A is: node 5 → node 6. Then node 6 is the node closest to node 5 in the message transmission path of pending request A, that is, node 6 is a peripheral node of node 5. Since the direction of the peripheral cache area relative to the peripheral node is: the corresponding direction of the target cache area relative to the target node, the peripheral cache area is the cache area corresponding to node 5 in the northbound transmission direction, that is, the fourth cache area on node 5 is the peripheral cache area of node 5.
[0134] when When , it means that the surrounding nodes have failed. From this we can conclude that when a peripheral node fails, The weight value of has a veto power, and the target node will not perform any action on the target cache request.
[0135] Indicates the amount of free space in the surrounding node's micro-slice cache. This parameter is introduced into the weight setting model to incorporate the overall NOC traffic load changes into the target node's arbitration strategy. This allows the target node to adjust its arbitration strategy based on the overall NOC traffic load changes, rather than solely referencing its own information. This allows for efficient traffic load control and optimized network resource utilization.
[0136] It represents the value reassigned based on the weight values of all pending requests in the target cache area during the previous arbitration cycle, excluding those with local destination addresses. The number of pending requests from the target node in the previous arbitration cycle, in the same output direction as the target cache request, is included in the target cache area. These two parameters are introduced into the weighting model to prevent a node in the on-chip network from obtaining transmission resources multiple times and to ensure uniform distribution of transmission resources.
[0137] Obviously, through the technical solution provided by this embodiment, the weight value corresponding to each request to be sent in the target set can be quantified, and then the weight value set can be obtained.
[0138] As a preferred embodiment, the arbitration method for on-chip network transmission information further includes:
[0139] In the previous arbitration cycle and in the target cache area, the weight values of the other pending requests on the target node excluding the pending requests with local destination addresses are obtained, and the pending requests are sorted in descending order of the weight values to obtain a first sequence;
[0140] If there are pending requests sent in the target output direction in the first sequence, then determining the pending requests sent in the target output direction in the first sequence to obtain a target screening set; the target output direction is the same as the output direction of the target cache request;
[0141] Determine the pending request corresponding to the largest weight value in the target screening set to obtain the target screening request, and filter out the pending requests that appear for the first time in different output directions from the first sequence to obtain a second sequence;
[0142] According to the order of the target filter requests in the second sequence Assign a value.
[0143] In this embodiment, in order to evenly distribute the transmission resources, it is necessary to set the weight setting model based on the weight values of the other pending requests in the target cache area except for the pending requests with local destination addresses in the previous arbitration cycle. The numerical value of .
[0144] Specifically, in the weight setting model When setting, the target node will obtain the weight values of other pending requests on the target node except for the pending requests with local destination addresses in the previous arbitration cycle and in the target cache area, and sort the pending requests in descending order according to the weight value to obtain a first sequence; then, the target node will determine whether there are pending requests sent by the target output direction in the first sequence; if there are pending requests sent by the target output direction in the first sequence, the target node will determine the pending requests sent by the target output direction in the first sequence to obtain a target filtering set, and will determine the pending request corresponding to the largest weight value in the target filtering set to obtain a target filtering request, and will also filter out the pending requests that appear for the first time in different output directions from the first sequence to obtain a second sequence, and will sort according to the arrangement order of the target filtering requests in the second sequence. If there is no request to be sent by the target output direction in the first sequence, the target node will assign a value to Perform forced assignment.
[0145] To illustrate this with an example, to calculate the weight corresponding to node 5's pending request A0 in the third cache area, request A0 is first parsed. Parsing reveals that the transmission path corresponding to request A0 is: node 5 → node 6. Node 6 is then a neighboring node of node 5, and the fourth cache area on node 6 is the neighboring cache area. Since request A0 is destined for node 6, node 6 is in the eastward output direction relative to node 5, so the target output direction is the eastward output direction.
[0146] Determine the weight value corresponding to the request A0 to be sent according to the weight setting model When , it is necessary to determine the weight setting model 、 、 and Assume that node 6 is not faulty and the number of free microslice size spaces in the fourth cache area on node 6 is 3, then 、 In order to determine The corresponding value needs to obtain the weight values corresponding to the requests to be sent by node 5 in the last arbitration cycle and in the third cache area except for the requests to be sent with the local destination address.
[0147] Assume that in the last arbitration cycle, node 5 has four pending requests A1, A2, A3 and A4 in the third cache area. These four pending requests A1, A2, A3 and A4 are pending requests in the westward output direction, northward output direction, westward output direction and eastward output direction, respectively. The weight values corresponding to these four pending requests A1, A2, A3 and A4 are 0.4, 0.3, 0.2 and 0.1, respectively. After sorting the pending requests A1, A2, A3 and A4 in descending order of weight value, the first sequence obtained is: A1, A2, A3 and A4. Since there are pending requests in the first sequence A1, A2, A3 and A4 with the same output direction as the pending request A0 (that is, there are pending requests in the east output direction in the first sequence A1, A2, A3 and A4), and there is only one pending request in the east output direction, A4 is the target screening request; then, the pending requests that appear for the first time in different output directions are filtered out from the first sequence A1, A2, A3 and A4 to obtain the second sequence, that is, the second sequence is A1, A2 and A4; finally, the target screening request A4 is sorted according to its arrangement order in the second sequence A1, A2 and A4. To assign a value, Assign a value of 3 and The value is 1.
[0148] In another application scenario, to calculate the weight corresponding to node 5's pending request B0 in the third cache area, B0 is first parsed. After parsing, it is found that the transmission path corresponding to B0 is: node 5 → node 4. Therefore, node 4 is a neighboring node of node 5, and the fourth cache area on node 4 is the neighboring cache area. Since the destination address of B0 is node 4, node 4 is in the westward output direction relative to node 5, so the target output direction is the westward output direction.
[0149] Determine the weight value corresponding to the request B0 to be sent according to the weight setting model When , it is necessary to determine the weight setting model 、 、 and Assume that node 4 is not faulty and the number of free microslice-sized spaces in the fourth cache area on node 4 is 4. Then 、 In order to determine The corresponding value needs to obtain the weight values corresponding to the requests to be sent by node 4 in the last arbitration cycle and in the third cache area except for the requests to be sent with the local destination address.
[0150] Assume that in the last arbitration cycle, node 5 has four pending requests B1, B2, B3 and B4 in the third cache area. These four pending requests B1, B2, B3 and B4 are pending requests in the westward output direction, northward output direction, westward output direction and eastward output direction, respectively. The weight values corresponding to these four pending requests B1, B2, B3 and B4 are 0.8, 0.7, 0.6 and 0.5, respectively. After sorting the pending requests B1, B2, B3 and B4 in descending order of weight value, the first sequence obtained is: B1, B2, B3 and B4. Since there are pending requests in the first sequence B1, B2, B3, and B4 with the same output direction as the pending request B0 (that is, there are pending requests in the first sequence B1, B2, B3, and B4 with a westward output direction), and there are two pending requests in the westward output direction in the first sequence, then the target screening request is the pending request corresponding to the two westward output directions with the highest weights in the first sequence, that is, the target screening request is B1. Then, the first pending requests that appear in different output directions are filtered out from the first sequence B1, B2, B3, and B4 to obtain the second sequence, that is, the second sequence is B1, B2, and B4; finally, the target screening request B1 is sorted according to the order in which it is arranged in the second sequence B1, B2, and B4. To assign a value, Assign a value of 1 and The value is 2.
[0151] In another application scenario, to calculate the weight corresponding to the pending request C0 in the second cache area of node 5, the pending request C0 is first parsed. After parsing, it is found that the transmission path corresponding to the pending request C0 is: node 5 → node 2. Node 2 is then a neighboring node of node 5, and the first cache area of node 2 is the neighboring cache area. Since the destination address of the pending request C0 is node 2, node 2 is in the northbound output direction relative to node 5, so the target output direction is the northbound output direction.
[0152] Determine the weight value corresponding to the request C0 to be sent according to the weight setting model When , it is necessary to determine the weight setting model 、 、 and Assume that there is no fault on node 2 and the number of free spaces of the micro-slice size in the first cache area on node 2 is 1, then 、 In order to determine The corresponding value needs to obtain the weight values corresponding to the requests to be sent by node 5 in the previous arbitration cycle and in the second cache area except for the requests to be sent with the local destination address.
[0153] Assume that in the last arbitration cycle, node 5 has four pending requests C1, C2, C3, and C4 in the second cache area. These four pending requests C1, C2, C3, and C4 are pending requests in the westward output direction, northward output direction, westward output direction, and eastward output direction, respectively. The weight values corresponding to these four pending requests C1, C2, C3, and C4 are 0.8, 0.6, 0.4, and 0.3, respectively. After sorting the pending requests C1, C2, C3, and C4 in descending order of weight value, the first sequence obtained is: C1, C2, C3, and C4. Since there is no pending request in the first sequence C1, C2, C3, and C4 with the same output direction as the pending request C0 (that is, there is no pending request in the northward output direction in the first sequence C1, C2, C3, and C4), then at this time, the pending requests can be sorted according to the target output direction. Perform forced assignment, and at this time .
[0154] Here we can pre-agreed that the requests to be sent in the east, south, west and north directions correspond to In other words, since there are no requests to be sent in the northbound output direction in the first sequence C1, C2, C3 and C4, The mandatory value is 7.
[0155] Obviously, through the technical solution provided by this embodiment, the weight values corresponding to the other requests to be sent in the target set except for the requests to be sent with local destination addresses can be accurately determined according to the weight setting model.
[0156] As a preferred embodiment, the arbitration method for on-chip network transmission information further includes:
[0157] Send target detection signals to surrounding nodes and determine whether surrounding nodes have faults based on the feedback signals returned by the surrounding nodes.
[0158] In this embodiment, in order to determine whether a peripheral node fails, the target node may further send a target detection signal to the peripheral nodes, and determine whether a peripheral node fails based on feedback signals returned by the peripheral nodes.
[0159] The target detection signal sent by the target node to the surrounding nodes is a special signal that can be used to test the connectivity and status of the links between the target node and the surrounding nodes. After receiving the target detection signal, the surrounding nodes return a feedback signal to the target node, which can contain information such as whether the target node has successfully reached the surrounding nodes and the link status between the target node and the surrounding nodes. Therefore, by analyzing the feedback signals returned by the surrounding nodes, the target node can determine whether the surrounding nodes have failed.
[0160] Obviously, through the technical solution provided by this embodiment, it is possible to accurately determine whether a peripheral node has a fault.
[0161] As a preferred embodiment, the above step of sending a target detection signal to a surrounding node and determining whether a surrounding node has a fault based on a feedback signal returned by the surrounding node includes:
[0162] Send a target detection signal to the surrounding nodes and determine whether the surrounding nodes can return a feedback signal corresponding to the target detection signal within a preset time;
[0163] If so, it is determined that no failure has occurred in the surrounding nodes;
[0164] If not, it is determined that the surrounding nodes are faulty.
[0165] It is understood that if the surrounding nodes are not faulty, then after the target node sends a target detection signal to the surrounding nodes, the surrounding nodes will definitely be able to return feedback signals corresponding to the target detection signal within a preset time. Therefore, in this embodiment, the above-mentioned attribute characteristics can be used to more quickly determine whether the surrounding nodes are faulty.
[0166] That is, after a target node sends a target detection signal to a surrounding node, if the surrounding node can return a feedback signal corresponding to the target detection signal within a preset time, it means that the surrounding node has not failed. If the surrounding node cannot return a feedback signal corresponding to the target detection signal within the preset time, it means that the surrounding node has failed.
[0167] Obviously, through the technical solution provided by this embodiment, it is possible to more quickly determine whether a peripheral node has a fault.
[0168] As a preferred embodiment, the arbitration method for on-chip network transmission information further includes:
[0169] When it is determined that a peripheral node fails, the destination node corresponding to the target cache request is determined, and the routing transmission path between the target node and the destination node corresponding to the target cache request is recalculated according to the routing algorithm.
[0170] In this embodiment, if a neighboring node is detected as faulty, it indicates that the target node is no longer able to send the target cache request to its corresponding destination address via the original routing path. In this case, to ensure that the target cache request can still accurately reach the destination node in the network-on-chip, the destination node corresponding to the target cache request must be determined and the routing path between the target node and the destination node corresponding to the target cache request must be recalculated using a routing algorithm.
[0171] Specifically, the routing algorithm can be set to Distance-Vector Routing, Link-State Routing, Path-Vector Routing, etc.
[0172] Obviously, the technical solution provided by this embodiment can further improve the fault tolerance of the on-chip network.
[0173] As a preferred embodiment, the above step of: before recalculating the routing transmission path between the target node and the target node corresponding to the target cache request according to the routing algorithm, further includes:
[0174] The target cache request is retained in the cache area where the target cache request is located, and the target cache request is added to the next round of arbitration cycle.
[0175] In this embodiment, before recalculating the routing path between the target node and its surrounding nodes based on the routing algorithm, the target cache request can be retained in the cache area where the target cache request resides and added to the next round of arbitration. This configuration not only prevents packet loss in the on-chip network but also ensures that the target node processes and responds to each pending request in an orderly manner.
[0176] Obviously, the technical solution provided by this embodiment can relatively improve the working performance and execution efficiency of the target node.
[0177] As a preferred embodiment, the arbitration method for on-chip network transmission information further includes:
[0178] The credit information of the surrounding nodes is obtained, and the number of free spaces of the size of the micro slices of the surrounding nodes in the surrounding cache area is determined according to the credit information of the surrounding nodes.
[0179] In this embodiment, the parameters in the weight setting model are determined. When the peripheral nodes are allocated, the credit information of the peripheral nodes can be obtained, and the number of free spaces of the micro-slice size of the peripheral nodes in the peripheral cache area can be determined according to the credit information of the peripheral nodes.
[0180] Because the credit information of the surrounding nodes can represent the buffer status of the surrounding nodes, generally speaking, the higher the credit value, the more free space of the micro-chip size in the surrounding cache area, and the lower the credit value, the fewer free space of the micro-chip size in the surrounding cache area. Therefore, in actual applications, the free space of the micro-chip size of the surrounding nodes in the surrounding cache area can be determined based on the credit information of the surrounding nodes.
[0181] Obviously, through the technical solution provided by this embodiment, the number of free spaces of the size of microchips of the peripheral nodes in the peripheral cache area can be accurately calculated.
[0182] As a preferred embodiment, the arbitration method for on-chip network transmission information further includes:
[0183] Pre-establishing a first table, a second table, and a third table;
[0184] Using the first table to record the fault status of the surrounding nodes;
[0185] Using the second table to record the number of free spaces of the size of the microchips of the surrounding nodes in the surrounding cache area;
[0186] The third table is used to record the weight values corresponding to the other pending requests of the target node in the previous arbitration cycle and in each cache area except the pending requests with local destination addresses.
[0187] In this embodiment, a first table, a second table and a third table can also be established in advance, and the first table can be used to record the fault status of the surrounding nodes, the second table can be used to record the number of free spaces of the micro-chip size of the surrounding nodes in the surrounding cache area, and the third table can be used to record the weight values corresponding to the target node's other pending requests in the previous arbitration cycle and in each cache area except for the pending requests with local destination addresses.
[0188] In this setting mode, when the target node determines the weight value corresponding to each pending request in the target set according to the weight setting model, it can directly call the values corresponding to each parameter in the weight setting model from the first table, the second table, and the third table. This can further improve the working performance of the target node and increase the response speed to each pending request. Figure 5 , Figure 5This is a schematic diagram of a target node calculating the weight value corresponding to each request to be sent in the target set according to the weight setting model by calling the data in the first table, the second table and the third table.
[0189] Obviously, the technical solution provided by this embodiment can further improve the working performance of the target node and increase the response speed to each request to be sent.
[0190] See Figure 6 , Figure 6 This is a structural diagram of an arbitration device for transmitting information in a network on chip provided by an embodiment of the present invention. The device is applied to a target node in a network on chip arranged in a grid shape, and the target node is any node in the network on chip, including:
[0191] The request acquisition module 21 is used to acquire the requests to be sent by the target node in each input direction in the current arbitration cycle to obtain a target set;
[0192] The weight calculation module 22 is configured to determine a weight value corresponding to each pending request in the target set based on the failure status of the node closest to the target node in the message transmission path of the target request, the number of free spaces in the flit size, the number of historical requests in the same transmission direction, and the historical weight value; the target request is any pending request in the target set; and the historical weight value is the weight value corresponding to each pending request sent by the target node in the previous arbitration cycle.
[0193] The request response module 23 is configured to sort the requests to be sent in the target set in descending order of weight value, so as to respond to the sorted requests to be sent using the crossbar switch on the target node.
[0194] In a specific embodiment of the present application, the request acquisition module 21 includes:
[0195] The request acquisition submodule is used to determine, with the position of the target node in the on-chip network as a reference point, requests to be sent by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the current arbitration cycle, and to determine requests to be sent that are retained by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the previous arbitration cycle, to obtain the target set.
[0196] In a specific embodiment of the present application, it also includes:
[0197] The data parsing submodule is configured to parse the data packet header flit of each to-be-sent request in the target set to determine the output direction of each to-be-sent request in the target set relative to the target node.
[0198] In a specific embodiment of the present application, the data parsing submodule includes:
[0199] The data parsing unit is configured to parse the header flit of each data packet of the to-be-sent request in the target set by using the router in the target node.
[0200] In a specific embodiment of the present application, the request acquisition submodule includes:
[0201] a request acquisition unit, configured to determine pending requests cached by the target node in the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area in a current arbitration cycle, and to determine pending requests cached by the target node in the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area in a previous arbitration cycle, to obtain a first subset, a second subset, a third subset, a fourth subset, and a fifth subset;
[0202] The target set includes the first subset, the second subset, the third subset, the fourth subset, and the fifth subset;
[0203] The first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area are cache areas corresponding to the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction, respectively;
[0204] The arbitration cycle is: starting from processing the pending requests cached and retained in the first cache area, the second cache area, the third cache area, the fourth cache area and the fifth cache area, and the time required to complete the processing of the pending requests cached and retained in the first cache area, the second cache area, the third cache area, the fourth cache area and the fifth cache area.
[0205] In a specific embodiment of the present application, the request response module 23 includes:
[0206] a request sorting submodule, configured to preferentially respond to pending requests with local destination addresses in the first subset, the second subset, the third subset, the fourth subset, and the fifth subset, and sort the pending requests in the first subset, the second subset, the third subset, the fourth subset, and the fifth subset, excluding the pending requests with local destination addresses, in descending order of weight to obtain a first request sequence, a second request sequence, a third request sequence, a fourth request sequence, and a fifth request sequence;
[0207] a request response submodule, configured to send the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence to the cross switch on the target node, so as to utilize the cross switch to respond to each of the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence.
[0208] In a specific embodiment of the present application, the request response submodule includes:
[0209] a first responding unit configured to, if there is only one to-be-sent request in each of the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, send the to-be-sent requests in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence to the crossbar switch on the target node, so that the crossbar switch responds to the to-be-sent requests in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence simultaneously;
[0210] a second response unit, configured to send the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence to the crossbar switch on the target node if there is more than one request to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, so that the crossbar switch responds to the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence based on a first-in-first-out principle.
[0211] In a specific embodiment of the present application, the weight calculation module 22 includes:
[0212] A weight calculation submodule, configured to determine, according to a weight setting model, weight values corresponding to the other requests to be sent in the target set excluding the requests to be sent with a local destination address;
[0213] Wherein, the expression of the weight setting model is:
[0214] ;
[0215] Where, Indicates a weight value corresponding to a target cache request in a target cache area, where the target cache area is any one of the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area; and the target cache request is any pending request in the target cache area except pending requests with a local destination address. is the fault state of the peripheral node, the peripheral node is the node closest to the target node in the message transmission path of the target cache request; when the peripheral node fails, , when the surrounding nodes are not faulty, ; Indicates the number of free spaces of the microchip size of the peripheral node in the peripheral cache area; the direction of the peripheral cache area relative to the peripheral node is: the corresponding direction of the target cache area relative to the target node; represents a value reassigned based on the weight values of all pending requests of the target node in the previous arbitration cycle and in the target cache area, excluding pending requests with local destination addresses; Indicates the preset coefficient; Indicates the number of pending requests in the target cache area with the same output direction as the target cache request in the previous round of arbitration cycle. If the target node does not have a pending request in the target cache area with the same output direction as the target cache request in the previous round of arbitration cycle, let .
[0216] In a specific embodiment of the present application, it also includes:
[0217] a first sorting unit, configured to obtain, in the target cache area and in the previous arbitration cycle, weight values of all pending requests on the target node excluding pending requests with local destination addresses, and to sort the pending requests in descending order of weight values to obtain a first sequence;
[0218] a request screening unit configured to, if there are pending requests sent in a target output direction in the first sequence, determine the pending requests sent in the target output direction in the first sequence to obtain a target screening set; the target output direction being the same as the output direction of the target cache request;
[0219] a second sorting unit, configured to determine the to-be-sent request corresponding to the largest weight value in the target screening set to obtain the target screening request, and to filter out the to-be-sent requests that appear for the first time in different output directions from the first sequence to obtain a second sequence;
[0220] A numerical value assignment unit is used to assign a value to the target filter request according to the order of arrangement of the target filter request in the second sequence. Assign a value.
[0221] In a specific embodiment of the present application, it also includes:
[0222] A forced assignment unit, configured to assign a value to the target output direction according to the target output direction if there is no request to be sent in the first sequence. Assign a value.
[0223] In a specific embodiment of the present application, it also includes:
[0224] The node determination unit is configured to send a target detection signal to the peripheral nodes and determine whether a fault occurs in the peripheral nodes according to feedback signals returned by the peripheral nodes.
[0225] In a specific implementation of the present application, the node determination unit includes:
[0226] a signal sending subunit, configured to send the target detection signal to the surrounding nodes, and determine whether the surrounding nodes can return a feedback signal corresponding to the target detection signal within a preset time;
[0227] a first determining subunit, configured to determine that no fault occurs in the peripheral node when the determination result of the signal sending subunit is yes;
[0228] The second determination subunit is configured to determine that a fault occurs in the peripheral node when the determination result of the signal sending subunit is negative.
[0229] In a specific embodiment of the present application, it also includes:
[0230] The path calculation unit is used to determine the destination node corresponding to the target cache request when it is determined that the peripheral node fails, and recalculate the routing transmission path between the target node and the destination node corresponding to the target cache request according to a routing algorithm.
[0231] In a specific embodiment of the present application, it also includes:
[0232] A request retention unit is used to retain the target cache request in the cache area where the target cache request is located before recalculating the routing transmission path between the target node and the target node corresponding to the target cache request according to the routing algorithm, and add the target cache request to the next round of arbitration cycle.
[0233] In a specific embodiment of the present application, it also includes:
[0234] The credit acquisition unit is configured to acquire the credit information of the peripheral node and determine the number of free spaces of the micro-slice size of the peripheral node in the peripheral cache area according to the credit information of the peripheral node.
[0235] In a specific embodiment of the present application, it also includes:
[0236] A table creation unit, configured to pre-create a first table, a second table, and a third table;
[0237] a first recording unit, configured to record the fault status of the peripheral nodes using the first table;
[0238] A second recording unit is configured to record the number of free spaces of the size of the flit of the peripheral node in the peripheral cache area using the second table;
[0239] The third recording unit is configured to use the third table to record the weight values corresponding to the other pending requests of the target node in the previous arbitration cycle and in each cache area except the pending requests with local destination addresses.
[0240] An arbitration device for network-on-chip (NOC) transmission information provided by an embodiment of the present invention has the beneficial effects of the aforementioned arbitration method for NOC transmission information.
[0241] See Figure 7 , Figure 7 This is a structural diagram of an arbitration device for transmitting information on a network on a chip provided by an embodiment of the present invention, the device comprising:
[0242] Memory 31, for storing computer programs;
[0243] The processor 32 is configured to execute the computer program to implement the steps of the arbitration method for transmitting information in a network on chip as disclosed above.
[0244] The arbitration device for transmitting information on a network on a chip provided in this embodiment may include, but is not limited to, a smart phone, a tablet computer, a laptop computer, or a desktop computer.
[0245] The processor 32 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 32 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 32 may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 32 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing content required to be displayed on the display screen. In some embodiments, the processor 32 may also include an artificial intelligence (AI) processor, which is used to handle computational operations related to machine learning.
[0246] The memory 31 may include one or more computer-readable storage media, which may be non-transitory. The memory 31 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 31 is at least used to store the following computer program 301, wherein, after the computer program is loaded and executed by the processor 32, it can implement the relevant steps of the arbitration method for on-chip network transmission information disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 31 may also include an operating system 302 and data 303, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 302 may include Windows, Unix, Linux, etc. The data 303 may include but is not limited to data involved in the arbitration method for on-chip network transmission information.
[0247] In some embodiments, the arbitration device for transmitting information on a network on chip may further include a display screen 33 , an input / output interface 34 , a communication interface 35 , a power supply 36 , and a communication bus 37 .
[0248] Those skilled in the art will understand that Figure 7The illustrated structure does not constitute a limitation on the arbitration device for transmitting information in a network on chip, and may include more or fewer components than shown in the figure.
[0249] It is understood that if the methods in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the current technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable ROM, a register, a hard drive, a removable disk, a CD-ROM, a magnetic disk, or an optical disk, and other media that can store program code.
[0250] An arbitration device for network-on-chip (NOC) transmission information provided by an embodiment of the present invention has the beneficial effects of the aforementioned arbitration method for NOC transmission information.
[0251] Accordingly, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the arbitration method for transmitting information in an on-chip network as disclosed above are implemented.
[0252] A computer-readable storage medium provided by an embodiment of the present invention has the beneficial effects of the aforementioned arbitration method for information transmission in a network on chip.
[0253] Accordingly, an embodiment of the present invention further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the arbitration method for transmitting information in an on-chip network as disclosed above.
[0254] A computer program product provided by an embodiment of the present invention has the beneficial effects of the aforementioned arbitration method for information transmission on a network on a chip.
[0255] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0256] The above describes in detail the arbitration method and related equipment for network-on-chip (NOC) information transmission provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core concept of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from the principles of the present invention, and such improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A method for arbitrating information transmitted by an on-chip network, characterized in that: Applied to a target node in a grid-arranged network on chip, the target node being any node of the network on chip, including: Determine the requests to be sent by the target node in each input direction in the current arbitration cycle to obtain a target set; Determining the requests to be sent by the target node in each input direction in the current arbitration cycle to obtain a target set includes: Taking the position of the target node in the on-chip network as a reference point, determining requests to be sent by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the current arbitration cycle, and determining requests to be sent that are retained by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the previous arbitration cycle, to obtain the target set; The determining of the requests to be sent by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the current arbitration cycle includes: Determine the pending requests cached by the target node in the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area in a current arbitration cycle; The first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area are cache areas corresponding to the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction, respectively; Determine a weight value corresponding to each pending request in the target set based on a failure status of a node closest to the target node in a message transmission path of the target request, the number of free spaces in the flit size, the number of historical requests in the same transmission direction, and a historical weight value; the target request is any pending request in the target set; and the historical weight value is a weight value corresponding to each pending request sent by the target node in a previous arbitration cycle; Determining weight values corresponding to the other requests to be sent in the target set excluding the requests to be sent whose destination addresses are local according to the weight setting model; Wherein, the expression of the weight setting model is: ; Where, Indicates a weight value corresponding to a target cache request in a target cache area, where the target cache area is any one of the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area; and the target cache request is any pending request in the target cache area except pending requests with a local destination address. is the fault state of the peripheral node, the peripheral node is the node closest to the target node in the message transmission path of the target cache request; when the peripheral node fails, , when the surrounding nodes are not faulty, ; Indicates the number of free spaces of the microchip size of the peripheral node in the peripheral cache area; the direction of the peripheral cache area relative to the peripheral node is: the corresponding direction of the target cache area relative to the target node; represents a value reassigned based on the weight values of all pending requests of the target node in the previous arbitration cycle and in the target cache area, excluding pending requests with local destination addresses; Indicates the preset coefficient; Indicates the number of pending requests in the target cache area with the same output direction as the target cache request in the previous round of arbitration cycle. If the target node does not have a pending request in the target cache area with the same output direction as the target cache request in the previous round of arbitration cycle, let ; The requests to be sent in the target set are sorted in descending order of weight value, so as to respond to the sorted requests to be sent by using the crossbar switch on the target node.
2. The arbitration method for information transmission in a network on chip according to claim 1, wherein: Also includes: The data packet header flit of each to-be-sent request in the target set is parsed to determine the output direction of each to-be-sent request in the target set relative to the target node.
3. The arbitration method for information transmission in a network on chip according to claim 2, wherein: The parsing of the data packet header flit of each to-be-sent request in the target set includes: The router in the target node is used to parse the header fragments of each data packet of the to-be-sent request in the target set.
4. The arbitration method for information transmission in a network on chip according to claim 1, wherein: The determining of the pending requests held by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the previous arbitration cycle includes: Determine pending requests of the target node that are respectively retained in the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area in a previous arbitration cycle, to obtain a first subset, a second subset, a third subset, a fourth subset, and a fifth subset; The target set includes the first subset, the second subset, the third subset, the fourth subset, and the fifth subset; The arbitration cycle is: starting from processing the pending requests cached and retained in the first cache area, the second cache area, the third cache area, the fourth cache area and the fifth cache area, and the time required to complete the processing of the pending requests cached and retained in the first cache area, the second cache area, the third cache area, the fourth cache area and the fifth cache area.
5. The arbitration method for information transmission in a network on chip according to claim 4, characterized in that: The step of sorting the requests to be sent in the target set in descending order of weight values, so as to respond to the sorted requests to be sent by using the crossbar switch on the target node, includes: Prioritizing responding to pending requests with local destination addresses in the first subset, the second subset, the third subset, the fourth subset, and the fifth subset, and sorting the pending requests in the first subset, the second subset, the third subset, the fourth subset, and the fifth subset, excluding the pending requests with local destination addresses, in descending order of weight to obtain a first request sequence, a second request sequence, a third request sequence, a fourth request sequence, and a fifth request sequence; The first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence are sent to the crossbar switch on the target node, so that the crossbar switch is used to respond to each of the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence.
6. The arbitration method for information transmission in a network on chip according to claim 5, characterized in that: The sending of the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence to the crossbar switch on the target node, so as to use the crossbar switch to respond to each of the to-be-sent requests in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, includes: If there is only one request to be sent in each of the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, sending the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence to the crossbar switch on the target node, so that the crossbar switch responds to the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence at the same time; If there is more than one request to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence, the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence are sent to the cross switch on the target node, so that the cross switch responds to the requests to be sent in the first request sequence, the second request sequence, the third request sequence, the fourth request sequence, and the fifth request sequence based on a first-in-first-out principle.
7. The arbitration method for information transmission in a network on chip according to claim 1, characterized in that: Also includes: In the previous arbitration cycle and in the target cache area, obtaining weight values of other pending requests on the target node except for pending requests with local destination addresses, and sorting the pending requests in descending order of weight values to obtain a first sequence; If there are pending requests sent by the target output direction in the first sequence, determining the pending requests sent by the target output direction in the first sequence to obtain a target screening set; The target output direction is the same as the target cache request output direction; Determine the to-be-sent request corresponding to the largest weight value in the target screening set to obtain the target screening request, and filter out the to-be-sent requests that appear for the first time in different output directions from the first sequence to obtain a second sequence; According to the order of arrangement of the target screening request in the second sequence Assign a value.
8. The arbitration method for information transmission in a network on chip according to claim 7, characterized in that: Also includes: If there is no request to be sent by the target output direction in the first sequence, Assign a value.
9. The arbitration method for information transmission in a network on chip according to claim 1, wherein: Also includes: A target detection signal is sent to the peripheral node, and whether the peripheral node fails is determined according to a feedback signal returned by the peripheral node.
10. The arbitration method for information transmission in a network on chip according to claim 9, characterized in that: The sending of the target detection signal to the surrounding node and determining whether the surrounding node fails according to the feedback signal returned by the surrounding node includes: Sending the target detection signal to the surrounding nodes, and determining whether the surrounding nodes can return a feedback signal corresponding to the target detection signal within a preset time; If so, it is determined that the peripheral node has not failed; If not, it is determined that the peripheral node fails.
11. The arbitration method for information transmission in a network on chip according to claim 10, characterized in that: Also includes: When it is determined that the peripheral node fails, the destination node corresponding to the target cache request is determined, and the routing transmission path between the target node and the destination node corresponding to the target cache request is recalculated according to a routing algorithm.
12. The arbitration method for information transmission in a network on chip according to claim 11, characterized in that: Before recalculating the routing transmission path between the target node and the target node corresponding to the target cache request according to the routing algorithm, the method further includes: The target cache request is retained in the cache area where the target cache request is located, and the target cache request is added to the next round of arbitration cycle.
13. The arbitration method for information transmission in a network on chip according to claim 1, characterized in that: Also includes: The credit information of the peripheral node is acquired, and the number of free spaces of the size of the microchip of the peripheral node in the peripheral cache area is determined according to the credit information of the peripheral node.
14. The arbitration method for information transmission in a network on chip according to claim 1, characterized in that: Also includes: Pre-establishing a first table, a second table, and a third table; Using the first table to record the fault status of the surrounding nodes; Using the second table to record the number of free spaces of the size of the flit of the peripheral nodes in the peripheral cache area; The third table is used to record the weight values corresponding to the other pending requests of the target node in the previous arbitration cycle and in each cache area except the pending requests with local destination addresses.
15. An arbitration device for transmitting information in a network on chip, characterized in that: Applied to a target node in a grid-arranged network on chip, the target node being any node of the network on chip, including: a request acquisition module, configured to acquire requests to be sent by the target node in each input direction in the current arbitration cycle, and obtain a target set; Determining the requests to be sent by the target node in each input direction in the current arbitration cycle to obtain a target set includes: Taking the position of the target node in the on-chip network as a reference point, determining requests to be sent by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the current arbitration cycle, and determining requests to be sent that are retained by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the previous arbitration cycle, to obtain the target set; The determining of the requests to be sent by the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction in the current arbitration cycle includes: Determine the pending requests cached by the target node in the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area in a current arbitration cycle; The first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area are cache areas corresponding to the target node in the east incoming direction, the west incoming direction, the south incoming direction, the north incoming direction, and the local incoming direction, respectively; A weight calculation module is configured to determine a weight value corresponding to each pending request in the target set based on the failure status of the node closest to the target node in the message transmission path of the target request, the number of free spaces in the flit size, the number of historical requests in the same transmission direction, and a historical weight value; the target request is any pending request in the target set; and the historical weight value is the weight value corresponding to each pending request sent by the target node in the previous arbitration cycle; Determining weight values corresponding to the other requests to be sent in the target set excluding the requests to be sent whose destination addresses are local according to the weight setting model; Wherein, the expression of the weight setting model is: ; Where, Indicates a weight value corresponding to a target cache request in a target cache area, where the target cache area is any one of the first cache area, the second cache area, the third cache area, the fourth cache area, and the fifth cache area; and the target cache request is any pending request in the target cache area except pending requests with a local destination address. is the fault state of the peripheral node, the peripheral node is the node closest to the target node in the message transmission path of the target cache request; when the peripheral node fails, , when the surrounding nodes are not faulty, ; Indicates the number of free spaces of the microchip size of the peripheral node in the peripheral cache area; the direction of the peripheral cache area relative to the peripheral node is: the corresponding direction of the target cache area relative to the target node; represents a value reassigned based on the weight values of all pending requests of the target node in the previous arbitration cycle and in the target cache area, excluding pending requests with local destination addresses; Indicates the preset coefficient; Indicates the number of pending requests in the target cache area with the same output direction as the target cache request in the previous round of arbitration cycle. If the target node does not have a pending request in the target cache area with the same output direction as the target cache request in the previous round of arbitration cycle, let ; The request response module is used to sort the requests to be sent in the target set in descending order of weight value, so as to respond to the sorted requests to be sent by using the crossbar switch on the target node.
16. An arbitration device for transmitting information on a network on a chip, characterized in that: include: memory for storing computer programs; A processor is configured to execute the computer program to implement the steps of the arbitration method for information transmission in an on-chip network as claimed in any one of claims 1 to 14.
17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the arbitration method for information transmission in an on-chip network are implemented as claimed in any one of claims 1 to 14.
18. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the arbitration method for information transmission in an on-chip network as claimed in any one of claims 1 to 14 are implemented.
Citation Information
Patent Citations
Routing device and routing method of network-on-chip
CN110620731A
Distributed network resource optimization scheduling method and system under load balancing strategy
CN119232741A