A hang-release circuit and hang-release method based on on-chip network

By adding channel buffers and deadlock management circuits to the ring path of the system-on-chip (SoC), the deadlock problem caused by the ring routing path is solved, and efficient data transmission is achieved without topology changes.

CN115914110BActive Publication Date: 2025-09-23STREAM COMPUTING INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110943339.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-17
Publication Date
2025-09-23
Estimated Expiration
2041-08-17

AI Technical Summary

Technical Problem

In system-on-chip (SoC), the hang phenomenon caused by circular routing paths frequently occurs. Existing technologies solve this problem by interrupting the circular path structure, but this leads to changes in the topology and increased data transmission delays.

Method used

By adding channel buffers and deadlock management circuits to the ring path, deadlock resolution can be achieved through signal interaction and arbitrator scheduling without changing the NoC topology and routing rules.

Benefits of technology

It effectively eliminates the deadlock phenomenon, reduces performance loss, and maintains the efficiency and balance of data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115914110B_ABST
    Figure CN115914110B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention discloses a deadlock release circuit and a deadlock release method based on an on-chip network. The deadlock release circuit of the embodiment of the present invention includes a plurality of block circuits, wherein the block circuit includes a synchronizer and a plurality of data buffers, and the plurality of block circuits are sequentially connected through a data channel to form a ring path; a channel buffer is provided on the ring path and is configured to cache data transmitted on the ring path in a cache state; a deadlock management circuit is connected to the block circuit and the channel buffer, and is used to receive the first signal and send cache enable information to the channel buffer; the cache enable information is used to instruct the channel buffer to switch from a pass-through state to a cache state, so as to provide data exchange space for the ring path to be unhooked. The deadlock management circuit and the channel buffer provided in the above circuit can release the deadlock caused by the ring path during data transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a deadlock release circuit and a deadlock release method based on a network on chip. Background Art

[0002] In a system on chip (SoC), a network-on-chip (NoC) is used to enable communication between subsystems. Specifically, a NoC architecture includes multiple subsystems (i.e., computing cores), which are interconnected through broadcast routers (BRs). The subsystems route data in the form of packets to the target subsystem.

[0003] Since system-on-chip (SoC) based on on-chip network (NoC) can better meet the needs of high-bandwidth and low-latency data transmission, NoC is the optimal interconnection mechanism for SoC. However, in the design of complex NoC network topology, during the data transmission process, hanging phenomena often occur due to circular routing paths (ring paths).

[0004] In summary, during the data transmission process, how to solve the deadlock phenomenon caused by the ring path is a problem that needs to be solved at present. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a deadlock release circuit and a deadlock release method based on an on-chip network, which can release the deadlock phenomenon caused by a ring path during data transmission.

[0006] In a first aspect, an embodiment of the present invention provides a hang-free circuit based on a network on chip, the circuit comprising:

[0007] Multiple block circuits, each block circuit including a synchronizer and multiple data buffers, wherein the multiple block circuits are sequentially connected via a data channel to form a ring path; the synchronizer is configured to send a first signal to a deadlock management circuit, wherein the first signal is configured to indicate that all data buffers of the block circuit have entered a deadlock state;

[0008] a channel buffer, provided on the ring path, and configured to cache data transmitted on the ring path in a cache state;

[0009] The deadlock management circuit is connected to the blocking circuit and the channel buffer, and is used to receive the first signal and send cache enable information to the channel buffer; the cache enable information is used to instruct the channel buffer to switch from a pass-through state to a cache state to provide data exchange space for the ring path to be unblocked.

[0010] Optionally, in response to all the data buffers in any block circuit on the ring path reaching a preset high watermark, the synchronizer of any block circuit sends the first signal to the deadlock management circuit; and / or, the data buffer enters a blocked state, indicating that the capacity of the data buffer reaches a preset high watermark.

[0011] Optionally, in response to the deadlock management circuit receiving a first signal sent by synchronizers of all the block circuits connected to it, the deadlock management circuit sends cache enable information to the channel buffer to switch the channel buffer from a pass-through state to a cache state.

[0012] Optionally, the block circuit includes an arbitrator;

[0013] The deadlock management circuit is further configured to send a deadlock release signal to the arbitrator located on the ring path in the block circuit;

[0014] After the arbitrator receives the hang-up release signal, the arbitrator switches from a fair scheduling state to a priority scheduling state, and preferentially schedules the data on the ring path.

[0015] Optionally, the deadlock management circuit is further configured to receive a second signal sent by the synchronizer of the block circuit, and send cache deactivation information to the channel buffer;

[0016] Among them, the second signal is used to indicate that all the data buffers in the blocking circuit enter the second state, and the second state indicates that the data buffer reaches a preset low watermark; the cache disable information is used to instruct the channel buffer to switch from the cache state to the pass-through state.

[0017] Optionally, in response to all the data buffers in any block circuit on the ring path reaching a preset low watermark, the synchronizer of any block circuit sends the second signal to the deadlock management circuit.

[0018] Optionally, in response to the deadlock management circuit receiving a second signal sent by the synchronizers of all the block circuits connected to it, the deadlock management circuit sends cache disable information to the channel buffer to switch the channel buffer from a cache state to a pass-through state.

[0019] Optionally, the deadlock management circuit is further configured to send a deadlock release signal to the arbitrator located on the ring path in the block circuit;

[0020] After the arbitrator receives the hang-up release signal, the arbitrator switches from the priority scheduling state to the fair scheduling state.

[0021] Optionally, the channel buffer is placed on the annular passage, specifically comprising:

[0022] The channel buffer is placed between any two of the block circuits; or,

[0023] The channel buffer is placed inside any of the block circuits.

[0024] Optionally, the annular path transmits data in a clockwise direction, or the annular path transmits data in a counterclockwise direction.

[0025] In a second aspect, an embodiment of the present invention provides a hang-free circuit based on a network on chip, the circuit comprising:

[0026] a plurality of block circuits, each block circuit including a plurality of data buffers, wherein the plurality of block circuits are sequentially connected via a data channel to form a ring path; the data buffer is configured to send a third signal to a deadlock management circuit, the third signal being configured to indicate that the data buffer has entered a deadlock state;

[0027] a channel buffer, provided on the ring path, and configured to cache data transmitted on the ring path in a cache state;

[0028] The deadlock management circuit is connected to the blocking circuit and the channel buffer, and is used to receive the third signal and send cache enable information to the channel buffer; the cache enable information is used to instruct the channel buffer to switch from a pass-through state to a cache state to provide data exchange space for the ring path to be unblocked.

[0029] In a third aspect, an embodiment of the present invention provides a method for resolving a deadlock based on an on-chip network, the method comprising:

[0030] receiving a plurality of first signals sent by synchronizers of all block circuits on a ring path, wherein the first signals are used to indicate that all the data buffers in the block circuits have entered a blocked state;

[0031] Sending cache enable information to the channel buffer to switch the channel buffer from a pass-through state to a cache state, and / or sending a release signal to the arbitrator on the ring path, wherein after receiving the release signal, the arbitrator switches from a fair scheduling state to a priority scheduling state, and preferentially schedules data on the ring path.

[0032] Optionally, a plurality of second signals sent by all the block circuits on the ring path are received, wherein the second signals are used to indicate that all the data buffers in the block circuits have entered a second state, and the second state indicates that the data buffers have reached a preset low watermark;

[0033] Sending cache disable information to the channel buffer to switch the channel buffer from a cache state to a pass-through state, and at the same time, sending a hang-up release signal to the arbitrator on the ring path, wherein after receiving the hang-up release signal, the arbitrator switches from the priority scheduling state to the fair scheduling state.

[0034] In a fourth aspect, an embodiment of the present invention provides a method for resolving a deadlock based on an on-chip network, the method comprising:

[0035] receiving a plurality of third signals sent by a plurality of data buffers of all block circuits on the ring path, wherein the third signals are used to indicate that the data buffers have entered a blocked state;

[0036] Sending cache enable information to the channel buffer to switch the channel buffer from a pass-through state to a cache state, and / or sending a release signal to the arbitrator on the ring path, wherein after receiving the release signal, the arbitrator switches from a fair scheduling state to a priority scheduling state, and preferentially schedules data on the ring path.

[0037] In a fifth aspect, an embodiment of the present invention provides an integrated circuit, comprising multiple cores, an on-chip network, and an on-chip network-based deadlock release circuit as described in the first aspect, any possible embodiment of the first aspect, or any one of the second aspects.

[0038] In a sixth aspect, an embodiment of the present invention provides a board card, which includes the integrated circuit of the fifth aspect.

[0039] In a seventh aspect, an embodiment of the present invention provides a server, comprising the board card according to the sixth aspect.

[0040] The embodiment of the present invention uses multiple block circuits, which include synchronizers and multiple data buffers. The multiple block circuits are connected in sequence through data channels to form a ring path; the synchronizer is used to send a first signal to the deadlock management circuit, and the first signal is used to indicate that all the data buffers of the block circuit have entered a blocked state; the channel buffer is set on the ring path and is configured to cache the data transmitted on the ring path in a cache state; the deadlock management circuit is connected to the block circuit and the channel buffer, and is used to receive the first signal and send cache enable information to the channel buffer; the cache enable information is used to instruct the channel buffer to switch from a pass-through state to a cache state, so as to provide data exchange space for the ring path to be unblocked. Through the above circuit, when all data buffers on the ring path enter a blocked state, that is, the ring path is dead, the channel buffer provides data exchange space for the ring path to be unblocked, thereby eliminating the deadlock phenomenon caused by the ring path during data transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0042] Figure 1 This is a schematic diagram of a network-on-chip (SOC) structure in the prior art;

[0043] Figure 2 This is a schematic diagram of a data transmission process in the prior art;

[0044] Figure 3 This is a schematic diagram of a deadlock release circuit in an embodiment of the present invention;

[0045] Figure 4 This is a schematic diagram of a deadlock release circuit in an embodiment of the present invention;

[0046] Figure 5 This is a schematic diagram of the internal circuit structure of a BR according to an embodiment of the present invention;

[0047] Figure 6 This is a schematic diagram of a deadlock release circuit in an embodiment of the present invention;

[0048] Figure 7 The present invention is a flowchart of a method for resolving a deadlock. DETAILED DESCRIPTION

[0049] The present disclosure is described below based on examples, but the present disclosure is not limited to these examples. Certain specific details are described in detail in the following detailed description of the present disclosure. A person skilled in the art can fully understand the present disclosure without these details. To avoid obscuring the essence of the present disclosure, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0050] Furthermore, persons of ordinary skill in the art will appreciate that the figures provided herein are for illustration purposes only and are not necessarily drawn to scale.

[0051] Unless the context clearly requires otherwise, words like “include”, “comprising” and the like throughout this application should be interpreted as including rather than exclusive or exhaustive; that is, as meaning “including but not limited to”.

[0052] In the description of the present disclosure, it should be understood that the terms "first," "second," etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance. In addition, in the description of the present disclosure, unless otherwise specified, "plurality" means two or more.

[0053] In the system on chip (SoC), a network on chip (NoC) is used to realize communication between subsystems. Specifically, the system on chip (SoC) includes multiple banks, each bank includes multiple computing cores (COREs). Assume that the system on chip (SoC) includes 4 banks, each bank includes 4 cores, as shown in the following example: Figure 1 As shown, the four banks are BANK0, BANK1, BANK2 and BANK3. Taking BANK0 as an example, its cores are CORE0, CORE1, CORE2 and CORE3. The names of the cores in other banks are as follows: Figure 1 Each bank also includes four broadcast routers (BRs): BR0, BR1, BR2, and BR3 in bank 0; BR4, BR5, BR6, and BR7 in bank 1; BR8, BR9, BR10, and BR11 in bank 2; and BR12, BR13, BR14, and BR15 in bank 3. Each bank also includes two asynchronous bridges (Async Bridges) for data connection with other banks.

[0054] The above four banks BR0, BR1, BR4, BR5, BR8, BR9, BR12, BR13 and the two asynchronous bridges included in each bank form a closed ring routing path, that is, a ring path. During data transmission, data can be transmitted clockwise on the above ring path, or counterclockwise on the above ring path. When the ring path is transmitting counterclockwise or clockwise, it generally has the characteristics of one-way cyclic transmission, which will cause the ring path to hang due to mutual waiting of resources. Take the clockwise transmission of data on the ring path as an example to illustrate the hang. Specifically, Figure 2 As shown in the figure, since data is transmitted clockwise, the buffer of each bank caches data sent to other banks. The buffer resources of bank 0 cache all data going to bank 3; the buffer resources of bank 1 cache all data going to bank 2; the buffer resources of bank 3 cache all data going to bank 0; and the buffer resources of bank 2 cache all data going to bank 1. Therefore, a resource waiting phenomenon occurs, that is, data going to bank 3 needs to wait for the buffer resource space of bank 1 to have resource space before it can continue to be routed forward; data going to bank 2 needs to wait for the buffer resource space of bank 3 to have resource space before it can continue to be routed forward; data going to bank 0 needs to wait for the buffer resource space of bank 2 to have resource space before it can continue to be routed forward; data going to bank 1 needs to wait for the buffer resource space of bank 0 to have resource space before it can continue to be routed forward. Due to the formation of a resource waiting dependency loop, there is no idle resource in the buffer of each bank, and thus the data in each bank cannot be routed forward, resulting in the clockwise data access path being hung. Figure 2 The number of buffers in the four banks described above is only exemplary and should be determined based on actual conditions.

[0055] In the prior art, in order to solve the problem of hanging, the ring path is usually interrupted in the structure, as shown above. Figure 2In the system shown, the physical channel between any two banks is disconnected. For example, the physical channel between BANK0 and BANK2 is disconnected. If BANK0 needs to send data to BANK2, it cannot send the data directly. Instead, the routing rules need to be modified. The original path from BANK0->BANK2 is replaced with BANK0->BANK1->BANK3->BANK2. If BANK2 needs to send data to BANK0, it cannot send the data directly. Instead, the routing rules need to be modified. The original path from BANK2->BANK0 is replaced with BANK2->BANK3->BANK1->BANK0. If BANK3 needs to send data to BANK0, the original path from BANK3->BANK2->BANK0 is replaced with BANK3->BANK1->BANK0. Disrupting the ring path structure not only changes the topology but also requires changing the routing method during data transmission, increasing the access delay between BANK0 and BANK2. The traffic between BANK3->BANK1->BANK0 is large, resulting in unbalanced traffic access, causing this path to become a performance bottleneck.

[0056] In summary, how to solve the problem of hanging in the ring passage and avoid various problems caused by interrupting the ring passage structure in the prior art is currently in need of solution.

[0057] In the embodiment of the present invention, in order to solve the deadlock problem in the ring path, a deadlock release circuit based on the on-chip network is proposed. Figure 3 As shown, Figure 3 This is a schematic diagram of a hang-free circuit based on a network on chip according to an embodiment of the present invention, which specifically includes: a plurality of block circuits 300 , a channel buffer 301 and a hang-free management circuit 302 .

[0058] For multiple block circuits 300, multiple block circuits are connected in sequence through data channels to form a ring path. The data channel is used for data transmission between block circuits. Each block circuit includes a synchronizer and multiple data buffers, and the data buffer is used to cache data transmitted on the ring path. After the synchronizer collects multiple data buffers in the block circuit where it is located and enters a blocked state, it is used to send a first signal to the deadlock management circuit. The first signal is used to indicate that all the data buffers of the block circuit have entered a blocked state. At this time, the block circuit where the synchronizer is located is blocked. Alternatively, each block circuit includes multiple data buffers, and the data buffer is used to send a third signal to the deadlock management circuit. The third signal is used to indicate that the data buffer has entered a blocked state. At this time, only the data buffer that sends the third signal is blocked. Since regardless of whether the block circuit has a synchronizer, the signal sent to the deadlock management circuit is sent from the block circuit, therefore, Figure 3 The connection line between the block circuit and the deadlock management module is drawn as an example. In actual operation, the connection line can be connected to the synchronizer in the block circuit, or to each data buffer in the block circuit, which will not be described here.

[0059] The channel buffer 301 is provided on the ring path and is configured to cache data transmitted on the ring path in a buffering state. The channel buffer has two states: a pass-through state and a cache state. When the channel buffer is in the pass-through state, it does not cache data on the ring path; when in the cache state, it can cache data on the ring path.

[0060] The deadlock management circuit 302 is connected to the block circuits and the channel buffers, and is configured to receive the first signal or the third signal, and, in response to the first signal received from the synchronizers of the plurality of block circuits or the third signal received from the data buffers of the plurality of block circuits, send cache enable information to the channel buffers. The cache enable information is configured to instruct the channel buffers to switch from a pass-through state to a cache state, thereby providing data exchange space for unblocking the ring path.

[0061] In the embodiment of the present invention, only a channel buffer needs to be added to the ring path, and the topology structure and routing rules of the NoC do not need to be changed. The implementation is simple, and the data transmission of other paths except the ring path will not be affected during the unhang process, resulting in less performance loss.

[0062] In one possible implementation, the block circuit includes a synchronizer and multiple data buffers. The synchronizer is connected to the multiple data buffers and the deadlock management module and is used to collect the status of the multiple data buffers. When all the data buffers enter the deadlock state, the synchronizer generates a first signal and sends it to the deadlock management module. When the amount of data in the data buffer reaches a preset high watermark, it indicates that the data buffer enters the deadlock state, that is, the data buffer is in a nearly full state. When all the data buffers on the ring path in the block circuit enter the deadlock state, it indicates that the block circuit is blocked on the ring path. In this embodiment, after collecting the nearly full status of all data buffers in the block circuit through the synchronizer of the block circuit, a unified signal is sent to the deadlock management module. At this time, the deadlock management module can determine that the block circuit that sent the signal is in a deadlock state upon receiving a signal.

[0063] In one possible implementation, the blocking circuit includes multiple data buffers, each of which is used to send a third signal to the deadlock management circuit, where the third signal is used to indicate that the amount of data in the data buffer has reached a preset high watermark.

[0064] In the embodiment of the present invention, the data buffer may be a first-in-first-out queue.

[0065] In the embodiment of the present invention, the number of channel buffers on the annular path may be one or more, depending on the actual situation. The greater the number of channel buffers, the faster the deadlock is resolved. Theoretically, only one channel buffer needs to be added to resolve the deadlock. The greater the number of channel buffers on the annular path, the faster the deadlock is resolved.

[0066] In one possible implementation, in response to the deadlock management circuit receiving a first signal sent by all the block circuits connected to it, the deadlock management circuit sends cache activation information to the channel buffer, switching the channel buffer from a pass-through state to a cache state. That is, when the deadlock management circuit determines that all the block circuits are in a deadlock state, the channel buffer on the ring path is activated.

[0067] In another possible implementation, if one or more channel buffers are set corresponding to each block circuit, when the deadlock management circuit receives the first signal of a certain block circuit, the deadlock management circuit can send cache enablement information to the channel buffer corresponding to the deadlocked block circuit.

[0068] In one possible implementation, the block circuit includes an arbiter; the deadlock management circuit is further configured to send a deadlock release signal to the arbiter located on the ring path within the block circuit. Upon receiving the deadlock release signal, the arbiter switches from a fair scheduling state to a priority scheduling state, thereby prioritizing data on the ring path. According to embodiments of the present invention, data on the ring path can be prioritized by not only enabling channel buffers on the ring path but also controlling the arbiter on the ring path.

[0069] In another feasible embodiment, the arbiter on the ring path can be switched from a fair scheduling state to a priority scheduling state before the channel buffer is activated. Specifically, after determining that the ring path is blocked, the deadlock management circuit first sends a deadlock release signal to the arbiter located on the ring path in the block circuit. After the arbiter switches from a fair scheduling state to a priority scheduling state, it feeds back an acknowledgment signal (ack) to the deadlock management circuit. The deadlock management circuit receives the acknowledgment signals from all arbiters located on the ring path in each block circuit and sends a cache activation message to the channel buffer. By restricting the activation order of the arbiter and channel buffer in this embodiment, the efficiency of unblocking the ring path can be guaranteed. This is because if the channel buffer is activated first while the arbiter is still in a fair scheduling state, data from other channels may be scheduled into the ring path, which is equivalent to data from other channels continuously flowing into the ring path. Even if the channel buffer is activated at this time, it may still be blocked and the deadlock cannot be released.

[0070] In another optional embodiment, the block circuit includes a synchronizer and multiple data buffers. The synchronizer can collect the status of the data buffers in the block circuit. When the status of all collected data buffers is the second state, a second signal is sent to the deadlock management module. At this time, the deadlock management module can determine that the block circuit has been unblocked upon receiving the second signal, that is, all data buffers located on the ring path in the block circuit have been unblocked.

[0071] Optionally, the blocking circuit includes multiple data buffers, each of which can send a fourth signal to the deadlock management circuit. The fourth signal indicates that the data buffer has reached a preset low watermark, i.e., the data buffer is nearly empty. The deadlock management circuit can determine that the blocking state of the blocking circuit has been released based on the fourth signals received from all data buffers of the blocking circuit.

[0072] In one possible implementation, the deadlock management circuit is further configured to send cache deactivation information to the channel buffer; the cache deactivation information is configured to instruct the channel buffer to switch from a cache state to a pass-through state. The deadlock management module sends the cache deactivation information to the channel buffer after the data cached in the channel buffer is cleared and the block circuit is in a deadlock-free state.

[0073] In one possible implementation, the deadlock management circuit is further configured to send a deadlock release signal to the arbitrator located on the ring path in the block circuit. After receiving the deadlock release signal, the arbitrator switches from the priority scheduling state to the fair scheduling state. That is, the ring path is unblocked, and data transmission returns to normal. Optionally, the deadlock management circuit may send a deadlock release signal to all arbitrators located on the ring path in each block circuit respectively. In another optional embodiment, the block circuit further includes a synchronizer, and the deadlock management circuit may send a deadlock release signal to the synchronizer of each block circuit, and then the synchronizer will synchronize the deadlock release signal to all arbitrators located on the ring path in the block circuit.

[0074] In the embodiment of the present invention, the Figure 3 The mid-channel buffer 301 is placed between any two of the block circuits.

[0075] Optionally, the channel buffer may also be placed inside any of the block circuits. Figure 4 As shown, the channel buffer is placed between two broadcast routers of any of the block circuits. In the embodiment of the present invention, the channel buffer can be multiple, Figure 3 and Figure 4 Only one example is used for illustration.

[0076] In the embodiment of the present invention, in order to express the data transmission of the ring path more clearly, the internal circuit structure of each BR is shown. The internal circuit structure of each BR is as follows: Figure 5As shown, each direction of the BR can receive and send data. The port for receiving data is represented by RX. Each BR can receive data from four directions, that is, each BR includes four RX ports, namely RX0, RX1, RX2 and RX3. The port for sending data is represented by TX. Each BR can send data in four directions, that is, each BR includes four TX ports, namely TX0, TX1, TX2 and TX3. Among them, each RX port in each direction is equipped with a splitter for sending data received in that direction to the TX ports in the other three directions. The BR also includes four splitters, namely Splitter0, Splitter1, Splitter2 and Splitter3. The splitter can also be called a distribution module. Each TX port in each direction is equipped with an arbiter for arbitrating data sent from the other three directions. The BR also includes four arbiters, namely Arbiter0, Arbiter1, Arbiter2 and Arbiter3. In addition, each TX port in each direction can receive data from three directions. Each data port received in each direction is equipped with a buffer, which is a first-in-first-out queue (FIFO). The BR also includes 12 FIFOs, with three FIFOs for each TX direction, used to buffer data from the other three directions. The connection between each FIFO and the arbiter for that direction forms a channel. Specifically, after data enters the BR, the data is transmitted as follows: it first passes through the splitter distribution module, which distributes the data to the FIFOs in the target direction based on the shortest path routing principle. Data that needs to be routed to the same TX direction is aggregated on the TX side. That is, the three FIFOs in the target direction aggregate the RX data from the three directions, with the data in each FIFO having the same RX direction and TX direction. The aggregated data is then scheduled by the arbiter in the target direction, and then sent to the TX port for output.

[0077] In the embodiment of the present invention, it is assumed that a ring path is formed by 4 BANKs, each BANK includes 4 COREs, and in the clockwise data transmission direction, the Figure 1 As shown in the figure, data is input from the port on one side of the BR in the bank and output from the port on the opposite side. If BR1, BR0, BR4, BR5, BR12, BR13, BR9, and BR8 all input data from RX3, then data is output from TX1. The specific diagram is as follows Figure 6As shown in the figure, the splitter2 corresponding to RX3 of BR1, BR0, BR4, BR5, BR12, BR13, BR9, and BR8, the arbiter1 on the TX1 side, and the FIFO between splitter2 and arbiter1 constitute the ring path of the four banks. Among them, the FIFO located on the ring path in the bank is the data buffer. Figure 6 In the design, a deadlock management circuit is connected to four banks, and four channel buffers are designed. Each channel buffer is between any two banks, for example, between TX1 of bank0 and RX3 of bank1. Figure 6 The dotted line in FIG. 1 represents a clockwise circular path.

[0078] In the embodiment of the present invention, the processing flow of the method for releasing a deadlock based on an on-chip network is as follows: Figure 7 As shown, the specific steps include:

[0079] Step S700: Receive multiple first signals sent by all block circuits on the ring path.

[0080] The first signal is used to indicate that all the data buffers in the block circuit have entered a blocked state. The data buffer enters a blocked state when the capacity of the data buffer reaches a preset high watermark.

[0081] In one possible implementation, the data buffer is a first-in-first-out queue. When the first-in-first-out queue reaches a preset high watermark, it indicates that the first-in-first-out queue will be unable to receive new data. For example, if the depth of the first-in-first-out queue in the BR is 8, the high watermark and low watermark of the FIFO on the ring path in the BR are set. Assume that the high watermark is 6 and the low watermark is 2; when the number of cached data in the FIFO exceeds the high watermark, the FIFO is considered to be in a nearly full state; when the number of cached data in the FIFO is less than the low watermark, the FIFO is considered to be in a nearly empty state.

[0082] In response to all the data buffers in any block circuit on the ring path reaching a preset high watermark, the any block circuit sends the first signal to the deadlock management circuit.

[0083] For example, Figure 6 In the BRO shown, part of the ring path is from RX3 of BR0 to TX1 of BR0, and the FIFO passed through is the gray FIFO in BR0 in the figure. When the gray FIFO reaches a preset high watermark or exceeds the preset high watermark, the BANK corresponding to the gray FIFO sends the first signal to the deadlock management circuit.

[0084] In a possible implementation, the Figure 6 The FIFO on the ring path can also be directly connected to the deadlock management circuit, which is not limited in the embodiment of the present invention. Figure 6 As shown, all eight FIFOs on the ring path send nearly full signals to the deadlock management module, indicating that all eight FIFOs on the ring path are in a nearly full state, which means that the ring path is deadlocked and data transmission is impossible.

[0085] Step S701: Sending cache activation information to the channel buffer to switch the channel buffer from a through state to a cache state. At the same time, sending a hang-up release signal to the arbiter on the ring path.

[0086] In this embodiment of the present invention, upon receiving a first signal transmitted by all of the block circuits connected thereto, the hang management circuit sends cache activation information to the channel buffer, switching the channel buffer from a pass-through state to a cache state. Assuming the channel buffer depth is set to 2, in the pass-through state, no data is cached in the channel buffer during data transmission. Only when the channel buffer switches from the pass-through state to the cache state, i.e., when the ring path hangs, is the channel buffer activated, data is cached in the channel buffer, and swap space is provided for unhanging data on the ring path.

[0087] Optionally, after the arbitrator receives the hang-up release signal, the arbitrator switches from a fair scheduling state to a priority scheduling state, and preferentially schedules data on the ring path.

[0088] Depend on Figure 6 It can be seen that each arbiter is connected to multiple FIFOs. Figure 6 When any arbiter in the ring path of any BR receives a contact hang signal sent by the hang management module, it will prioritize scheduling data in the gray FIFO. When receiving the contact hang signal, if data in other channels is being transmitted, it will complete the transmission of one granularity of data in the channel before switching channels.

[0089] In the embodiment of the present invention, the FIFO on the ring path may be directly connected to the deadlock management module, or may be connected to the deadlock management module through the BANK where the FIFO is located, which is not limited in the embodiment of the present invention.

[0090] Step S702: Receive multiple second signals sent by all block circuits on the ring path.

[0091] The second signal is used to indicate that all the data buffers in the blocking circuit enter a second state, and the second state indicates that the data buffer reaches a preset low watermark.

[0092] In a possible implementation, the second state may also be referred to as an empty state.

[0093] In one possible implementation, the ring path is unblocked due to the activation of the channel buffer, the data cached in the first-in-first-out queue on the ring path falls below the low watermark, and the FIFO is in a near-empty state. In response to all data buffers in any block circuit on the ring path reaching the preset low watermark, any block circuit sends the second signal to the deadlock management circuit. If any block circuit on the ring path sends the second signal to the deadlock management circuit via its synchronizer, or if all FIFOs on the ring path directly send near-empty signals to the deadlock management circuit, the ring path is unblocked.

[0094] Step S703: Sending cache deactivation information to the channel buffer, switching the channel buffer from a cache state to a pass-through state, and at the same time, sending a hang-up release signal to the arbiter on the ring path.

[0095] An embodiment of the present invention provides an integrated circuit, comprising a plurality of cores, an on-chip network, and a deadlock release circuit based on the on-chip network.

[0096] An embodiment of the present invention provides a board card, which includes the integrated circuit.

[0097] An embodiment of the present invention provides a server, which includes the board card.

[0098] As will be appreciated by those skilled in the art, various aspects of embodiments of the present invention may be implemented as systems, methods, or computer program products. Thus, various aspects of embodiments of the present invention may take the form of a complete hardware implementation, a complete software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software aspects with hardware aspects, which may all be generally referred to herein as a "circuit," "module," or "system." Additionally, various aspects of embodiments of the present invention may take the form of a computer program product implemented in one or more computer-readable media having computer-readable program code implemented thereon.

[0099] Any combination of one or more computer-readable media can be utilized. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples (non-exhaustive enumeration) of computer-readable storage media will include the following: an electrical connection with one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the context of an embodiment of the present invention, a computer-readable storage medium can be any tangible medium that can contain or store a program used by an instruction execution system, device, or apparatus, or a program used in conjunction with an instruction execution system, device, or apparatus.

[0100] A computer-readable signal medium may include a propagated data signal having computer-readable program code embodied therein, such as in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or apparatus.

[0101] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0102] The computer program code for performing the operations for various aspects of the embodiments of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, C++, etc., and conventional procedural programming languages ​​such as "C" programming language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer as a stand-alone software package; partially on the user's computer and partially on a remote computer; or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0103] The flowchart legends and / or block diagrams of the methods, devices (systems) and computer program products according to embodiments of the present invention described above describe various aspects of embodiments of the present invention. It will be understood that each block of the flowchart legends and / or block diagrams and the combination of blocks in the flowchart legends and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that the instructions (executed by the processor of the computer or other programmable data processing device) create a device for implementing the function / action specified in the flowchart and / or block diagram block or block.

[0104] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing device, or other apparatus to operate in a particular manner, so that the instructions stored in the computer-readable medium produce an article of manufacture that includes instructions for implementing the functions / actions specified in the flowchart and / or block diagram blocks or blocks.

[0105] The computer program instructions may also be loaded onto a computer, other programmable data processing device or other apparatus to cause a series of operable steps to be performed on the computer, other programmable device or other apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in the flowchart and / or block diagram block or blocks.

[0106] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A hang-up release circuit based on a network on chip, characterized in that: The circuit includes: Multiple block circuits, each block circuit including a synchronizer and multiple data buffers, wherein the multiple block circuits are sequentially connected via a data channel to form a ring path; the synchronizer is configured to send a first signal to a deadlock management circuit, wherein the first signal is configured to indicate that all data buffers of the block circuit have entered a deadlock state; a channel buffer, provided on the ring path, and configured to cache data transmitted on the ring path in a cache state; The deadlock management circuit is connected to the blocking circuit and the channel buffer, and is used to receive the first signal and send cache enable information to the channel buffer; the cache enable information is used to instruct the channel buffer to switch from a pass-through state to a cache state to provide data exchange space for the ring path to be unblocked.

2. The circuit according to claim 1, wherein In response to the deadlock management circuit receiving a first signal sent by synchronizers of all the block circuits connected thereto, the deadlock management circuit sends cache enable information to the channel buffer to switch the channel buffer from a pass-through state to a cache state.

3. The circuit according to claim 1 or 2, characterized in that The blocking circuit includes an arbiter; The deadlock management circuit is further configured to send a deadlock release signal to the arbitrator located on the ring path in the block circuit; After receiving the hang-up release signal, the arbitrator switches from a fair scheduling state to a priority scheduling state to give priority to scheduling the data on the ring path.

4. The circuit according to any one of claims 1 or 2, characterized in that The deadlock management circuit is further configured to receive a second signal sent by the synchronizer of the block circuit, and send cache deactivation information to the channel buffer; The second signal is used to indicate that all the data buffers in the block circuit enter a second state, and the second state indicates that the data buffer reaches a preset low watermark; The cache disable information is used to instruct the channel buffer to switch from a cache state to a pass-through state.

5. The circuit according to claim 4, wherein In response to the deadlock management circuit receiving the second signal sent by the synchronizers of all the block circuits connected thereto, the deadlock management circuit sends cache disable information to the channel buffer to switch the channel buffer from a cache state to a pass-through state.

6. The circuit according to claim 3, wherein: The deadlock management circuit is further configured to send a deadlock release signal to the arbiter located on the ring path in the block circuit; After the arbitrator receives the hang-up release signal, the arbitrator switches from the priority scheduling state to the fair scheduling state.

7. The circuit according to claim 1, wherein: The channel buffer is arranged on the annular passage, and specifically includes: The channel buffer is placed between any two of the block circuits; or, The channel buffer is placed inside any of the block circuits.

8. A hang-up release circuit based on a network on chip, characterized in that: The circuit includes: a plurality of block circuits, each block circuit including a plurality of data buffers, wherein the plurality of block circuits are sequentially connected via a data channel to form a ring path; the data buffer is configured to send a third signal to a deadlock management circuit, the third signal being configured to indicate that the data buffer has entered a deadlock state; a channel buffer, provided on the ring path, and configured to cache data transmitted on the ring path in a cache state; The hang management circuit is connected to the blocking circuit and the channel buffer, and is used to receive the third signal, and send cache enable information to the channel buffer by receiving the third signal of all data buffers of the blocking circuit; the cache enable information is used to instruct the channel buffer to switch from a pass-through state to a cache state, so as to provide data exchange space for the ring path to be unblocked.

9. A method for resolving a deadlock based on a network on chip, characterized in that: The method includes: receiving a plurality of first signals sent by synchronizers of all block circuits on a ring path, wherein the ring path is formed by sequentially connecting the plurality of block circuits via data channels, the first signals being used to indicate that all data buffers of the block circuits have entered a blocked state; the blocked state of the data buffer indicates that the capacity of the data buffer has reached a preset high watermark; Sending cache enablement information to the channel buffer to switch the channel buffer from a pass-through state to a cache state, and / or sending a hang-release signal to the arbitrator on the ring path, wherein after receiving the hang-release signal, the arbitrator switches from a fair scheduling state to a priority scheduling state to preferentially schedule data on the ring path; the channel buffer is set on the ring path.

10. The method according to claim 9, wherein The method further includes: receiving a plurality of second signals sent by all the block circuits on the ring path, wherein the second signals are used to indicate that all the data buffers of the block circuits have entered a second state; and the data buffer entering the second state indicates that the data buffer has reached a preset low watermark; Sending cache disable information to the channel buffer to switch the channel buffer from a cache state to a pass-through state, and / or sending a hang-up release signal to the arbitrator on the ring path, wherein after receiving the hang-up release signal, the arbitrator switches from the priority scheduling state to the fair scheduling state.

11. A method for resolving a deadlock based on a network on chip, characterized in that: The method includes: receiving a plurality of third signals sent by a plurality of data buffers of all block circuits on a ring path, wherein the ring path is formed by sequentially connecting the plurality of block circuits via data channels, the third signals being used to indicate that the data buffers have entered a blocked state; Sending cache enablement information to the channel buffer to switch the channel buffer from a pass-through state to a cache state, and / or sending a hang-release signal to the arbitrator on the ring path, wherein after receiving the hang-release signal, the arbitrator switches from a fair scheduling state to a priority scheduling state to preferentially schedule data on the ring path; the channel buffer is set on the ring path.

12. An integrated circuit, characterized in that: The integrated circuit includes a plurality of on-chip network-based deadlock release circuits according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Network on chip and hedge suspension relieving method

    CN108632172A

  • Method and system for controlling processing

    CN110322390A