Many-core chip and routing method using the same

By introducing a second routing node and processing core into the many-core chip, the routing path is reconstructed, which solves the problem of chip unusability caused by faulty processing cores, and achieves cost reduction while maintaining software versatility.

CN116028424BActive Publication Date: 2026-03-24WUXI LINGXI BRAIN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Due to manufacturing processes and other reasons, many-core chips contain fault handling cores, which can render the chip unusable or require software modifications to avoid faulty cores, resulting in high manufacturing costs and reduced software versatility.

Method used

By introducing a second routing node and processing core into the many-core chip, the routing path is reconstructed, and the task is transferred to the second processing core for processing, avoiding modification of the mapping software and using the fault handling core to continue executing the task.

Benefits of technology

By maximizing the use of chips with faulty cores, manufacturing and usage costs are reduced, while maintaining software versatility and eliminating the need to replace or discard chips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116028424B_ABST
    Figure CN116028424B_ABST
Patent Text Reader

Abstract

The present disclosure provides a many-core chip and a routing method using the same. The many-core chip comprises: a plurality of first processing cores; a plurality of first routing nodes, which are directly or indirectly connected with each other, and each of the first routing nodes is connected with one of the first processing cores; at least one second routing node, each of the second routing nodes is connected with at least part of the first routing nodes; and at least one second processing core, each of the second processing cores is connected with one of the second routing nodes, and the second processing core is used to replace a failed first processing core to execute a task to be processed. The embodiments of the present disclosure can maximize the utilization of the many-core chip with a failed core.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of chips, and in particular, to a many-core chip and a routing method using the same. BACKGROUND

[0002] With the development of artificial intelligence technology, users have higher and higher requirements for the processing capacity of chips. Since the processing capacity of a single-core chip is limited, a many-core chip can distribute processing tasks to multiple processing cores for parallel execution, significantly improving the processing capacity of the chip. Therefore, research on many-core chips is increasingly valued. SUMMARY

[0003] The present disclosure provides a many-core chip and a routing method using the same.

[0004] In a first aspect, the present disclosure provides a many-core chip, comprising:

[0005] a plurality of first processing cores;

[0006] a plurality of first routing nodes, the plurality of first routing nodes being directly or indirectly connected to each other by signals, and each first routing node being connected to a first processing core by signals;

[0007] at least one second routing node, each second routing node being connected to at least part of the first routing nodes by signals;

[0008] at least one second processing core, each second processing core being connected to a second routing node by signals, and the second processing core being used to replace a failed first processing core to execute a to-be-processed task.

[0009] In a second aspect, the present disclosure provides a routing method based on a many-core chip, comprising:

[0010] determining a target processing core corresponding to a to-be-processed task based on routing information, wherein the routing information is path information for transmitting the to-be-processed task to a target routing node corresponding to the target processing core, the target processing core is a processing core in the plurality of first processing cores, and the target routing node is a routing node in the plurality of first routing nodes;

[0011] determining whether the target processing core is a failed core;

[0012] in a case where it is determined that the target processing core is a failed core, the first routing node corresponding to the target processing core reroutes the to-be-processed task to a second routing node, and the to-be-processed task is processed by a second processing core.

[0013] In a third aspect, the present disclosure provides a many-core chip, comprising:

[0014] a plurality of first processing cores;

[0015] a first topology network comprising a plurality of first routing nodes, the plurality of first routing nodes being directly or indirectly signal connected to each other, each of the first routing nodes being signal connected to one of the first processing cores;

[0016] a second topology network comprising at least one second routing node, each of the second routing nodes being signal connected to at least part of the first routing nodes in the first topology network;

[0017] at least one second processing core, each of the second processing cores being signal connected to one of the second routing nodes;

[0018] a to-be-processed task reaches a target processing core through the first topology network or the second topology network, the target processing core being used to process the to-be-processed task.

[0019] In a fourth aspect, the present disclosure provides a routing method based on a many-core chip, the routing method comprising:

[0020] determining a routing strategy of the many-core chip based on a comparison of routing steps of a routing task in the first topology network and the second topology network, a comparison of loads of the first topology network and the second topology network, a comparison of delays of the first topology network and the second topology network, and / or a congestion situation of the first topology network.

[0021] The many-core chip provided by the embodiments of the present disclosure, since the second routing nodes are signal connected to at least part of the first routing nodes, when a first processing core fails, routing information can be reconstructed through the first routing nodes and the second routing nodes, a to-be-processed task can be transmitted to a second processing core, and the second processing core can be used to replace the failed first processing core to execute the to-be-processed task, without the need to modify mapping software, so that the many-core chip can continue to be used to execute the to-be-processed task, and more, the many-core chip does not need to be replaced or discarded, the many-core chip with a failed core is maximally utilized, and the manufacturing cost of the many-core chip and the use cost of a user are reduced.

[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0023] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of the specification, illustrate embodiments of the present disclosure and are used to explain the present disclosure together with the specification. The accompanying drawings, which should not be taken as limiting the present disclosure, illustrate the following embodiments:

[0024] Figure 1 A partial structure schematic diagram of a many-core chip provided for an embodiment of the present disclosure is shown in FIG. 1;

[0025] Figure 2 A partial structure schematic diagram of another many-core chip provided for an embodiment of the present disclosure is shown in FIG. 2;

[0026] Figure 3 A partial structure schematic diagram of still another many-core chip provided for an embodiment of the present disclosure is shown in FIG. 3;

[0027] Figure 4 A flowchart of a routing method provided for an embodiment of the present disclosure is shown in FIG. 4;

[0028] Figure 5 A flowchart of a routing method provided for an embodiment of the present disclosure is shown in FIG. 4; DETAILED DESCRIPTION

[0029] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, the description below omits the description of well-known functions and structures.

[0030] The embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict, if necessary.

[0031] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0032] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. "Coupled" or "connected" or similar terms are not restricted to physical or mechanical connections or associations, but can also include electrical connections, whether direct or indirect.

[0033] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly literal or overly formal sense unless expressly so defined herein.

[0034] The many-core chip includes a plurality of processing cores (or computing engines) and routing nodes, each of the routing nodes corresponding to a processing core, and the plurality of processing cores are signal-connected with the respective routing nodes. In the production process of the chip, due to manufacturing process and other reasons, some processing cores inevitably have faults, and the processing cores with faults cannot execute algorithms (or cannot map algorithms to the processing cores with faults), so that the chip cannot be used, and even the many-core chip of the whole multi-chip architecture has to be discarded, which finally leads to high manufacturing cost of the many-core chip. In addition, in the use process of the many-core chip, due to overheating, errors and other reasons, some processing cores cannot be normally used temporarily, and if the chip with the processing cores with faults is to be continuously used, the software needs to be modified to avoid the processing cores with faults in the mapping. However, this will reduce the software universality. Therefore, it is necessary to provide a many-core chip and a routing method, which can reconfigure the routing to maximize the use of the chip with the processing cores with faults.

[0035] Figure 1 A partial structure schematic diagram of a many-core chip provided by an embodiment of the present disclosure is shown in FIG. 1. Referring to FIG. 1, Figure 1 The many-core chip provided by the embodiment of the present disclosure includes:

[0036] a plurality of first processing cores 11; wherein the first processing core 11 can be used to execute computing, storage, reading and other processing tasks.

[0037] a plurality of first routing nodes 12, the plurality of first routing nodes 12 are directly or indirectly signal-connected with each other, and each first routing node 12 is signal-connected with a first processing core 11.

[0038] In some embodiments, two adjacent first routing nodes 12 are connected by a direct connection, while two non-adjacent first routing nodes 12 are connected by an indirect connection. An indirect connection means that one or more first routing nodes 12 can be configured between two non-adjacent first routing nodes 12.

[0039] In this embodiment of the disclosure, multiple first routing nodes 12 are directly or indirectly connected to each other to form a first topology network, and the task to be processed can be transmitted to the target processing core through the first topology network.

[0040] In this embodiment, the many-core chip further includes at least one second routing node 22, each second routing node being signal-connected to at least a portion of the first routing nodes. Multiple first routing nodes 12 and second routing nodes 22 constitute a second topology network, through which tasks to be processed can be transmitted to the target processing core.

[0041] At least one second processing core 21, each second processing core being signal-connected to a second routing node, the second processing core being used to replace the first processing core.

[0042] In this embodiment of the disclosure, the function of the second processing core 21 is the same as that of the first processing core 11, that is, the second processing core 21 can be used to perform calculation, storage, reading and waiting processing tasks.

[0043] The many-core chip provided in this disclosure embodiment, since the second routing node is signal-connected to at least part of the first routing node, can reconstruct routing information through the first and second routing nodes when the first processing core fails, transmit the task to be processed to the second processing core, and use the second processing core to replace the failed first processing core to execute the task. The many-core chip can continue to be used to execute the task without modifying the mapping software, and there is no need to replace or discard the many-core chip. This maximizes the utilization of the many-core chip with the faulty core, reducing the manufacturing cost of the many-core chip and the user's usage cost.

[0044] In some embodiments, the number of second routing nodes and second processing cores is one, and the second routing node is signal-connected to all or part of the first routing nodes.

[0045] like Figure 1 As shown, the many-core chip includes a second routing node 22 and a second processing core 21. The second processing core 21 is directly signal-connected to the second routing node 22. Moreover, the second routing node 22 can be signal-connected to all the first routing nodes 12 in the many-core chip. The second processing core 21 can be directly signal-connected to any of the first routing nodes 12 through the second routing node 22.

[0046] In this embodiment of the disclosure, the second routing node 22 is configured with an address identifier, which can be used to clearly locate the routing node.

[0047] When mapping tasks to be processed to the many-core chip, the second processing core 21 does not map any tasks. When the first processing core 11 executing the task in the many-core chip fails, the failed first processing core 11 is identified as the target processing core, and the second processing core 12 is used to execute the task in place of the failed target processing core. For example, the identifier of the first routing node 12 corresponding to the target processing core is replaced with the identifier of the second routing node 22, and the address of the target processing core is replaced with the address of the second processing core 21, so that the task to be processed can be transmitted to the second processing core 21 for execution.

[0048] like Figure 2 As shown, the many-core chip includes a second routing node 22 and a second processing core 21. The second processing core 21 is directly signal-connected to the second routing node 22, and the second routing node 22 is also signal-connected to some of the first routing nodes 12 in the many-core chip. For example, the second routing node 22 is directly signal-connected to the first routing nodes 12 in the first column, and also directly signal-connected to the first routing nodes 12 in the first and fourth rows. The second routing node 22 is indirectly signal-connected to other first routing nodes 12 in the many-core chip, meaning that the second routing node 22 can indirectly signal-connect to other first routing nodes 12 through the directly connected first routing nodes 12.

[0049] In some embodiments, the first routing node 12 is divided into multiple routing node groups according to a preset step size. For example, if the preset step size is 5, then any first routing node 12 can reach the target processing core within every 5 steps by passing through the second routing node 22. Grouping in this way can reduce the complexity of the topology network, thereby reducing the latency, congestion and power consumption of the topology network.

[0050] In other embodiments, the multiple first routing nodes are divided into multiple routing node groups according to their regional locations. That is, the multiple first routing nodes are divided into multiple routing node groups based on the physical location of the first routing node 12. For example, ten adjacent first routing nodes 12 are divided into one routing node group.

[0051] In some embodiments, the number of second routing nodes and second processing cores is two or more. Multiple first routing nodes are divided into multiple routing node groups, and the number of routing node groups is the same as the number of second routing nodes. Each second routing node is signal-connected to all or some of the first routing nodes within a routing node group.

[0052] likeFigure 3 As shown, the many-core chip includes two second routing nodes 22 and two second processing cores 21, with each second processing core 21 and second routing node 22 connected to each other in a one-to-one signal connection. In some embodiments, the two second routing nodes 22 are each connected to a portion of the first routing nodes 12. The first routing nodes 12 connected to the two second routing nodes 22 are not duplicated; that is, each first routing node 12 is connected to only one of the second routing nodes 22. In other embodiments, the two second routing nodes 22 cover all the first routing nodes 12, that is, all the first routing nodes in the many-core chip are divided into two groups, with each group of first routing nodes connected to one of the second routing nodes 22.

[0053] In some embodiments, multiple second routing nodes 22 may be directly or indirectly connected by signals.

[0054] In some embodiments, a plurality of first processing cores are arranged in an array, with each row and column of first processing cores intersecting to define a plurality of routing areas; wherein, a routing area refers to an area where first routing nodes can be configured. A plurality of first routing nodes are arranged in an array, with each first routing node configured in one routing area.

[0055] like Figures 1 to 3 As shown, the first processing core 11 and the first routing node 12 form a processing core array and a routing node array, respectively. The two arrays are interleaved. That is, in the processing core array, the intersection area of ​​each row and each column of the first processing cores defines multiple routing areas, and each routing area can be configured with one first routing node 12. The intersection area of ​​each row and each column of the first routing node 12 defines multiple processing core areas, and each processing core area can be configured with one first processing core 11. Multiple first routing nodes 12 are spaced apart in multiple routing areas, and multiple first processing cores 11 are spaced apart in multiple processing core areas, forming a cross-array of processing core array and routing node array.

[0056] In some embodiments, each of the first routing nodes 12 in the array is signal-connected to the second routing node 22, such as Figure 1 As shown. In some embodiments, only a portion of the first routing node 12 in the array is signal-connected to the second routing node 22, such as... Figure 2 As shown. In some embodiments, the first routing node 12 configured in the array can be signal-connected to one second routing node 22, or it can be signal-connected to multiple second routing nodes, such as... Figure 3 As shown.

[0057] In some embodiments, where some first routing nodes are signal-connected to second routing nodes, the first routing nodes are signal-connected to second routing nodes at predetermined intervals in the row direction and / or column direction.

[0058] The number of first routing nodes 12 at intervals can be preset. For example, the preset number can be 3. Then, in the row direction, every 3 first routing nodes 12 are signal-connected to the second routing node 22. In the column direction, every 3 first routing nodes 12 are signal-connected to the second routing node 22. The interval between the first routing nodes 12 and the second routing nodes 22 can reduce the complexity of many-core chips and allow the task to be processed to reach the target processing core from any first processing core 11 within a preset step size.

[0059] In some embodiments, the many-core chip includes a substrate and one or more functional layers stacked sequentially, with a plurality of first processing cores 11 and at least one second processing core 21 disposed on the same or different layers of the many-core chip.

[0060] In some embodiments, the many-core chip includes one or more network layers, and a plurality of first routing nodes and at least one second routing node are interconnected through one or more network layers.

[0061] For example, multiple first routing nodes and second routing nodes are configured in one network layer, and the multiple first routing nodes and second routing nodes are interconnected through one network layer. Alternatively, multiple first routing nodes are configured in one network layer, and second routing nodes are configured in another network layer, and the multiple first routing nodes and second routing nodes are interconnected through two network layers.

[0062] Secondly, embodiments of this disclosure also provide a routing method that, by reconstructing routes, can avoid faulty processing cores, thereby maximizing the utilization of chips with faulty cores.

[0063] Figure 4 A flowchart illustrating a routing method provided in an embodiment of this disclosure. Figure 4 As shown, this routing method is based on a many-core chip provided in this disclosure embodiment, and includes:

[0064] Step S401: Determine the target processing core corresponding to the task to be processed based on the routing information.

[0065] The routing information refers to the path information that transmits the task to be processed to the target routing node corresponding to the target processing core. The target processing core is a processing core among multiple first processing cores, and the target routing node is a routing node among multiple first routing nodes.

[0066] In some embodiments, after receiving the routing information, the first routing node parses the routing information to obtain the destination address of the task to be processed. The first routing node determines whether the address of the first processing core is the target processing core (local processing core). If not, it continues to send the task to be processed to the next first routing node. If it is, step S402 is executed.

[0067] Step S402: Determine whether the target processing core is a faulty core. If yes, proceed to step S403; otherwise, proceed to step S404.

[0068] This disclosure does not limit the method of determining the faulty core. Any determination method in the related technology can be used to determine whether the target processing core is a faulty core.

[0069] If the target processing core is determined not to be a faulty core, then the target processing core will handle and execute the pending task. If the target processing core is determined to be a faulty core, then step S403 will be executed.

[0070] In step S403, if it is determined that the target processing core is a faulty core, the first routing node corresponding to the target processing core will route the task to be processed to the second routing node, and the second processing core will process it.

[0071] In this process, the first routing node, acting as a relay routing node, only needs to modify the storage address within the target processing core to the storage address within the second processing core. It does not need to modify the mapping relationship of the many-core chips. In other words, when generating the mapping relationship, the target processing core with a fault can still perform the compilation task normally to generate the mapping relationship without any additional processing.

[0072] In some embodiments, the second routing node is identified using its identifier, and the storage address within the second processing core is a unique address for the second processing core. By replacing the storage address within the target processing core with the storage address within the second processing core, the first routing node can forward the task to be processed to the second routing node, and the second processing core can process the task in place of the target processing core.

[0073] In this embodiment, since only the first routing node needs to modify the storage address within the processing core, there is no need to modify the mapping software to avoid the faulty core, which improves the versatility of the mapping software. Moreover, even if there is a faulty core in the many-core chip, it can still be used, thereby maximizing the utilization of the chip with the faulty core, and thus reducing the manufacturing cost of the many-core chip and the user's usage cost.

[0074] Step S404: The task to be processed is sent to the second routing node.

[0075] The routing method provided in this embodiment determines the target processing core based on routing information, and then determines whether the target processing core is a faulty core. If it is a faulty core, the task to be processed is routed to a second routing node, and the second processing core processes the task. This way, the mapping software does not need to be modified, and the many-core chip can continue to be used to execute the task. Furthermore, there is no need to replace or discard the many-core chip, which maximizes the utilization of the many-core chip with faulty cores and reduces the manufacturing cost of the many-core chip and the user's usage cost.

[0076] In some embodiments, the task to be processed is a read / write task, wherein a read task refers to reading data from the target processing core, and a write task refers to writing data to the target processing core.

[0077] If the target processing core is determined to be a faulty core, the first routing node corresponding to the target processing core will route the task to be processed to the second routing node, where the second processing core will process it. This includes: if the target processing core is determined to be a faulty core, the first routing node corresponding to the target processing core will replace the storage address in the target processing core with the storage address in the second processing core; the task to be processed will be routed to the second routing node, and the second processing core will complete the read / write task according to the storage address.

[0078] In some embodiments, the first routing node modifies the storage address in the target processing core to the storage address in the second processing core, thereby routing the task to be processed to the second routing node, and the second processing core completes the task (such as a read / write task) on behalf of the target processing core.

[0079] In some embodiments, when the target processing core is determined to be a faulty core, before the first routing node corresponding to the target processing core replaces the storage address in the target processing core with the storage address in the second processing core, the method further includes: the first routing node creating an address replacement correspondence between the storage address in the faulty core and the storage address in the second processing core.

[0080] In this embodiment of the disclosure, each first routing node can set an address replacement correspondence. When the first processing core corresponding to the first routing node fails, the first routing node can determine the storage address in the second processing core according to the address replacement correspondence, and modify the storage address in the target processing core to the storage address in the second processing core. The first routing node can forward the task to be processed to the second routing node so that the second processing core can execute the task.

[0081] The routing method provided in this embodiment replaces the storage address in the faulty core with the storage address in the second processing core, and enables the second processing core to perform the task to be processed in place of the faulty core. It can continue to use the many-core chip to perform the task to be processed without modifying the mapping software, and there is no need to replace or discard the many-core chip. It maximizes the utilization of the many-core chip with faulty core, and reduces the manufacturing cost of the many-core chip and the user's usage cost.

[0082] Thirdly, embodiments of this disclosure also provide a many-core chip. (Refer to...) Figures 1 to 3 The many-core chip provided in this disclosure includes:

[0083] Multiple first processing cores 11, wherein the first processing cores 11 can be used to perform calculation, storage, and read waiting processing tasks.

[0084] The first topology network (the network indicated by solid lines) includes multiple first routing nodes 12, which are directly or indirectly connected to each other via signaling. Each first routing node 12 is connected to a first processing core 11 via signaling.

[0085] In some embodiments, two adjacent first routing nodes 12 are connected by a direct connection, while two non-adjacent first routing nodes 12 are connected by an indirect connection. An indirect connection means that one or more first routing nodes 12 can be configured between two non-adjacent first routing nodes 12.

[0086] The second topology network (shown by dashed lines) includes at least one second routing node 22, each second routing node 22 being signal-connected to at least a portion of the first routing node 12.

[0087] At least one second processing core 21, each second processing core 21 being signal-connected to a second routing node 22.

[0088] The task to be processed reaches the target processing core through the first topology network or the second topology network. The target processing core is used to process the task to be processed. In other words, the task to be processed can reach the target processing core through the first topology network or through the second topology network.

[0089] In some embodiments, the first topology network and the second topology network are configured at the same network layer or at different network layers. When the first topology network and the second topology network are configured at different network layers, the first topology network and the second topology network can be interconnected through the first routing node 12 and the second routing node 22.

[0090] In some embodiments, a plurality of first processing cores 11 are arranged in an array, with each row and column of first processing cores 11 intersecting to define a plurality of first routing areas; a plurality of first routing nodes 12 are arranged in an array, with each first routing node configured in one first routing area. At least one second processing core 21 is arranged in an array, with each row and column of second processing cores 21 intersecting to define a plurality of second routing areas; at least one second routing node 22 is arranged in an array, with each second routing node 22 configured in one second routing area.

[0091] For example, the first processing core 11 and the first routing node 12 form a processing core array and a routing node array, respectively. The two arrays are interleaved. That is, in the processing core array, the intersection area of ​​each row and each column of the first processing cores defines multiple routing areas, and each routing area can be configured with one first routing node 12. The intersection area of ​​each row and each column of the first routing node 12 defines multiple processing core areas, and each processing core area can be configured with one first processing core 11. Multiple first routing nodes 12 are spaced apart in multiple routing areas, and multiple first processing cores 11 are spaced apart in multiple processing core areas, forming a cross-array of processing core array and routing node array.

[0092] The second processing core 21 and the second routing node 22 form a processing core array and a routing node array, respectively. These two arrays are interleaved. Specifically, in the processing core array, the intersection of each row and column of the second processing core 21 defines multiple routing areas, and each routing area can have one second routing node 22. Similarly, the intersection of each row and column of the second routing node 22 defines multiple processing core areas, and each processing core area can have one second processing core 21. By interleaving multiple second routing nodes 22 across multiple routing areas and multiple second processing cores 21 across multiple processing core areas, a cross-array of processing core and routing node arrays is formed.

[0093] In some embodiments, the number of second routing nodes and second processing cores is one, and the second routing node is signal-connected to all or part of the first routing nodes; or, the number of second routing nodes and second processing cores is two or more, the multiple first routing nodes are divided into multiple routing node groups, and the number of routing node groups is the same as the number of second routing nodes, and each second routing node is signal-connected to all or part of the first routing nodes in a routing node group.

[0094] In this embodiment of the disclosure, the many-core chip includes a first topology network and a second topology network. The task to be processed can reach the target processing core through either the first topology network or the second topology network. When the first topology network is congested, the second topology network can be selected. When the transmission efficiency of the first topology network is lower than that of the second topology network, the second topology network can be selected. When the load of the first topology network exceeds a preset load threshold, the second topology network can be selected.

[0095] Fourthly, this disclosure also provides a routing method that uses a many-core chip provided in this disclosure, which can reconstruct routes based on routing conditions and improve the processing efficiency of the many-core chip.

[0096] The routing method provided in this disclosure includes: determining the routing strategy of the many-core chip based on a comparison of the routing step size of the task to be processed in the first topology network and the second topology network, a comparison of the load of the first topology network and the second topology network, a comparison of the delay of the first topology network and the second topology network, and / or the congestion status of the first topology network.

[0097] In this context, the routing step size refers to the number of routing nodes a task needs to traverse to reach its target node. In some embodiments, the task is transmitted hop-by-hop between routing nodes, with each pair of adjacent routing nodes constituting one routing step size. Load refers to the data transmission status of the first and second topology networks; a higher load indicates more data transmission in both networks, and vice versa. Latency refers to the time taken for data to be transmitted in the first and second topology networks, such as the time from when data enters the first and / or second topology networks to when it leaves. Congestion refers to a sustained overload in the first and / or second topology networks, resulting in a decrease in data transmission capacity. Congestion can be localized or network-wide.

[0098] The routing method provided in this disclosure selects a routing strategy based on the routing step size, load, delay, and congestion of the task to be processed in the first topology network and the second topology network. For example, the task to be processed can be transmitted through either the first topology network or the second topology network, which can improve routing efficiency and thus improve the processing efficiency of many-core chips.

[0099] Figure 5 A flowchart illustrating a routing method provided in an embodiment of this disclosure. Figure 5 As shown, this routing method is based on a many-core chip provided in this disclosure embodiment, and includes:

[0100] Step S501: Based on the first topology network, the task to be processed is transmitted to the target routing node by the first routing step.

[0101] Step S502: Determine the second routing step size for transmitting the task to be processed to the target routing node based on the second topology network.

[0102] Step S503: Calculate the difference between the first route step size and the second route step size to obtain the route step size difference.

[0103] Step S504: Determine whether the routing step size difference is greater than the preset step size threshold.

[0104] Step S505: If the routing step size difference is greater than the preset step size threshold, the task to be processed is transmitted using the second topology network.

[0105] When the first routing node determines that the task to be processed will be transmitted by the second topology network based on the routing step size difference, it modifies the address of the next hop to the routing node in the second topology network, that is, the task to be processed will be transmitted by the second topology network, so as to shorten the step size of the task to be processed.

[0106] Step S506: If the routing step size difference is less than or equal to a preset step size threshold, the task to be processed is transmitted using the first topology network.

[0107] For example, by parsing routing information and looking up the node table in the packet header, the first routing step size for transmitting the task to the target routing node can be calculated, as well as the second routing step size. Assume the first routing step size is 5 and the second routing step size is 3. The difference between the first and second routing step sizes is calculated, resulting in a routing step size difference of 2. If the preset step size threshold is 3, the routing step size difference is less than the preset threshold, and the task to be processed is transmitted to the target routing node through the first topology network. If the preset step size threshold is 1, the routing step size difference is greater than the preset threshold, and the task to be processed is transmitted to the target routing node through the second topology network. That is, compared to the first topology network, the second topology network can transmit the task to the target routing node faster, saving resources in the first topology network, thereby reducing the probability of congestion in the first topology network and improving transmission efficiency.

[0108] For example, when the routing step size is 10 and the second routing step size is 2, the difference between the first and second routing step sizes is calculated, resulting in a routing step size difference of 8. If the preset step size threshold is 3, and the routing step size difference is greater than the preset step size threshold, then the transmission efficiency of the task to be processed through the second topology network to the target routing node is much greater than the transmission efficiency through the first topology network to the target routing node.

[0109] In this embodiment of the disclosure, the first routing node can determine the difference between the first routing step size and the second routing step size based on the routing information, and determine whether the task to be processed is transmitted by the first topology network or the second topology network based on the routing step size difference and the preset step size threshold, so as to shorten the transmission path of the task to be processed and improve the processing efficiency of the many-core chip.

[0110] When the number of first processing cores in a many-core chip is large and many-to-many data interactions are performed, the routing of tasks to be processed is more complex, and the amount of data on some paths is large. The second topology network can simplify the routing and effectively reduce the amount of data in the first topology network, especially the amount of data on paths with large data volumes, thereby improving the operating efficiency of the first topology network and thus improving the processing efficiency of the many-core chip.

[0111] In some embodiments, when the second routing node establishes a signal connection with all the first routing nodes, any one of the first routing nodes can be reached through the second routing node. Therefore, when the amount of data in the first topology network is large, the task to be processed can be quickly transmitted to the target routing node (a routing node in the first topology network) corresponding to the target processing core through the second routing node, thereby reducing the amount of data in the first topology network.

[0112] In this embodiment of the disclosure, by comparing the first routing step size and the second routing step size, and based on the difference between the first routing step size and the second routing step size and a preset step size threshold, the first topology network or the second topology network is selected to transmit the task to be processed. This simplifies the routing, reduces the resource consumption of the first topology network, reduces the latency and congestion probability of the first topology network, and improves the transmission rate of the task to be processed.

[0113] In some embodiments, the routing method includes: step S61, obtaining a first network load of a first topology network and a second network load of a second topology network; step S62, comparing the first network load and the second network load; and step S63, if the first network load exceeds the second network load, using the second topology network to transmit the task to be processed.

[0114] The embodiments disclosed herein do not limit the method of obtaining the first network load and the second network load. Any method of measuring / monitoring network load in the relevant field can be used to obtain the first network load and the second network load.

[0115] For example, when the first network load of the first topology network is large, especially when it exceeds the second network load of the second topology network, the task to be processed can be transmitted through the second topology network to reduce the load of the first topology network and thus avoid congestion in the first topology network.

[0116] In this embodiment of the disclosure, based on the first network load and the second network load, selecting the network with less load to transmit the task to be processed can not only improve the transmission rate of the task to be processed, but also reduce the load of the first topology network, thereby improving the transmission efficiency of the first topology network and reducing the latency of the first topology network.

[0117] In some embodiments, the routing method includes: step S71, obtaining a first network delay of a first topology network and a second network delay of a second topology network; step S72, comparing the first network delay and the second network delay; and step S73, if the first network delay exceeds the second network delay, using the second topology network to transmit the task to be processed.

[0118] For example, when one or more first routing nodes in a many-core chip send and receive data at inconsistent rates, and the on-chip network of the many-core chip can only temporarily store a limited amount of data, it can easily cause a first routing node to wait for data from another first routing node before proceeding to the next operation. This will prolong the routing time of the first topology network. In this embodiment of the present disclosure, when the latency of the first network is greater than the latency of the second network, the second topology network is selected to transmit the task to be processed. That is, the second routing node is selected as a relay routing node, and data is received and buffered at the rate at which the processing core sends data, thus alleviating the latency of the first topology network. Then, the task to be processed is quickly transmitted to the target processing core.

[0119] In this embodiment of the disclosure, by comparing the first network latency and the second network latency, and selecting the topology network with the smaller latency to transmit the task to be processed, not only can the latency of the topology network be alleviated, but the task to be processed can also be quickly transmitted to the target processing core.

[0120] In some embodiments, the first topology network and the second topology network have the same bandwidth. In other embodiments, the first topology network and the second topology network have different bandwidths, and the bandwidth of the second topology network is greater than that of the first topology network.

[0121] In some embodiments, the routing method further includes: obtaining the data transmission rate of the sending core and the data reception rate of the receiving core; wherein the sending core and the receiving core are processing cores among a plurality of first processing cores; and in the case that the data reception rate and the data transmission rate are inconsistent, using a second processing core to relay the data transmission and reception.

[0122] When the data transmission rate of the sending core is inconsistent with the data reception rate of the receiving core (mismatch), it indicates congestion in the routing of the first topology network. Either the sending or receiving core needs to wait for data from the other, which increases routing time and reduces the processing efficiency of the many-core chip. If a second processing core is used as a relay core, the sending core sends data to the second processing core, which then forwards the data to the receiving core. This avoids the extended routing time caused by congestion, thereby improving the processing efficiency of the many-core chip.

[0123] In some embodiments, the second processing core receives data at the rate at which the sending core sends data, and / or sends data at the rate at which the receiving core receives data.

[0124] When the second processing core is signal-connected to all the first routing nodes in the first topology network through the second routing node, the second processing core can send data to any one of the first processing cores in the many-core chip, that is, the second processing core can act as a relay core between any two first processing cores.

[0125] In some embodiments, the tasks to be processed include, but are not limited to, storage tasks, reading tasks, and computing tasks. The first processing core and the second processing core have the same function and can both execute the tasks to be processed.

[0126] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0127] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.

Claims

1. A many-core chip, characterized in that, include: Multiple first processing cores; Multiple first routing nodes are directly or indirectly signal-connected to each other, and each first routing node is signal-connected to a first processing core. At least one second routing node, each second routing node being signal-connected to at least a portion of the first routing node; At least one second processing core, each second processing core being signal-connected to a second routing node, the second processing core being used to replace the faulty first processing core to perform the task to be processed; The method of replacing the faulty first processing core with the second processing core includes: The identifier of the first routing node corresponding to the faulty first processing core is replaced with the identifier of the second routing node corresponding to the second processing core, and the address of the faulty first processing core is replaced with the address of the second processing core. The method of signal connection between the second routing node and the first routing node includes: connecting the first routing node to the second routing node at intervals according to a preset number.

2. The many-core chip according to claim 1, characterized in that, The number of the second routing node and the second processing core is one, and the second routing node is signal-connected to all or part of the first routing nodes; Alternatively, the number of the second routing node and the second processing core is two or more, the plurality of first routing nodes are divided into a plurality of routing node groups, and the number of the routing node groups is the same as the number of the second routing nodes, and each second routing node is signal-connected to all or part of the first routing nodes in one of the routing node groups.

3. The many-core chip according to claim 2, characterized in that, The plurality of first routing nodes are divided into plurality of routing node groups according to region or preset step size.

4. The many-core chip according to any one of claims 1-3, characterized in that, The plurality of first processing cores are arranged in an array, with each row of first processing cores and each column of first processing cores intersecting to define multiple routing areas; The plurality of first routing nodes are arranged in an array, and each first routing node is located in one of the routing areas.

5. The many-core chip according to claim 4, characterized in that, In cases where some of the first routing nodes are signal-connected to the second routing nodes, the first routing nodes are signal-connected to the second routing nodes at predetermined intervals in the row direction and / or column direction.

6. The many-core chip according to claim 1, characterized in that, The plurality of first routing nodes and the at least one second routing node are interconnected through one or more network layers.

7. A routing method based on the many-core chip according to any one of claims 1 to 6, characterized in that, include: The target processing core corresponding to the task to be processed is determined based on the routing information, wherein the routing information is the path information for transmitting the task to be processed to the target routing node corresponding to the target processing core, the target processing core is the processing core among the plurality of first processing cores, and the target routing node is the routing node among the plurality of first routing nodes; Determine whether the target processing core is a faulty core; If the target processing core is determined to be a faulty core, the first routing node corresponding to the target processing core will route the task to be processed to the second routing node, where it will be processed by the second processing core.

8. The routing method according to claim 7, characterized in that, The task to be processed is a read / write task; The first routing node corresponding to the target processing core routes the task to be processed to the second routing node, including: The first routing node corresponding to the target processing core replaces the storage address in the target processing core with the storage address in the second processing core; The task to be processed is routed to the second routing node, and the second processing core completes the read and write tasks according to the storage address.

9. The routing method according to claim 8, characterized in that, Before the first routing node corresponding to the target processing core replaces the storage address in the target processing core with the storage address in the second processing core, the method further includes: The first routing node creates an address replacement correspondence between the storage address in the fault core and the storage address in the second processing core.

10. A many-core chip, characterized in that, include: Multiple first processing cores; A first topology network, comprising a plurality of first routing nodes, wherein the plurality of first routing nodes are directly or indirectly signal-connected to each other, and each first routing node is signal-connected to a first processing core; A second topology network, the second topology network including at least one second routing node, each second routing node being signal-connected to at least a portion of the first routing nodes in the first topology network; At least one second processing core, each second processing core being signal-connected to a second routing node, the second processing core being used to replace the faulty first processing core to perform the task to be processed; The task to be processed arrives at the target processing core through the first topology network or the second topology network, and the target processing core is used to process the task to be processed; The method of replacing the faulty first processing core with the second processing core includes: The identifier of the first routing node corresponding to the faulty first processing core is replaced with the identifier of the second routing node corresponding to the second processing core, and the address of the faulty first processing core is replaced with the address of the second processing core. The method of signal connection between the second routing node and the first routing node includes: connecting the first routing node to the second routing node at intervals according to a preset number.

11. The many-core chip according to claim 10, characterized in that, The first topology network and the second topology network are configured in the same network layer or in different network layers.

12. The many-core chip according to claim 10, characterized in that, The plurality of first processing cores are arranged in an array, with each row of first processing cores and each column of first processing cores intersecting to define a plurality of first routing areas; The plurality of first routing nodes are arranged in an array, and each first routing node is located in a first routing area; And / or, The at least one second processing core is arranged in an array, with each row of the second processing core and each column of the second processing core intersecting to define multiple second routing areas; The at least one second routing node is configured in an array, with each second routing node configured in a second routing area.

13. The many-core chip according to claim 10, characterized in that, The number of the second routing node and the second processing core is one, and the second routing node is signal-connected to all or part of the first routing nodes; Alternatively, the number of the second routing node and the second processing core is two or more, the plurality of first routing nodes are divided into a plurality of routing node groups, and the number of the routing node groups is the same as the number of the second routing nodes, and each second routing node is signal-connected to all or part of the first routing nodes in one of the routing node groups.

14. A routing method, characterized in that, The routing method is based on the many-core chip according to any one of claims 10 to 13, and the routing method includes: The routing strategy of the many-core chip is determined based on the comparison of the routing step size of the task to be processed in the first topology network and the second topology network, the comparison of the load of the first topology network and the second topology network, the comparison of the delay of the first topology network and the second topology network, and / or the congestion status of the first topology network.

15. The routing method according to claim 14, characterized in that, include: Based on the first topology network, determine the first routing step size for transmitting the task to be processed to the target routing node; The second routing step size for transmitting the task to be processed to the target routing node is determined based on the second topology network. Calculate the difference between the first route step size and the second route step size to obtain the route step size difference; If the routing step size difference is greater than a preset step size threshold, the task to be processed is transmitted using the second topology network; If the routing step size difference is less than or equal to a preset step size threshold, the task to be processed is transmitted using the first topology network.

16. The routing method according to claim 14, characterized in that, include: Obtain the first network load of the first topology network and the second network load of the second topology network; Compare the first network load and the second network load; If the load on the first network exceeds the load on the second network, the task to be processed is transmitted using the second network topology.

17. The routing method according to claim 14, characterized in that, include: Obtain the first network latency of the first topology network and the second network latency of the second topology network; Compare the first network latency and the second network latency; If the first network latency exceeds the second network latency, the task to be processed is transmitted using the second topology network.

18. The routing method according to claim 14, characterized in that, include: The data transmission rate of the sending core and the data reception rate of the receiving core are obtained; wherein, the sending core and the receiving core are processing cores among the plurality of first processing cores; If the rate of receiving data and the rate of sending data are inconsistent, the second processing core is used to relay the sending data and the receiving data.

19. The routing method according to claim 18, characterized in that, The second processing core receives data at the rate at which the sending core sends data, and / or sends data at the rate at which the receiving core receives data.

Citation Information

Patent Citations

  • Many-core chip, manufacturing method of many-core chip, route reconstruction method and device

    CN115422119A