Wafer level processor chip, routing unit and routing method

By introducing a computing engine and control forwarding components into the routing unit of a wafer-level processor chip, data processing and traffic aggregation are achieved on the network, solving the problem that the routing unit in traditional wafer-level processor chips does not have data processing capabilities, thus improving network efficiency and reducing latency.

CN120892385APending Publication Date: 2025-11-04BEIJING TSINGMICRO INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510864517.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

The routing unit of a traditional wafer-level processor chip does not have data processing capabilities, which means that the data processing of the collective communication operator needs to be pushed down to the computing unit, introducing communication latency and resource consumption, and the long communication path hinders communication efficiency.

Method used

By adding computing engine components and control forwarding components to the routing unit of the processor unit, data on-network computation and traffic aggregation can be realized. The routing unit completes the aggregated communication operation and offloads computing tasks.

Benefits of technology

It improves the effective bandwidth utilization of 2D Mesh networks, reduces communication latency and network congestion, and optimizes the communication bottleneck for AI training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892385A_ABST
    Figure CN120892385A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of processor chips, and particularly relates to a wafer-level processor chip, a routing unit and a routing method. According to the invention, a calculation engine component and a control forwarding component are added in a routing unit in a processor unit, bidirectional linkage is carried out between the control forwarding component and the calculation engine component, and furthermore, the control forwarding component controls a data operation process in the calculation engine component according to data channel associated control information; data operation of an aggregate communication operator can be unloaded into a routing unit, so that aggregate communication in-network computing capability is introduced into a wafer-level processor chip, and congestion and delay of an aggregate communication network in a network are effectively reduced; moreover, the problem that a traditional wafer-level processor chip does not support the correlation operation of a set communication operator is solved, the utilization rate of the effective bandwidth of the 2D Mesh network is improved, and the communication delay of the 2D Mesh network is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of processor chip technology, specifically relating to a wafer-level processor chip, a routing unit, and a routing method. Background Technology

[0002] To address the bottlenecks in training large AI models, such as limited computing power and memory bandwidth, various novel processor architectures have been proposed by academia and industry. Among them, wafer-level processors (WLPs), due to their unique architectural design, possess significantly higher computing power and memory bandwidth than traditional AI chips. This has made them stand out among numerous new processor architectures, attracting widespread research and application in the industry. The main characteristic of a wafer-level processor is that it is manufactured on a single wafer. This processor chip contains numerous processor units arranged in a grid to form a wafer-level processor chip. The processor units within the wafer-level processor chip are interconnected through a 2D mesh network, enabling communication between the various processor units within the chip.

[0003] Collective communication is a type of communication operator that has achieved high-performance AI training and inference. These operators allow efficient data exchange and collaborative work between computing nodes according to specific rules. Therefore, the communication efficiency of collective communication operators directly affects the computational efficiency of AI clusters. Collective communication operators have been implemented in traditional wafer-level processor chips. However, because the routing units in traditional wafer-level processor chips lack data processing capabilities, the data processing of their collective communication operators needs to be offloaded to the computing units within the processor units. On the one hand, offloading data from the routing units to the computing units not only introduces additional communication latency but also consumes computing resources within the computing units. On the other hand, due to the large array size of processor units within wafer-level processor chips, the communication path of collective communication operators typically needs to traverse multiple processor units to complete. This characteristic significantly hinders the communication efficiency of traditional wafer-level processor chips. Summary of the Invention

[0004] To address at least one of the problems mentioned in the background section, this invention proposes a wafer-level processor chip, a routing unit, and a routing method.

[0005] According to a first aspect of the present invention, the present invention provides a routing unit for a wafer-level processor chip. Specifically, the routing unit is used to receive, process, and forward data. The routing unit comprises a computing engine component, which is used to perform on-network computation on the data. The on-network computation includes at least one of the following computation methods: logical operation and numerical operation. The wafer-level processor chip on which the routing unit is located includes a computing array composed of N rows and M columns of processor units. The processor units in the computing array are interconnected by a 2D Mesh network. Each processor unit includes a computing unit and a routing unit.

[0006] The computing engine component is used to perform on-network computation on target data according to the target computation method. The target data includes part or all of the data carried in the data packet sent by the upstream unit. The data packet also carries data along-path control information. The target data and the target computation method are specified in the data along-path control information.

[0007] Traditional wafer-level processor chips lack data processing capabilities in their routing units, requiring data computation for aggregated communication operators to be offloaded to computation units within the processor unit. This invention proposes adding a computation engine component to the routing unit of the processor unit, thus solving the problem of traditional wafer-level processor chips not supporting operations related to aggregated communication operators. This allows the proposed wafer-level processor chip to not only offload common aggregated communication operations to the routing unit, thereby improving the utilization of effective bandwidth in 2D Mesh networks and reducing communication latency, but also to provide the computational foundation for performing aggregated communication operator functions.

[0008] The routing unit further includes an input buffer component, an output buffer component, and a control forwarding component. The input buffer component is used to temporarily store data packets sent by the upstream unit, the data packets carrying data and data-as-a-path control information. The output buffer component is used to temporarily store data packets sent to the downstream unit, the data packets sent to the downstream unit carrying data calculated on the network and the data-as-a-path control information, the downstream unit determining based on the data-as-a-path control information.

[0009] The routing unit further includes a control and forwarding component, which is used to perform at least one of the following functions:

[0010] The downstream unit is determined based on the data-as-a-path control information; target data is determined based on the data-as-a-path control information, the target data including part or all of the data carried in the data packet sent by the upstream unit; a target calculation method is determined based on the data-as-a-path control information; and a data packet is composed based on the data-as-a-path control information and sent to the downstream unit.

[0011] The wafer-level processor chip proposed in this invention adds a control forwarding component to the routing unit in the processor unit. The control forwarding component realizes the routing and forwarding function of control data packets. This allows the wafer-level processor chip proposed in this invention to perform aggregation processing on network traffic in advance according to the aggregated communication characteristics. This can not only increase the bandwidth of effective traffic in the network, but also significantly reduce the probability of network congestion.

[0012] According to a second aspect of the present invention, the present invention also provides a wafer-level processor chip, which includes a computing array composed of N rows and M columns of processor units, and the processor units in the computing array are interconnected by a 2D Mesh network. The processor units include computing units and any of the routing units described in the first aspect of the present invention, wherein the computing units are used for data computation.

[0013] According to a third aspect of the present invention, the present invention also provides a routing method applicable to any of the routing units described in the first aspect of the present invention, the routing method comprising:

[0014] The computing engine component of the routing unit performs on-network computation on the received data. The on-network computation includes at least one of the following computation methods: logical operation and numerical operation. The wafer-level processor chip on which the routing unit is located contains a computing array composed of N rows and M columns of processor units. The processor units in the computing array are interconnected by a 2D Mesh network. Each processor unit contains a computing unit and a routing unit.

[0015] The computing engine component of the routing unit performs on-network calculations on the received data, including:

[0016] The computing engine component of the routing unit performs on-network computation on the target data according to the target computation method. The target data includes some or all of the data carried in the data packet sent by the upstream unit. The data packet also carries data along-path control information. The target data and the target computation method are specified in the data along-path control information.

[0017] In-network computing is a network communication acceleration technology in the field of distributed parallel computing. By offloading the operations related to aggregate communication operators to the network, it enables the aggregation and processing of network data on the network, eliminates redundancy in communication data and shortens the transmission path, which can significantly reduce the latency of aggregate communication, reduce network congestion, and optimize and alleviate the communication bottleneck problem in AI training. At present, in-network computing technology is mainly used in network devices such as network cards and switches, and there is no such technology application in wafer-level processor chips.

[0018] This invention proposes to add a computing engine component and a control forwarding component to the routing unit of the processor unit in a wafer-level processor chip, and to establish a bidirectional link between the control forwarding component and the computing engine component. Furthermore, the control forwarding component controls the data operation process in the computing engine component according to the data along-path control information, which can offload the data operation of the aggregated communication operator to the routing unit, thereby introducing aggregated communication on-network computing capability into the wafer-level processor chip, and effectively reducing aggregated communication network congestion and latency in the network. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the wafer-level processor chip proposed in this invention;

[0020] Figure 2 This is a schematic diagram of the routing unit in the wafer-level processor chip proposed in this invention. Detailed Implementation

[0021] To make the objectives and features of this invention more apparent and understandable, the technical solution will be described in detail below through embodiments and in conjunction with the accompanying drawings.

[0022] like Figure 2 As shown, the present invention provides a routing unit for a wafer-level processor chip. Specifically, the routing unit is used to receive, process, and forward data. The routing unit comprises a computing engine component, which is used to perform on-network computation on the data. The on-network computation includes at least one of the following computation methods: logical operation and numerical operation. The wafer-level processor chip on which the routing unit is located contains a computing array composed of N rows and M columns of processor units. Furthermore, the processor units in the computing array are interconnected by a 2D Mesh network. Each processor unit contains a computing unit and a routing unit.

[0023] The computing engine component is used to perform on-network computation on target data according to the target computation method. The target data includes part or all of the data carried in the data packet sent by the upstream unit. The data packet also carries data along-path control information. The target data and the target computation method are specified in the data along-path control information.

[0024] Traditional wafer-level processor chips lack data processing capabilities in their routing units, requiring data computation for aggregated communication operators to be offloaded to computation units within the processor unit. This invention proposes adding a computation engine component to the routing unit of the processor unit, thus solving the problem of traditional wafer-level processor chips not supporting operations related to aggregated communication operators. This allows the proposed wafer-level processor chip to not only offload common aggregated communication operations to the routing unit, thereby improving the utilization of effective bandwidth in 2D Mesh networks and reducing communication latency, but also to provide the computational foundation for performing aggregated communication operator functions.

[0025] The routing unit further includes an input buffer component, an output buffer component, and a control forwarding component. The input buffer component is used to temporarily store data packets sent by the upstream unit, the data packets carrying data and data-as-a-path control information. The output buffer component is used to temporarily store data packets sent to the downstream unit, the data packets sent to the downstream unit carrying data calculated on the network and the data-as-a-path control information, the downstream unit determining based on the data-as-a-path control information.

[0026] The routing unit's connection method includes: a unidirectional link between the input side of the input cache component and the upstream unit; a unidirectional link between one end of the output side of the input cache component and the input side of the control forwarding component; a unidirectional link between the other end of the output side of the input cache component and the input side of the computing engine component; a unidirectional link between the output side of the control forwarding component and one end of the input side of the output cache component; a unidirectional link between the output side of the computing engine component and the other end of the input side of the output cache component; a unidirectional link between the output side of the output cache component and the downstream unit; and a bidirectional link between the control forwarding component and the computing engine component.

[0027] The routing unit further includes a control and forwarding component, which is used to perform at least one of the following functions:

[0028] The downstream unit is determined based on the data-as-a-path control information; target data is determined based on the data-as-a-path control information, the target data including part or all of the data carried in the data packet sent by the upstream unit; a target calculation method is determined based on the data-as-a-path control information; and a data packet is composed based on the data-as-a-path control information and sent to the downstream unit.

[0029] The wafer-level processor chip proposed in this invention adds a control forwarding component to the routing unit in the processor unit. The control forwarding component realizes the routing and forwarding function of control data packets. This allows the wafer-level processor chip proposed in this invention to perform aggregation processing on network traffic in advance according to the aggregated communication characteristics. This can not only increase the bandwidth of effective traffic in the network, but also significantly reduce the probability of network congestion.

[0030] Figure 1 In this context, PE represents a processor unit, CE represents a computing unit within the processor unit, and Router represents a routing unit within the processor unit. Figure 1 As shown, this embodiment of the invention also provides a wafer-level processor chip, which includes a computing array composed of N rows and M columns of processor units. The processor units in the computing array are interconnected by a 2D Mesh network. Each processor unit includes a computing unit and any of the routing units described in this embodiment of the invention. The computing unit is used for data computation.

[0031] This invention also provides a routing method applicable to any of the routing units described in this invention, the routing method comprising:

[0032] The computing engine component of the routing unit performs on-network computation on the received data. The on-network computation includes at least one of the following computation methods: logical operation and numerical operation. The wafer-level processor chip on which the routing unit is located contains a computing array composed of N rows and M columns of processor units. The processor units in the computing array are interconnected by a 2D Mesh network. Each processor unit contains a computing unit and a routing unit.

[0033] By way of example and not limitation, logical and numerical operations may include operations such as summation, logical operations, multiplication and division, maximum value, minimum value, etc. Furthermore, the computing engine module unit supports fixed-point and floating-point operations of different data formats, which is not limited in this invention.

[0034] Alternatively, data formats such as INT8, INT16, FP8, FP16, BP16, and FP32 may be used.

[0035] The computing engine component of the routing unit performs on-network calculations on the received data, including:

[0036] The computing engine component of the routing unit performs on-network computation on the target data according to the target computation method. The target data includes some or all of the data carried in the data packet sent by the upstream unit. The data packet also carries data along-path control information. The target data and the target computation method are specified in the data along-path control information.

[0037] In-network computing is a network communication acceleration technology in the field of distributed parallel computing. By offloading the operations related to aggregate communication operators to the network, it enables the aggregation and processing of network data on the network, eliminates redundancy in communication data and shortens the transmission path, which can significantly reduce the latency of aggregate communication, reduce network congestion, and optimize and alleviate the communication bottleneck problem in AI training. At present, in-network computing technology is mainly used in network devices such as network cards and switches, and there is no such technology application in wafer-level processor chips.

[0038] This invention proposes to add a computing engine component and a control forwarding component to the routing unit of the processor unit in a wafer-level processor chip, and to establish a bidirectional link between the control forwarding component and the computing engine component. Furthermore, the control forwarding component controls the data operation process in the computing engine component according to the data along-path control information, which can offload the data operation of the aggregated communication operator to the routing unit, thereby introducing aggregated communication on-network computing capability into the wafer-level processor chip, and effectively reducing aggregated communication network congestion and latency in the network.

[0039] In this embodiment, the data-as-you-go control information includes the data packet format and the data processing method.

[0040] One implementation method is as follows: the data packet format and the target data have a fixed mapping relationship. The fixed format of the data packet means that the data in the data packet needs to perform on-net operations on the data in the data packet according to the established rules.

[0041] For example, data packet format A means that the data in the data packet needs to be processed in the network according to the established rules and the data operation method M.

[0042] Another implementation is that the data packet format indicates how the data in the data packet is split, and the data along-path control information also includes additional information to indicate which data needs to be processed on the network.

[0043] As an example and not a limitation, the routing method includes:

[0044] The input buffer component of the routing unit temporarily stores data packets sent by the upstream unit, the data packets carrying data and data-as-route control information; the output buffer component of the routing unit temporarily stores data packets sent to the downstream unit, the data packets sent to the downstream unit carrying data calculated on the network and the data-as-route control information, the downstream unit determining based on the data-as-route control information.

[0045] Optionally, the upstream unit can be either a computing unit within the same processor unit or another processor unit adjacent to that processor unit.

[0046] By way of example and not limitation, the routing method also includes at least one of the following steps:

[0047] The control and forwarding component of the routing unit determines the downstream unit based on the data-as-route control information;

[0048] The control and forwarding component of the routing unit determines the target data based on the data-as-you-go control information. The target data includes some or all of the data carried in the data packets sent by the upstream unit.

[0049] The control and forwarding component of the routing unit determines the target computation method based on the data-as-route control information.

[0050] The control and forwarding component of the routing unit assembles data packets and sends them to the downstream unit based on the data-as-you-go control information.

[0051] Optionally, the downstream unit can be either a computing unit within the same processor unit or another processor unit adjacent to that processor unit.

[0052] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

[0053] It should be noted that the sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

Claims

1. A routing unit for a wafer-level processor chip, characterized in that, The routing unit is used to receive, process, and forward data. The routing unit comprises a computing engine component, which is used to perform on-network computation on the data. The on-network computation includes at least one of the following computation methods: logical operation and numerical operation. The wafer-level processor chip on which the routing unit is located contains a computing array composed of N rows and M columns of processor units. The processor units in the computing array are interconnected by a 2D Mesh network. Each processor unit contains a computing unit and a routing unit.

2. The routing unit according to claim 1, characterized in that, The computing engine component is used to perform on-network computation on target data according to the target computation method. The target data includes part or all of the data carried in the data packet sent by the upstream unit. The data packet also carries data along-path control information. The target data and the target computation method are specified in the data along-path control information.

3. The routing unit according to claim 1, characterized in that, The routing unit further includes an input buffer component, an output buffer component, and a control forwarding component. The linking method of the routing unit includes: The input side of the input buffer component is unidirectionally linked to the upstream unit; one end of the output side of the input buffer component is unidirectionally linked to the input side of the control forwarding component; the other end of the output side of the input buffer component is unidirectionally linked to the input side of the computing engine component; the output side of the control forwarding component is unidirectionally linked to one end of the input side of the output buffer component; the output side of the computing engine component is unidirectionally linked to the other end of the input side of the output buffer component; the output side of the output buffer component is unidirectionally linked to the downstream unit; and a bidirectional link is established between the control forwarding component and the computing engine component.

4. The routing unit according to claim 1, characterized in that, The routing unit further includes an input buffer component and an output buffer component; The input buffer component is used to temporarily store data packets sent by the upstream unit, and the data packets carry data and data-as-a-path control information; The output buffer component is used to temporarily store data packets sent to downstream units. The data packets sent to downstream units carry data calculated on the network and data-as-a-path control information. The downstream units determine the data based on the data-as-a-path control information.

5. The routing unit according to claim 4, characterized in that, The routing unit further includes a control and forwarding component, which is used to perform at least one of the following functions: The downstream unit is determined based on the data-driven control information; The target data is determined based on the data-as-you-go control information, and the target data includes some or all of the data carried in the data packets sent by the upstream unit. The target calculation method is determined based on the aforementioned data-driven control information. The data packet is composed of the data along-path control information and sent to the downstream unit.

6. A wafer-level processor chip, characterized in that, The wafer-level processor chip includes a computing array consisting of N rows and M columns of processor units, and the processor units in the computing array are interconnected by a 2D Mesh network. Each processor unit includes a computing unit and a routing unit as described in any one of claims 1 to 5, wherein the computing unit is used for data computation.

7. A routing method applicable to the routing unit of a wafer-level processor chip, characterized in that, include: The computing engine component of the routing unit performs on-network computation on the received data. The on-network computation includes at least one of the following computation methods: logical operation and numerical operation. The wafer-level processor chip on which the routing unit is located contains a computing array composed of N rows and M columns of processor units. The processor units in the computing array are interconnected by a 2D Mesh network. Each processor unit contains a computing unit and a routing unit.

8. The method according to claim 7, characterized in that, The computing engine component of the routing unit performs on-network calculations on the received data, including: The computing engine component of the routing unit performs on-network computation on the target data according to the target computation method. The target data includes some or all of the data carried in the data packet sent by the upstream unit. The data packet also carries data along-path control information. The target data and the target computation method are specified in the data along-path control information.

9. The method according to claim 7, characterized in that, The method further includes: The input buffer component of the routing unit temporarily stores data packets sent by the upstream unit, and the data packets carry data and data-as-route control information; The output buffer component of the routing unit temporarily stores the data packets sent to the downstream unit. The data packets sent to the downstream unit carry the data calculated on the network and the data-as-route control information. The downstream unit determines the data based on the data-as-route control information.

10. The method according to claim 9, characterized in that, The method further includes at least one of the following steps: The control and forwarding component of the routing unit determines the downstream unit based on the data-as-route control information; The control and forwarding component of the routing unit determines the target data based on the data-as-you-go control information. The target data includes some or all of the data carried in the data packets sent by the upstream unit. The control and forwarding component of the routing unit determines the target computation method based on the data-as-route control information. The control and forwarding component of the routing unit assembles data packets and sends them to the downstream unit based on the data-as-you-go control information.