Network-on-chip divergent multicast routing system and implementation method and device

By optimizing the multicast routing transmission architecture and data packet structure of the on-chip network, a multicast routing system with one-time transmission and multiple-point reception is realized, which solves the problems of high transmission latency and easy congestion, and improves the transmission efficiency and system performance of the on-chip network.

CN120880965APending Publication Date: 2025-10-31SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510981353.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing multicast transmission technologies suffer from high transmission latency and congestion in on-chip networks, especially performing poorly in applications with high data volume and high real-time requirements, resulting in low transmission efficiency.

Method used

Design an on-chip network-based distributed multicast routing system. By optimizing the multicast routing transmission architecture, data packet structure, and routing calculation unit logic, and using five independent routing calculation units (East, South, West, North, and Local) and a multicast data packet structure, data transmission can be achieved by sending once and receiving at multiple points.

Benefits of technology

It significantly improves the transmission efficiency of on-chip networks, reduces latency and network congestion, enhances system reliability and availability, and adapts to on-chip network systems of different sizes and needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120880965A_ABST
    Figure CN120880965A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of integrated circuit system-on-chip communication, and relates to a network-on-chip divergent multicast routing system and an implementation method and device. The system comprises a multicast routing transmission architecture, a multicast data packet structure and a routing calculation unit, wherein the multicast routing transmission architecture comprises five independent routing calculation units which respectively correspond to five directions of east, south, west, north and local, and each routing calculation unit is connected with the input module and the output module; the multicast data packet is composed of a head flit, a data flit and a tail flit; and the routing calculation unit dynamically generates a multicast routing direction combination based on a comparison result of the target node coordinate and the current node coordinate. According to the invention, by designing a divergent multicast routing transmission architecture, one-time transmission and multi-point receiving of data are realized, and the problem of repeated transmission in a traditional unicast mode is effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of integrated circuit system-on-chip communication technology, and to an on-chip network divergent multicast routing system and its implementation method and apparatus. Background Technology

[0002] In the current field of integrated circuit design, with the rapid development of high-tech applications such as artificial intelligence and big data processing, the amount of data processed by systems is showing an exponential growth trend. This change places higher demands on the communication technology in integrated circuit systems on a chip (SoC), especially in network-on-chip (NoC) environments, where efficient and reliable data transmission between cores has become a key factor restricting the overall performance of the system.

[0003] Traditional unicast transmission mode is increasingly revealing its limitations in high-data-volume scenarios. Unicast transmits data point-to-point, which not only easily leads to link congestion during data transmission but also causes significant transmission delays, thus affecting the overall system's response speed and throughput. Especially in scenarios where the same data needs to be sent to multiple destinations simultaneously, unicast requires repeated transmissions, greatly wasting bandwidth resources and reducing transmission efficiency.

[0004] To address this challenge, multicast transmission mode emerged. Multicast technology, through a mechanism of sending once and receiving at multiple points, effectively alleviates the load pressure on NoC communication links and improves data transmission efficiency and reliability. However, existing multicast transmission technologies still have many shortcomings in their application in NoCs.

[0005] On the one hand, some NoC systems simulate multicast effects by using multiple unicasts. Although this method is simple, it consumes a lot of valuable bandwidth resources and causes serious redundancy problems due to repeated data transmission, which further exacerbates network congestion.

[0006] On the other hand, existing multicast transmission technologies generally face problems such as high transmission latency and congestion when handling large-scale data transmission. These problems not only reduce the transmission performance of NoCs but also limit their widespread deployment in applications with high real-time and high bandwidth requirements. Therefore, designing an efficient and reliable multicast routing implementation method has become the key to improving NoC performance.

[0007] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0008] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0009] This disclosure provides an on-chip network-based divergent multicast routing system, implementation method, and apparatus, aiming to solve the problems of high transmission latency and congestion in existing multicast technologies. By optimizing the multicast routing transmission architecture, data packet composition, and implementation logic of the routing calculation unit, the transmission efficiency and overall performance of the on-chip network are significantly improved, meeting the growing demand for high-performance computing.

[0010] In some embodiments, the system includes:

[0011] Multicast routing transmission architecture, multicast data packet structure, and routing calculation unit;

[0012] The multicast routing transmission architecture includes five independent routing calculation units, corresponding to the five directions of east, south, west, north, and local, respectively. Each routing calculation unit is connected to an input module and an output module.

[0013] The multicast data packet consists of a header fragment, a data fragment, and a tail fragment.

[0014] The routing calculation unit dynamically generates multicast route direction combinations based on the comparison results between the destination node coordinates and the current node coordinates.

[0015] Preferably, the header micro-piece includes the following fields: multicast flag, identifying unicast or multicast transmission; destination node row mask, indicating the row number where the destination node exists; coordinates of all target nodes, the set of coordinates of the multicast destination node; packet size, data packet length information.

[0016] Preferably, the bit width of the destination node row mask is equal to the number of on-chip network rows, with each bit corresponding to one row; when a bit is 1, it indicates that there is at least one destination node coordinate in that row.

[0017] Preferably, the routing calculation unit performs the following steps:

[0018] Determine whether the input micro-piece is a header micro-piece;

[0019] If it is a header fragment and the multicast flag is 1, then the multicast route direction flag is generated based on the round-robin comparison of the destination node coordinates and the current node coordinates; if the multicast flag is 0, then the unicast route direction is calculated using the dimension-order XY routing algorithm.

[0020] The multicast output direction is determined based on the combination of direction markers.

[0021] Preferably, the direction flag consists of 5 bits, representing the five directions of east, south, west, north, and local, respectively, with a bit value of 1 indicating that data needs to be transmitted in that direction.

[0022] Preferably, the coordinate comparison logic of the local routing calculation unit includes:

[0023] Compare the y-coordinates of the destination node in the same row; if the y-coordinate value of the destination node is greater than the current y-coordinate value, set the east direction flag; if the y-coordinate value of the destination node is less than the current y-coordinate value, set the west direction flag.

[0024] Compare the x-coordinates of the destination nodes in the same column; if the x-coordinate of the destination node is greater than the current x-coordinate, set the south direction flag; if the x-coordinate of the destination node is less than the current x-coordinate, set the north direction flag.

[0025] Preferably, the coordinate comparison logic of the northbound routing calculation unit includes:

[0026] Only select target nodes whose x-coordinate is less than or equal to the current x-coordinate for round-robin comparison;

[0027] Compare the y-coordinates of nodes in the same row; if the y-coordinate value of the destination node is less than the current y-coordinate value, set the west direction flag; if the y-coordinate value of the destination node is greater than the current y-coordinate value, set the east direction flag.

[0028] Compare the x-coordinates in the same column; if the x-coordinate value of the destination node is equal to the x-coordinate value of the current node, set the local direction flag; if the x-coordinate value of the destination node is greater than the x-coordinate value of the current node, set the south direction flag.

[0029] The coordinate comparison logic of the southbound routing calculation unit includes:

[0030] Only select target nodes whose x-coordinate is less than or equal to the current x-coordinate for round-robin comparison;

[0031] Compare the y-coordinates of nodes in the same row; if the y-coordinate value of the destination node is less than the current y-coordinate value, set the west direction flag; if the y-coordinate value of the destination node is greater than the current y-coordinate value, set the east direction flag.

[0032] Compare the x-coordinates in the same column; if the x-coordinate value of the destination node is less than the x-coordinate value of the current node, set the north direction flag; if the x-coordinate value of the destination node is equal to the x-coordinate value of the current node, set the local direction flag.

[0033] Preferably, the coordinate comparison logic of the eastward routing calculation unit includes:

[0034] Only select target nodes in the current row whose y-coordinate is less than or equal to the current y-value for round-robin comparison;

[0035] Compare the y-coordinates of the same row; if y is less than the current y value, set the west direction flag; if y is equal to the current y value, set the local direction flag.

[0036] The coordinate comparison logic of the westward route calculation unit includes:

[0037] Only select target nodes in the current row whose y-coordinate is greater than or equal to the current y-value for round-robin comparison;

[0038] Compare the y-coordinates of the same line; if y is greater than the current y value, set the east direction flag; if y is equal to the current y value, set the local direction flag.

[0039] In some embodiments, the on-chip network divergent multicast routing implementation method includes the following steps:

[0040] Construct a multicast data packet by splitting the data packet into a header fragment, a data fragment, and a tail fragment; wherein, the header fragment includes: a 1-bit multicast flag, a destination node row mask, a set of coordinates of all destination nodes, and data packet size information;

[0041] The multicast routing direction is dynamically calculated. When a data packet is input to the routing calculation unit, it is identified whether the current micro-fragment is a header micro-fragment. If it is a header micro-fragment and the multicast flag is 1, the multicast routing direction flag is generated based on the round-robin comparison between the destination node coordinates and the current node coordinates. If the multicast flag is 0, the dimensional XY routing algorithm is used to calculate the unicast routing direction.

[0042] Divergent multicast transmission copies data packets simultaneously to all defined output directions, enabling transmission with a single send and multiple receivers.

[0043] In some embodiments, the apparatus includes a processor and a memory storing program instructions, wherein the processor is configured to execute the on-chip network divergent multicast routing implementation method when the program instructions are executed.

[0044] The on-chip network divergent multicast routing system, implementation method, and apparatus provided in this disclosure can achieve the following technical effects:

[0045] This invention, through the design of a divergent multicast routing transmission architecture, achieves one-time data transmission and multi-point reception, effectively avoiding the problem of repeated transmission in traditional unicast mode. Simultaneously, by optimizing the routing calculation unit implementation logic in five directions (north, south, east, west, and local), it reduces data waiting time and processing latency during transmission, thereby significantly improving the transmission efficiency of the on-chip network. For example, in scenarios requiring simultaneous data transmission to multiple destinations, this method can route data to all target nodes at once, avoiding the time overhead of multiple unicast transmissions.

[0046] Existing multicast technologies often suffer from high transmission latency and congestion when handling large-scale data transmission, leading to a decline in NoC transmission performance. This invention effectively distributes data flows through precise routing calculations and direction flag settings, avoiding excessive concentration of data on specific paths and thus reducing network congestion. This helps maintain the stable operation of the on-chip network, improving system reliability and availability.

[0047] This invention innovates in packet structure by introducing mechanisms such as multicast flags and destination node row masks, reducing the logical design complexity in multicast routing implementation. For example, the use of destination node row masks allows for quick location of the row containing the destination coordinates instead of comparing all destination coordinates one by one when calculating the route direction, thus simplifying the routing calculation process and saving hardware resources.

[0048] The multicast routing implementation method of this invention has good scalability and flexibility. As the scale of the on-chip network increases or application requirements change, it can adapt to new scenarios by simply adjusting the logic of the routing calculation unit or adding new directional routing calculation units. This flexibility enables this method to be widely applied to on-chip network systems of different sizes and requirements.

[0049] In summary, the divergent multicast routing implementation method of this invention can significantly improve the transmission efficiency of on-chip networks, reduce latency, reduce congestion, save resources, and enhance the scalability and flexibility of the system. These improvements work together to effectively enhance the overall performance of the on-chip network system, providing more efficient and reliable communication support for high-performance computing applications such as artificial intelligence and big data processing.

[0050] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description

[0051] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein:

[0052] Figure 1 This is a diagram of a multicast routing transmission architecture.

[0053] Figure 2 This is a diagram showing the structure of a data packet.

[0054] Figure 3 Flowchart for the Local routing calculation unit.

[0055] Figure 4 Flowchart for the North routing calculation unit implementation.

[0056] Figure 5 Flowchart for the South routing calculation unit implementation.

[0057] Figure 6 Flowchart for the East routing calculation unit implementation.

[0058] Figure 7 Flowchart for the West routing calculation unit implementation.

[0059] Figure 8 This is a structural diagram of a directional sign.

[0060] Figure 9 A schematic diagram of the device structure is provided for embodiments of the invention. Detailed Implementation

[0061] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0062] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0063] Unless otherwise stated, the term "multiple" means two or more.

[0064] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0065] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0066] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.

[0067] like Figures 1-8As shown, an on-chip network-based distributed multicast routing system includes: a multicast routing transmission architecture, a multicast data packet structure, and a routing calculation unit;

[0068] The multicast routing transmission architecture includes five independent routing calculation units, corresponding to the five directions of east, south, west, north, and local, respectively. Each routing calculation unit is connected to an input module and an output module.

[0069] The multicast data packet consists of a header fragment (Head_Flit), a data fragment (Body_Flit), and a tail fragment (Tail_Flit);

[0070] The routing calculation unit dynamically generates multicast route direction combinations based on the comparison results between the destination node coordinates and the current node coordinates.

[0071] As a refinement of the above embodiments, the appendix Figure 1 The multicast routing transmission architecture is described in the figure, with a routing calculation unit configured for each input module. During multicast routing transmission, data from the input module is transmitted to the corresponding output module according to the routing direction obtained by the routing calculation unit. The multicast routing direction is a free combination of five directions: East, South, West, North, and Local, such as Southeast, Northeast, East Local, Southeast West, etc. When input data from the Local_input module has a routing direction of Southeast, Northwest, or Eastwest according to the routing calculation module, the input data can be transmitted to the East_output, West_output, South_output, and North_output modules simultaneously. When input data from the North_input module has a routing direction of Southeast West Local, the input data can be transmitted to the East_output, West_output, South_output, and Local_output modules simultaneously.

[0072] As a refinement of the above embodiments, the appendix Figure 2The structure of the data packet is described in detail. As shown in the figure, Head_Flit contains parameters such as multicast flag, destination node row mask, coordinates of all destination nodes, and packet size. The multicast flag is used to initiate multicast transmission, and the destination node coordinates are used to calculate the direction of the route. The multicast flag uses 1 bit to represent the data; when the multicast flag is 1, multicast transmission is performed; when the multicast flag is 0, unicast transmission is performed. The destination node row mask is used to indicate whether a certain row has destination coordinates. For example, in a 4x4 mesh NoC structure, a destination node row mask of 1100 indicates that only the third and fourth rows have destination coordinates, and the other rows do not. This avoids retrieving data from the second and third rows, thus reducing the complexity of the logical design. Body_Flit and Tail_Flit are only used for data transmission.

[0073] As a refinement of the above embodiments, the appendix Figure 3 This demonstrates the implementation logic of the Local routing calculation unit. Upon inputting "Flit", the system first checks "Head_Flit". If it is "Head_Flit", the multicast flag is retrieved; otherwise, routing calculation is not performed. When the multicast flag is 1, multicast routing direction calculation begins. If the multicast flag is 0, the commonly used dimension-order XY routing algorithm is used to calculate the unicast routing direction. During multicast routing calculation, all destination coordinates are extracted based on the destination node's row mask and compared with the current node's coordinates in a round-robin fashion. First, the y-values ​​of the row coordinates are compared. If the y-value is greater than the current node's y-value, the East direction flag is set to 1; otherwise, it is set to 0. If the y-value is less than the current node's y-value, the West direction flag is set to 1; otherwise, it is set to 0. Then, the x-values ​​of the column coordinates are compared. If the x-value is greater than the current node's x-value, the South direction flag is set to 1; otherwise, it is set to 0. If the x-value is less than the current node's x-value, the North direction flag is set to 1; otherwise, it is set to 0. Finally, the routing direction is obtained based on these direction flags. The direction flags are attached. Figure 8 As shown, when the value is 01011, the routing direction is east, west, and north.

[0074] As a refinement of the above embodiments, the appendix Figure 4The implementation logic of the North routing calculation unit is demonstrated, and its implementation process is largely the same as that of the Local routing calculation unit. Specifically, during multicast route calculation, only all destination coordinates with x-values ​​less than or equal to the current coordinate are retrieved and compared with the coordinates of the current node in a round-robin fashion. For example, if the current node coordinates are (1,2), then only the coordinates from the destination node coordinate lists in rows 1 and 2 of Head_Flit need to be retrieved. If the destination node row mask is 0001, only row 1 needs to be retrieved. Furthermore, when comparing x-values ​​in the same column, comparisons are not performed for values ​​less than the current node's x-value; instead, comparisons are performed for values ​​equal to the current node's x-value. If the x-value is equal to the current node's coordinate, the Local direction flag is set to 1; otherwise, it is set to 0.

[0075] As a refinement of the above embodiments, the appendix Figure 5 The implementation logic of the South routing calculation unit is shown. The difference is that it retrieves all destination coordinates whose x-values ​​are greater than or equal to the current coordinate and compares them with the coordinates of the current node. In addition, when comparing the x-values ​​of coordinates in the same column, it does not compare those greater than the current node's x-value, but rather compares those equal to the current node's x-value.

[0076] As a refinement of the above embodiments, the appendix Figure 6 The implementation logic of the East route calculation unit is shown. It only needs to extract all destination coordinates in the current row whose y-value is less than or equal to the current coordinate and poll them. Then, it only compares the y-values ​​of the coordinates in the same row. If the y-value is less than the current node's coordinate y-value, the West direction flag is set to 1, otherwise it is set to 0. If the y-value is equal to the current node's coordinate y-value, the Local direction flag is set to 1, otherwise it is set to 0.

[0077] As a refinement of the above embodiments, the appendix Figure 7 The implementation logic of the West route calculation unit is shown. It only needs to retrieve all destination coordinates in the current row whose y-value is greater than or equal to the current coordinate and poll them. Then, it only compares the y-values ​​of the coordinates in the same row. When the y-value is greater than the current node's coordinate y-value, the East direction flag is set to 1, otherwise it is set to 0. When the y-value is equal to the current node's coordinate y-value, the Local direction flag is set to 1, otherwise it is set to 0.

[0078] In summary, this invention improves the transmission efficiency and overall performance of the on-chip network by designing a multicast routing transmission architecture, defining the composition of data packets, implementing the logic of five-directional routing calculation units, obtaining the transmission routing direction of multicast routes, and realizing multicast transmission.

[0079] Example 2

[0080] A method for implementing on-chip network distributed multicast routing includes the following steps:

[0081] S1: Construct a multicast data packet by splitting the data packet into a header fragment, a data fragment, and a tail fragment; wherein, the header fragment includes: a 1-bit multicast flag, a destination node row mask, a set of coordinates of all target nodes, and data packet size information;

[0082] S2: Dynamically calculate the multicast routing direction. When a data packet is input to the routing calculation unit, it identifies whether the current micro-fragment is a header micro-fragment. If it is a header micro-fragment and the multicast flag is 1, a multicast routing direction flag is generated based on the polling comparison between the destination node coordinates and the current node coordinates. If the multicast flag is 0, the dimensional XY routing algorithm is used to calculate the unicast routing direction.

[0083] S3: Divergent multicast transmission, which simultaneously copies data packets to all defined output directions, enabling transmission with a single send and multiple receivers.

[0084] Combination Figure 9 As shown, this disclosure provides an on-chip network divergent multicast routing implementation device 300, including a processor 304 and a memory 301. Optionally, the device may further include a communication interface 302 and a bus 303. The processor 304, communication interface 302, and memory 301 can communicate with each other via the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call logical instructions in the memory 301 to execute the on-chip network divergent multicast routing implementation method of the above embodiment.

[0085] Furthermore, the logic instructions in the aforementioned memory 301 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0086] The memory 301, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 304 executes functional applications and data processing by running the program instructions / modules stored in the memory 301, thereby implementing the on-chip network divergent multicast routing implementation method in the above embodiments.

[0087] The memory 301 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 301 may include high-speed random access memory and may also include non-volatile memory.

[0088] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to execute the above-described on-chip network divergent multicast routing implementation method.

[0089] The aforementioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

[0090] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code; it can also be a transient storage medium.

[0091] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0092] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0093] The methods and products (including but not limited to devices and equipment) disclosed in the embodiments herein can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection shown or discussed between each other may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0094] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. An on-chip network-based distributed multicast routing system, characterized in that, include: Multicast routing transmission architecture, multicast data packet structure, and routing calculation unit; The multicast routing transmission architecture includes five independent routing calculation units, corresponding to the five directions of east, south, west, north, and local, respectively. Each routing calculation unit is connected to an input module and an output module. The multicast data packet consists of a header fragment, a data fragment, and a tail fragment. The routing calculation unit dynamically generates multicast route direction combinations based on the comparison results between the destination node coordinates and the current node coordinates.

2. The on-chip network divergent multicast routing system according to claim 1, characterized in that, The header micro-chip contains the following fields: multicast flag, indicating unicast or multicast transmission; destination node row mask, indicating the row number where the destination node exists; coordinates of all target nodes, the set of coordinates of the multicast destination node; packet size, data packet length information.

3. The on-chip network divergent multicast routing system according to claim 2, characterized in that, The bit width of the destination node row mask is equal to the number of on-chip network rows, with each bit corresponding to one row; when a bit is 1, it indicates that there is at least one destination node coordinate in that row.

4. The on-chip network divergent multicast routing system according to claim 1, characterized in that, The routing calculation unit performs the following steps: Determine whether the input micro-piece is a header micro-piece; If it is a header fragment and the multicast flag is 1, then a multicast routing direction flag is generated based on a round-robin comparison of the destination node coordinates and the current node coordinates. If the multicast flag is 0, the unicast route direction is calculated using the dimension-order XY routing algorithm; The multicast output direction is determined based on the combination of direction markers.

5. The on-chip network divergent multicast routing system according to claim 4, characterized in that, The direction flag consists of 5 bits, representing the five directions: east, south, west, north, and local. A bit value of 1 indicates that data needs to be transmitted in that direction.

6. The on-chip network divergent multicast routing system according to claim 4, characterized in that, The coordinate comparison logic of the local routing calculation unit includes: Compare the y-coordinates of the destination node in the same row; if the y-coordinate value of the destination node is greater than the current y-coordinate value, set the east direction flag; if the y-coordinate value of the destination node is less than the current y-coordinate value, set the west direction flag. Compare the x-coordinates of the destination nodes in the same column; if the x-coordinate of the destination node is greater than the current x-coordinate, set the south direction flag; if the x-coordinate of the destination node is less than the current x-coordinate, set the north direction flag.

7. The on-chip network divergent multicast routing system according to claim 4, characterized in that, The coordinate comparison logic of the northbound routing calculation unit includes: Only select target nodes whose x-coordinate is less than or equal to the current x-coordinate for round-robin comparison; Compare the y-coordinates of nodes in the same row; if the y-coordinate value of the destination node is less than the current y-coordinate value, set the west direction flag; if the y-coordinate value of the destination node is greater than the current y-coordinate value, set the east direction flag. Compare the x-coordinates in the same column; if the x-coordinate value of the destination node is equal to the x-coordinate value of the current node, set the local direction flag; if the x-coordinate value of the destination node is greater than the x-coordinate value of the current node, set the south direction flag. The coordinate comparison logic of the southbound routing calculation unit includes: Only select target nodes whose x-coordinate is less than or equal to the current x-coordinate for round-robin comparison; Compare the y-coordinates of nodes in the same row; if the y-coordinate value of the destination node is less than the current y-coordinate value, set the west direction flag; if the y-coordinate value of the destination node is greater than the current y-coordinate value, set the east direction flag. Compare the x-coordinates in the same column; if the x-coordinate value of the destination node is less than the x-coordinate value of the current node, set the north direction flag; if the x-coordinate value of the destination node is equal to the x-coordinate value of the current node, set the local direction flag.

8. The on-chip network divergent multicast routing system according to claim 4, characterized in that, The coordinate comparison logic of the eastward route calculation unit includes: Only select target nodes in the current row whose y-coordinate is less than or equal to the current y-value for round-robin comparison; Compare the y-coordinates of the same row; if y is less than the current y value, set the west direction flag; if y is equal to the current y value, set the local direction flag. The coordinate comparison logic of the westward route calculation unit includes: Only select target nodes in the current row whose y-coordinate is greater than or equal to the current y-value for round-robin comparison; Compare the y-coordinates of the same line; if y is greater than the current y value, set the east direction flag; if y is equal to the current y value, set the local direction flag.

9. A method for implementing divergent multicast routing in an on-chip network, characterized in that, Includes the following steps: Construct a multicast data packet by splitting the data packet into a header fragment, a data fragment, and a tail fragment; wherein, the header fragment includes: a 1-bit multicast flag, a destination node row mask, a set of coordinates of all destination nodes, and data packet size information; The multicast routing direction is dynamically calculated. When a data packet is input to the routing calculation unit, it is identified whether the current micro-piece is a header micro-piece. If it is a header micro-piece and the multicast flag is 1, the multicast routing direction is calculated using the routing calculation unit of the on-chip network divergent multicast routing system as described in any of claims 1-8. If the multicast flag is 0, the unicast routing direction is calculated using the dimension-order XY routing algorithm. Divergent multicast transmission copies data packets simultaneously to all defined output directions, enabling transmission with a single send and multiple receivers.

10. An on-chip network divergent multicast routing implementation device, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the on-chip network divergent multicast routing implementation method as described in claim 8 when running the program instructions.