Network-on-chip design method and device, electronic equipment and storage medium
By performing system-level modeling before designing an on-chip network, decomposing the routing node processing flow to a streamlined level and building a timing model, the problems of high design risks and high cost in the existing technology are solved, and efficient and low-cost on-chip network design is achieved.
Patent Information
- Application Number
- CN202510887661.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-27
AI Technical Summary
In the prior art, on-chip network design relies on designer experience, resulting in high design risks and high costs, making it difficult to quickly verify whether it meets expectations.
By determining the performance requirements of the on-chip network, decomposing the routing node processing flow to the pipeline level, building a timing model and verifying it, ensuring that the on-chip network is designed after the verification is passed.
Improves the efficiency of on-chip network design, reduces design costs, reduces rework risks, and ensures that the design meets performance requirements.
Smart Images

Figure CN120373231A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a method and apparatus for on-chip network design, an electronic device, and a storage medium. Background Art
[0002] In the context of continuous breakthroughs in integrated circuit technology, the system-on-chip is evolving towards a highly heterogeneous integration mode, and more heterogeneous components with different functions need to be integrated on the system-on-chip. As the core communication architecture of the on-chip network, due to this trend of complexity, the communication scale inside the chip has increased exponentially, thus posing higher requirements for the design of the on-chip network.
[0003] In the prior art, the design of the on-chip network usually relies on the experience of designers. Designers establish the on-chip network architecture based on their understanding of the technology, and judge whether the on-chip network architecture meets the expectations. In the case of serious non-compliance with the expectations, the previous work needs to be overthrown, resulting in high risks and costs. Summary of the Invention
[0004] Embodiments of the present application provide a method and apparatus for on-chip network design, an electronic device, and a storage medium, which can reduce the risks and costs in the process of designing the on-chip network architecture and improve the efficiency of designing the on-chip network architecture.
[0005] In a first aspect, embodiments of the present application disclose a method for on-chip network design, the method including: Determine the performance requirements of the on-chip network to be designed; According to the performance requirements, determine the architecture parameters corresponding to the on-chip network; For each routing node of the on-chip network, decompose the processing flow of the routing node into at least one pipeline stage, and determine the number of clock cycles corresponding to each pipeline stage; According to the architecture parameters and the number of clock cycles corresponding to all routing nodes, construct a first timing model; Verify the first timing model, and in the case of successful verification, design the on-chip network according to the first timing model.
[0006] Optionally, the constructing a first timing model according to the architecture parameters and the number of clock cycles corresponding to all routing nodes includes: For each routing node, construct a timing model of a single routing node according to the number of clock cycles corresponding to each pipeline stage; According to the data transmission cycle corresponding to each link in the on-chip network, construct a link timing model; According to the topological structure of the on-chip network, the timing models corresponding to each routing node, and the link timing models corresponding to each link, construct a first timing model.
[0007] Optionally, the verification of the first timing model includes: Obtain the data traffic of the first timing model in at least one data mode; For each data mode, inject the data traffic into the first timing model for simulation to determine the current first network performance of the first timing model; the model parameters of the first timing model adopt a first configuration parameter group; According to the first network performance, adjust the first configuration parameter group and enter the next round of simulation until the first network performance meets the preset conditions, and then stop the simulation.
[0008] Optionally, the adjustment of the first configuration parameter group according to the first network performance includes: Determine the first traffic corresponding to each link in the first timing model; When the distribution of the first traffic on each link is uneven, adjust the routing paths for communication between each routing node by adjusting the hash seed in the routing allocation algorithm.
[0009] Optionally, the adjustment of the first configuration parameter group according to the first network performance includes: Determine the input traffic of each routing node in the first timing model; According to the input traffic and a preset threshold, determine at least one first routing node; For each first routing node, adjust the first configuration parameter corresponding to the first routing node.
[0010] Optionally, the determination of the architecture parameters corresponding to the network-on-chip according to the performance requirement includes: According to the performance requirement, determine the preset range corresponding to each architecture parameter; the architecture parameters include at least one of a topology structure, a flow control mechanism, and an internal structure of a routing node; According to the influence degree of the topology structure, the flow control mechanism, and the internal structure of the routing node on the performance requirement respectively, select a first topology structure, a first flow control mechanism, and a first internal structure of a routing node from the preset range.
[0011] In a second aspect, an embodiment of the present application discloses a network-on-chip design device, and the device includes: A requirement determination module, configured to determine the performance requirement of the network-on-chip to be designed; A parameter determination module, configured to determine the architecture parameters corresponding to the network-on-chip according to the performance requirement; A decomposition module, configured to decompose the processing flow of each routing node of the on-chip network into at least one pipeline stage, and determine the number of clock cycles corresponding to each pipeline stage; A construction module, configured to construct a first timing model according to the architecture parameters and the number of clock cycles corresponding to all routing nodes; A design module, configured to verify the first timing model, and design the on-chip network according to the first timing model when the verification is passed.
[0012] Optionally, the construction module includes: A first construction sub-module, configured to construct a timing model of a single routing node for each routing node according to the number of clock cycles corresponding to each pipeline stage; A second construction sub-module, configured to construct a link timing model according to the data transmission period corresponding to each link in the on-chip network; A third construction sub-module, configured to construct a first timing model according to the topology structure of the on-chip network, the timing models corresponding to each routing node, and the link timing models corresponding to each link.
[0013] In a third aspect, an embodiment of the present application discloses an electronic device, which includes a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store executable instructions, and the executable instructions cause the processor to execute the foregoing on-chip network design method.
[0014] In a fourth aspect, an embodiment of the present application discloses a readable storage medium. When the instructions in the readable storage medium are executed by a processor of an electronic device, the electronic device can execute the foregoing on-chip network design method.
[0015] The embodiments of the present application have the following advantages: A method for designing a network-on-chip provided by an embodiment of the present application can determine the performance requirements of the network-on-chip to be designed; determine the architecture parameters corresponding to the network-on-chip according to the performance requirements; for each routing node of the network-on-chip, decompose the processing flow of the routing node into at least one pipeline stage, and determine the number of clock cycles corresponding to each pipeline stage; construct a first timing model according to the architecture parameters and the number of clock cycles corresponding to all routing nodes; verify the first timing model, and when the verification is passed, design the network-on-chip according to the first timing model. By performing system-level modeling before actually establishing the network-on-chip architecture, simulating the system behavior of each component in the entire network-on-chip and the number of clock cycles consumed by the behavior, it is possible to quickly determine whether there is a deviation between the performance of the network-on-chip architecture and the expected target, and modify and verify the working parameters of the network-on-chip architecture to obtain a network-on-chip model that meets the performance requirements, improving the efficiency of designing the network-on-chip and reducing the cost of designing the network-on-chip. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0017] Figure 1 is a flowchart of the steps of an embodiment of a method for designing a network-on-chip of the present invention; Figure 2 is a structural block diagram of a network-on-chip design device of the present invention; Figure 3 is a structural block diagram of an electronic device provided by an example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] The terms "first", "second", etc. in the description and claims of the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the description and claims is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. In the embodiments of the present invention, the term "a plurality" refers to two or more, and other quantifiers are similar.
[0020] Method embodiment Refer to Figure 1 , a step flowchart of an embodiment of a network-on-chip design method of the present invention is shown. The method may specifically include the following steps: Step 101, determine the performance requirements of the network-on-chip to be designed; Step 102, determine the architecture parameters corresponding to the network-on-chip according to the performance requirements; Step 103, for each routing node of the network-on-chip, decompose the processing flow of the routing node into at least one pipeline stage, and determine the number of clock cycles corresponding to each pipeline stage; Step 104, construct a first timing model according to the architecture parameters and the number of clock cycles corresponding to all routing nodes; Step 105, verify the first timing model. If the verification is passed, design the network-on-chip according to the first timing model.
[0021] The on-chip network design method provided by the embodiments of this application performs system-level modeling on the on-chip network before Register Transfer Level (RTL) design, quickly validates the feasibility and performance of the on-chip network architecture through the system-level model, and reduces the risk of rework in the later stage. System-level modeling mainly describes the overall behavior of the system, such as data flow and performance metrics. RTL design mainly describes the data transmission and logical operations between registers, and the connections of clocks, registers, and combinational logic need to be clarified. System-level modeling is used to verify the feasibility of the on-chip network architecture model, such as whether the topology is reasonable and whether the routing algorithm is deadlock-free. System-level modeling uses high-level languages to describe transaction-level behaviors, such as C language, SystemC, and Python language. While RTL design uses hardware description languages to describe register-level behaviors, such as VHDL and Verilog. The role of system-level modeling is architecture verification, and the role of RTL design is hardware implementation. System-level modeling is a rehearsal for RTL-level design. It quickly validates the feasibility and performance of the designed on-chip network architecture through high-level languages and reduces the risk of rework in the later stage. RTL design is the implementation, which transforms the abstract system-level behavior into hardware description. Among them, after completing the RTL code, modifying the topology or routing algorithm of the architecture may require rewriting a large amount of code, with a high cost. While system-level modeling using high-level languages can quickly enumerate and verify multiple architecture schemes through parametric design. Therefore, performing system-level modeling before RTL-level design and validating through the model obtained from system-level modeling can quickly verify the feasibility and performance of the designed on-chip network architecture, improve the efficiency of designing the on-chip network, and reduce the cost required in the process of designing the on-chip network.
[0022] Among them, performance requirements design the fundamental goals and constraints of the on-chip network. Performance requirements define the quantitative metrics regarding data transmission capabilities that the on-chip network needs to meet and are proposed by upper-layer applications. The upper-layer applications can be multi-core processors, AI accelerators, System on a Chip (SoC). Specifically, the performance requirements can include at least one of the following: throughput, latency, and power consumption. Different application scenarios have different performance requirements for the on-chip network. For example, in high-performance computing scenarios such as scientific computing and tensor computing communication of AI chips, the performance requirements for the on-chip network are high throughput and low latency; supporting parallel data-intensive tasks. In real-time system scenarios such as for autonomous driving chips, the performance requirements for the on-chip network are to limit the upper limit of latency and high reliability to ensure the determinacy of event response.
[0023] Architectural parameters are the parameters corresponding to the topology and functional components of the network-on-chip, which are used to construct the physical and logical frameworks of the network-on-chip. Selecting different parameter combinations will directly affect the performance, area, and power consumption of the network-on-chip. Architectural parameters can include at least one of the following parameters: topology, routing algorithm, flow control mechanism, router microarchitecture, link characteristics, and network interface. Among them, the topology defines the interconnection method of IP modules on the chip, and the IP modules can be computing cores, memories, input / output systems (I / O), etc.; the flow control mechanism refers to the strategy for buffer and link resource allocation between routing nodes to prevent congestion and deadlocks; the router microarchitecture refers to the functional unit design and workflow inside a single routing node.
[0024] Before designing the network-on-chip, the target application scenario can be deeply analyzed, the bandwidth required for each critical path or the entire network can be quantified, and the latency requirements for critical communications can be defined. Set the power consumption budget and energy efficiency goals, determine the scalability requirements, and evaluate the reliability / fault tolerance requirements. The performance requirements can be recorded in a network-on-chip performance requirements specification.
[0025] Decomposing the processing flow of routing nodes into at least one pipelining stage is to decompose the tasks required for a router to process a data flit into logical steps. Map each logical step to a specific pipelining stage. A logical step may take one or more clock cycles, or multiple logical steps may be combined into one cycle. Perform logic synthesis and timing analysis on each pipelining stage to determine the number of target clock cycles required for each pipelining stage to complete its function. In most cases, the goal is that the cycles per instruction of the pipeline stage (CPI) is 1 (i.e., CPI = 1). However, in some low-power or high-frequency designs, complex stages with long critical paths may require multiple clock cycles (CPI > 1). Integrate the architectural parameters and the detailed pipelining timing information of each routing node to construct a first timing model that can predict the timing behavior of network communications. The first timing model can simulate the behavior of each component in the network-on-chip in each clock cycle. Establishing the first timing model can include the following aspects: establishing the topology of the entire network-on-chip, modeling the latency of the network-on-chip, such as calculating the single-hop latency and the total latency of multi-hop paths, and modeling the latency of the network-on-chip. For example, the throughput of the entire network-on-chip is equal to the product of the sum of the link bandwidths and the bandwidth utilization rate.
[0026] Verify the first timing model, mainly verifying the routing algorithm, timing, power consumption, and throughput. Verify whether the routing algorithm is correct and whether the timing meets the performance requirements. For example, whether the delay between ends is less than or equal to the expected target value, and evaluate the dynamic or static power consumption on the network-on-chip through a power analysis tool. Specifically, drive the first timing model with the target workload and measure the performance metrics of the key first timing model, such as end-to-end delay, bandwidth utilization, critical path delay, buffer occupancy, power consumption estimation, etc. Strictly compare the measurement results with the performance requirement specifications. If all key requirements are met in the simulation, it is considered that the verification is passed.
[0027] In the embodiment of the present application, the performance requirements of the network-on-chip to be designed can be determined; according to the performance requirements, the architecture parameters corresponding to the network-on-chip are determined; for each routing node of the network-on-chip, the processing flow of the routing node is decomposed into at least one pipeline stage, and the number of clock cycles corresponding to each pipeline stage is determined; according to the architecture parameters and the number of clock cycles corresponding to all routing nodes, a first timing model is constructed; the first timing model is verified, and in the case of successful verification, the network-on-chip is designed according to the first timing model. By performing system-level modeling before actually establishing the network-on-chip architecture, simulating the system behavior of each component in the entire network-on-chip and the number of clock cycles consumed by the behavior, it is possible to quickly determine whether there is a deviation between the performance of the network-on-chip architecture and the expected target, and modify and verify the working parameters of the network-on-chip architecture to obtain a network-on-chip model that meets the performance requirements, improving the efficiency of designing the network-on-chip and reducing the cost of designing the network-on-chip.
[0028] Optionally, the constructing the first timing model according to the architecture parameters and the number of clock cycles corresponding to all routing nodes includes: Step S11: For each routing node, construct a timing model of a single routing node according to the number of clock cycles corresponding to each pipeline stage; Step S12: Construct a link timing model according to the data transmission cycle corresponding to each link in the network-on-chip; Step S13: Construct a first timing model according to the topology structure of the network-on-chip, the timing models corresponding to each routing node, and the link timing models corresponding to each link.
[0029] Specifically, the tasks required for a router to process a flit are decomposed into logical steps. Typical steps include: Buffer Write (BW), Route Calculation (RC), Virtual Channel Allocation (VA), Switch Allocation (SA), and Switch Traversal (ST). Determine the number of clock cycles consumed by route calculation, switch arbitration, link transmission, buffer read / write respectively.
[0030] The timing model of a single node needs to clarify the total delay of data passing through this node. According to the number of pipeline stages and CPI per stage, calculate the minimum or typical hop latency (HopLatency) required for a flit to cross a router (from input to output), and establish a routing node delay model. For example, for a 4-stage pipeline with CPI = 1 per stage, the minimum crossing delay is 4 clock cycles (excluding link delay). According to the link length and physical characteristics (or a simple estimated value of link delay per hop, such as 1 or 2 cycles), calculate the time (in clock cycles) required for a flit to be transmitted on the link, and establish a link timing model. It is also possible to consider the time that a flit waits in the input buffer due to flow control (such as credit mechanism, VC allocation competition) and network congestion, and establish a queuing delay model. And combine the source / destination network interface delay, the number of router hops on the path, the router delay per hop, the link delay per hop, and the queuing delay on the path to estimate the total delay of a data packet from the source IP to the destination IP, and establish an end-to-end delay model. Based on the link bandwidth, according to the router crossbar bandwidth and flow control mechanism, estimate the maximum saturated throughput of the network and the available bandwidth of a specific path, and establish a throughput / bandwidth model.
[0031] In the embodiment of the present application, for each routing node, according to the number of clock cycles corresponding to each pipeline stage, construct a timing model of a single routing node; according to the data transmission cycle corresponding to each link in the network-on-chip, construct a link timing model; according to the topology of the network-on-chip, the timing models corresponding to each routing node, and the link timing models corresponding to each link, construct a first timing model. By establishing the timing models corresponding to the routing nodes and links respectively, and finally establishing the first timing model comprehensively according to the topology, a first timing model closer to the real application scenario can be established.
[0032] Optionally, the verification of the first timing model includes: Step S21, obtain the data traffic of the first timing model under at least one data mode; Step S22: For each data pattern, inject the data traffic into the first timing model for simulation to determine the current first network performance of the first timing model; the model parameters of the first timing model adopt the first configuration parameter group; Step S23: Adjust the first configuration parameter group according to the first network performance, and enter the next round of simulation until the simulation is stopped when the first network performance meets the preset conditions.
[0033] Among them, the data pattern refers to the typical data transmission patterns that the chip will face in actual application scenarios. For example, matrix operations, memory read and write, and video processing. The first timing model is designed according to the project to be processed. The project to be processed defines the application scenario and core functions of the target chip. According to the project requirements, the typical data transmission characteristics of the actual chip in different working scenarios are obtained, and these characteristics are abstracted into data patterns. For example, if the project to be processed is to design an autonomous driving chip, the data patterns may include: sensor fusion mode, neural network inference mode, and emergency control mode. The data traffic refers to the transmission characteristics of data packets under a specific data pattern. For example, the traffic distribution between nodes, the size of data packets, bursty traffic or periodic traffic. The data traffic can be obtained by statistical analysis based on actual application scenarios or theoretical modeling. For example, a traffic generation tool generates a data packet sequence. After obtaining the data traffic, the traffic is injected into the input port of the first timing model through a network interface to simulate the communication behavior of a real computing unit.
[0034] After injecting the data traffic into the first timing model, the traffic simulation is run through a simulation tool to evaluate the network performance of the on-chip network under a specific configuration parameter group. The first network performance refers to a key metric for measuring the overall efficiency of the first timing model, rather than the local performance in the first timing model. The first network performance may include at least one of the following: latency, throughput, degree of long-tail effect, energy consumption, and degree of congestion. It should be noted that during the construction of the first timing model, some parameters are set to default according to the requirements of the project to be processed. For example, according to the computing power requirements of the project to be processed, the number of CPU cores and the capacity of the cache are set to fixed values. The first configuration parameter group refers to the set of parameters in the first timing model that can be adjusted, mainly including the parameters that can affect the communication ability of the first timing model. The first configuration parameter group may include: network scale, link width, buffer depth, number of virtual channels, priority scheduling rules, etc.
[0035] It should be noted that adjusting the configuration parameter group according to network performance is a multi-objective optimization problem. It is necessary to combine the correlation between performance indicators and parameters, determine the configuration parameters with a high degree of correlation with the performance indicators, and adjust the configuration parameters. For example, there is a high degree of correlation between latency and buffer size, and a high degree of correlation between throughput and link bandwidth; after modifying the parameters, re-simulate to verify whether the performance has improved and avoid local optimality. In the case of conflicts among multiple network performance indicators, such as reducing latency resulting in increased energy consumption, make a trade-off according to the expected goals and core requirements of the project to be processed. Moreover, when adjusting each parameter in the configuration parameter group, it can be adjusted within the preset parameter range. For example, if the preset number of virtual channels is between 2 and 14, then when adjusting the number of virtual channels, only select values between 2 and 14, and the preset number of virtual channels can be set according to experience.
[0036] Inject the data traffic of each data mode into the first timing model in sequence, and perform independent simulation for each data mode. For each data mode, adjust the configuration parameter group of the first timing model at least once, and verify whether the network performance of the model reaches the expected goal corresponding to this data mode after each adjustment. When the network performance meets the preset conditions, stop the simulation to obtain the second configuration parameter group corresponding to this data mode, that is, the optimal configuration parameter group under limited conditions. Among them, each data mode has its own corresponding preset conditions, and the preset conditions refer to the requirements for the network performance of the model under this data mode, such as throughput being greater than or equal to the preset throughput threshold, and latency being less than or equal to the preset time threshold. It should be noted that the expected goals defined by the project to be processed include the preset conditions for each data mode.
[0037] After obtaining the optimal configuration parameter group of the first timing model for each data mode, weigh the priorities of each data mode in the project to be processed, determine the final configuration parameter group according to the priorities, and design the chip according to the final configuration parameter group. Among them, the configuration parameter group under a certain data mode can be determined as the configuration parameter group for reference when finally designing the chip, or the configuration parameter groups under multiple data modes can be comprehensively considered to obtain a configuration parameter group with a compromise configuration, and this configured configuration parameter group can be determined as the configuration parameter group for reference when finally designing the chip.
[0038] In an embodiment of the present application, an initial first model is constructed according to performance requirements, and data traffic corresponding to multiple data patterns is obtained. The data traffic corresponding to each data pattern is sequentially input into the first time-series model, and each data pattern is independently simulated. During the simulation process, the network performance of the first time-series model is evaluated, and the configuration parameter group in the first time-series model is adjusted until the network performance meets the preset conditions corresponding to the data pattern, and the optimal parameter combination under this data pattern is obtained. Compared with engineers determining the design parameters of the on-chip network based on experience, the limitations of traditional design methods are solved through data-driven and system-level verification, and it is possible to handle the on-chip network design for rich and complex scenarios, and design an on-chip network that better meets the requirements of real application scenarios.
[0039] Optionally, the adjusting the first configuration parameter group according to the first network performance includes: Step S31: Determine the first traffic corresponding to each link in the first time-series model; Step S32: In the case where the distribution of the first traffic on each link is uneven, adjust the routing path for communication between each routing node by adjusting the hash seed in the routing allocation algorithm.
[0040] Among them, a link refers to a communication channel connecting two adjacent routing nodes, which is responsible for transmitting data between nodes. The routing allocation algorithm determines the logical rule of the path that the data packet passes from the source routing node to the destination routing node. For example, in multi-path routing based on hashing, the data packet is mapped to multiple paths through a hash function. The routing path is the sequence of links that the data packet passes from the source routing node to the destination routing node, and the link is the basic unit of the path.
[0041] In the first time-series model, a traffic monitoring module is deployed for each link to record in real time the amount of data passing through each link per unit time, and the traffic on the link can also be visualized through a heat map or a line graph. For example, draw a traffic change curve according to the order of the routing path. If the curve shows "monotonically decreasing", it means that there is a gradient decrease in the traffic of each link on the routing path. The essence of the gradient decrease in the traffic on each link is the uneven path selection of the routing allocation algorithm, resulting in overload of some links and underutilization of subsequent links. In the embodiment of the present application, the path distribution is adjusted by changing the hash seed to balance the link load.
[0042] A hash seed is an initial parameter of a hash function. Different seeds will change the arrangement of hash results. By adjusting the hash seed, the hash results can be evenly distributed across multiple paths, thereby enabling traffic to be dispersed to different links. In the first timing model, multiple hash seed values are traversed, the hash results are recalculated for each hash seed, and the traffic distribution of each link is statistically analyzed. The smaller the standard deviation of the traffic on each link, the smaller the gradient coefficient. The gradient coefficient refers to the ratio of the link with the highest traffic to the link with the lowest traffic, and the more evenly distributed the traffic is. Multiple hash seed values are traversed, and the seed that minimizes the traffic standard deviation and has a gradient coefficient closest to 1 is selected.
[0043] In the embodiments of the present application, a traffic monitoring module is deployed for each link to record in real time the amount of data passing through each link per unit time, and determine the first traffic corresponding to each link in the first timing model; in the case where the distribution of the first traffic on each link is uneven, by adjusting the hash seed in the routing allocation algorithm, the routing paths for communication between each routing node are adjusted, and the path distribution is changed by adjusting the hash seed to balance the link load.
[0044] Optionally, the adjusting the first configuration parameter group according to the first network performance includes: Step S41: Determine the input traffic of each routing node in the first timing model; Step S42: Determine at least one first routing node according to the input traffic and a preset threshold; Step S43: For each first routing node, adjust the first configuration parameter corresponding to the first routing node.
[0045] It should be noted that the first timing model in the embodiments of the present application uses a lossless on-chip network. A lossless on-chip network is an interconnection architecture designed specifically for internal chip communication, and its core goal is to completely avoid data packet loss during data transmission. In the steady-state transmission without data loss, that is, when the buffer is not saturated and there is no backpressure signal, the long-term average input traffic of a routing node is equal to the output traffic. When the buffer of the output port or downstream node of a certain routing node is full, the node will send a "pause" signal to the upstream node to prevent data from continuing to flow in and avoid packet loss caused by buffer overflow. The input traffic refers to the amount of data received and forwarded by each routing node, such as the number of data packets and the number of bytes.
[0046] Exemplarily, the traffic load of the first routing node is significantly higher than that of the routing nodes at the average level of the first timing model. For example, the preset threshold is 2 times the average traffic. By real-time monitoring the input traffic of each routing node in the first timing model, when the input traffic of a routing node is greater than or equal to the preset threshold, this routing node is determined as the first routing node. Among them, the preset threshold can be static or dynamic. For example, the preset threshold can be set according to experience and remains unchanged when adjusting the first configuration parameter; it can also be adjusted in real time according to the input traffic of each routing node in the first timing model. Usually, the first routing node is located in the middle area of the entire first timing model, that is, on the middle common link.
[0047] After injecting the data traffic corresponding to the data pattern into the first timing model, the performance metrics of each routing node can also be monitored and analyzed. The performance metrics can include: buffer occupancy rate and power consumption. The buffer occupancy rate refers to the proportion of the used space in the input buffer of the routing node. For example, the buffer size is 512 bytes, and the current occupancy is 400 bytes, so the buffer occupancy rate is 78%; the power consumption refers to the dynamic power consumption generated by the data processing of the routing node, such as routing calculation and buffer read and write. The preset threshold can be set according to the performance target. For example, the preset threshold can be a power consumption threshold of 20 mW, and the routing node with a power consumption greater than 20 mW is determined as the first routing node.
[0048] Adjusting the first configuration parameter corresponding to the first routing node means adjusting the network physical structure parameter corresponding to the first routing node, controlling the traffic passing through the first routing node, adjusting the transmission path of the data packet, reducing the traffic load of the first routing node, and adjusting the network resources allocated to the first routing node. Exemplarily, when the queuing time of the data packet in the buffer of the first routing node is too long, that is, when the delay exceeds the standard, the input buffer of the first routing node can be enlarged; or the routing calculation can be split into multiple-level pipelines to shorten the single-level delay. For example, the address resolution and virtual channel selection in the routing calculation process are each divided into one-level pipelines. When the power consumption of the first routing node exceeds the standard due to frequent data processing, the complex adaptive routing algorithm can be replaced with a simpler fixed routing to reduce the calculation power consumption.
[0049] In an embodiment of the present application, during the simulation process, the input traffic of each routing node in the first timing model is monitored in real time, and a routing node with an input traffic greater than or equal to a preset threshold is determined as the first routing node; wherein, the preset threshold can be static and fixed, or can be dynamically adjusted according to the real-time traffic in the on-chip network model. For each first routing node, the network resources allocated to the first routing node are adjusted, for example, the network physical structure and the transmission path of the data packet. High-fidelity traffic monitoring and threshold strategy can accurately perceive the traffic of each routing node, and adjust the parameters considering both the immediate effect and long-term stability to avoid waste of global resources caused by local hotspots.
[0050] Optionally, determining the architecture parameters corresponding to the on-chip network according to the performance requirements includes: Step S51: Determine the preset range corresponding to each architecture parameter according to the performance requirements; the architecture parameters include at least one of a topology structure, a flow control mechanism, and an internal structure of a routing node; Step S52: Select a first topology structure, a first flow control mechanism, and a first internal structure of a routing node from the preset range according to the influence degrees of the topology structure, the flow control mechanism, and the internal structure of the routing node on the performance requirements.
[0051] It should be noted that the preset range corresponding to each structure parameter can be determined according to performance requirements, historical experience, and relevant theories. For example, the mapping relationship between architecture parameters and performance requirements can be determined through literature. According to experience or literature, it is known that virtual channels can solve the deadlock problem of wormhole routing but increase the area; the backpressure mechanism can avoid congestion by stopping data transmission but may increase the delay. Based on the mapping relationship, architecture parameters that obviously do not meet the requirements are excluded. For example, if the performance requirement is "high throughput (1TB / s)", then the ring topology structure with fewer links (fewer links) is excluded from the topology structure, and the Mesh topology structure with more links is retained.
[0052] Moreover, the topology structure can affect the communication mode and scale, and thus the topology structure can be selected according to the communication mode, scale, delay / bandwidth requirements, and area constraints. For example, uniform communication is suitable for Mesh, and those with strong locality may use Cluster or Hierarchical structures, and those with high bandwidth requirements may use Fat Tree. The routing algorithm can also be selected according to the topology, deadlock avoidance requirements, and load balancing requirements. Deterministic routing (such as XY) is commonly used for simple topologies, and adaptive routing is used for complex or load balancing requirements. The flow control mechanism can affect the delay and throughput, and thus the flow control mechanism can be selected according to the delay, throughput, buffer cost, and deadlock avoidance requirements. For example, wormhole flow control with multiple virtual channels is usually adopted for low delay.
[0053] In the embodiments of the present application, the preset range corresponding to each structural parameter can be determined according to performance requirements, historical experience, and relevant theories. For example, the mapping relationship between architecture parameters and performance requirements can be determined through literature or simulation tools. By determining the topology structure, flow control mechanism, and internal structure of the routing node within the preset range, the number of parameters that need to be traversed when designing the on-chip network is reduced, and the efficiency of designing the on-chip network is improved.
[0054] In summary, a method for designing an on-chip network provided by the embodiments of the present application can determine the performance requirements of the to-be-designed on-chip network; determine the architecture parameters corresponding to the on-chip network according to the performance requirements; for each routing node of the on-chip network, decompose the processing flow of the routing node into at least one pipeline stage, and determine the number of clock cycles corresponding to each pipeline stage; for each routing node, construct a timing model of a single routing node according to the number of clock cycles corresponding to each pipeline stage; construct a link timing model according to the data transmission period corresponding to each link in the on-chip network; construct a first timing model according to the topology structure of the on-chip network, the timing models corresponding to each routing node, and the link timing models corresponding to each link. Verify the first timing model. During the verification process, the configuration parameters adopted by the first timing model can be modified multiple times to determine the network performance of the on-chip network when each group of configuration parameters is adopted, and then determine the target configuration parameter combination that can achieve the expected goal. Design the on-chip network according to the first timing model adopting the target configuration parameter combination, and the on-chip network architecture that meets the requirements can be accurately and quickly determined. By performing system-level modeling before actually establishing the on-chip network architecture, simulating the system behavior of each component in the entire on-chip network and the number of clock cycles consumed by the behavior, it can be quickly judged whether there is a deviation between the performance of the on-chip network architecture and the expected goal, and the working parameters of the on-chip network architecture can be modified and verified to obtain an on-chip network model that meets the performance requirements, improving the efficiency of designing the on-chip network and reducing the cost of designing the on-chip network.
[0055] Device embodiments Refer to Figure 2 , which shows a structural block diagram of an on-chip network design device of the present application. The device may specifically include: A requirement determination module 210, configured to determine the performance requirements of the to-be-designed on-chip network; A parameter determination module 220, configured to determine the architecture parameters corresponding to the on-chip network according to the performance requirements; A decomposition module 230, configured to, for each routing node of the on-chip network, decompose the processing flow of the routing node into at least one pipeline stage, and determine the number of clock cycles corresponding to each pipeline stage; A construction module 240, configured to construct a first timing model according to the architecture parameters and the number of clock cycles corresponding to all routing nodes; A design module 250, configured to verify the first timing model, and design an on-chip network according to the first timing model when the verification is passed.
[0056] Optionally, the construction module includes: A first construction sub-module, configured to construct a timing model of a single routing node for each routing node according to the number of clock cycles corresponding to each pipeline stage. A second construction sub-module, configured to construct a link timing model according to the data transmission period corresponding to each link in the on-chip network. A third construction sub-module, configured to construct a first timing model according to the topological structure of the on-chip network, the timing models corresponding to each routing node, and the link timing models corresponding to each link.
[0057] Optionally, the design module includes: An acquisition module, configured to acquire the data traffic of the first timing model in at least one data mode. A simulation module, configured to inject the data traffic into the first timing model for simulation for each data mode, and determine the current first network performance of the first timing model; the model parameters of the first timing model adopt a first configuration parameter group. An adjustment module, configured to adjust the first configuration parameter group according to the first network performance, and enter the next round of simulation until the first network performance meets a preset condition, and then stop the simulation.
[0058] Optionally, the adjustment module includes: A link traffic determination module, configured to determine the first traffic corresponding to each link in the first timing model. A first adjustment sub-module, configured to adjust the routing path for communication between each routing node by adjusting the hash seed in the routing allocation algorithm when the distribution of the first traffic on each link is uneven.
[0059] Optionally, the adjustment module includes: A node traffic determination module, configured to determine the input traffic of each routing node in the first timing model. A node determination module, configured to determine at least one first routing node according to the input traffic and a preset threshold. A second adjustment sub-module, configured to adjust the first configuration parameter corresponding to each first routing node.
[0060] Optionally, the parameter determination module includes: A range determination module for determining a preset range corresponding to each structural parameter according to the performance requirements; the architecture parameters include at least one of a topology structure, a flow control mechanism, and an internal structure of a routing node; A selection module for selecting a first topology structure, a first flow control mechanism, and a first internal structure of a routing node from the preset ranges according to the influence degrees of the topology structure, the flow control mechanism, and the internal structure of the routing node on the performance requirements respectively.
[0061] In summary, an on-chip network design device provided by an embodiment of the present application can determine the performance requirements of an on-chip network to be designed; determine the architecture parameters corresponding to the on-chip network according to the performance requirements; for each routing node of the on-chip network, decompose the processing flow of the routing node into at least one pipelining stage, and determine the number of clock cycles corresponding to each pipelining stage; for each routing node, construct a timing model of a single routing node according to the number of clock cycles corresponding to each pipelining stage; construct a link timing model according to the data transmission period corresponding to each link in the on-chip network; construct a first timing model according to the topology structure of the on-chip network, the timing models corresponding to each routing node respectively, and the link timing models corresponding to each link respectively. Verify the first timing model. During the verification process, the configuration parameters adopted by the first timing model can be modified multiple times to determine the network performance of the on-chip network when each group of configuration parameters is adopted respectively, and then determine the target configuration parameter combination that can achieve the expected goal. Design the on-chip network according to the first timing model adopting the target configuration parameter combination, and can accurately and quickly determine the on-chip network architecture that meets the requirements. By performing system-level modeling before actually establishing the on-chip network architecture, simulating the system behaviors of each component in the entire on-chip network and the number of clock cycles consumed by the behaviors, it can be quickly judged whether there is a deviation between the performance of the on-chip network architecture and the expected goal, and the working parameters of the on-chip network architecture can be modified and verified to obtain an on-chip network model that meets the performance requirements, improving the efficiency of designing the on-chip network and reducing the cost of designing the on-chip network.
[0062] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the related parts, refer to the partial description of the method embodiment.
[0063] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0064] Regarding the processor in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0065] Refer toFigure 3 , which is a structural block diagram of an electronic device for on-chip network design provided by an embodiment of the present invention. As Figure 3 shown, the electronic device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store executable instructions, and the executable instructions cause the processor to execute the on-chip network design method of the foregoing embodiment.
[0066] The processor may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other editable devices, transistor logic devices, hardware components, or any combination thereof. The processor may also be a combination that implements a computing function, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0067] The communication bus may include a path for transmitting information between the memory and the communication interface. The communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0068] The memory may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0069] An embodiment of the present invention further provides a non - temporary computer - readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device (server or terminal), the processor is enabled to execute Figure 1 the on - chip network design method shown.
[0070] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0071] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk memories, CD - ROMs, optical memories, etc.) containing computer - usable program code.
[0072] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the method, terminal device (system), and computer program product according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general - purpose computer, a special - purpose computer, an embedded processor, or other programmable data - processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data - processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0073] These computer program instructions can also be stored in a computer - readable memory that can guide a computer or other programmable data - processing terminal devices to work in a predictive manner, so that the instructions stored in the computer - readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data - processing terminal devices, so that a series of operation steps are executed on the computer or other programmable terminal devices to generate a computer - implemented process. Thus, the instructions executed on the computer or other programmable terminal devices provide for implementing the functions specified in Figure 1 one process or multiple processes and / or blocksFigure 1 Steps of the functions specified in one or more boxes.
[0075] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0076] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.
[0077] The above has introduced in detail a method, device, electronic device and readable storage medium for on-chip network design provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for on-chip network design, characterized in that, The method includes: Determining the performance requirements of the Network-on-Chip (NoC) to be designed; Determining the architecture parameters corresponding to the NoC according to the performance requirements; For each routing node of the NoC, decomposing the processing flow of the routing node into at least one pipeline stage and determining the number of clock cycles corresponding to each pipeline stage; Constructing a first timing model according to the architecture parameters and the number of clock cycles corresponding to all routing nodes; Verifying the first timing model, and when the verification is passed, designing the NoC according to the first timing model.
2. The method according to claim 1, wherein The constructing a first timing model according to the architecture parameters and the number of clock cycles corresponding to all routing nodes includes: For each routing node, constructing a timing model of a single routing node according to the number of clock cycles corresponding to each pipeline stage; Constructing a link timing model according to the data transmission period corresponding to each link in the NoC; Constructing a first timing model according to the topology structure of the NoC, the timing models corresponding to each routing node, and the link timing models corresponding to each link.
3. The method according to claim 1, wherein The verifying the first timing model includes: Obtaining the data traffic of the first timing model in at least one data mode; For each data mode, injecting the data traffic into the first timing model for simulation to determine the current first network performance of the first timing model; the model parameters of the first timing model adopt a first configuration parameter group; Adjusting the first configuration parameter group according to the first network performance and entering the next round of simulation until the simulation is stopped when the first network performance meets the preset conditions.
4. The method according to claim 3, wherein The adjusting the first configuration parameter group according to the first network performance includes: Determining the first traffic corresponding to each link in the first timing model; When the distribution of the first traffic on each link is uneven, adjusting the routing paths for communication between each routing node by adjusting the hash seed in the routing allocation algorithm.
5. The method according to claim 3, characterized in that, The adjusting the first configuration parameter group according to the first network performance includes: Determining the input traffic of each routing node of the first timing model; Determining at least one first routing node according to the input traffic and a preset threshold; For each first routing node, adjusting the first configuration parameter corresponding to the first routing node.
6. The method according to claim 1, characterized in that, The determining the architecture parameters corresponding to the NoC according to the performance requirements includes: Determining the preset range corresponding to each architecture parameter according to the performance requirements; the architecture parameters include at least one of topology structure, flow control mechanism, and internal structure of the routing node; Selecting a first topology structure, a first flow control mechanism, and a first internal structure of the routing node from the preset range according to the influence degree of the topology structure, the flow control mechanism, and the internal structure of the routing node on the performance requirements.
7. A Network-on-Chip design device, characterized in that, The device includes: A requirement determination module for determining the performance requirements of the Network-on-Chip (NoC) to be designed; A parameter determination module for determining the architecture parameters corresponding to the NoC according to the performance requirements; A decomposition module, which is used to decompose the processing flow of each routing node of the network-on-chip into at least one pipeline stage for each routing node of the network-on-chip, and determine the number of clock cycles corresponding to each pipeline stage; A construction module, which is used to construct a first timing model according to the architecture parameters and the number of clock cycles corresponding to all routing nodes; A design module, which is used to verify the first timing model, and design the network-on-chip according to the first timing model when the verification is passed.
8. The device according to claim 7, wherein The construction module includes: A first construction sub-module, which is used to construct a timing model of a single routing node for each routing node according to the number of clock cycles corresponding to each pipeline stage; A second construction sub-module, which is used to construct a link timing model according to the data transmission period corresponding to each link in the network-on-chip; A third construction sub-module, which is used to construct a first timing model according to the topology of the network-on-chip, the timing models corresponding to each routing node, and the link timing models corresponding to each link.
9. An electronic device, characterized in that, The electronic device includes a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store executable instructions, and the executable instructions cause the processor to execute the network-on-chip design method according to any one of claims 1 to 6.
10. A readable storage medium, characterized in that, When the instructions in the readable storage medium are executed by the processor of the electronic device, the processor can execute the network-on-chip design method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Layering and reconfigurable on-chip network modeling and simulation system
CN103970939A
Network-on-chip setting method and network-on-chip setting structure
CN114385547A
Customized network-on-chip topological structure generation method and system
CN118153241A
Network-on-chip design method and device, electronic equipment and readable storage medium
CN119378461A
Network-on-chip simulation method and device, electronic equipment and storage medium
CN119397720A