Chip design method and device, electronic equipment and readable storage medium

By performing data-driven simulation and parameter adjustment of the on-chip network model, the optimal configuration parameter group is determined, which solves the problem of relying on experience in chip design and designs a chip that meets the needs of real applications.

CN120373230APending Publication Date: 2025-07-25BEIJING INSTITUTE OF OPEN SOURCE CHIP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510887653.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, artificial intelligence chip design relies on designer experience and cannot reasonably verify whether the design parameters meet actual needs during the early design process.

Method used

By obtaining the on-chip network model of the pending project and its data traffic in the data mode, performing simulation, adjusting the configuration parameter group until the preset conditions are met, and the optimal configuration parameter group is determined for designing the target chip.

Benefits of technology

Based on data-driven and system-level verification, a chip that is more in line with the needs of real application scenarios is designed, which solves the limitations of traditional design methods and adapts to rich and complex application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373230A_ABST
    Figure CN120373230A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a chip design method and device, electronic equipment and a readable storage medium. The method comprises the steps that an on-chip network model corresponding to a to-be-processed item and data traffic of the on-chip network model in at least one data mode are acquired; for each data mode, injecting the data traffic into the network-on-chip model for simulation, and determining the current first network performance of the network-on-chip model; model parameters of the network-on-chip model adopt a first configuration parameter group; adjusting the first configuration parameter group according to the first network performance, entering a next round of simulation, stopping simulation until the first network performance meets a preset condition, and determining the current first configuration parameter group as a second configuration parameter group corresponding to the data mode; and designing the target chip according to the second configuration parameter group corresponding to the at least one data mode. Different data modes are simulated, the configuration parameter set of the model is adjusted, and a chip better meeting application requirements is designed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a chip design method, apparatus, electronic device, and readable storage medium. Background Art

[0002] In recent years, deep learning algorithms such as recurrent neural networks and convolutional neural networks have shown increasing influence in various different fields and are applied in all aspects of social life. With the continuous development of deep learning algorithms, especially the proposal and application of large language models, a large number of tasks of large language models are implemented by artificial intelligence chips, and the current demand for artificial intelligence chips in the cloud and edge is very strong. Therefore, the demand for hardware has gradually increased, posing new challenges to chip design.

[0003] Currently, in the design process of artificial intelligence chips, the design parameters of the chips are completely determined by the experience of designers, and whether the design parameters are reasonable and meet the actual requirements cannot give reasonable evidence in the early design process. Summary of the Invention

[0004] Embodiments of this application provide a chip design method, apparatus, electronic device, and readable storage medium, which can design chips that better meet the actual application requirements.

[0005] In a first aspect, embodiments of this application disclose a chip design method, the method including: Obtain an on-chip network model corresponding to a project to be processed, and the data traffic of the on-chip network model in at least one data mode; For each data mode, inject the data traffic into the on-chip network model for simulation, and determine the current first network performance of the on-chip network model; the model parameters of the on-chip network model adopt a first configuration parameter group; According to the first network performance, adjust the first configuration parameter group and enter the next round of simulation. Stop the simulation until the first network performance meets a preset condition, and determine the current first configuration parameter group as the second configuration parameter group corresponding to the data mode; Design a target chip according to the second configuration parameter groups corresponding to the at least one data mode.

[0006] Optionally, the adjusting the first configuration parameter group according to the first network performance includes: Determine the input traffic of each routing node of the on-chip network model; Determine at least one first routing node according to the input traffic and a preset threshold; For each first routing node, adjust the first configuration parameter corresponding to the first routing node.

[0007] Optionally, for each first routing node, adjusting the first configuration parameter corresponding to the first routing node includes: Determining a second routing node from the remaining first routing nodes according to the connection relationship between the first routing node and the remaining first routing nodes; data transmission between the second routing node and the first routing node does not need to be forwarded by the remaining routing nodes; Determining the link between the first routing node and the second routing node as the target link; Increasing the number of virtual channels corresponding to the target link.

[0008] Optionally, for each first routing node, adjusting the first configuration parameter corresponding to the first routing node includes at least one of the following: Increasing the storage capacity of the cache queue in the first routing node; Adjusting the transmission priority of data packets in the first routing node; Adjusting the topological structure of the area corresponding to the first routing node.

[0009] Optionally, determining the current first network performance of the on-chip network model includes: For each routing node, determining the amount of data transmitted by the routing node; Determining the port rate of the routing node according to the amount of data and the simulation time; Determining the first network performance according to the port rates corresponding to each routing node.

[0010] Optionally, designing a target chip according to the second configuration parameter group corresponding to the at least one data pattern includes: Determining a target data pattern according to the usage ratio of each data pattern; Determining a target configuration parameter group according to the second configuration parameter group corresponding to the target data pattern; Designing a target chip according to the target configuration parameter group.

[0011] In a second aspect, an embodiment of the present application discloses a chip design device, and the device includes: An acquisition module, configured to acquire an on-chip network model corresponding to a to-be-processed project and the data traffic of the on-chip network model in at least one data pattern; A simulation module, configured to, for each data pattern, inject the data traffic into the on-chip network model for simulation to determine the current first network performance of the on-chip network model; the model parameters of the on-chip network model adopt a first configuration parameter group; An adjustment module, configured to adjust the first configuration parameter group according to the first network performance, and enter the next round of simulation. When the first network performance meets a preset condition, the simulation is stopped, and the current first configuration parameter group is determined as the second configuration parameter group corresponding to the data mode; A design module, configured to design a target chip according to the second configuration parameter group corresponding to the at least one data mode.

[0012] Optionally, the adjustment module includes: A traffic determination module, configured to determine the input traffic of each routing node of the on-chip network model; A node determination module, configured to determine at least one first routing node according to the input traffic and a preset threshold; An adjustment sub-module, configured to adjust the first configuration parameter corresponding to each first routing node for each first routing node.

[0013] In a third aspect, an embodiment of the present application discloses an electronic device, which includes a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store executable instructions, and the executable instructions cause the processor to execute the foregoing chip design method.

[0014] In a fourth aspect, an embodiment of the present application discloses a readable storage medium. When the instructions in the readable storage medium are executed by a processor of an electronic device, the electronic device can execute the foregoing chip design method.

[0015] The embodiments of the present application have the following advantages: Obtain the on-chip network model corresponding to the item to be processed and the data traffic of the on-chip network model in at least one data mode; for each data mode, inject the data traffic into the on-chip network model for simulation to determine the current first network performance of the on-chip network model; according to the first network performance, adjust the first configuration parameter group and enter the next round of simulation until the first network performance meets the preset conditions, then stop the simulation and determine the current first configuration parameter group as the second configuration parameter group corresponding to the data mode; design the target chip according to the second configuration parameter groups corresponding to at least one data mode. By constructing the on-chip network model and performing independent simulations for multiple data modes respectively, the optimal configuration parameter groups of the on-chip network model in each data mode are determined under limited conditions. Further, according to multiple optimal configuration parameter groups, the final configuration parameter group is determined. The chip is designed according to the final configuration parameter group. Compared with engineers determining the chip design parameters based on experience, the limitations of traditional design methods are solved through data-driven and system-level verification, enabling the chip design to handle rich and complex scenarios and design a chip that better meets the requirements of real application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 is a flowchart of the steps of an embodiment of a chip design method of the present invention; Figure 2 is a structural block diagram of a chip design device of the present invention; Figure 3 is a structural block diagram of an electronic device provided by an example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0019] The terms "first", "second", etc. in the description and claims of the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the description and claims is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. In the embodiments of the present invention, the term "a plurality of" refers to two or more, and other quantifiers are similar.

[0020] Method embodiment Refer to Figure 1 , a step flowchart of an embodiment of a chip design method of the present invention is shown. The method may specifically include the following steps: Step 101, obtain an on-chip network model corresponding to the item to be processed, and the data traffic of the on-chip network model in at least one data mode; Step 102, for each data mode, inject the data traffic into the on-chip network model for simulation, and determine the current first network performance of the on-chip network model; the model parameters of the on-chip network model adopt the first configuration parameter group; Step 103, adjust the first configuration parameter group according to the first network performance, and enter the next round of simulation until the first network performance meets the preset conditions, then stop the simulation, and determine the current first configuration parameter group as the second configuration parameter group corresponding to the data mode; Step 104, design a target chip according to the second configuration parameter groups corresponding to the at least one data mode.

[0021] The chip design method provided by the embodiments of the present application can be used to design complex chip systems that require efficient and flexible on-chip communication. For example, artificial intelligence or machine learning chips, autonomous driving and edge computing chips, 5G or 6G communication baseband chips, data center and cloud computing chips.

[0022] Among them, the item to be processed refers to the current chip project that needs to be designed or optimized. The item to be processed can specify the specific application scenario and design objectives of the target chip. For example, the application scenario is the design of a deep learning accelerator chip or the design of a 5G communication chip; the design objective refers to the expected objective for performance, such as low power consumption and high bandwidth. The on-chip network model refers to a model that simulates the communication architecture between multiple computing units inside the chip, and is composed of components such as routers, communication links, and network interfaces; the computing units include CPU cores, GPU cores, caches, peripherals, etc. The data pattern refers to the typical data transfer pattern that the chip will face in the actual application scenario. For example, matrix operations, memory read and write, and video processing. The item to be processed defines the application scenario and core functions of the target chip. According to the project requirements, the typical data transfer characteristics of the actual chip in different working scenarios are obtained, and these characteristics are abstracted into data patterns. For example, if the item to be processed is to design an autonomous driving chip, the data patterns can include: sensor fusion mode, neural network inference mode, and emergency control mode. The data traffic refers to the transmission characteristics of data packets under a specific data pattern. For example, the traffic distribution between nodes, the size of data packets, bursty traffic or periodic traffic. The data traffic can be obtained by statistical analysis based on the actual application scenario or generated by theoretical modeling. For example, a traffic generation tool generates a data packet sequence. After obtaining the data traffic, the traffic is injected into the input port of the on-chip network model through the network interface to simulate the communication behavior of real computing units.

[0023] After injecting the data traffic into the on-chip network model, traffic simulation is run through a simulation tool to evaluate the network performance of the on-chip network under a specific set of configuration parameters. The first network performance refers to a key metric for measuring the overall efficiency of the on-chip network model, rather than the local performance in the on-chip network model. The first network performance can include at least one of the following: latency, throughput, degree of long-tail effect, energy consumption, and degree of congestion. It should be noted that during the process of constructing the on-chip network model, according to the requirements of the item to be processed, some parameters are set as default values. For example, according to the computing power requirements of the item to be processed, the number of CPU cores and the capacity of the cache are set as fixed values. The first set of configuration parameters refers to the set of parameters that can be adjusted in the on-chip network model, mainly including the parameters that can affect the communication ability of the on-chip network model. The first set of configuration parameters can include: network scale, link width, buffer depth, number of virtual channels, priority scheduling rules, etc.

[0024] It should be noted that adjusting the configuration parameter group according to network performance is a multi-objective optimization problem. It is necessary to combine the correlation between performance indicators and parameters, determine the configuration parameters with high correlation with performance indicators, and adjust the configuration parameters. For example, there is a high degree of correlation between latency and buffer size, and a high degree of correlation between throughput and link bandwidth. After modifying the parameters, re-simulate to verify whether the performance has improved and avoid local optima. If there are conflicts among multiple network performance indicators, for example, reducing latency leads to an increase in energy consumption, make a trade-off according to the expected goals and core requirements of the project to be processed. Moreover, when adjusting each parameter in the configuration parameter group, the adjustment can be made within a preset parameter range. For example, if the preset number of virtual channels is between 2 and 14, then when adjusting the number of virtual channels, only select values between 2 and 14, and the preset number of virtual channels can be set according to experience.

[0025] Inject the data traffic of each data mode into the on-chip network model in sequence, and perform independent simulations for each data mode. For each data mode, adjust the configuration parameter group of the on-chip network model at least once, and verify whether the network performance of the model reaches the expected goal corresponding to this data mode after each adjustment. When the network performance meets the preset conditions, stop the simulation to obtain the second configuration parameter group corresponding to this data mode, that is, the optimal configuration parameter group under limited conditions. Among them, each data mode has its own corresponding preset conditions, and the preset conditions refer to the requirements for the network performance of the model under this data mode, such as throughput being greater than or equal to the preset throughput threshold, and latency being less than or equal to the preset time threshold. It should be noted that the expected goals defined by the project to be processed include the preset conditions for each data mode.

[0026] After obtaining the optimal configuration parameter group of the on-chip network model for each data mode, weigh the priorities of each data mode in the project to be processed, determine the final configuration parameter group according to the priorities, and design the chip according to the final configuration parameter group. Among them, the configuration parameter group under a certain data mode can be determined as the configuration parameter group for reference when finally designing the chip, or the configuration parameter groups under multiple data modes can be comprehensively considered to obtain a configuration parameter group with a compromise configuration, and this configured configuration parameter group can be determined as the configuration parameter group for reference when finally designing the chip.

[0027] In the embodiments of the present application, an initial on-chip network model is constructed according to the computing power requirements of the item to be processed, and the data traffic corresponding to multiple possible data patterns in the item to be processed is obtained. The data traffic corresponding to each data pattern is sequentially input into the on-chip network model, and independent simulation is performed on each data pattern. During the simulation process, the network performance of the on-chip network model is evaluated, and the configuration parameter group in the on-chip network model is adjusted until the network performance meets the preset conditions corresponding to the data pattern, and the optimal parameter combination under this data pattern is obtained. After obtaining the optimal parameter combinations corresponding to each data pattern respectively, the priorities corresponding to each data pattern in the item to be processed are weighed. Furthermore, according to the priorities corresponding to each data pattern, a chip is designed based on multiple groups of second configuration parameter groups. By constructing an on-chip network model and performing independent simulation on multiple data patterns respectively, under limited conditions, the optimal configuration parameter group of the on-chip network model under each data pattern is determined. Further, according to multiple optimal configuration parameter groups, the final configuration parameter group is determined. According to the final configuration parameter group, a chip is designed. Compared with the engineers determining the design parameters of the chip according to experience, the limitations of the traditional design method are solved through data-driven and system-level verification, and it can cope with the chip design in rich and complex scenarios, and a chip that more meets the requirements of the real application scenario can be designed.

[0028] Optionally, the adjusting the first configuration parameter group according to the first network performance includes: Step S11: Determine the input traffic of each routing node of the on-chip network model; Step S12: Determine at least one first routing node according to the input traffic and a preset threshold; Step S13: For each first routing node, adjust the first configuration parameter corresponding to the first routing node.

[0029] It should be noted that the on-chip network model in the embodiments of the present application adopts a lossless on-chip network. A lossless on-chip network is an interconnection architecture specially designed for internal chip communication, and its core goal is to completely avoid packet loss during data transmission. In the steady-state transmission without data loss, that is, when the buffer is not saturated and there is no backpressure signal, the long-term average input traffic of the routing node is equal to the output traffic. When the buffer of the output port or the downstream node of a certain routing node is full, the node will send a "pause" signal to the upstream node to prevent data from flowing in continuously, avoiding packet loss caused by buffer overflow. The input traffic refers to the amount of data received and forwarded by each routing node, such as the number of data packets and the number of bytes.

[0030] The routing node with a traffic load significantly higher than the average level of the on-chip network model, for example, the preset threshold is 2 times the average traffic. By monitoring the input traffic of each routing node in the on-chip network model in real time, when the input traffic of the routing node is greater than or equal to the preset threshold, the routing node is determined as the first routing node. Among them, the preset threshold can be static or dynamic. For example, the preset threshold can be set according to experience and remains unchanged when adjusting the first configuration parameter; it can also be adjusted in real time according to the input traffic of each routing node in the on-chip network model. Usually, the first routing node is located in the middle area of the entire on-chip network model, that is, on the middle common link.

[0031] Adjusting the first configuration parameter corresponding to the first routing node refers to adjusting the network physical structure parameter corresponding to the first routing node, controlling the traffic passing through the first routing node, adjusting the transmission path of the data packet, reducing the traffic load of the first routing node, and adjusting the network resources allocated to the first routing node.

[0032] In the embodiment of the present application, during the simulation process, the input traffic of each routing node in the on-chip network model is monitored in real time, and the routing node with the input traffic greater than or equal to the preset threshold is determined as the first routing node; among them, the preset threshold can be static and unchanged, or can be dynamically adjusted according to the real-time traffic in the on-chip network model. For each first routing node, the network resources allocated to the first routing node are adjusted, for example, the network physical structure and the transmission path of the data packet. High-fidelity traffic monitoring and threshold strategy can accurately perceive the traffic of each routing node, and balance the immediate effect and long-term stability to adjust the parameters to avoid waste of global resources caused by local hotspots.

[0033] Optionally, the adjusting the first configuration parameter corresponding to each first routing node includes: Step S21, determining a second routing node from the remaining first routing nodes according to the connection relationship between the first routing node and the remaining first routing nodes; data transmission between the second routing node and the first routing node does not need to be forwarded by the remaining routing nodes; Step S22, determining the link between the first routing node and the second routing node as the target link; Step S23, increasing the number of virtual channels corresponding to the target link.

[0034] Among them, the first routing node refers to the node with excessive input traffic, the second routing node refers to the routing node directly connected to the first routing node, and the communication between the two nodes does not need to be forwarded through an intermediate node. The target link refers to the physical direct link connecting the first routing node and the second routing node. A virtual channel (VC) refers to a logical channel divided on a single physical link, which is used to isolate traffic with different priorities or types.

[0035] For example, in a Mesh topology, (X, Y) is used to identify the position of a routing node in the topology structure, where X represents the position of the routing node in the horizontal direction and Y represents the position of the routing node in the vertical direction. The first routing node is the O routing node, and the position of the O routing node is (2, 2). The routing nodes directly connected to the O routing node are A(2, 3), B(2, 1), C(3, 2), and D(1, 2) respectively. Assuming that the C routing node (3, 2) and the A routing node (2, 3) are also identified as the first routing nodes, then the C routing node and the A routing node are the second routing nodes corresponding to the first routing node. The target links connecting the first routing node and the second routing node are L1: (2, 2) ↔ (3, 2) and L2: (2, 2) ↔ (2, 3) respectively. Assuming that the number of virtual channels corresponding to L1 is 2, namely VC0 and VC1, and the number of virtual channels allocated to L1 is increased to 4, namely VC0, VC1, VC2, and VC3.

[0036] In the embodiment of the present application, according to the connection relationship between the first routing node and the remaining first routing nodes, the remaining first routing nodes directly connected to the first routing node are determined as the second routing nodes, the physical link between the first routing node and the second routing node is determined as the target link, and the number of virtual channels allocated to the target link is increased. Through refined resource allocation, resources are concentrated in hot spots, such as virtual channels, to specifically solve local congestion problems and achieve significant performance improvement at a small cost.

[0037] Optionally, for each first routing node, adjusting the first configuration parameter corresponding to the first routing node includes at least one of the following: Step S31: Increase the storage capacity of the cache queue in the first routing node; Step S32: Adjust the transmission priority of the data packets in the first routing node; Step S33: Adjust the topology structure of the area corresponding to the first routing node.

[0038] Among them, the input traffic of the first routing node is relatively high, resulting in frequent overflows of the buffer queue, causing packet loss or retransmission, and increasing latency. By increasing the depth of the buffer queue or the number of buffer queues in the first routing node, the traffic tolerance of the first routing node is improved, and overflows are reduced. Exemplarily, the buffer queue depth of each input port of the first routing node is 4, that is, Buffer Depth = 4. During the traffic peak, the buffer occupancy rate reaches 100% and the overflow rate is 15%. The depth of the buffer queue of each input port of the first routing node is increased from 4 to 8, that is, Buffer Depth = 8.

[0039] Since the first routing node processes multiple types of traffic simultaneously, such as high-priority real-time control instructions and computing data, resource contention causes the latency of critical traffic to exceed the standard. Priority scheduling can be used to ensure the priority transmission of high-importance packets and reduce the latency of high-importance packets. Exemplarily, initially, all packets in the first routing node are not distinguished, and the default priority is 0; adjusting the transmission priority of packets in the first routing node means that different priorities can be set for different instructions according to the importance of the instructions, and the packets are transmitted in order of priority. For example, the real-time control instruction is assigned a priority of 3, the computing data has a priority of 1, and the storage access has a priority of 0.

[0040] It should be noted that the first routing node may also become a bottleneck due to physical path limitations, such as excessive hop counts or insufficient link bandwidth. Through local topology optimization, the transmission efficiency of the area where the first routing node is located is improved. Specifically, the topology structure of the area where the first routing node is located can be adjusted by inserting bypass links (Express Link), increasing the link width (Link Width), or changing to a high-dimensional topology (such as changing from Mesh to Torus). Adjusting the topology structure can include not only the connection relationship between each routing node but also the physical changes to the link path itself. Exemplarily, the on-chip network model is a 4x4 Mesh structure, and the node N7 (coordinates 2, 2) is connected to adjacent nodes through a 1GHz, 64-bit-wide link. Inserting a bypass link means adding a direct link between N7 and N10 (coordinates 3, 3), and the hop count is reduced from 2 to 1. Increasing the link width means increasing the link width between N7 and the adjacent node N8 from 64 bits to 128 bits. Changing the 3x3 area centered on N7 to a small-scale Fat-Tree optimizes global communication.

[0041] Optionally, in the embodiments of the present application, congestion awareness can also be performed in real time, such as monitoring the buffer occupancy rate of each node, dynamically adjusting the data transmission path, bypassing congested nodes, that is, bypassing the first routing node, and reducing the traffic load of the first routing node.

[0042] In the embodiments of the present application, the storage capacity of the cache queue in the first routing node is increased to improve the traffic tolerance of the first routing node, reduce overflows, increase throughput, and reduce latency. The transmission priorities of the data packets in the first routing node are adjusted to control the transmission of each data packet in the first routing node in the order of priority to avoid conflicts and congestion. The topological structure of the area corresponding to the first routing node is adjusted. For example, by inserting a bypass link, that is, adding a direct link to the link that originally needed to be redirected through other routing nodes, the number of hops is reduced, the latency is reduced, the link width is increased to increase the throughput, or a high-dimensional topology is used to improve the network performance of the entire on-chip network model through local optimization.

[0043] Optionally, determining the current first network performance of the on-chip network model includes: Step S41: For each routing node, determine the amount of data transmitted by the routing node; Step S42: Determine the port rate of the routing node according to the amount of data and the simulation time; Step S43: Determine the first network performance according to the port rates corresponding to each routing node.

[0044] Among them, within each round of simulation cycle, the total amount of data sent, received, and forwarded by each routing node is usually measured in bits (bit) or the number of data packets (Packets). The input / output data amount of each routing node can be recorded by a simulation tool, such as BookSim. The port rate refers to the data transmission rate of each physical port of the routing node, with the unit of Gb / s. During the entire simulation process, the routing node continuously transmits data. Therefore, the port rate = (the amount of data transmitted by the node × the average size of the data packet) / simulation time. Example: Node A transmits 500 data packets within a simulation time of 100 ns, and the average packet size = 64 B (512 bit), then the port rate = (500 × 512 bit) / 100 ns = 2560 Gb / s.

[0045] The port rates corresponding to each routing node are summarized to determine the first network performance of the entire on-chip network model. For example, the latency of the first on-chip network model can be represented by the average / maximum transmission time of the data packet from the source to the destination, and the throughput refers to the total amount of data successfully transmitted by the on-chip network model per unit time.

[0046] In the embodiments of the present application, for each routing node, the amount of data transmitted by the routing node is determined; according to the amount of data and the simulation time, the port rate of the routing node is determined; according to the port rates corresponding to the respective routing nodes, the first network performance is determined. By separately monitoring the amount of data transmitted by each routing node, the port rate of each routing node is calculated, and then, based on the port rates of each routing node, a summary is made to determine the first network performance of the entire on-chip network model. By monitoring the data transmission volume of a single routing node to evaluate the network performance of the overall model, it essentially utilizes the economy of local observation and the correlation of the network topology to balance among hardware resources, monitoring latency, and accuracy, solves the problem of high overhead of traditional global monitoring, and at the same time provides data support for local dynamic routing optimization and resource allocation.

[0047] Optionally, the designing of the target chip according to the second configuration parameter group corresponding to the at least one data pattern includes: Step S51: Determine the target data pattern according to the usage ratio of each data pattern; Step S52: Determine the target configuration parameter group according to the second configuration parameter group corresponding to the target data pattern; Step S53: Design the target chip according to the target configuration parameter group.

[0048] In the embodiments of the present application, sorting is performed according to the usage ratio of each data pattern, and at least one target data pattern is determined for the data patterns greater than or equal to the preset ratio threshold. Alternatively, the target data pattern can be screened according to the core function of the chip and the importance of the pattern. Specifically, the data pattern with the highest usage ratio can be determined as the target data pattern, and the second configuration parameter group corresponding to the target data pattern can be determined as the target configuration parameter group. According to the target configuration parameter group, the hardware parameters in the process of designing the target chip are determined. For example, the target configuration parameter group can be directly determined as the hardware parameters of the target chip, or the target configuration parameter group can be fine-tuned to meet specific application requirements. Alternatively, multiple data patterns can be determined as the target data patterns, the second configuration parameter groups of the respective target data patterns are extracted, and the target configuration parameter group is obtained through weighted or constrained combination.

[0049] Exemplarily, there are three data modes for the autonomous driving chip, namely the sensor fusion mode (60% of the time), the path planning mode (30% of the time), and the emergency control mode (10% of the time). The target data mode is determined according to the usage ratio and the important mode, and the sensor fusion mode and the emergency control mode are determined as the target data modes. The second configuration parameter group corresponding to the sensor fusion mode is that the topology is 4x4 Mesh and bypass links, the adaptive routing algorithm, 3 priority channels, and the buffer depth is 8. The second configuration parameter group corresponding to the emergency control mode is the static XY routing algorithm and dedicated low-latency paths, with 20% bandwidth reservation. The target configuration parameter group is to retain the bypass links and the adaptive routing for sensor fusion coverage, configure dedicated paths and bandwidth reservation for emergency control, and set the buffer depth to 8 to balance the requirements of both data modes.

[0050] In the embodiments of the present application, according to the usage ratio of the data mode, the data mode with the highest usage ratio is determined as the target data mode, or the target data mode is determined according to the usage ratio, core functions, and importance of the mode. The second configuration parameter group of the target data mode with the highest usage ratio is directly determined as the target configuration parameter group, or the second configuration parameter groups corresponding to at least two target data modes are weighted and summed, and constraint merged to balance the performance of multiple data modes and determine the target configuration parameter group. According to the target configuration parameter group, the target chip is designed, which can optimize resources and save costs, and improve network performance and efficiency. Parameter fusion can also be used to make the chip adapt to diverse workloads and avoid the limitations of single-mode optimization.

[0051] In summary, the chip design method provided by the embodiments of the present application can construct an initial on-chip network model according to the computing power requirements of the project to be processed, and obtain the data traffic corresponding to multiple data patterns that may exist in the project to be processed. The data traffic corresponding to each data pattern is sequentially input into the on-chip network model, and each data pattern is independently simulated. During the simulation process, the network performance of each routing node in the on-chip network model is evaluated, and the routing node with a relatively large transmission traffic in the entire on-chip network model is determined as a hot spot. For the configuration parameters related to the hot spot, for example, increasing the storage capacity of the cache queue of the hot spot, increasing the number of virtual channels of the link where the hot spot is located, etc., to improve the throughput of the on-chip network model and reduce the latency of the on-chip network model, that is, to improve the network performance of the on-chip network model until the network performance meets the preset conditions corresponding to the data pattern, and obtain the optimal parameter combination under this data pattern. After obtaining the optimal parameter combination corresponding to each data pattern, the priorities corresponding to each data pattern in the project to be processed are weighed. Furthermore, according to the priorities corresponding to each data pattern, the chip is designed according to multiple sets of second configuration parameter groups. By constructing an on-chip network model and independently simulating multiple data patterns, under limited conditions, the optimal configuration parameter group of the on-chip network model in each data pattern is determined. Further, according to multiple optimal configuration parameter groups, the final configuration parameter group is determined. According to the final configuration parameter group, the chip is designed. Compared with engineers determining the design parameters of the chip based on experience, the limitations of the traditional design method are solved through data-driven and system-level verification, and it can handle the chip design of rich and complex scenarios, and can design a chip that more meets the requirements of the real application scenario.

[0052] Device embodiment Referring to Figure 2 , a structural block diagram of a chip design device according to the present invention is shown. The device may specifically include: An acquisition module 210, configured to acquire an on-chip network model corresponding to a project to be processed, and the data traffic of the on-chip network model in at least one data pattern; A simulation module 220, configured to inject the data traffic into the on-chip network model for simulation for each data pattern, and determine the current first network performance of the on-chip network model; the model parameters of the on-chip network model adopt a first configuration parameter group; An adjustment module 230, configured to adjust the first configuration parameter group according to the first network performance, and enter the next round of simulation until the first network performance meets the preset conditions, stop the simulation, and determine the current first configuration parameter group as the second configuration parameter group corresponding to the data pattern; A design module 240, configured to design a target chip according to the second configuration parameter groups corresponding to the at least one data pattern.

[0053] Optionally, the adjustment module includes: A traffic determination module, configured to determine the input traffic of each routing node of the on-chip network model; A node determination module, configured to determine at least one first routing node according to the input traffic and a preset threshold; An adjustment sub-module, configured to adjust the first configuration parameter corresponding to each first routing node for each first routing node.

[0054] Optionally, the adjustment sub-module includes: A node determination sub-module, configured to determine a second routing node from the remaining first routing nodes according to the connection relationship between the first routing node and the remaining first routing nodes; data transmission between the second routing node and the first routing node does not need to be forwarded by the remaining routing nodes; A link determination module, configured to determine the link between the first routing node and the second routing node as a target link; A first increase module, configured to increase the number of virtual channels corresponding to the target link.

[0055] Optionally, the adjustment sub-module includes: A second increase module, configured to increase the storage capacity of the cache queue in the first routing node; A first adjustment unit, configured to adjust the transmission priority of data packets in the first routing node; A second adjustment unit, configured to adjust the topological structure of the area corresponding to the first routing node.

[0056] Optionally, the simulation module includes: A data volume determination module, configured to determine the data volume transmitted by each routing node for each routing node; A port rate determination module, configured to determine the port rate of the routing node according to the data volume and the simulation time; A performance determination module, configured to determine the first network performance according to the port rates corresponding to each routing node.

[0057] Optionally, the design module includes: A mode determination module, configured to determine a target data mode according to the usage ratio of each data mode; A configuration parameter group determination module, configured to determine a target configuration parameter group according to the second configuration parameter group corresponding to the target data mode; A design sub-module, configured to design a target chip according to the target configuration parameter group.

[0058] In summary, the chip design device provided by the embodiments of the present application can construct an initial on-chip network model according to the computing power requirements of the project to be processed, and obtain the data traffic corresponding to multiple data patterns that may exist in the project to be processed. The data traffic corresponding to each data pattern is sequentially input into the on-chip network model, and each data pattern is independently simulated. During the simulation process, the network performance of each routing node in the on-chip network model is evaluated, and the routing node with a relatively large transmission traffic in the entire on-chip network model is determined as the hot spot. For the configuration parameters related to the hot spot, for example, increasing the storage capacity of the cache queue of the hot spot, increasing the number of virtual channels of the link where the hot spot is located, etc., to improve the throughput of the on-chip network model and reduce the latency of the on-chip network model, that is, to improve the network performance of the on-chip network model until the network performance meets the preset conditions corresponding to the data pattern, and the optimal parameter combination under this data pattern is obtained. After obtaining the optimal parameter combinations corresponding to each data pattern, the priorities corresponding to each data pattern in the project to be processed are weighed. Furthermore, according to the priorities corresponding to each data pattern, the chip is designed according to multiple groups of second configuration parameter groups. By constructing the on-chip network model and independently simulating each of the multiple data patterns, under limited conditions, the optimal configuration parameter groups of the on-chip network model under each data pattern are determined. Further, according to multiple optimal configuration parameter groups, the final configuration parameter group is determined. According to the final configuration parameter group, the chip is designed. Compared with the engineer determining the design parameters of the chip based on experience, the limitations of the traditional design method are solved through data-driven and system-level verification, and it can handle the chip design in rich and complex scenarios, and can design a chip that more meets the requirements of the real application scenario.

[0059] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, please refer to the partial description of the method embodiment.

[0060] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0061] Regarding the processor in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0062] Refer to Figure 3 , which is the structural block diagram of an electronic device for chip design provided by the embodiments of the present invention. As Figure 3As shown in the figure, the electronic device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus. The memory is used to store executable instructions, and the executable instructions cause the processor to execute the chip design method of the foregoing embodiment.

[0063] The processor may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable devices, transistor logic devices, hardware components, or any combination thereof. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0064] The communication bus may include a path for transmitting information between the memory and the communication interface. The communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 3 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0065] The memory may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), magnetic tape, a floppy disk, and optical data storage devices, etc.

[0066] An embodiment of the present invention also provides a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device (server or terminal), the processor can execute Figure 1 the chip design method shown.

[0067] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0068] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0069] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0070] These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing terminal devices to work in a predictive manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal devices, so that a series of operation steps are executed on the computer or other programmable terminal devices to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable terminal devices provide for implementing the functions in Figure 1 one flow or multiple flows and / or blocksFigure 1 Steps of functions specified in one or more boxes.

[0072] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

[0073] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.

[0074] The above has introduced in detail a chip design method, device, electronic device and readable storage medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A chip design method, characterized in that, The method includes: Obtaining an on-chip network model corresponding to a project to be processed, and data traffic of the on-chip network model in at least one data mode; For each data mode, injecting the data traffic into the on-chip network model for simulation, and determining the current first network performance of the on-chip network model; the model parameters of the on-chip network model adopt a first configuration parameter group; Adjusting the first configuration parameter group according to the first network performance, and entering the next round of simulation. Stop the simulation until the first network performance meets the preset conditions, and determine the current first configuration parameter group as the second configuration parameter group corresponding to the data mode; Designing a target chip according to the second configuration parameter groups corresponding to the at least one data mode.

2. The method according to claim 1, wherein The adjusting the first configuration parameter group according to the first network performance includes: Determining the input traffic of each routing node of the on-chip network model; Determining at least one first routing node according to the input traffic and a preset threshold; For each first routing node, adjusting the first configuration parameter corresponding to the first routing node.

3. The method according to claim 2, wherein The adjusting the first configuration parameter corresponding to the first routing node for each first routing node includes: Determining a second routing node from the remaining first routing nodes according to the connection relationship between the first routing node and the remaining first routing nodes; data transmission between the second routing node and the first routing node does not need to be forwarded by the remaining routing nodes; Determining the link between the first routing node and the second routing node as the target link; Increasing the number of virtual channels corresponding to the target link.

4. The method according to claim 2, wherein The adjusting the first configuration parameter corresponding to the first routing node for each first routing node includes at least one of the following: Increasing the storage capacity of the cache queue in the first routing node; Adjusting the transmission priority of data packets in the first routing node; Adjusting the topological structure of the area corresponding to the first routing node.

5. The method according to claim 1, wherein The determining the current first network performance of the on-chip network model includes: For each routing node, determining the amount of data transmitted by the routing node; Determining the port rate of the routing node according to the amount of data and the simulation time; Determining the first network performance according to the port rates corresponding to each routing node.

6. The method according to claim 1, wherein The designing a target chip according to the second configuration parameter groups corresponding to the at least one data mode includes: Determining a target data mode according to the usage ratio of each data mode; Determining a target configuration parameter group according to the second configuration parameter group corresponding to the target data mode; Designing a target chip according to the target configuration parameter group.

7. A chip design device, characterized in that, The device includes: An obtaining module, configured to obtain an on-chip network model corresponding to a project to be processed, and data traffic of the on-chip network model in at least one data mode; A simulation module, configured to, for each data mode, inject the data traffic into the on-chip network model for simulation, and determine the current first network performance of the on-chip network model; the model parameters of the on-chip network model adopt a first configuration parameter group; An adjustment module, configured to adjust the first configuration parameter group according to the first network performance, and enter the next round of simulation. The simulation is stopped until the first network performance meets a preset condition, and the current first configuration parameter group is determined as the second configuration parameter group corresponding to the data mode; A design module, configured to design a target chip according to the second configuration parameter group corresponding to the at least one data mode.

8. The device according to claim 7, characterized in that, The adjustment module includes: A traffic determination module, configured to determine the input traffic of each routing node of the on-chip network model; A node determination module, configured to determine at least one first routing node according to the input traffic and a preset threshold; An adjustment sub-module, configured to adjust the first configuration parameter corresponding to the first routing node for each first routing node.

9. An electronic device, characterized in that, The electronic device includes a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store executable instructions, and the executable instructions cause the processor to execute the chip design method according to any one of claims 1 to 6.

10. A readable storage medium, characterized in that, When the instructions in the readable storage medium are executed by the processor of the electronic device, the processor is enabled to execute the chip design method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Simulation model generation method and device of network-on-chip, electronic equipment and computer readable storage medium

    CN115114755A

  • Network-on-chip simulation system for multi-core-particle combined chip

    CN115460128A

  • Network-on-chip simulation method and device, electronic equipment and storage medium

    CN119397720A

  • Traffic scheduling simulation method and device, equipment and storage medium

    CN119402369A

  • Transactional traffic specification for network-on-chip design

    US20150358211A1