A method and device for performance monitoring and performance optimization of a network on chip
By setting three thresholds to monitor and analyze the real-time data transmission status of the bus system, we can judge the congestion cause of the AI chip bus system, and formulate an optimization plan to solve the problem of inaccurate positioning of congestion in the existing technology, and realize the relief of data transmission pressure without reducing performance and ensuring high-performance computing of the AI chip.
Patent Information
- Application Number
- CN202510649156.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The prior art cannot accurately locate the causes of local transmission congestion in the AI chip bus system, resulting in reducing the chip's computing power when alleviating data transmission pressure.
By setting three thresholds, the real-time data transmission status of the bus system is monitored, the transmission performance parameters of the master component, slave component and NoC routes are analyzed, the causes of congestion are judged, and the optimization plan is formulated, including adjusting the frequency and priority of data requests to alleviate stress.
Accurately locate the causes of bus system congestion, formulate reasonable optimization plans, alleviate transmission pressure without reducing system performance, and ensure the high-performance computing capabilities of AI chips.
Smart Images

Figure CN120186099B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of AI chips, and particularly relates to a method and device for monitoring and optimizing the performance of a network on chip. Background Art
[0002] In recent years, with the rapid development of semiconductor processes and technologies, the application scenarios related to artificial intelligence (AI) technology have been continuously enriched, and the demand for the high-performance computing capabilities of AI chips has been increasing day by day. A network on chip (NoC) is the core of information interaction in an AI chip, and data interaction between chip components needs to be transmitted and forwarded through the NoC. Therefore, a high-performance NoC bus system is the key to achieving high-bandwidth and low-latency data processing in an AI chip, which can significantly improve the high-performance computing capabilities of the AI chip. Usually, when an AI chip performs high-performance computing, a large number of data processing requests are sent by multiple core components inside it. The NoC needs to transmit and forward the data request operations. The instantaneous high-density data requests often easily cause excessive bus data transmission pressure in the local routing of the NoC, resulting in an increase in the data transmission delay and a decrease in the bandwidth of the NoC, causing the problem of data transmission congestion on the bus system, and ultimately seriously affecting the real-time computing performance of the AI chip.
[0003] Currently, the mainstream methods for monitoring and optimizing the performance of the bus system are to monitor the number of outstanding commands, real-time bandwidth, and latency on the interface between the master component (main device component) in the system and the system bus, compare the bandwidth or latency with the transmission threshold set by the system, and perform backpressure processing on the master component after exceeding the threshold. By restricting the send request operations of the master component, the data processing pressure on the bus system is reduced. While this operation reduces the data transmission pressure on the bus system, it also reduces the computing capabilities of the AI chip. This solution does not deeply analyze the problem of local transmission congestion in the bus system and does not determine the cause of the local transmission congestion problem in the bus system.
[0004] In summary, there is an urgent need for a new network on chip performance monitoring and performance optimization solution to monitor the transmission pressure of the bus system, accurately locate whether the transmission congestion in the bus system comes from the system master component, the NoC local routing, or the system slave component (slave device component), and at the same time, formulate a reasonable bus performance optimization solution to relieve the transmission pressure of the bus system without reducing the performance of the AI chip. Summary of the Invention
[0005] In view of the above problems, the object of the present invention is to provide a method and device for performance monitoring and performance optimization of a network-on-chip, aiming to solve the technical problem that in the existing technical solution, the bus system of the AI chip cannot accurately locate the local transmission congestion of the bus system.
[0006] The present invention adopts the following technical solutions:
[0007] On the one hand, the method for performance monitoring and performance optimization of the network-on-chip includes the following steps:
[0008] Step S1: Configure the data flow operation request sent by the master component to the specified slave component according to the current application scenario of the bus system;
[0009] Step S2: Configure the priority parameter of the data flow operation request sent by the master component on the NoC route according to the current application scenario of the bus system;
[0010] Step S3: Configure the data flow transmission request frequency parameters of the master component, slave component and NoC route according to the current application scenario of the bus system, and start the simulation;
[0011] Step S4: Monitor the real-time data flow on the bus interfaces of the master component and the slave component, and obtain the real-time transmission performance parameters of the master component and the slave component;
[0012] Step S5: Analyze and judge whether the bus transmission performance of the master component of the current data request operation meets the first threshold of the current application scenario according to the monitored real-time transmission performance parameters. If it meets, continue to monitor;
[0013] Step S6: If it does not meet the first threshold, judge whether the bus transmission performance of the slave component of the current data request operation meets the second threshold of the current application scenario;
[0014] Step S7: If it still does not meet the second threshold of the slave component, determine that one of the reasons for the bus system transmission congestion comes from the slave component, adjust the operation request of the current master component to the slave component whose bus transmission performance is less than the set second threshold, and then return to Step S4 to continue monitoring;
[0015] Step S8: If it meets the second threshold of the slave component, calculate the path transmission performance parameter of the data request operation on the NoC path according to the real-time transmission performance parameters of the current data request operation;
[0016] Step S9: Judge whether the path transmission performance parameter meets the third threshold of the current application scenario;
[0017] Step S10: If the third threshold is not met, it is determined that one of the reasons for the bus system transmission congestion comes from the NoC routing. Configure to increase the data transmission priority of the master component in the on-chip network, and then return to step S4 to continue monitoring;
[0018] Step S11: If the third threshold is met, it is determined that one of the reasons for the bus system transmission congestion comes from the master component. Configure the data flow transmission request frequency parameters of the master component, slave component, and NoC routing, and increase the transmission clock frequency of the data request on the chip bus system to achieve overclock transmission of the data flow in the on-chip network.
[0019] On the other hand, the on-chip network performance monitoring and performance optimization device is applied to the bus system. The bus system has NoC routings connected by multiple paths, and several master components and several slave components are mounted on the bus system. The device includes:
[0020] A parameter configuration module, which is used to configure the data flow operation request sent by the master component to the specified slave component according to the current application scenario of the bus system, configure the priority parameter of the data flow operation request sent by the master component on the NoC routing, and configure the data flow transmission request frequency parameters of the master component, slave component, and NoC routing, and start the simulation;
[0021] A performance monitoring module, which is used to monitor the real-time data flow on the bus interfaces of the master component and the slave component, and obtain the real-time transmission performance parameters of the master component and the slave component;
[0022] An analysis and optimization module is used to analyze and determine whether the bus transmission performance of the master component in the current data request operation meets the first threshold of the current application scenario according to the monitored real-time transmission performance parameters. If it meets, continue to monitor; if it does not meet the first threshold, then determine whether the bus transmission performance of the slave component in the current data request operation meets the second threshold of the current application scenario; if it still does not meet the second threshold of the slave component, it is determined that one of the reasons for the bus system transmission congestion comes from the slave component, adjust the operation request of the current master component to the slave component whose bus transmission performance is less than the set second threshold, and then return to continue monitoring; if it meets the second threshold of the slave component, calculate the path transmission performance parameter of the data request operation on the NoC path according to the real-time transmission performance parameters of the current data request operation; determine whether the path transmission performance parameter meets the third threshold of the current application scenario; if it does not meet the third threshold, it is determined that one of the reasons for the bus system transmission congestion comes from the NoC routing, configure to improve the data transmission priority of the master component in the on-chip network, and then return to continue monitoring; if it meets the third threshold, it is determined that one of the reasons for the bus system transmission congestion comes from the master component, configure the data stream transmission request frequency parameters of the master component, slave component and NoC routing, increase the transmission clock frequency of the data request on the chip bus system, and realize the overclock transmission of the data stream in the on-chip network.
[0023] The beneficial effects of the present invention are as follows: By setting three thresholds for the master component, slave component and NoC routing, and analyzing the real-time data transmission status of the bus system to obtain the corresponding transmission performance parameters, and then through logical judgment and analysis, comparing with various set thresholds in sequence, it can accurately locate whether the bus system transmission congestion comes from the master component, slave component or NOC local routing. At the same time, a reasonable transmission performance optimization scheme is formulated to relieve the bus system transmission performance pressure without reducing the system performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flowchart of the method for monitoring and optimizing the performance of the on-chip network provided by the embodiment of the present invention;
[0025] Figure 2 is a topological diagram of an application scenario example provided by the embodiment of the present invention;
[0026] Figure 3 is a bus request and response timing diagram;
[0027] Figure 4 is a data request and response timing diagram;
[0028] Figure 5 It is a structural schematic diagram of a performance monitoring and performance optimization device for a network-on-chip provided by an embodiment of the present invention. Detailed implementation manners
[0029] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0030] In order to illustrate the technical solutions described in the present invention, the following will be described through specific embodiments.
[0031] Embodiment 1:
[0032] As Figure 1 shown, the performance monitoring and performance optimization method for the network-on-chip provided in this embodiment includes the following steps:
[0033] Step S1: Configure the data stream operation requests sent by the master component to the specified slave component according to the current application scenario of the bus system.
[0034] Step S2: Configure the priority parameters of the data stream operation requests sent by the master component on the NoC routing according to the current application scenario of the bus system.
[0035] Step S3: Configure the data stream transmission request frequency parameters of the master component, slave component and NoC routing according to the current application scenario of the bus system, and start the simulation.
[0036] The above steps S1 - S3 are the parameter configuration phase.
[0037] The application scenario described in this embodiment refers to a high-performance application scenario in which the chip bus system is performing a performance limit test, including several master components, slave components and NoC routings, etc. Each master component is sequentially connected to each slave component through several NoC routings, and a NoC path is formed between the master component and the slave component. By configuring the working parameters of the relevant master component, NoC routing and slave component, the limit performance test of the chip core components (master component and slave component) and the local NoC routing is realized to verify the performance monitoring and performance optimization functions of the chip bus system.
[0038] As Figure 2 shown, an application scenario example shows a specific connection relationship between the core components and the local NoC routing.
[0039] In the illustration, M0, M1, and M2 are master components, S0, S1, S2, and S3 are slave components, and Router0, Router1, Router2, Router3, Router4, and Router5 are local NoC routers. During the high-performance test of the chip, the master components will send a large number of data request operations to the slave components through the NoC routers. For example, if most of the data request operations are concentrated on S1 and S2, as the number of data request operations of M0, M1, and M2 continues to increase, the path delays of the total line paths P0->P3->P6->P9, P1->P4->P10, P0->P3->P7->P9, and P2->P5-P10 will gradually increase, and at the same time, the response delays of S1 and S2 will also gradually increase. If the performance of the bus system of the chip is not optimized, as the chip performance test time increases, it is very easy to have a transmission congestion problem in the local bus system, and ultimately lead to a serious decline in the performance of the chip.
[0040] Therefore, when implementing monitoring and optimization in this embodiment, first, it is necessary to configure the request relationship between the data operations of the master components and the corresponding target slave components, then configure the priority of the data flow operations, and then configure the transmission request frequency.
[0041] Step S4: Monitor the real-time data flow on the bus interfaces of the master components and the slave components, and obtain the real-time transmission performance parameters of the master components and the slave components.
[0042] This step is the performance monitoring stage, monitoring the real-time transmission performance parameters on the bus-related interfaces, mainly including real-time bandwidth, the number of outstanding commands, and transmission delay, etc., which can be used as the criteria to measure the speed of on-chip network data transmission. The calculation methods of each performance parameter are as follows:
[0043] The calculation method of the number of outstanding commands Outstanding is: Outstanding is the number of data request commands that have not received a return command among the sent data request commands. That is, for each sent data request, Outstanding is incremented by 1, and for each received data response, Outstanding is decremented by 1. When the sent data request and the data response are received simultaneously, Outstanding remains unchanged, and Outstanding is dynamically counted.
[0044] The calculation method of the bus data transmission delay Latency is: Latency = request_interval_time + transaction_response_time. Here, request_interval_time is the request interval time, and transaction_response_time is the request transaction response time. For example Figure 3In the bus request and response timing diagram shown, taking the calculation of the Latency of bus request ID1 as an example, in each data request sent on the AXI bus interface, when calculating the Latency of bus request ID1, the completion time of ID0 request sending is used as the start time of the ID1 request interval time, the start time of ID1 request sending is used as the end time of the ID1 request interval time and the start time of the ID1 request transaction response time, the response received by the bus for ID1 is used as the end time of the ID1 request transaction response time, and the sum of the request interval time and the request transaction response time is calculated as the transmission delay Latency of the ID1 data request transaction.
[0045] The calculation method of the real-time bandwidth Bandwidth is: Bandwidth = transaction_data / Latency, where transaction_data is the number of bits of each data request, and transaction_data = (len + 1) * 2 size , len is the length of each data request, and size is the data bit width of each data request. In the bus system, the length len of the data requests sent by the master component is usually not a fixed value, which will cause the calculated real-time bandwidth to fluctuate greatly due to the change of the length len, and it is impossible to quantitatively judge whether the bus system meets the performance requirements through the real-time bandwidth parameter. Therefore, in this embodiment, the concept of real-time average bandwidth is added on the basis of the real-time bandwidth. The real-time average bandwidth Average_Bandwidth = Bandwidth / (len + 1), that is, 2 size / Latency, which eliminates the influence of the length len in the data request operation on the real-time bandwidth.
[0046] Step S5: According to the monitored real-time transmission performance parameters, analyze and judge whether the bus transmission performance of the master component of the current data request operation meets the first threshold of the current application scenario. If it meets, continue to monitor.
[0047] In this step, the judgment criterion for whether it meets the first threshold is: whether both the transaction delay and the real-time average bandwidth of the master component sending a request to the bus meet the corresponding first thresholds, that is, specifically, these two performance parameters of transaction delay and real-time average bandwidth are used for judgment.
[0048] First analyzed here is that the master component sends a data request to the bus, and the bus responds. The corresponding transaction latency and real-time average bandwidth of the master component are calculated according to the calculation method shown in step S4, which will not be elaborated here. There are two first thresholds, corresponding to the transaction latency and the real-time average bandwidth respectively. When at least one of the two performance parameters monitored does not meet the first threshold condition, it is determined that the first threshold of the current application scenario is not met. If both performance parameters meet the corresponding first threshold conditions, it is determined that the first threshold of the current application scenario is met.
[0049] Step S6: If the first threshold is not met, determine whether the bus transmission performance of the slave component for the current data request operation meets the second threshold of the current application scenario.
[0050] If the first threshold is not met, further determine whether the performance parameters of the slave component meet the second threshold of the previous application scenario. Similarly, after the slave component receives the bus response, the slave component returns the content corresponding to the data request to the bus, and is also calculated according to the calculation method shown in step S4, which will not be elaborated here either, and the corresponding transaction latency and real-time average bandwidth of the slave component can also be obtained. These two performance parameters are compared with the two set second thresholds.
[0051] Step S7: If the second threshold of the slave component is still not met, it is determined that one of the reasons for the bus system transmission congestion comes from the slave component. Adjust the operation request of the current master component to the slave component whose bus transmission performance is less than the set second threshold, and then return to step S4 to continue monitoring.
[0052] If at least one of the two performance parameters does not meet the second threshold condition, it is determined that the second threshold of the current application scenario is not met. At this time, it can be located that one of the reasons for the bus system transmission congestion comes from the slave component. Then, during optimization, first, according to the method of step S1, adjust the operation request of the current master component to the slave component whose bus transmission performance is less than the set second threshold, and then continue to return to step S4 for monitoring.
[0053] Step S8: If the second threshold of the slave component is met, calculate the path transmission performance parameter of the data request operation on the NoC path according to the real-time transmission performance parameter of the current data request operation.
[0054] If the condition of the second threshold is met, further calculate the path transmission performance parameter of the data request operation on the NoC path. As Figure 4 shown in the complete data request and response timing diagram, the data request is sent from the master component and forwarded through the NoC router to the slave component to obtain the data response.
[0055] In the illustration, the request data of the master component has an ID number and an address. Through the ID number and the address, the ID number and the address of the response data of the slave component bus interface can be matched, so that the path transmission performance parameters of the data response operation of the slave component corresponding to the same data request operation of the master component on the bus system can be obtained. In this embodiment, it includes the path delay NoC_latency and the delay average bandwidth NoCAverage_Bandwidth.
[0056] As can be seen from the figure, the transaction delay of the master component is Tmaster_latency, and the transaction delay of the slave component for data response is Tslave_latency. The path delays of the data request transaction on the NoC are the NoC request delay and the NoC response delay respectively. Therefore, the path delay NoC_latency of the master component data request on the NoC routing is NoC_latency = Tmaster_latency - Tslave_latency. Referring to the above calculation method of the average real-time bandwidth, the delay average bandwidth NoCAverage_Bandwidth of the data request on the NoC routing can be calculated as NoCAverage_Bandwidth = 2 size / NoC_latency.
[0057] Step S9: Determine whether the path transmission performance parameters meet the third threshold of the current application scenario.
[0058] Similarly, the judgment criterion for whether it meets the third threshold is: whether both the path delay and the delay average bandwidth of the request data on the NoC routing meet the corresponding third thresholds.
[0059] Step S10: If the third threshold is not met, it is determined that one of the reasons for the transmission congestion of the bus system comes from the NoC routing. Configure to increase the data transmission priority of the master component in the on-chip network, and then return to step S4 to continue monitoring.
[0060] If the third threshold is not met, that is, at least one of the two performance parameters does not meet the corresponding third threshold, it can be located that one of the reasons for the transmission congestion of the bus system comes from the NoC routing. When optimizing, first configure to increase the data transmission priority of the master component on the NoC routing according to the method of step S2, and then continue to return to step S4 for monitoring.
[0061] Step S11: If the third threshold is met, it is determined that one of the reasons for the transmission congestion in the bus system comes from the master component. Configure the data stream transmission request frequency parameters of the master component, slave component, and NoC routing, increase the transmission clock frequency of data requests on the chip bus system, and achieve overclock transmission of the data stream on the on-chip network.
[0062] If the condition of the third threshold is met, one of the reasons for the congestion on the bus system can be located as coming from the master component. When optimizing, first set the data stream transmission request frequency parameters of the master component, slave component, and NoC according to the method in Step S3, increase the transmission clock frequency of data requests on the bus system, achieve overclock transmission of the data stream on the NoC, and relieve the transmission pressure of the bus system in high-performance application scenarios.
[0063] The above Steps S5 - S11 are the analysis, judgment, and optimization adjustment stages based on the monitoring results. Through the three-level comparison result of the real-time transmission bandwidth of the bus and the set threshold, the real-time data transmission state of the bus system can be truly reflected, and it can accurately locate whether the transmission congestion of the bus system comes from the master component, the NOC local routing, or the slave component. At the same time, a reasonable transmission performance optimization plan is formulated to relieve the transmission performance pressure of the bus system without reducing the system performance. In addition, it should be noted that the first threshold, second threshold, and third threshold here are the optimal performance parameter values measured when a single master component sends data requests under the empty load of the bus system, and 60% of the optimal performance parameter value is taken as the corresponding threshold for performance judgment.
[0064] Embodiment 2:
[0065] As Figure 5 shown, this embodiment provides a performance monitoring and performance optimization device for an on-chip network, which is applied to a bus system. The bus system has multiple paths connected to NoC routing, and the bus system is mounted with several master components (illustrated as M0, M1 in the figure) and several slave components (illustrated as S0 - S3 in the figure). The device includes:
[0066] A parameter configuration module 101, configured to configure the data stream operation requests sent by the master component to the specified slave component, configure the priority parameters of the data stream operation requests sent by the master component on the NoC routing, and configure the data stream transmission request frequency parameters of the master component, slave component, and NoC routing, and start the simulation;
[0067] The performance monitoring module 102 is used to monitor the real-time data stream on the bus interfaces of the master component and the slave component, and obtain the real-time transmission performance parameters of the master component and the slave component;
[0068] The analysis and optimization module 103 is used to analyze and determine whether the bus transmission performance of the master component of the current data request operation meets the first threshold of the current application scenario according to the monitored real-time transmission performance parameters. If it meets, continue to monitor; if it does not meet the first threshold, then determine whether the bus transmission performance of the slave component of the current data request operation meets the second threshold of the current application scenario; if it still does not meet the second threshold of the slave component, it is determined that one of the reasons for the bus system transmission congestion comes from the slave component, and the operation request of the current master component is adjusted to the slave component whose bus transmission performance is less than the set second threshold, and then return to continue monitoring; if it meets the second threshold of the slave component, calculate the path transmission performance parameter of the data request operation on the NoC path according to the real-time transmission performance parameter of the current data request operation; determine whether the path transmission performance parameter meets the third threshold of the current application scenario; if it does not meet the third threshold, it is determined that one of the reasons for the bus system transmission congestion comes from the NoC routing, configure to improve the data transmission priority of the master component in the on-chip network, and then return to continue monitoring; if it meets the third threshold, it is determined that one of the reasons for the bus system transmission congestion comes from the master component, configure the data stream transmission request frequency parameters of the master component, the slave component and the NoC routing, increase the transmission clock frequency of the data request on the chip bus system, and realize the overclock transmission of the data stream in the on-chip network.
[0069] The three functional modules 101-103 of the device in this embodiment correspondingly implement the three stages in Embodiment 1. Specifically, module 101 correspondingly implements steps S1-S3, module 102 correspondingly implements step S4, and module 103 correspondingly implements steps S5-S11.
[0070] The parameter configuration module 101 is mainly used to configure working parameters. The performance monitoring module 102 respectively monitors and obtains the real-time performance parameters of the master component and the slave component on the NoC bus interface, and then transfers the real-time performance parameters of the master component and the slave component to the bus performance analysis and optimization module 103. The analysis and optimization module 103 caches the real-time performance parameters, and matches the real-time performance parameters corresponding to the data requests of the master component and the slave component according to the ID number and address information in the data request, so as to obtain the real-time bus performance parameters of the same data request transaction on the bus system, and calculates the real-time bus performance parameters of the data request transaction on the NoC path. The analysis and optimization module 103 then analyzes and judges the bus transmission performance, and sequentially judges whether the delays and real-time average bandwidths on the master component, the slave component and the NoC router meet the thresholds set by the bus system. For the data request transactions whose performance parameters do not meet the thresholds set by the bus system, analyze the reasons for bus congestion, and judge whether the congestion is caused by the NoC router, the slave component, or the master component.
[0071] After locating the specific congestion location, the parameter configuration module 101 generates corresponding configuration parameters for the specified functional component or NoC router, and completes the reconfiguration of the simulation parameters to implement the NoC transmission performance optimization operation, and then continues the aforementioned performance monitoring and performance optimization operations until the transmission congestion problem on the chip bus system is finally solved, so as to relieve the transmission performance pressure of the bus system without reducing the system performance and ensure the high-performance computing ability of the chip.
[0072] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for performance monitoring and performance optimization of a network on chip, characterized in that, The method includes the following steps: Step S1: Configure the data stream operation request sent by the master component to the specified slave component according to the current application scenario of the bus system; Step S2: Configure the priority parameter of the data stream operation request sent by the master component on the NoC route according to the current application scenario of the bus system; Step S3: Configure the data stream transmission request frequency parameters of the master component, slave component, and NoC route according to the current application scenario of the bus system, and start the simulation; Step S4: Monitor the real-time data stream on the bus interfaces of the master component and slave component, and obtain the real-time transmission performance parameters of the master component and slave component; Step S5: Analyze and judge whether the bus transmission performance of the master component of the current data request operation meets the first threshold of the current application scenario according to the monitored real-time transmission performance parameters. If it meets, continue to monitor; Step S6: If the first threshold is not met, judge whether the bus transmission performance of the slave component of the current data request operation meets the second threshold of the current application scenario; Step S7: If the second threshold of the slave component is still not met, determine that one of the reasons for the bus system transmission congestion comes from the slave component, adjust the operation request of the current master component to the slave component whose bus transmission performance is less than the set second threshold, and then return to Step S4 to continue monitoring; Step S8: If the second threshold of the slave component is met, calculate the path transmission performance parameter of the data request operation on the NoC path according to the real-time transmission performance parameters of the current data request operation; Step S9: Judge whether the path transmission performance parameter meets the third threshold of the current application scenario; Step S10: If the third threshold is not met, determine that one of the reasons for the bus system transmission congestion comes from the NoC route, configure to increase the data transmission priority of the master component in the on-chip network, and then return to Step S4 to continue monitoring; Step S11: If the third threshold is met, determine that one of the reasons for the bus system transmission congestion comes from the master component, configure the data stream transmission request frequency parameters of the master component, slave component, and NoC route, increase the transmission clock frequency of the data request on the chip bus system, and realize the overclock transmission of the data stream in the on-chip network.
2. The method for performance monitoring and performance optimization of the network-on-chip according to claim 1, characterized in that The application scenario is a high-performance application scenario in which the chip bus system is performing a performance limit test. Each master component is sequentially connected to each slave component through several NoC routes, and a NoC path is formed between the master component and the slave component.
3. The method for performance monitoring and performance optimization of the network-on-chip according to claim 2, wherein The real-time transmission performance parameters include real-time bandwidth, number of advanced commands, and transmission delay, where, The calculation method of the outstanding command number is as follows: Outstanding is the number of data request commands sent but not yet received with a response. That is, for each data request sent, Outstanding is incremented by 1, and for each data response received, Outstanding is decremented by 1. When both data requests and data responses are sent and received simultaneously, Outstanding remains unchanged, and Outstanding is dynamically calculated. The calculation method of the transmission latency is as follows: Latency = request_interval_time + transaction_response_time, where request_interval_time is the request interval time and transaction_response_time is the request transaction response time. The calculation method of real-time bandwidth is: Bandwidth = transaction_data / Latency. Here, transaction_data is the amount of bits for each data request, and transaction_data = (len + 1) * 2 size , where len is the length of each data request, and size is the data bit width of each data request.
4. The method for performance monitoring and performance optimization of the network-on-chip according to claim 3, wherein The real-time transmission performance parameter further includes the real-time average bandwidth. The real-time average bandwidth Average_Bandwidth = Bandwidth / (len + 1), that is, 2 size / Latency.
5. The method for performance monitoring and performance optimization of the network-on-chip according to claim 4, wherein The path transmission performance parameters described in step S8 include path delay and average delay bandwidth. The path delay NoC_latency = Tmaster_latency - Tslave_latency, where Tmaster_latency is the transaction delay of the master component and Tslave_latency is the transaction delay of the slave component; the average delay bandwidth NoCAverage_Bandwidth = 2 size / NoC_latency.
6. The method for performance monitoring and performance optimization of the network-on-chip according to claim 5, wherein In step S5, the criterion for determining whether the first threshold is met is whether both the transaction latency and the real-time average bandwidth when the master component sends a request to the bus meet the corresponding first threshold. In step S6, the criterion for determining whether the second threshold is met is whether both the transaction latency and the real-time average bandwidth when the slave component responds to the request to the bus meet the corresponding second threshold. In step S9, the criterion for determining whether the third threshold is met is whether both the path latency and the latency average bandwidth of the data on the NoC route meet the corresponding third threshold.
7. The method for performance monitoring and performance optimization of the network-on-chip according to claim 6, wherein, The first threshold, the second threshold, and the third threshold are the optimal performance parameter values measured when a single master component sends a data request under a no-load condition of the bus system, and 60% of the optimal performance parameter value is taken as the corresponding threshold for performance judgment.
8. A performance monitoring and performance optimization device for a network-on-chip, characterized in that, The device is applied to a bus system. The bus system has multiple NoC routes connected by multiple paths and is equipped with several master components and several slave components. The device includes: A parameter configuration module, which is used to configure the data stream operation requests sent by the master component to the specified slave component, configure the priority parameters of the data stream operation requests sent by the master component on the NoC route, and configure the data stream transmission request frequency parameters of the master component, the slave component, and the NoC route, and start the simulation. A performance monitoring module, which is used to monitor the real-time data stream on the bus interfaces of the master component and the slave component and obtain the real-time transmission performance parameters of the master component and the slave component. An analysis and optimization module is used to analyze and determine whether the bus transmission performance of the master component in the current data request operation meets the first threshold of the current application scenario according to the monitored real-time transmission performance parameters. If it meets the requirement, continue monitoring; if it does not meet the first threshold, then determine whether the bus transmission performance of the slave component in the current data request operation meets the second threshold of the current application scenario; if it still does not meet the second threshold of the slave component, it is determined that one of the reasons for the bus system transmission congestion comes from the slave component, and the operation request of the current master component is adjusted to the slave component whose bus transmission performance is less than the set second threshold, and then return to continue monitoring; if it meets the second threshold of the slave component, calculate the path transmission performance parameters of the data request operation on the NoC path according to the real-time transmission performance parameters of the current data request operation; determine whether the path transmission performance parameters meet the third threshold of the current application scenario; if they do not meet the third threshold, it is determined that one of the reasons for the bus system transmission congestion comes from the NoC routing, configure to increase the data transmission priority of the master component in the on-chip network, and then return to continue monitoring; if it meets the third threshold, it is determined that one of the reasons for the bus system transmission congestion comes from the master component, configure the data flow transmission request frequency parameters of the master component, slave component and NoC routing, increase the transmission clock frequency of the data request on the chip bus system, and realize the overclock transmission of the data flow in the on-chip network.
Citation Information
Patent Citations
Performance verification method and device of system on chip
CN115952074A
Conversion system from I2C to AXImaster
CN216956942U